Hacker News new | past | comments | ask | show | jobs | submit
I spent the last three days (off and on) using Gemini to configure my edge router 4 with my iOS devices on a vpn and it's been awesome. In the past I'd do a google search and read a few sources of documentation, do another google search and read another set of documentation. Now, Gemini aggregates multiple pages together so all of the work of reading source docs from multiple locations is now n a single step.

Oh, I should mention though. There was no advertising at all. They didn't make any money off me. It was 100% Gemini which I recognize as not long-term feasible.

All the information Gemini surfaced was created with human effort and published on the internet with the expectation that humans would visit the website and the creator would get some reward - advertising dollars, bragging rights, popularity, subscribers or whatever else.

If the only visitors to websites are now LLM training bots then what incentive is there to publish anything new? For how long can we continue to rely on pre-2024 non-AI generated content?

I've seen websites put up some draconian measures to try and get a grip on the scraping. So much for the sub-second loading experience when you have Cloudflare, Google, Anubis, and all these other captcha services trying to see if you're a human. It's made the web browsing experience so much worse.

Some of the proposals to address this include charging bots for access to web resources, but they will also have repercussions for regular users. I don't see how you solve this cleanly.

> Cloudflare, Google, Anubis, and all these other captcha services trying to see if you're a human.

Yep. IMO, this is so far the biggest AI-inflicted damage to the web. A bit of anecdata - wikipedia (and all other wikimedia sites) are blocking my Firefox since about a week, with a "please respect our bot policy" message. Outright block, not even a captcha.

It took me a while to figure out they don't like me disabling some SSL ciphers, so now "JA4 browser fingerprint" is not matching user-agent. Funnily enough curl (what I would imagine a bot would use) pulls exact same URLs from exact same client IP, just fine.

Sure - it sucks, unfortunately the alternative is the sites going away entirely. When the load from scraper bots is constantly knocking the site offline the choices are literally to allow it to remain inaccessible for much of the time, put up a layer of defenses with all the user-annoyance compromises that entails, or just give up and unpublish the site.
The alternative is simple.. Go dark. VPN tech is known from like 30 years. Pretty much everyone can use it (VPN providers). But instead using it to browse net, build VPN overlay networks of interest for people. Gaming networks, R&D networks, Retro Networks. People will peer to PoP and use resources. Bad actor? BAN it from network. You have control. This could be done in Internet, but big corpos and big money won the battle. Just wake F*ing up...
I think you'd struggle to keep LLM bots off the network unfortunately.

If it had any real value, anyway.

Small, truly private communities could be an interesting thing though.

Continuing on your suggestion.

There could be open source tooling to create custom private "closednets", with

- trust ring mechanism to allow invitations, flagging, banning, and banning those that invite people who were banned

- the rules of the closednet

- search engine with opt-in scraping

- portal (remember the 80s?) with all the registered nodes, perhaps by service category such as public git repo hosts, web sites etc.

etc.

The first closednet could be Hacker News.

It doesn't have to be an IP-layer network. A website that you need to log in to view works just as well.
Cloudflare specifically has a block for LLM and AI training bots now.

Not sure of the effectiveness but it's there.

Minimal. I'm behind Cloudflare and 90% of the traffic is still scrapers. I don't think they're serious about the long tail.

I think the main thing Cloudflare is trying to do is block direct traffic from frontier labs and then start charging them for access. They might end up shooting themselves in the foot, as this simply empowers sketchy residential-proxy outfits to undercut Cloudflare and sell the data to labs for less.

I think the other thing they're trying to do is get most of the internet to send them all of their cleartext traffic. Expect in 2040 the PRISM2 docs will get leaked by some Eduardo Rainedon and we'll find out Cloudflare was the NSA all along.
It still only blocks "well-behaved" bots that have proper User-Agents and respect robots.txt, so it's largely pointless.

The problematic bots are all disguising themselves as Chrome and sending requests from millions of residential proxy IPs, and the only real solution to those is some sort of captcha or PoW page on first visit.

Me, I'm just scraping the parts of the internet I like, toying with local LLMs… ready really to just shove off.
> If the only visitors to websites are now LLM training bots then what incentive is there to publish anything new?

I write because I have ideas I want to share, and whether that happens with LLMs as an intermediary isn't important to me.

The danger that's concerning people (rightly or wrongly) isn't that LLMs are going to be an intermediary to your website. It's that they'll be the only thing reading it. No one will ever read your post or know what you wrote. The only consumers will be LLMs, they'll train on a version that strips out you as the author (probably more due to expedience than any sort of malice; it's not like you're famous, are you?), and your idea might get embedded into a set of model weights somewhere. No human will see a byte of it.

Are you actually saying you'd be OK with that?

loading story #49256246
Yes, that would be fine. I write primarily communicate ideas, not for credit or fame.

Empirically, however, LLMs don't strip out the author: the big models know a lot about what I've written even with search disabled. Ex: https://claude.ai/share/8cbcdf88-a360-421a-8c06-ae7b7992e866

Sure they know about you if you ask, but generally they won't credit you if they cite an idea from their latent space that came from you.
Not the OP, but I suspect no one goes to my website anyway. (I write nonetheless.)
I do! You invented the crayon picker in MacOS..?
Your level of thinking is defined by ICD-10.

Yeah, authors don't want to be recognized as authors, they don't want any reward for their work, they don't want to amass pool of loyal readers, interact with them, etc.

All they want is for halucinating AI to take excerpts of their work and compile it with random sh!t.

GENIUS

Thanatic drive masquerading as transhumanist virtue signalling.
I mainly shared my projects for learning, discussion and bragging rights.

LLMs just use everything, generate similar code with no attribution and keep users from visiting, so no bragging rights or attention.

Worse, there are some PRs that seem fully generated ...

So i mostly stopped sharing and started pulling my old repos offline.

At this pace, i don't want to compete with a clone of myself in the future that will do my work for much cheaper.

There are different types of writing. If we depend on people writing because it's enjoyable at some level, we're going to lose writing that's important but also a bit tedious.
Of course! I do think we'd lose a lot of great writing if it went amateur-only. But my parent seemed to be saying the incentive would entirely disappear, so I wanted to give my perspective.
> I write because I have ideas I want to share, and whether that happens with LLMs as an intermediary isn't important to me.

Sure. But you can see that for some people (myself included), writing for peers is part of the joy? And that if instead a megacorp places an opaque computer program between the author and the readers, that joy might be ruined?

loading story #49256290
> "humans would visit the website and the creator would get some reward"

That expectation is a problem, has always been a problem, and Tim Berners Lee never mentioned anything about a reward structure when coming up with the WWW.

Your thinking too narrowly about the reward. Sometimes, it's just about the getting the knowledge out there that's motivating the creator, not anything tangible for themselves.
> it's just about the getting the knowledge out there that's motivating the creator

In that case the creator should welcome AIs with open arms; a human reader will forget eventually, but the AI will preserve the knowledge forever.

"the AI will preserve the knowledge forever"

no, only some mangled form of it

> Your thinking

Should be "You're thinking".

> Should be "You're thinking".

This should be "This should be 'You're thinking'." don't you think? Why bother correcting someone's grammar with a sentence fragment? You're just trading one mistake for another. I'm hoping someone finds a grammar error in my post, because continuing this would be hilarious.

> This should be "This should be 'You're thinking'." don't you think?

Reflexively, I think it should be more like ...

  javascript: `This should be "You're thinking".` ; 
  // to preserve the original character use and to avoid '...'...' parse foos
  // however `"...".` also possibly deserves a [sic] to critique the original 
  // i.e. ~grammar police say the period belongs within the quote marks, no?
... but then that's just me, in [my] quirks mode.
Did the operators of HN mention anything about a reward structure for posting comments here? I'm sure you can see how that's still attractive to some.
Well in the example above the manufacturer still has incentive to provide the manual's and guides that describe how to use their products, and if that is subsequently served by an LLM that's totally fine. The only sites that LLM's would have a negative effect on are those that are only hosting content for the ad views.
Manuals don't always well explain how to use their products with everybody else's products because there are too many to do that. But there are lots of people trying things out and might figure out the fine details on how to make various things work. They then publish these how-to pieces (which exist no where else) to the internet, or at least they used to when there were incentives to do so.
Easy to fix a well documented router now. Difficult to fix a non documented router in five years time because no one has been contributing to the web about its bug fixes.
At least LLMs almost always transform the original - it usually isn’t as straightforward as “Here’s the original but without the ads that pay for it”.

But we already have the latter case that exists - ad blockers. Ad blockers literally serve up the word-for-word original content minus the ads.

Gemini can just consume the device documents. There's an incentive for device makers to publish this content.
There had always been some incentive for manufactures to publish device documentation, and yet it has often been quite lacking either in quality or overall existence. I doubt LLM/agents being the readers will change that at all. What I expect AI scraping and using without credit will impact is people publishing their own unofficial help and guidance, and the affect there is likely to be negative. It won't stop all of them, but enough to be noticeable. Another possible negative is the manufactures documentation being AI generated without sufficient review, so possibly more erroneous than before, or intentionally not producing full documentation at all and expecting AI to fill the gap (MS seems to be heading this way: pushing "ask copilot" all over Azure instead of links direct to good reference material). All this would add up to a situation that is somewhere between "a little worse than pre-AI" and "an absolute shit show".
Exactly. What’s problematic about comments like your parent is the absence of mid-to-long term thinking.

It’s like bragging about a new highly addictive psychedelic drug that a dealer gave you a taste of for free. The effects are awesome today, you feel so fun and free! Never mind that it’s destroying your body and that the dealer will eventually charge you or demand you pay in other ways, that’s a problem for another day. Weeee!

I wonder the same thing. I only imagine that what comes next is worse: AI companies using vast resources to develop new training data, in house, locked down. They are already doing this with developers and code at Meta. Information will become locked away behind AI paywalls and chatbots.
{"deleted":true,"id":49251597,"parent":49251574,"time":1786406974,"type":"comment"}
That is next earnings quarters problem is the approach being taken
> incentive is there to publish anything new? For how long can we continue to rely on pre-2024 non-AI generated content?

obviously new content still has value because it remains the source layer for LLM agents. it just wont be ads giving you revenues thats all.

People write and create regardless of profit motive, it has been that way for thousands of years.
loading story #49256425
People are way less likely to write if there is no one to read it. And blog were also monkey see monley do - people seen other peoples blogs and got inspired.

When people wont see others blogs, they wont start writing own. When there will bw no ome to actually read it, they will go to do something else.

Actually no. The original copyright laws were created in part because of the realities of needed profit motive to have high value writing done, time consuming compilation work done. It was even titled "An Act for the Encouragement of Learning". The thousands of years writing you are talking about was often funded by patrons, who kept the output in their private libraries to show off (and maybe lend out) for prestige. It was a horrible limitation of knowledge and ideas. Much worse than the profit motive, copyright based system that came after that spawned a new age of knowledge in which everyone had cheap access, and those that didn't had access to the (no longer just private) libraries.

I'm sure the billionaire class would love a return to patronage based libraries, NDAs on authors of books, and the elitism they would feel with a return to private libraries locking away all kinds of knowledge that would happen if patronage become the only way authors could make money (such as with AI just regurgitating their works, or if the stupid 'do away with copyright' people got their way).

Actually it grew out of censorship and monopolies: https://en.wikipedia.org/wiki/Statute_of_Anne#Background

Cheap access came from the invention of cheap printing . The laws were passed to restrict it.

It's not copyright that caused cheap access. The printing machine allowed for cheaper publications, that's what spawned the new age of knowledge. The raw materials and the duplication of knowledge was the bottleneck. With digital systems this cost is minuscule, but still there.
Well companies are starting to put hidden ads in text content if the user agent is an AI crawler
Funny, I published information in the hopes that humans would benefit from it. If it happens to be through collective intelligence of LLMs I'm ok with that--even more so if through open models.
> Funny, I published information in the hopes that humans would benefit from it.

Sure, humans would benefit.

It took them searching, reading themselves, maybe even understanding something in the process, to complete a 360° revolution of their squirrel cages in time T.

Now they can omit searching, skip reading to the regurgitated answer, throw away understanding, and complete a full revolution in T/N, where N is a heuristic value directly proportional to the amount of skin in the AI hype.

But the catch is that the squirrel cage must run non-stop still.

That works now because there are human made sources that the AI can find and summarize for you. But now there are no incentives at all for humans to write anything on the internet and if they do the content will be buried by hallucinated content someone else posted at a larger scale.

I had the similar experience to yours yesterday and it lead nowhere. Funnily enough I was also trying to configure a vpn on a router, google didn't return anything useful (besides a blog post clearly written by AI and with absolutely no information in it). Claude managed to give some interesting pointers, but its suggestions were not working and I also noticed that it started to hallucinate badly about ipv6 and gave me some suggestions that were just plain untrue. Claude Opus is smart, usually when it gets so convinced about something is after researching the internet and not just based on its training data. I wonder where it got so convinced about it. Maybe reading some other hallucinated blog post like the one I stumbled upon?

{"deleted":true,"id":49255589,"parent":49255114,"time":1786441677,"type":"comment"}
That only works when there's an abundance of documentation for your specific device. The moment you're on a more recent version of something and have a weird issue, all hope is lots. You get stuck in loops because all LLMs keep recycling old advice that no longer applies. This problem will only get worse and worse as people no longer as questions on public forums, so answers are not publicly available either.
Yes. When my kids wonder why I don't hate AI the way they do I tell them it's because I hate Google even more.

To be more precise, I hate the SEO shithole the internet has become, that Google serves up, that Google facilitated, indirectly created.

(I really don't have any tears to shed if there is a death of the Corporate Internet™.)

AI seems to be on the same trajectory? Search was very useful in the start also, until it became entrenched. Then search placement became a target, and they are just focusing on extracting rents. All way paying the content providers zero or near-zero. And with years of that dynamic, we end up where we are now. It was the same with "social media". The same will happen with AI. AI is a power for more enshittification - being currently less shit than Google is (mosy likely) temporary.
It's definitely temporary. Remember there are no ad blockers for LLMs.
You're right, this is the best it will ever be. But there likely will be ad blockers for them -- local LLMs that filter for any brand placement, etc.
They're probably going to figure out how to aggressively monetize and enshittify LLMs at some point.

We're likely at the "golden age" of LLM-assisted web searching and summarization.

Hopefully open models keep it cracked open, but expecting enshittification is always the safe bet these days.

For the free to use LLMs, for sure. But if I pay $20 monthly, why should they enshittify it?
$20 a month is not nearly enough. The current finances require AI companies to make far more money than that to not implode.
Because that’s what always happens? Because their raison d’etre is to squeeze you dry?
look at every paid streaming service...
To be fair those enshittified because of copyright owners, not because of paid streaming services.
The amount of money you pay is irrelevant if you don't leave the platform once they starting adding revenue streams that degrade your experience.
Is there some search engine that you think could have become popular and not ended up with SEO optimization?
One you pay for yourself !

Very happy with Kagi personally

>> Is there some search engine that you think could have become popular and not ended up with SEO optimization?

> One you pay for yourself !

SEO is the practice done by webmasters of optimizing a website to improve its visibility and ranking in search engine results.

If you are paying to use your search engine, does that mean webmasters are no longer incentivized to/will not try to improve their visibility/ranking in your search results?

loading story #49256266
I don’t agree with the parent that the solution for SEO spam is paying for the search engine (I think there are other good reasons to do this though), especially since people have been doing things like naming their company “AAA Auto Repair” to be first in the phone book since before computers ever existed. But the person you originally replied to does have a point in blaming Google for the problem. Most SEO spam sites make their money from ads, and Google are the ones who run the ad network, which means Google are the ones funding them and creating an incentive for them to exist.
{"deleted":true,"id":49254165,"parent":49253837,"time":1786430191,"type":"comment"}
> SEO is the practice done by webmasters of optimizing a website to improve its visibility and ranking in search engine results.

They do this because they benefit from their site being visited or the information they are providing being noticed.

> If you are paying to use your search engine, does that mean webmasters are no longer incentivized to/will not try to improve their visibility/ranking in your search results?

I see no reason it would have that effect. It does, however, create different incentives for the search provider to improve the signals indicating page relevance since the user is the priority instead of advertisers.

Ah, the argument is that Google intentionally avoids showing you the most relevant results, or at least avoids "solving" the problem of webmasters who attempt to 'game' the system.

>>>> Is there some search engine that you think could have become popular and not ended up with SEO optimization?

>>> One you pay for yourself !

>> SEO is the practice done by webmasters of optimizing a website to improve its visibility and ranking in search engine results.

> They do this because they benefit from their site being visited or the information they are providing being noticed.

>> If you are paying to use your search engine, does that mean webmasters are no longer incentivized to/will not try to improve their visibility/ranking in your search results?

> I see no reason it would have that effect.

Agreed.

Google did not really lost to optimization, they enshittified to make you spwnd more time on search (and see more ads)
Ah, the argument is that Google intentionally avoids showing you the best results, so that you spend more time searching (and being exposed to ads)?
I'd frame it more that Google is incentivised to show you the results which are most lucrative for them to display, rather than the ones which are most beneficial for you to see.
Incentives. Which is why platforms (such as search) should never be allowed to be in the same company with things built on top of them (such as ads), if the combined company is significant for an important market.

We used to know better, Standard Oil vertical integration was dismantled.

While I understand that you feel good as you got the device configured, you would have been better off without gemini. The 'old' way as you call it would have resulted in you knowing the backgrounds and the inner workings of your router, making maintaining it a breeze and helps you actually understand your setup. Besides that, it would probably trigger you to rethink some of the things you now blindly have implemented because gemini did not show alternatives nor reasoning behind it. (something that most definitely would have been documented on the source pages)
Yep - it's awesome, the problem is that Google isn't sharing the revenue with the context creators anymore - over time, unless fixed, this will decimate the knowledge base it feeds on.

> Oh, I should mention though. There was no advertising at all. They didn't make any money off me

Are you sure about that? Even if you didn't see ads ( remember people pay even if you don't click - just like a billboard ) - they are still profiling you to better sell you ads in the future, and using your interaction as free training data.

I've run into major problems with LLMs as I maintain my home Linux systems. If I just copied commands they list, I would have a near 100% failure rate, as most of the information they have has been gleaned from forum posts that are years out of date.

They're a good jumping off point, but I need to delve into the original sources just like I did when I used Google.

I'm not sure. I run a homelab tailscale/k3s setup with gitops, dozens of services, VMs, backups, etc, all vibed by claude, and works just fine. Didn't write a single line of code for this. I don't know kubernetes and never will.
> I don't know kubernetes and never will.

He said, proud of his own ignorance.

Proud? No, but I am not ashamed of it also.

I would be the first to admit my ignorance on the absolute majority of topics. There is a limited number of things I can learn in life, and kubernetes won't be one of them - I'm just not interested in it (and all the other infra stuff, to be honest), as long as it works.

What on earth LLM are you using and do you tell it which distro/version of Linux they are supposed to be working with + give them access to web search or man pages? I haven't had this issue in like a year and a half.
By default I use Leo, and I do specify the distro. I tried to qualify by version, but when I do that I lose out on a lot of correct information for things that haven't changed in awhile.
I've had the opposite experience giving claude code SSH/ADB access to my devices. Not sure if it's the model itself or the harness, but I haven't had to manually do a sysadmin task in months.
Have a look at https://github.com/ThorOdinson246/whatisit-nl2sh

It seems to get shell scripts right most of the time.

I hate to break it to you, but we are at the start of enshittification circle. Once Gemini has monopoly you will reminisce with joy google search results.
I've been mostly enjoying Gemini, but it also clearly and definitively told me something i was trying to do was not possible with the library im using, so i wrote a different implementation, an hour later to discover that the library does in fact do precisely what i wanted in exactly the way i wanted with less headache. If I'd just gone right to the documentation instead, it actually would have saved me time.
Believe me they're making money off you one way or the other
And I've spent the last few days irritated that everything I ask Gemini is answered with something that's blatantly wrong and I'm not even a subject matter expert. A quick Google search for the same questions gives plenty of results that counter what the LLM gave me.

YMMV.

Confidently wrong summaries are the bane of Google search, and unfortunately I'm finding the AI seems to bleed into the actual search results too now, often turning up pages that back up what the summary is (wrongly) suggesting instead of surfacing actually relevant results for what I'm asking.
its pretty good with logs too. i mostly paste a few pages i suspect hace a problem in them and let it go to town. its almost always correct and is way faster than me
loading story #49256464
You deserve what's coming for you.
I built a simple Claude Code container on my homelab to do the same. I can just SSH in and get tech support when I need it.
AI is on the whole bad but it’s impossible to forget that the ostensibly human-curated Web as presented by modern search engines is bad too. Invasive ads, popups, and worst of all the substantive content has a nine inch frame of SEO filler and a three inch picture (what you were after). And some things are not even high-tech slop stolen. It is just old-school verbatim copied from another website and repackaged with another frame.

But yeah, text remix machines are not a long-term solution to that problem.

Imagine someone making an indie router and sell it too you then the AI not be able to surface their doc
no, no, no. you are supposed to romanticize the hunt for correct information /s
Even with the /s I think you misunderstand. If you just want to configure your router, then AI is neat. If you want to learn and understand what's going on in the router as you set it up, then being given the answers basically teaches you nothing.

It's the same reason why we don't give students the answers to things, we teach them to find the answers.

There will be new ways and incentives for content creators to be compensated. Many AI search startups are already talking about this or have created programs that help incentivize content creation.
Are you sure the lack of advertising isn't long term feasible? It seems to me that AI models have proven themselves to be something consumers ARE willing to pay a subscription - or even pay per use/token for.
Are they profiting from these subscriptions yet?
I've been using Gemini free tier exclusively for all my AI "needs" and would not be willing to pay for it
It took you 4 days because you used Gemini. Gemini is the worst AI model I ever used. It is way behind even open models. It looks like Google just reached its AOL moment.
Gemini isn't great as a model, googles search and ability to cite textbooks down to the paragraph make it better than every other model for human in the loop tasks.

I end up using the gemini api for with search enabled for the cases that I don't have access to good grounding data even in agentic tasks.

Come now, copilot is worse in every way
Guessing this is the relevant part: "worst AI model I ever used"
Copilot can use any model like Sol or Opus and is just a harness so not sure why people say this.
Personally it's because i used it when it was brand new and work paid for it and i had no idea what model it was using, my boss just turned it on for automatic PR summaries and code reviews and it was universally dogshit, and the auto complete in my IDE was awful as well.