Hacker News new | past | comments | ask | show | jobs | submit
Indeed, but what is the difference between crawl data and model data but decay rate? Models are trained on previous crawl data, but if an LLM provider engages a search engine to get "live data," that data isn't live but previously crawled as well (and perhaps not yet integrated into models as crawl data).

So, why would you use Google as a tool or search target when you can, in some combination, go direct to the website (or whatever the target data endpoint is) yourself as an LLM provider to retrieve the most recent data or rely on your own "hot cache" of that data that was crawled recently but said data is not stale enough warranting a live web crawl to retrieve and present to the user or AI agent? Is this capability to perform retrieval from a data source in real time not similar to an AI agent?

Broadly speaking, I'm just spitballing on the concept of "You must use a search engine for an LLM to return 'live-ish' results" as I think we're directionally headed to where that isn't the case.