Hacker News new | past | comments | ask | show | jobs | submit
I'm surprised neither OpenAI nor Anthropic made this move first. The Chinese open weight models are pulling ahead and commoditizing their value proposition.

Baking models onto silicon would've been the next logical move to get a moat.

Google is already doing this and has an experimental project on top of already having TPUs and cramming their quantized flash onto individual TPUs for inference.

loading story #49206037
Personally I think Apple should have acquired them. if you could burn a gemma4 class model into an iphone and actually get extremely low latency and low battery usage it would feel like the future IMO. even if it means you wont get frontier intelligence, there might actually be incentive to buy a new mobile device every year again.
loading story #49203553
I don't think this works out from a cost/silicon perspective. Small models already run pretty well in software (since the weights fit in cache) and big models require silicon area proportional to the size of weights. On a mobile device putting a chip like this is competing directly in BOM and power against a whole lot more l3 cache, and the l3 cache makes everything faster
loading story #49204604
loading story #49203581
loading story #49204401
loading story #49203604
That's actually a really good point... There's currently zero incentive to buying more hardware, and that's one very good reason do have a new one.
loading story #49203910
From what I remember, these chips are not mobile size yet
A small model would be. I think that’s more the point. It’s definitely not SOTA but it’s fast and energy efficient and local.
> A small model would be [mobile size]

A ~30mm side for the HC1 tech for an 8b model (still unclear the planned HC2)?

Is that analogue or are they baking floating points into the silicon?
loading story #49204444
Nope, a small model would be larger than the whole iPhone SoC.
loading story #49204056
loading story #49203558
Considering the rate of model development and rail hopping, seems like baking models into silicon is speed-running obsolescence.
loading story #49205990
If you’re only running models for frontier capabilities, yeah. For tasks where current models are smart enough, running them 100x faster is the most impactful improvement you can make. Consider all the things you could use a model for, but don’t, because the latency is just a bit too high.
loading story #49203665
loading story #49205194
"seems like baking models into silicon is speed-running obsolescence"

Now maybe. When models are flying passenger aircraft, other prerogatives will assert themselves. When a 50TB ROM means you can impulse purchase a ChatGPT 6.3 xhigh that runs on batteries, yet more use cases will be apparent.

loading story #49203559
loading story #49204375
I could see this making sense when model development start to settle down ... it's going to settle down, right? ...
Not sure. You can fix the transistors but leave the connections between them open for flexibility, so you only need to change the manufacturing process for the upper masks for every new model.
Surely that added flexibility negatively impacts the density/parameter count of the model you could etch?
Compute the cost of producing n of them devices, imagine a fair price based on that, and see if that local, blazing fast card* can be an asset that could be replaced periodically.

*(It's local: private files managing firm oriented. It's blazing fast: it can be placed into recursive, intensive local workflows.)

Which is exactly what companies and shareholders want to increase sales.
Look at it the other way: compared to the cost of training a model, the cost of making a custom ASIC is trivial.
loading story #49203664
loading story #49205616
You need to find customers for several-generations-ago models before this makes any sense. AMD is a lot more incentivized to look than mr vanilla llm is
ASICs is what took over Bitcoin mining, cheaper in all ways, and lasts longer than Nvidia GPUs for inference.
loading story #49203915
Isn’t that kind of useless for the stock? It sounds complicated, unlike having number of CPUs go up.

It’s like talking about anything else than Megapixels when everyone was convinced that megapixels must go up in certain periods of the smartphone boom.

I’m surprised Nvidia hasn’t partnered to make a Claude chip yet. It’s a win/win you can license them out, sell them when they become obsolete, etc.
Apparently Anthropic is moving that way: https://arstechnica.com/ai/2026/08/anthropic-confirms-plans-...
loading story #49203772
loading story #49204600
Just to see how fast it is try chatjimmy.ai
loading story #49205033
loading story #49204142
loading story #49205113
Because Openai and anthropic are not hardware companies. They outsource that to Broadcom and AWS' Annapurna labs.
OpenAI and Anthropic are both designing ASICs.
loading story #49203703
A model can't be updated, and a chip that is only relevant for 6 months at max?
Depends what you mean by relevant. If you use AI primarily as a search/knowledge engine, it makes no sense. If it's your capable assistant that has a lot of general knowledge, can do tool calls, and has a big context window, very doable.

Indeed, for some kinds of applications involving secure/legal data etc. I can see the consistency of silicon winning out, because it combines performance with immutability and guardrails in hardware. Some chips have write-once PROMs to store password hashes and similar, you could do the same thing with prompt hashing to absolutely force or forbid certain behaviors. A model that can't be updated is also a model that can't be hacked.

People already buy new phones every year, this just creates even more reason to do so
Outside of this website I've never met a person who buys a new phone every year. It's closer to every 3-4 years for most people.
loading story #49203988
Your location/income bias is showing. Most people do not buy new phones every year.
loading story #49205220
loading story #49204797
One of these chips smart enough to take orders at a drive-thru would be relevant for a decade, minimum.
It's a terrible moat. You etch the silicon then nobody wants to run it in 6 months because models have advanced that much further.
Not so if it's embedded in something smart enough for its intended purpose.

Think vision, spatial reasoning, speech synthesis, even some speech analysis. Think self-driving cars (and drones) that need 10x less power for the brain, and can think at 10x situation per second.

loading story #49203521
{"deleted":true,"id":49203320,"parent":49203171,"time":1786054264,"type":"comment"}
If a model is good enough today, it's still gonna be good enough in a year. Except you'll be able to serve it 1/100 of the price. Or 100x the speed.
loading story #49203811
It googles models suck