Hacker News new | past | comments | ask | show | jobs | submit
I'm surprised neither OpenAI nor Anthropic made this move first. The Chinese open weight models are pulling ahead and commoditizing their value proposition.

Baking models onto silicon would've been the next logical move to get a moat.

Google is already doing this and has an experimental project on top of already having TPUs and cramming their quantized flash onto individual TPUs for inference.

Personally I think Apple should have acquired them. if you could burn a gemma4 class model into an iphone and actually get extremely low latency and low battery usage it would feel like the future IMO. even if it means you wont get frontier intelligence, there might actually be incentive to buy a new mobile device every year again.
That's actually a really good point... There's currently zero incentive to buying more hardware, and that's one very good reason do have a new one.
I don't think this works out from a cost/silicon perspective. Small models already run pretty well in software (since the weights fit in cache) and big models require silicon area proportional to the size of weights. On a mobile device putting a chip like this is competing directly in BOM and power against a whole lot more l3 cache, and the l3 cache makes everything faster
From what I remember, these chips are not mobile size yet
A small model would be. I think that’s more the point. It’s definitely not SOTA but it’s fast and energy efficient and local.
loading story #49203496
loading story #49203526
Considering the rate of model development and rail hopping, seems like baking models into silicon is speed-running obsolescence.
loading story #49203421
"seems like baking models into silicon is speed-running obsolescence"

Now maybe. When models are flying passenger aircraft, other prerogatives will assert themselves. When a 50TB ROM means you can impulse purchase a ChatGPT 6.3 xhigh that runs on batteries, yet more use cases will be apparent.

Compute the cost of producing n of them devices, imagine a fair price based on that, and see if that local, blazing fast card* can be an asset that could be replaced periodically.

*(It's local: private files managing firm oriented. It's blazing fast: it can be placed into recursive, intensive local workflows.)

Which is exactly what companies and shareholders want to increase sales.
Not sure. You can fix the transistors but leave the connections between them open for flexibility, so you only need to change the manufacturing process for the upper masks for every new model.
Surely that added flexibility negatively impacts the density/parameter count of the model you could etch?
I could see this making sense when model development start to settle down ... it's going to settle down, right? ...
Look at it the other way: compared to the cost of training a model, the cost of making a custom ASIC is trivial.
loading story #49203992
loading story #49203510
Isn’t that kind of useless for the stock? It sounds complicated, unlike having number of CPUs go up.

It’s like talking about anything else than Megapixels when everyone was convinced that megapixels must go up in certain periods of the smartphone boom.

loading story #49203522
loading story #49204136
I’m surprised Nvidia hasn’t partnered to make a Claude chip yet. It’s a win/win you can license them out, sell them when they become obsolete, etc.
A model can't be updated, and a chip that is only relevant for 6 months at max?
loading story #49203508
loading story #49203691
People already buy new phones every year, this just creates even more reason to do so
Because Openai and anthropic are not hardware companies. They outsource that to Broadcom and AWS' Annapurna labs.
It's a terrible moat. You etch the silicon then nobody wants to run it in 6 months because models have advanced that much further.
It's a fantastic business opportunity to sell local models though. Sell a proprietary PCIe inference ASIC that accepts a ROM daughterboard holding a model. You don't actually need the compute of 8 flagship GPUs for inference. The name of the game has been the terabytes of memory. Move to an old node like 28nm and you sidestep supply chain issues and have a massively affordable product. Gut tells me you'd be able to turn a tidy margin selling those daughterboards at a couple hundred bucks.

Update the model? Customers have to buy the new daughterboard, giving you a persistent income stream. Update the meta-architecture? You sell a new ASIC. Congratulations, you now dropped the capex for trillion-parameter models down from the price of a house, to the price of a normal computer peripheral.

Not so if it's embedded in something smart enough for its intended purpose.

Think vision, spatial reasoning, speech synthesis, even some speech analysis. Think self-driving cars (and drones) that need 10x less power for the brain, and can think at 10x situation per second.

If a model is good enough today, it's still gonna be good enough in a year. Except you'll be able to serve it 1/100 of the price. Or 100x the speed.
It googles models suck
loading story #49204720
The demo: https://chatjimmy.ai/
I know it's a relatively tiny model, but damn, is that thing fast.

It also mostly passes the "schlong" test

https://pastes.io/YcxSi8Fp

I read the paste, it got the etymology wrong, no? Schlong comes from shlang (snake), not shlemp (is this even a word? I don't speak Yiddish but couldn't find it on Google).

Oxford also claim that its first recorded use was from the 60s, not the 20s; https://www.oed.com/dictionary/schlong_n?tl=true

It did get it wrong but it also got a lot farther than much more recent, but worse models like 6.7GB on disk size ternary bonsai. It at least knows it's from Yiddish. The "schlemp" appears to be a total hallucination or it's confusing it with schlep, which is not related to schlong. One of the reasons why I said it "mostly" passes the test. Something much larger on the size of qwen 3.5 122B, deepseek v4 flash or similar that runs in 120GB to 190GB of RAM in my experience will answer perfectly unless it has been ruined by something like Q2 quantization.
I didn't realize there was a SchlongBench™ (but of course there is). What's it test? (asking seriously)
There isn't SchlongBench(TM) yet, it's a specific question I've been asking of differently sized models as a randomly chosen gauge of how much less commonly used knowledge is perma-baked into it. In this case a question about a specific yiddish origin slang term. Small/bad models don't know it's from middle high german or Yiddish and get its origin and meaning totally wrong (or it runs into model censorship related to slang related to the male anatomy).

It's also a question I have found will cause models that don't know what it is to go off quickly in a direction of hallucination trying to explain it, so the hallucination is evident very quickly starting from the first ever prompt issued with 0 context fill. Example: I had a model write four detailed supposedly-accurate sounding, grammatically correct paragraphs saying its origin is from AAVE (African American Vernacular English), which it most certainly is not

You could do the same by picking any topic that is very rarely discussed in conversation, some esoteric and narrow piece of knowledge and asking the model about it.

Oh ya, this is like the approach from the Incompressible Knowledge Probes [0] paper - smart!

[0] Incompressible Knowledge Probes: Estimating Black-Box LLM Parameter Counts via Factual Capacity [https://arxiv.org/abs/2604.24827]

I freakin' love this demo. It feels magical.
I had the same reaction but then I showed it to my partner. She completely didn't get it, in her words "how can it be thinking of a good answer when it's that quick?"

I tried to explain but I fear were probably going to be adding artificial sleeps to these things to convince the masses it's doing something clever.

I asked it some old hardware command line questions I'd recently asked Gemini, it hallucinated parts of the answer.

The characters in the 3-act Shakespearean play had very little depth, many of the names were similar, and they were not very smart, but the simple plot was cohesive.

to be fair, the model used for Chat Jimmy is not very smart, but the world where it is smart is very interesting.

It’s going to be really crazy when the bottle neck for agents is the speed of the tool calls rather than the speed of inference. Imagine an agent interacting with the terminal near instantly…

loading story #49204944
loading story #49204199
Wow, feels like Google web search in 1999.
If you still want the experience, go and browse McMaster Carr. Wizards designed that website.
loading story #49204390
loading story #49204097
loading story #49203887
OMFG this thing is fast.
its fast but try to get it to give you pi to 50 decimal places. it didnt go well for me.
I think the same exact model running on CPU-only and RAM, or a small GPU, would do about the same? It's quite an old model now and small, you could throw a GGUF into llama-server or something for a side by side comparison.

https://huggingface.co/meta-llama/Llama-3.1-8B

As I remember just about any english language model from mid 2024 and earlier didn't even do well if you asked it to count sequentially from 0 to 100, nevermind calculating stuff.

loading story #49205005
I understand the appeal due to the speed
{"deleted":true,"id":49203177,"parent":49202316,"time":1786053485,"type":"comment"}
loading story #49203762
loading story #49204685
It doesn’t believe it’s running on that chip, it’s arguing with me
It's running a very small, non-reasoning model at the moment. But more generally, almost all LLMs argue on the hardware/model they are/are on.
What would tokens/sec performance look like for a reasoning model? An order of magnitude slower?
Reasoning models are the same speed. They’re just post trained with RL to do CoT inside tags like <thinking></thinking> before a tag like <response></response>

There’s no difference in the inference implementation, parameter count, or speed.

Which model? Or how many active parameters?
Llama 3.1 8B model
loading story #49203053
{"deleted":true,"id":49202817,"parent":49202775,"time":1786051660,"type":"comment"}
loading story #49204302
loading story #49205726
loading story #49203451
loading story #49203424
loading story #49205219
This is probably a win-win. The team gets paid, and we get greater assurance that their best ideas and architectures -- which are truly impressive -- are going to see the light of day in actual products.
They were too small for this to be a meaningfully sized purchase for AMD, there's real risk they get sucked into a team that ultimately delivers sqat, not to mention the chances of anything being delivered in an even remotely consumer-priced bracket are definitely out the window
loading story #49205534
loading story #49205123
loading story #49205535
loading story #49205102
loading story #49205330
Well so much for that dream.

Guess we can look forward to picking these up ex-enterprise on ebay for under $5k a pop in a decade or two

Is there any LLM from exactly one year ago that would be worth running?

In Aug 2025 you had

- OpenAI o3

- Opus 4.1

- Gemini 2.5 Pro

- Grok 4

Even if those were almost free to run, you'd be way better off with Deepseek flash 0731 or GPT 5.6 Luna, which already are almost free.

Other than for things where the t/s are critical, it seems like a bad idea to etch a model into silicon.

loading story #49203829
loading story #49203778
AMD could have saved their money and used their own hardware! I've got a language model doing 60k tok/s on AMD hardware already, a Xilinx Kria K26 SOM, with the weights baked into URAM/BRAM with zero DRAM in the token loop. Same thesis as Taalas: single-stream decode is bandwidth bound, so stop fetching weights from far away.

Caveats stacked high, obviously. It's 3.16M parameters (tinystories, and I also have a kevin-speak lemmatised version), the tokens are characters, and the 60k record is 16 streams that each remember exactly one token of context, so it's blisteringly fast at saying nothing. The honest build with full context and KV caching still does ~19k tok/s on one stream though.

I keep messing with the blogpost with the live demo, but I'm planning on flipping it to live in the next day or two

Well, technically it is their hardware now...
And their team, if they treat them well.
loading story #49204484
loading story #49203576
Really hoped to see their hw out in the wild one day
loading story #49205217
loading story #49205342
loading story #49203797
While this design is self-limiting I think its a good approach. It doesn't take an entirely new architecture or infinite memory to produce significant performance improvement.
Toronto Canada startup btw.
Works well, I remember driving by the ATI building as a kid.
loading story #49202746
loading story #49205020
loading story #49204331
Can anyone comment on the economics and likely turnaround times of this process, when it’s more mature?

Would it be realistic for a frontier lab to deploy this or would the turnaround time mean the model is always too out of date?

Assuming the weights and architecture are eventually stable, how much cheaper would this end up being?

There are always uses for outdated models.

Claude Code is still using haiku 4.5 from ages ago for explore subagents for instance. Not to mention production uses like customer service that only need to be "good enough"

Customer service has really degraded huh. 4 years ago they expected opus performance out of human call center agents

I guess losing some customers due to poor customer service is ok if the price of customer service is right.

Just looked this up, no longer true. Explore subagents inherit whatever model the parent is. And you can of course make other subagent configs.
loading story #49203240
loading story #49203264
2 to 3 months optimistically assuming everything goes smoothly and is fully automated.

6 months or even a year if something goes wrong in the fabrication process and you need to update things.

If they do more standard asic design, it could be a lot longer as the design needs to be validated on an FPGA cluster, which would necessarily need to be very big for something like a LLM. Easily up to 2 years.

There's a reason chatjimmy isn't demonstrating newer models and why they only show of an 8B model.

I mean even if it take a few months, it'll still be out of date. But there was a hypothetical when it came up in Feb, would you want Qwen 3.5 at like 10k tokens per second.

At the time people were no doubt saying yes but now 3.8 is out, is that still desirable?

There's soooo much stuff that such a model is still capable of doing in the pursuit of getting a better overall answer. Imagine a powerful research agent that blasts out dozens of the small, cheap models to fetch and summarize one page each. Then the beefy researcher model performs the final analysis.
Honestly, this is starting to make more and more sense. SOTA models are starting to converge to certain architecture and capabilities. I wouldn’t be surprised we end up with a base model ASIC + “fine tune” card where it’s a physical LoRA style adapter.
loading story #49202698
loading story #49203239
loading story #49202875
loading story #49202663
loading story #49202563
loading story #49203001
loading story #49202615
loading story #49204735
loading story #49203512
loading story #49203919
Didn't even give them a chance to launch the hardware.
loading story #49203672
loading story #49203975
loading story #49204017
loading story #49204204
With web search and tool call a decent current generation model at the speed of the chatjimmy could do a lot. People saying it would be out of date are missing the point. It’s not going to make much sense for frontier companies that’s chasing the SOTA. But for a lot of business use cases if someone can put GLM 5.2 and sell it as a box, it would make so much sense.

My partner has been asking for a “completely private” model for doing research and shifting through volumes of data that can’t leave the office and $$$ for the current hardware makes no sense. It would be an easy sell if someone walks in with a black box that contains “ChatGPT”.

100% agree - you don't need the most up-to-date model to have something that's useful in agentic contexts. They could even produce chips with weights that make all the decision making/logical reasoning and have it delegate to other specialized agents. If it becomes cheap enough to print a run of custom chips, releasing a batch for each major advancement does not seem unreasonable for SOTA companies.
There are so many use cases for supremely fast offline models. The first thing that comes to my mind is for real-time video processing or other non-textual content in real time.
loading story #49203837
loading story #49204764
loading story #49205069
loading story #49204773
loading story #49204777
I feel like NAND process tech could become useful at solving some of these problems. A GPU where you can update the weights a few thousand times may be sufficient.
The basis of Taalas is "compute in memory" electronics - past Von Neumann's separation of processor and memory.

You need to be able to add|mul where the data (the weights) are stored.

NAND hasn't been scaling great lately. It seems like PCM or MRAM would both be better fits.
FPGA model storage?
loading story #49205337
Token quantity will have a quality all its own.
so qwen3.x-27b on hardware? or better deepseek-v4-flash on hardware .
loading story #49202473
It must be a “super model”. What will be if new model released? New chips?
loading story #49203999
Yeah, so https://chatjimmy.ai/ ... the model is crap, but the speed is amazing. Worth checking out.
Imagine the size of chip needed to 'etch' something like Qwen 3.6 27B in size.
Interesting thought, because it's a yield question. How tolerant are models today to a few broken weights.

If tolerant, they could churn out many cheaper chips, some perhaps with slight abnormal tendencies ;)

> How tolerant are models today to a few broken weights.

Extremely! You can remove entire layers and the model will still work just fine, with barely perceptible capability losses.

I've cut/bypassed ~15% of total parameters out of Gemma 4 31B on a pod once. Still got perfectly coherent responses out of it. Certain layers are a lot more important than others, particularly early and late ones; but it's honestly astonishing how much can be cut out from the middle without destroying the model's coherence.

I didn't run any meaningful benchmarks, so I have no idea what the capability loss looks like exactly. But "produce coherent and sensible English in response to a wide variety of prompts" was definitely not among the things the model unlearned.

I wonder if you had a few percent of problems in the yield, if it would be functionally equivalent to the difference between a unsloth-published Q6 standard size GGUF vs. the nearly perfect precision of an unsloth Q8-K-XL. Or more like Q4 vs Q8 where a lot is lost.
Not too dissimilar to the first HC1 (6nm 815mm² 53B Transistors embedding an 8b LLM):

> Our second model, still based on Taalas’ first-generation silicon platform (HC1), will be a mid-sized reasoning LLM

If someone has that sort of knowledge; how big a chip would be required? Is it possible?
Well, given the data above, roughly a 220b transistors chip for the HC1 tech.
loading story #49203484
loading story #49204198
loading story #49205073