Hacker News new | past | comments | ask | show | jobs | submit
>> We have an internal eval that measures performance on tasks for a handful of embedded systems repos for our mmWave radios

Okay but the parent said real-world usage, presumably meaning coding tasks.

We have a whole bunch of complex evals that Haiku 4.5 passes. That doesn't mean it is a good model for coding.

They literally stated in their first sentence that it was coding tasks.
Yes, these are coding tasks in the embedded systems domain (I mentioned Rust and C).