Hacker News new | past | comments | ask | show | jobs | submit

From the creator of Redis; run LLM locally with ds4

https://dwarfstar.sh/

  small native inference engine optimized first for DeepSeek V4 Flash (including the experimental vision model), DeepSeek V4.1 Flash (Metal, and text inference on CUDA), and additionally GLM 5.2 and 5.3, GLM 5.3 Flash and DeepSeek V4 PRO, and Qwen3.8 Flash Next (Metal and CUDA)
This is local targeting high end consumer hardware like DGX Spark or AMD Ryzen AI Halo.

For our mere mortals that were kids not long ago and can't really believe we've got our hands on a x090 series targeting Qwen3.8 27b, https://github.com/noonghunna/club-3090 is the way to go.

I'm maintaining a web frontend for this, trying to at least. You can follow it here: https://github.com/gchamon/club-3090-server

I maintain a fork of ds4 as shared libraries and thus can be used with other languages via FFI, along with public builds/binaries [1]. I made ds4go [2] against ds4 using techniques inspired by yzma.

In addition to the library bindings, we have a small library of tools (workspace for view/edit, scratchpad for persistence) and making your own is registering a Go function. And in recent weeks, I added the Vision and Qwen support, as ds4 added them.

Even if you don't use the Go library, the ds4go binary makes it really easy to download the libraries off of HuggingFace with a TUI available vie Homebrew.

Here's some TUI toy screenshots, sorry I still haven't released that code; it's of different quality than the others. [3]

EDIT: add ds4go TUI screenshot gist [4]

[1] https://github.com/NimbleMarkets/ds4/releases/tag/v0.8.20260...

[2] https://github.com/nimblemarkets/ds4go#install

[3] https://gist.github.com/neomantra/ae47422c8daf7a458212c93992...

[4] https://gist.github.com/neomantra/40180ade13df93290250ce8c6d...

https://github.com/antirez/ds4

The project GitHub page is a much better introduction for the hn crowd.

Nothing comparable but inspired from DwarfStar I wrote a little inference engine for Intel Xe-LP (no XMX) 32GB laptops. The only model supported right now is a quantized Gemma-4, but I don't exclude in the future to support other MoE of similar size. Too bad we have no Qwen 3.8 35B-A3B yet.

I'm also looking into expanding the protocol and the engine to support various steering techniques.

https://github.com/simoneiacomino/xenolith

loading story #49939079
It is pretty nifty. I spend some time over last weekend implementing fused TQ to allow for 1m context lengths on a 128 gb MacBook M5 Max when using Qwen 3.8 flash next (https://github.com/antirez/ds4/pull/1115 if you are interested). If I get bored I might port over the Metal kernels from oMLX -- the speed increase they have for the v0.7.0 release is amazeballs.
How is this different from other LLM runners?
loading story #49939270
loading story #49939224
curious, why did antirez go with C instead of something like Rust ?
loading story #49939049
loading story #49938951
the problem is the dsv4 checkpoint so quantized isn't very good
loading story #49938976
loading story #49938856
loading story #49938917
Random comment but the name is funny to me, reminds me of Silicon Valley.

What are we going to name the company, how about Dwarfism 2.0? What happened to 1.0 Jared?

loading story #49939018
loading story #49938901