Hacker News new | past | comments | ask | show | jobs | submit
I've got a working recipe to run this model on Dual DGX Spark: https://github.com/volfco/spark-vllm-docker/blob/main/recipe...

Averages ~25-35tok/s which isn't bad for a first attempt.