Hacker News new | past | comments | ask | show | jobs | submit
Given that it apparently defaults to 'xhigh', this is probably the answer.

Granted, it's still much lower tokens/s than you'll get out of many MoE models.

Edit: Even set to medium or low there's still a lot of second guessing, less consistency, lower 'acceptable response' rate, and slower/more token churn vs gemma4:26b-a3b. I think gemma4 is just a better 'general purpose' model.