Hacker News new | past | comments | ask | show | jobs | submit
> "confidence": 0

OP and the linked page talk about the confidence score and using it as an action threshold, so it looks like an appropriate total response to me.

Right, but that's not the same thing as reporting a benchmark across a test set. It doesn't help me determine how well the model does across a decently-large sample size of commands. It doesn't tell me with what reliability the confidence will be below a given threshold when it should be, above that threshold when it should be, etc.