* Models like DeepSeek V4.1 Flash are much cheaper on DeepSeek their API directly because of the cache handeling is better. Neuralwatt can hit up to 98% but DeepSeek can do 99.x... That may not sound like a big difference but it quickly widens the gap on long tasks to grow 2x a 3x in price. DeepSeek their cache handeling is S-tier (with a ton of features, for instance 24h caching).
* The same issue is also present if you compare GLM 5.3 Flash with z.ai vs Neuralwatt. Its just way more cheaper from the source, then from Neuralwatt.
* The energy numbers from Neuralwatt are ... to be taken with a ton of salt. Past energy numbers had the same models (for instance) GLM 5.2 up to 6x cheaper in energy usage, then after they "fixed" issues with the energy numbers. In reality, those energy numbers are just a different form of billing, but not a actual representation of the energy usage of AI models. Things like profits are inside those energy numbers. So seeing 4kWH used for a model, does not mean that it uses 4Kwh.
Edit: That are some interesting downvotes ...
To answer the questions. It was stated by the CEO himself in one of the video blogs that the energy prices inc their profit margins. Regarding their published numbers ... I like to point out that this is the same company that had up to 6x cheaper energy numbers at the start of the year until they got updated. Again, its in one of those video blogs the CEO did. Its around the same time when they increased the price from $5/1kwh to $10/1kwh.
Yes, DeepSeek API is cheaper then Neuralwatt. I have done way too many comparisons between NW and other providers, regarding their prices. Over long sessions, that gap grows because of the differences in caching. You need to use the NW Flex option to reduce the impact but then your constantly waiting on responses (good for overnight work, not great in prime time).
Edit 2: I am getting a little bit fed up with the people who downvote and do not give their reasons for the downvotes.