Meanwhile the GB300 used by hosted llms:
GPU Memory Bandwidth: 7.1 TB/s Interconnect Bandwidth: 900 GB/s bidirectional
https://pi3g.com/nvidia-gb300-specifications-including-memor...
If you think M7 will hit even 15% of these speeds you're very optimistic.
A hosted instance serves multiple customers at a time. A local model only one.
loading story #49297621