1 article
The Helios system is already running GPT-3 scale models with sub-millisecond latency, and that's a direct shot at Nvidia's inference throne.