
- The liquid-cooled LPX rack contains 256 Samsung-manufactured Groq chips and is designed for low-latency AI inference.
- Nvidia says the rack delivers 3,400 tokens per second, giving it a quantified performance advantage in the source's comparison with Cerebras-powered inference.
- The deployment commercializes technology and talent obtained through Nvidia's $20 billion purchase of Groq assets, announced in December.
- Adoption remains concentrated around one named customer, Nebius, and the product is specialized for inference rather than general-purpose AI workloads.
Quotes
“3,400 tokens per second”