Nvidia brings Groq inference racks to production

Yahoo Finance · 13 days ago
Nvidia brings Groq inference racks to production
  • The liquid-cooled LPX rack contains 256 Samsung-manufactured Groq chips and is designed for low-latency AI inference.
  • Nvidia says the rack delivers 3,400 tokens per second, giving it a quantified performance advantage in the source's comparison with Cerebras-powered inference.
  • The deployment commercializes technology and talent obtained through Nvidia's $20 billion purchase of Groq assets, announced in December.
  • Adoption remains concentrated around one named customer, Nebius, and the product is specialized for inference rather than general-purpose AI workloads.