AI Inference Chips & ASICs Beyond NVIDIA (Aug 2026)
Groq, Cerebras, SambaNova, AWS Trainium/Inferentia, Google TPU, Intel Gaudi — the AI-accelerator battlefield outside NVIDIA. Compared by architecture, memory type (large SRAM vs HBM) and deployment, highlighting differentiated paths like deterministic dataflow and wafer-scale single-chip.
| Vendor / chip | Architecture | Workload | Memory | Deployment | Perf (vendor claim) |
|---|---|---|---|---|---|
| Groq LPU | Deterministic dataflow; weights in SRAM; ultra-low latency | Inference-only | 500MB SRAM/chip (no HBM) | Cloud/API (GroqCloud) | 1,000 tok/s/user (rack) |
| Cerebras WSE-3 | Wafer-scale single chip; 900k AI cores | Training + inference | 44GB on-chip SRAM (no HBM) | Buy (CS-3) + cloud | Llama4 400B at 2,500 tok/s/user |
| SambaNova SN40L | Reconfigurable RDU; three-tier memory | Dual (inference-first) | 520MB SRAM + 64GB HBM + 1.5TB DDR | Buy + cloud/API | Llama 405B 129 tok/s; 8B 1,042 tok/s |
| AWS Trainium2 | Custom systolic ASIC; NeuronLink | Training + inference | 96GB HBM3e/chip (~2.9TB/s) | AWS cloud only | 667 dense BF16 TFLOP/s/chip |
| AWS Inferentia2 | Inference ASIC (NeuronCore-v2) | Inference-only | 32GB HBM/chip | AWS cloud only | up to 190 TFLOPS FP16/chip |
| Google TPU v7 Ironwood | Systolic array; built for inference era | Inference-first (also training) | 192GB HBM3e/chip (7.37TB/s) | Google Cloud only | 4,614 TFLOPS/chip; pod 42.5 ExaFLOPS |
| Intel Gaudi 3 | Dual die + on-die 24×200GbE Ethernet | Training + inference | 128GB HBM2e (3.7TB/s) + 96MB SRAM | Buy + cloud | 1.8 PFLOPS FP8/BF16; 900W |
| AMD MI355X (ref.) | CDNA4 general GPU (not inference ASIC) | Dual (inference-first) | 288GB HBM3e (8TB/s) | Buy + cloud | 10 PFLOPS FP4/FP6 |
Method & sources
Architecture/memory/deployment are from each vendor's site or technical white paper (Aug 2026), retrieved 2026-08-26. **All performance figures are VENDOR CLAIMS** (tokens/s, PFLOPS are self-reported, not independent third-party benchmarks; only AMD MI355X's MLPerf is a public leaderboard, still an AMD submission). **Classes**: inference-only = Groq, Inferentia2; dual-use inference-first = SambaNova, TPU Ironwood, MI355X; dual-use = Cerebras, Trainium2, Gaudi 3. **Memory profile**: large SRAM, no HBM = Groq (500MB/chip), Cerebras (44GB on-chip); SambaNova is SRAM+HBM+DDR three-tier; the rest are HBM-based. **Note**: Aug-2026 searches repeatedly show odd "NVIDIA Groq 3 LPX" branding of uncertain provenance, so Groq's latest-gen figures are lower-confidence; Cerebras specs are official numbers via tech-media relay. AMD MI355X is strictly a general GPU, not an inference ASIC — included for comparison only. Sources: Google TPU https://blog.google/innovation-and-ai/infrastructure-and-cloud/google-cloud/ironwood-tpu-age-of-inference/ ; AWS Trainium2 https://aws.amazon.com/ec2/instance-types/trn2/ ; SambaNova https://sambanova.ai/blog/sn40l-chip-best-inference-solution ; Intel Gaudi3 https://cdrdv2-public.intel.com/817486/gaudi-3-ai-accelerator-white-paper.pdf ; Groq https://groq.com/lpu-architecture .
Retrieved: 2026-08-26
FAQ
- What does "AI Inference Chips & ASICs Beyond NVIDIA (Aug 2026)" cover?
- It covers Groq LPU, Cerebras WSE-3, SambaNova SN40L, AWS Trainium2, AWS Inferentia2, Google TPU v7 Ironwood, Intel Gaudi 3, AMD MI355X (ref.), compared across: Vendor / chip, Architecture, Workload, Memory, Deployment, Perf (vendor claim).
- What are the sources and methodology?
- Architecture/memory/deployment are from each vendor's site or technical white paper (Aug 2026), retrieved 2026-08-26.
- When was this data last updated?
- The data was retrieved/updated on 2026-08-26.