AI Inference Chips & ASICs Beyond NVIDIA (Aug 2026)

Groq, Cerebras, SambaNova, AWS Trainium/Inferentia, Google TPU, Intel Gaudi — the AI-accelerator battlefield outside NVIDIA. Compared by architecture, memory type (large SRAM vs HBM) and deployment, highlighting differentiated paths like deterministic dataflow and wafer-scale single-chip.

Vendor / chipArchitectureWorkloadMemoryDeploymentPerf (vendor claim)
Groq LPUDeterministic dataflow; weights in SRAM; ultra-low latencyInference-only500MB SRAM/chip (no HBM)Cloud/API (GroqCloud)1,000 tok/s/user (rack)
Cerebras WSE-3Wafer-scale single chip; 900k AI coresTraining + inference44GB on-chip SRAM (no HBM)Buy (CS-3) + cloudLlama4 400B at 2,500 tok/s/user
SambaNova SN40LReconfigurable RDU; three-tier memoryDual (inference-first)520MB SRAM + 64GB HBM + 1.5TB DDRBuy + cloud/APILlama 405B 129 tok/s; 8B 1,042 tok/s
AWS Trainium2Custom systolic ASIC; NeuronLinkTraining + inference96GB HBM3e/chip (~2.9TB/s)AWS cloud only667 dense BF16 TFLOP/s/chip
AWS Inferentia2Inference ASIC (NeuronCore-v2)Inference-only32GB HBM/chipAWS cloud onlyup to 190 TFLOPS FP16/chip
Google TPU v7 IronwoodSystolic array; built for inference eraInference-first (also training)192GB HBM3e/chip (7.37TB/s)Google Cloud only4,614 TFLOPS/chip; pod 42.5 ExaFLOPS
Intel Gaudi 3Dual die + on-die 24×200GbE EthernetTraining + inference128GB HBM2e (3.7TB/s) + 96MB SRAMBuy + cloud1.8 PFLOPS FP8/BF16; 900W
AMD MI355X (ref.)CDNA4 general GPU (not inference ASIC)Dual (inference-first)288GB HBM3e (8TB/s)Buy + cloud10 PFLOPS FP4/FP6

Method & sources

Architecture/memory/deployment are from each vendor's site or technical white paper (Aug 2026), retrieved 2026-08-26. **All performance figures are VENDOR CLAIMS** (tokens/s, PFLOPS are self-reported, not independent third-party benchmarks; only AMD MI355X's MLPerf is a public leaderboard, still an AMD submission). **Classes**: inference-only = Groq, Inferentia2; dual-use inference-first = SambaNova, TPU Ironwood, MI355X; dual-use = Cerebras, Trainium2, Gaudi 3. **Memory profile**: large SRAM, no HBM = Groq (500MB/chip), Cerebras (44GB on-chip); SambaNova is SRAM+HBM+DDR three-tier; the rest are HBM-based. **Note**: Aug-2026 searches repeatedly show odd "NVIDIA Groq 3 LPX" branding of uncertain provenance, so Groq's latest-gen figures are lower-confidence; Cerebras specs are official numbers via tech-media relay. AMD MI355X is strictly a general GPU, not an inference ASIC — included for comparison only. Sources: Google TPU https://blog.google/innovation-and-ai/infrastructure-and-cloud/google-cloud/ironwood-tpu-age-of-inference/ ; AWS Trainium2 https://aws.amazon.com/ec2/instance-types/trn2/ ; SambaNova https://sambanova.ai/blog/sn40l-chip-best-inference-solution ; Intel Gaudi3 https://cdrdv2-public.intel.com/817486/gaudi-3-ai-accelerator-white-paper.pdf ; Groq https://groq.com/lpu-architecture .

Source: https://blog.google/innovation-and-ai/infrastructure-and-cloud/google-cloud/ironwood-tpu-age-of-inference/

Retrieved: 2026-08-26

FAQ

What does "AI Inference Chips & ASICs Beyond NVIDIA (Aug 2026)" cover?
It covers Groq LPU, Cerebras WSE-3, SambaNova SN40L, AWS Trainium2, AWS Inferentia2, Google TPU v7 Ironwood, Intel Gaudi 3, AMD MI355X (ref.), compared across: Vendor / chip, Architecture, Workload, Memory, Deployment, Perf (vendor claim).
What are the sources and methodology?
Architecture/memory/deployment are from each vendor's site or technical white paper (Aug 2026), retrieved 2026-08-26.
When was this data last updated?
The data was retrieved/updated on 2026-08-26.

Related datasets