PidokuInfra

Other Accelerators

Expert Advanced 50 min Difficulty 3/5 Topic 02 of 04

Prerequisites 01, IV.01

The idea in one minute#

NVIDIA GPUs are the default for AI, but not the only option. Alternatives fall into a few families: other GPUs (AMD, Intel), cloud vendors’ own chips (Google TPU, AWS, Microsoft, Meta), inference-specialized designs (Groq, Cerebras, SambaNova), and unified-memory systems-on-chip (Apple M-series). Each makes a different bet about which constraint matters most.

You can evaluate any of them with the tools you already have: capacity, bandwidth, arithmetic, interconnect, power — plus one new factor that often decides the matter: software.

A picture#

flowchart TB
  Q{"What limits you?"}
  Q -->|"Software risk, flexibility"| NV["NVIDIA GPU<br/>largest ecosystem"]
  Q -->|"Memory capacity per dollar"| AMD["AMD Instinct<br/>large HBM, ROCm"]
  Q -->|"Cost at cloud scale"| ASIC["Cloud ASICs<br/>TPU, Trainium, Maia"]
  Q -->|"Tokens per second per user"| SPEC["Specialized inference<br/>Groq, Cerebras"]
  Q -->|"Local, private, low power"| APL["Unified memory SoC<br/>Apple M-series"]
  class Q queue
  class NV compute
  class AMD,ASIC io
  class SPEC memory
  class APL neutral

How it really works#

The families#

FamilyExamplesThe betTrade-off
General GPUsNVIDIA; AMD Instinct (MI300/MI350 series; MI455X in “Helios” racks); IntelFlexibility: any model, training and inferenceNot optimal for any single workload
Cloud ASICsGoogle TPU (Ironwood); AWS Trainium3 / Inferentia; Microsoft Maia 200; Meta MTIAOwn the design, cut cost at huge scale for known workloadsOnly in that cloud; software tied to it
Wafer-scaleCerebrasOne enormous chip: massive bandwidth, no inter-chip communicationExotic system; model support is narrower
Deterministic inference chipsGroq LPU — now also sold by NVIDIA as Groq 3 LPX racksVery high tokens/s per user via on-chip memory and a static scheduleLimited memory per chip; many chips per model
Reconfigurable dataflowSambaNovaHardware shaped to the model’s data flowNiche tooling
Unified-memory SoCsApple M-series; some laptop/edge chipsCPU and GPU share one large memory pool: no PCIe copiesLower bandwidth and FLOPs than data-center parts

The field in October 2026#

Details date quickly; the pattern they show is the point.

  • AMD went rack-scale. The Helios rack (72 Instinct MI455X GPUs, each with a reported 432 GB of HBM4) began shipping to large customers in the second half of 2026 — AMD’s first answer to NVIDIA’s rack systems, and still a bet on memory per device.
  • Cloud ASICs are specializing by workload. Google’s seventh-generation TPU (Ironwood) is generally available and aimed at inference, and its announced eighth generation splits into separate training and inference chips. AWS’s Trainium3 is generally available; Microsoft’s Maia 200 (January 2026) is described as an inference accelerator.
  • The specialist was absorbed by the incumbent. NVIDIA licensed Groq’s inference technology in December 2025 (a reported $20 billion) and announced the Groq 3 LPX rack at GTC 2026 as a low-latency companion to its GPU racks. The lever Groq attacked — per-user token speed — turned out to matter enough for the GPU vendor to sell it.
  • Wafer-scale kept going. Cerebras runs production inference for large customers and presented its next systems at Hot Chips 2026.

Every one of these is a vendor claim until you run your model on it.

Why “FLOPs per dollar” is not enough#

A faster, cheaper chip is worthless if your model does not run on it, or runs through an immature kernel that reaches 20% of the hardware’s potential. CUDA’s advantage is fifteen years of tuned libraries (III.05), and every framework, optimization and new research technique lands on CUDA first.

Questions to ask of any alternative:

  1. Does my exact model and serving stack run on it today, at full speed?
  2. Are the optimizations I rely on (quantization formats, attention kernels, batching server) supported?
  3. How do I debug and profile it?
  4. Can I get capacity, and from more than one supplier?

For large, stable workloads the answers can justify the switch — that is why the largest operators build their own chips. For a small team iterating quickly, software maturity usually outweighs hardware advantages.

Where alternatives genuinely win#

  • AMD GPUs frequently offer more memory per dollar. For capacity-bound inference (V.03) that is exactly the right lever, and the major inference servers support them.
  • Specialized inference chips deliver several times the per-user token rate of a GPU. They attack the batch-1 memory-bound limit (II.03) directly, by putting weights in much faster memory. Valuable when latency per user is the product.
  • Cloud ASICs win on cost for the cloud’s own high-volume models, and are offered to customers at attractive prices for supported frameworks.
  • Apple silicon makes large models usable on a laptop because the GPU can address all of system memory. Bandwidth (roughly 100–800 GB/s depending on the chip) limits speed, but a 64 GB model simply runs, with no PCIe in the way.

CPUs still count#

A modern server CPU with wide vector instructions and many memory channels runs small models, embeddings and classical ML perfectly well, with no accelerator to schedule. From IV.01: if the workload is memory-bound and small, the CPU’s bandwidth may be enough.

Code#

A comparison harness using the course’s own model: fit, memory-bound speed, and cost.

Go
// compare.go — evaluate any accelerator with the same five questions.
package main

import "fmt"

type Accel struct {
	Name                string
	MemGB, BandwidthGBs float64
	DollarsPerHour      float64
	Mature              bool // does your stack run on it today, unmodified?
}

func main() {
	const modelGB = 40.0 // e.g. a 70B model at ~4.5 bits per weight
	candidates := []Accel{
		{"Data-center GPU A", 80, 3350, 2.50, true},
		{"Data-center GPU B (more memory)", 192, 5300, 2.20, true},
		{"Laptop SoC, unified memory", 128, 546, 0.10, true},
		{"Server CPU, 12 memory channels", 512, 300, 0.80, true},
		{"Specialized inference system", 64, 20000, 6.00, false},
	}
	fmt.Println("accelerator                        fits   tok/s (1 user)   $ per M tokens (1 user)  software")
	for _, a := range candidates {
		if modelGB > a.MemGB*0.9 {
			fmt.Printf("%-34s no\n", a.Name)
			continue
		}
		tps := a.BandwidthGBs / modelGB
		cost := a.DollarsPerHour / (tps * 3600) * 1e6
		sw := "ready"
		if !a.Mature {
			sw = "verify first"
		}
		fmt.Printf("%-34s yes  %12.0f  %22.2f   %s\n", a.Name, tps, cost, sw)
	}
	fmt.Println("\nSingle-user figures. Batching changes the cost column by 10-100x on GPUs.")
}

All rows are illustrative. Note the last line of output: single-user cost flatters hardware that cannot batch and penalizes hardware that can. Decide which case is yours before comparing.

Remember this#

  • Alternatives differ in which constraint they attack: capacity, bandwidth, per-user speed, cost at scale, or locality.
  • Software maturity is frequently the deciding factor.
  • Evaluate with capacity, bandwidth, arithmetic, interconnect, power — then cost per unit of your work.
  • CPUs and unified-memory laptops are legitimate inference hardware for the right model size.

Try it#

  1. Run compare.go. Add a Batch field and a compute limit, and recompute cost per million tokens at batch 32.
  2. Pick one alternative and find out whether a model you use runs on it with your serving stack. How long did it take you to find out? That time is part of the cost.
  3. Why does unified memory remove a whole class of problems from module III?

Check yourself#

  1. Name three families of non-NVIDIA accelerators and each one’s bet.
  2. Why can a slower chip with mature software beat a faster one without?
  3. Which hardware lever do specialized inference chips attack?

↑↓ navigate↵ openesc close