The idea in one minute#
A tensor core is a circuit inside each SM that does one thing: multiply two small matrices and add the result to a third, in a single operation. Because it is built for that one job and works on small number formats (16, 8 or even 4 bits per number), it does 10–30x more arithmetic per second than the GPU’s general-purpose units.
Nearly all of a modern GPU’s headline FLOP figure comes from tensor cores, and you only get it if your numbers are in a format the tensor cores accept.
An analogy#
A general arithmetic unit is a skilled worker with a calculator: any operation, one at a time.
A tensor core is a stamping press. It produces one specific part — a small matrix product — in a single stroke. It cannot do anything else, but nothing matches it at that part. And the thinner the sheet metal (fewer bits per number), the more parts per stroke.
A picture#
flowchart LR
subgraph IN["Inputs"]
A["Matrix A<br/>small tile"]
B["Matrix B<br/>small tile"]
C["Matrix C<br/>accumulator"]
end
TC["Tensor core<br/>D = A x B + C<br/>one operation"]
D["Matrix D"]
A --> TC
B --> TC
C --> TC
TC --> D
class A,B,C,D memory
class TC computeHow it really works#
Number formats#
A floating-point number stores a sign, an exponent and a fraction. Fewer bits means less precision and a smaller range — and half the memory, half the bandwidth, and simpler circuits.
| Format | Bits | Rough precision | Typical use |
|---|---|---|---|
| FP64 | 64 | 15–16 digits | Science, finance |
| FP32 | 32 | 7 digits | Classic default |
| FP16 | 16 | 3–4 digits, max ≈ 65,504 | AI inference and training |
| BF16 | 16 | 2–3 digits, same range as FP32 | AI training (safer range) |
| FP8 | 8 | 1–2 digits | Recent AI inference/training |
| INT8 | 8 | 256 levels | Quantized inference |
| FP4 / INT4 | 4 | 16 levels | Aggressively quantized inference |
Neural networks tolerate low precision remarkably well: their weights are noisy estimates to begin with, so three digits are usually enough.
What the tensor core does to the FLOP count#
H100, vendor figures:
| Path | Format | Peak |
|---|---|---|
| General units | FP64 | ~34 TFLOP/s |
| General units | FP32 | ~67 TFLOP/s |
| Tensor cores | FP16 / BF16 | ~990 TFLOP/s |
| Tensor cores | FP8 / INT8 | ~1,980 TFLOP/s |
Same chip. A 30x spread depending only on the number format and whether the work is expressed as matrix multiplication.
The conditions for using them#
- The operation must be a matrix multiply (or a convolution, which libraries turn into
one). An elementwise
a + bnever touches a tensor core. - The data must be in a supported format. FP32 inputs on most GPUs fall back to the slow path (some GPUs offer a reduced-precision “TF32” mode as a middle ground).
- Dimensions should be friendly, typically multiples of 8, so tiles fill completely.
You rarely call a tensor core yourself. Libraries (cuBLAS, cuDNN) and frameworks choose them automatically when the conditions hold. Your job is to make the conditions hold.
Mixed precision#
Low precision is safe for the bulk multiplies but risky for a few steps, such as summing
thousands of small numbers or computing exp() of large values. Mixed precision runs the
heavy matrix work in FP16/BF16/FP8 and keeps the fragile steps in FP32. It gets nearly all the
speed and nearly all the accuracy.
Code#
Go has no 16-bit float type, which makes it a good place to see what the format does. This program converts FP32 values to FP16 bits and back, by hand, and shows the two ways it bites: rounding and overflow.
// half.go — what FP16 keeps and what it loses.
package main
import (
"fmt"
"math"
)
// roundTrip converts a value to IEEE 754 half precision (round to nearest) and back.
func roundTrip(f float32) float32 {
if math.IsNaN(float64(f)) {
return f
}
const maxHalf = 65504
if f > maxHalf {
return float32(math.Inf(1))
}
if f < -maxHalf {
return float32(math.Inf(-1))
}
if f == 0 {
return 0
}
// FP16 has an 11-bit significand: keep 11 significant bits of the value.
exp := math.Floor(math.Log2(math.Abs(float64(f))))
exp = math.Max(exp, -14) // below this, half precision loses bits (subnormals)
step := math.Pow(2, exp-10)
return float32(math.Round(float64(f)/step) * step)
}
func main() {
for _, v := range []float32{1.0, 0.1, 3.14159265, 1000.123, 65504, 70000, 1e-7} {
h := roundTrip(v)
fmt.Printf("fp32 %-12g -> fp16 %-12g error %.2e\n", v, h, math.Abs(float64(v-h)))
}
// The classic trap: adding a small number to a big one.
sum32, sum16 := float32(2048), float32(2048)
for i := 0; i < 1000; i++ {
sum32 += 0.5
sum16 = roundTrip(sum16 + 0.5)
}
fmt.Printf("\n2048 + 1000 x 0.5: fp32 = %g, fp16 = %g\n", sum32, sum16)
}The last line is the important one. In FP16, numbers near 2048 are spaced 2 apart, so adding 0.5 does nothing — a thousand times. This is exactly why accumulations stay in FP32 under mixed precision.
Remember this#
- Tensor cores do
D = A × B + Con small matrix tiles in one operation. - They provide most of a GPU’s peak arithmetic — but only for matmul, in low-precision formats.
- Fewer bits = less memory, less bandwidth, more FLOP/s, less accuracy.
- Mixed precision: heavy work in low precision, fragile steps in FP32.
Try it#
- Run
half.go. Which input overflows? Which loses the most relative accuracy? - Change the accumulation to start at 1.0 instead of 2048. Does FP16 now keep up? Why?
- An 8-billion-parameter model: how many GB in FP32, FP16, FP8 and INT4?
Check yourself#
- What single operation does a tensor core perform?
- Give two conditions for work to run on tensor cores.
- Why is summation kept in FP32 under mixed precision?