Deep-Learning Floating-Point Formats: FP32 / TF32 / FP16 / BF16 / FP8
Bit layout (sign/exponent/mantissa) and use of common DL float formats; exponent bits set range, mantissa bits set precision.
| Format | Total bits | Sign | Exponent | Mantissa | Notes |
|---|---|---|---|---|---|
| FP32 (single) | 32 | 1 | 8 | 23 | IEEE 754; training precision baseline |
| TF32 (NVIDIA) | 19 | 1 | 8 | 10 | Ampere+ Tensor Cores; FP32 range, ~FP16 precision |
| FP16 (half) | 16 | 1 | 5 | 10 | IEEE 754; narrow range, often needs loss scaling |
| BF16 (bfloat16) | 16 | 1 | 8 | 7 | FP32 dynamic range, lower precision; common for training |
| FP8 E4M3 | 8 | 1 | 4 | 3 | Higher precision; forward/weights |
| FP8 E5M2 | 8 | 1 | 5 | 2 | Wider range; gradients |
Method & sources
Compiled from IEEE 754 (FP32/FP16), NVIDIA mixed-precision docs (TF32), Google's bfloat16 explainer, and the FP8 (E4M3/E5M2) interchange spec. Bit counts are definitional.
Source: https://docs.nvidia.com/deeplearning/performance/mixed-precision-training/index.html
Retrieved: 2026-09-05
FAQ
- What does "Deep-Learning Floating-Point Formats: FP32 / TF32 / FP16 / BF16 / FP8" cover?
- It covers FP32 (single), TF32 (NVIDIA), FP16 (half), BF16 (bfloat16), FP8 E4M3, FP8 E5M2, compared across: Format, Total bits, Sign, Exponent, Mantissa, Notes.
- What are the sources and methodology?
- Compiled from IEEE 754 (FP32/FP16), NVIDIA mixed-precision docs (TF32), Google's bfloat16 explainer, and the FP8 (E4M3/E5M2) interchange spec.
- When was this data last updated?
- The data was retrieved/updated on 2026-09-05.