Deep-Learning Floating-Point Formats: FP32 / TF32 / FP16 / BF16 / FP8

Bit layout (sign/exponent/mantissa) and use of common DL float formats; exponent bits set range, mantissa bits set precision.

FormatTotal bitsSignExponentMantissaNotes
FP32 (single)321823IEEE 754; training precision baseline
TF32 (NVIDIA)191810Ampere+ Tensor Cores; FP32 range, ~FP16 precision
FP16 (half)161510IEEE 754; narrow range, often needs loss scaling
BF16 (bfloat16)16187FP32 dynamic range, lower precision; common for training
FP8 E4M38143Higher precision; forward/weights
FP8 E5M28152Wider range; gradients

Method & sources

Compiled from IEEE 754 (FP32/FP16), NVIDIA mixed-precision docs (TF32), Google's bfloat16 explainer, and the FP8 (E4M3/E5M2) interchange spec. Bit counts are definitional.

Source: https://docs.nvidia.com/deeplearning/performance/mixed-precision-training/index.html

Retrieved: 2026-09-05

FAQ

What does "Deep-Learning Floating-Point Formats: FP32 / TF32 / FP16 / BF16 / FP8" cover?
It covers FP32 (single), TF32 (NVIDIA), FP16 (half), BF16 (bfloat16), FP8 E4M3, FP8 E5M2, compared across: Format, Total bits, Sign, Exponent, Mantissa, Notes.
What are the sources and methodology?
Compiled from IEEE 754 (FP32/FP16), NVIDIA mixed-precision docs (TF32), Google's bfloat16 explainer, and the FP8 (E4M3/E5M2) interchange spec.
When was this data last updated?
The data was retrieved/updated on 2026-09-05.