QF8 (38.1 dB) vs FP8 E4M3 (31.5 dB) on N(0,1) — a 4.6× reduction in quantization noise power. The advantage is consistent (+6.3 to +6.8 dB) across Gaussian, log-normal, Laplace, and 90%-sparse distributions, and across matrix sizes from 16×32×16 to 128×256×128. This is structural: 16 vs. 8 levels per octave.
QF8 (QuakeFloat8) is an 8-bit log-domain number format that provides 16 representable levels per octave (factor-of-2 range), compared to 8 levels per octave for IEEE FP8 E4M3 (the industry-standard 8-bit floating-point format used in NVIDIA H100, AMD MI300, etc.). Both formats use the same effective storage: 8.25 bits per element (8 per-element bits + shared block exponent).
This entry presents benchmark results showing that QF8's 2× finer per-octave resolution translates to a consistent SQNR (signal-to-quantization-noise ratio) advantage of +6.6 dB — equivalent to a 4.6× reduction in quantization noise power.
| Format | SQNR on N(0,1) | Gap from QF8 |
|---|---|---|
| QF8 (log-uniform, 256 levels) | 38.1 dB | — |
| FP8 E4M3 (IEEE) | 31.5 dB | 6.6 dB worse |
This advantage is predicted by the minimax NMSE theorem and confirmed empirically below.
Comparing at equal levels-per-octave (8), log-uniform still slightly wins due to truly constant relative cell width vs. FP8's piecewise-constant structure:
| Metric | IEEE E4M3 | Log-uniform (8/oct) | Ratio |
|---|---|---|---|
| Average MSRE | 6.510×10⁻⁴ | 6.256×10⁻⁴ | 1.041 |
| Peak relative error | 0.0625 | 0.04427 | 1.41 |
| Worst-case MSRE (HR) | 1.302×10⁻³ | 6.256×10⁻⁴ | 2.08 |
But QF8's main advantage is structural: 16 levels/octave vs. 8.
Measured vs. float64 ground truth on random matrices:
| Size | bfloat16 | float32 | FP8 E4M3 (block) | QF8 |
|---|---|---|---|---|
| 16×32×16 | 52.3 dB | 139.0 dB | 28.6 dB | 35.3 dB |
| 64×128×64 | 52.5 dB | 133.6 dB | 28.4 dB | 35.1 dB |
| 128×256×128 | 52.6 dB | 130.8 dB | 28.5 dB | 35.1 dB |
The advantage is consistent across matrix sizes, confirming the dimension-free property.
| Distribution | FP8 E4M3 | QF8 | Advantage |
|---|---|---|---|
| N(0, 0.02²) — typical weights | 31.4 dB | 38.2 dB | +6.7 dB |
| N(0, 1) — post-LayerNorm activations | 31.6 dB | 37.9 dB | +6.3 dB |
| LogN(0, 1) — gradient-like | 31.5 dB | 38.3 dB | +6.8 dB |
| Laplace(0, 0.02) — sparse weights | 31.5 dB | 38.0 dB | +6.5 dB |
| 90% sparse + N(0,1) | 31.7 dB | 38.3 dB | +6.6 dB |
The +6.3 to +6.8 dB advantage is stable across all tested distributions. This is expected from the minimax theorem: log-uniform quantization achieves the same NMSE regardless of input distribution, so the advantage over FP8 is structural rather than distribution-specific.