Repository navigation
RFE: Add support for more floating point low-precision ML data types (bfloat16, fp8, nvfp4) #930
Description
Activity
I support adding new data types to unlock improved performance and memory reduction (no doubt, goodness there), but we also want to judiciously weigh it with the fact that WebNN has many different backends of differing capabilities, and that these newer microscaled data types are a hot research area in flux, whereas Web standards should be durable and last for years ⚖️.
So focusing on the broadly supported types improves compatibility and reduces fragmentation (granted, the queryable
opSupportLimitsdoes enable more opt-in with graceful fallback). Of those listed above, thefloat16m7e8s1_ttype (brain float) is the most mature and longest-lived new kid on the block. I worry about these microscaled formats though with so many potential variants. I mean, there isn't simply one float8, as I've counted at least 8 in the wild! 🔍float8m3e4s1_t- mantissa: 3 bits, exponent: 4 bits, sign: 1 bit, infinities, NaNsfloat8m2e5s1_t- mantissa: 2 bits, exponent: 5 bits, sign: 1 bit, infinities, NaNsfloat8m3e4s1fnuz_t- mantissa: 3 bits, exponent: 4 bits, sign: 1 bit, no infinities, NaN as -0float8m2e5s1fnuz_t- mantissa: 2 bits, exponent: 5 bits, sign: 1 bit, no infinities, NaN as -0float8m4e3s1_t- mantissa: 4 bits, exponent: 3 bits, sign: 1 bit, infinities, NaN'sfloat8m0e8s0fn_t- mantissa: 8 bits, exponent: 0 bits, sign: 0 bits, no infinities, NaN via all ones (purely range scaling)float8m3e4b11s1fnuz_t- mantissa: 3 bits, exponent: 4 bits with bias 11, sign: 1 bit, no infinities, NaN as -0float8m3e4s1fn_t- mantissa: 3 bits, exponent: 4 bits, sign: 1 bit, infinities, NaNs
Additionally, many of these are not just the element by itself in a tensor but also entangle scaling factors and zero points wrapped into various block formats, which I have definitely lost count of now... 😵💫
So, whatever we add, we would want to pick the ones:
- supported widely by various hardware.
- available in backends CoreML/TFLite/ORT... (granted there may be some chicken-or-egg-first considerations, because WebNN can't use a feature not found in the backends, but then backends might be more inclined to add support if WebNN had it...)
- likely to stand the test of time (unlike say LSTM and GRU operators which had their highlight, but are now superseded by other operators, but add a lot of complexity to the API).
Reacted by Uday Kakade, Markus Tavenrath, mwyrzykowski and shiyiReacted by mwyrzykowskiFor fp8 I'd like follow the OCP standard (https://github.com/opencomputeproject/FP8) defined by NVIDIA, Intel, ARM, Google, AMD, and Meta, which defines float8m4e3 and float8m5e2.
Further research told me that float8m5e2 is used for training due to its higher range and and float8me3 is used for inference due to its higher precision. Since WebNN targets inference only the logical choice would be to support only float8m4e3 as described in the OCP specification.
Any specification for fp4 might be to early at the moment.
Reacted by mwyrzykowski and Dwayne RobinsonMy suggestion from the working group meeting this morning was that implementations could support the specific
constant -> castsubgraph where the constant type is not supported by the underlying framework as long as the output of the cast is. The value would be cast by the implementation before being provided to the underlying framework. This would maintain the benefit of reduced download size while keeping the site in control of the compute precision.RESOLUTION: Survey the existing backends' support for low-precision floating-point data types
For fp8 I'd like follow the OCP standard (https://github.com/opencomputeproject/FP8) ... Since WebNN targets inference only the logical choice would be to support only float8m4e3
@mtavenrath: So concretely that would be this type, right?
struct float8m3e4s1fn { uint8_t mantissa : 3; uint8_t exponent : 4; // Bias = 7 (pow(2, 4) - 1) uint8_t sign : 1; };- NaN = yes (all one's in mantissa and exponent)
- Infinity = no
- Signed and unsigned zero = yes
- Subnormal/Denormal numbers = yes
Let the backend survey commence! 👀🔍🙂
... implementations could support the specific
constant -> castsubgraph where the constant type is not supported by the underlying framework as long as the output of the cast is. The value would be cast by the implementation before being provided to the underlying framework. This would maintain the benefit of reduced download size ...@reillyeon: 🤔 You know, that could be convenient for WebNN callers for dynamic inputs too, not just constants, for cases like CoreML graphs too where uint8 is supported internally as a data type, but not as a graph input type (even though images are often uint8 per channel, making it desirable to feed it directly).
I took a look at fp8 and bfloat16 in current graphics + compute APIs.
- Vulkan has bfloat16 and fp8 (OCP) extensions.
- Apple has limited support for bfloat16. MLX & PyTorch upcast fp8 to fp16 on Apple devices.
- DirectX doesn't support bfloat16. It specifies fp8_e4m3 and fp8_e5m2 in the new linalg extension, but doesn't tell us anything about the exact variant. The documentation also states that fp8 being emulated if it's unsupported by hardware. (https://microsoft.github.io/DirectX-Specs/d3d/D3D12LinearAlgebraRuntimeFeatureSupport.html)
Reacted by Dwayne Robinson@reillyeon: 🤔 You know, that could be convenient for WebNN callers for dynamic inputs too, not just constants, for cases like CoreML graphs too where uint8 is supported internally as a data type, but not as a graph input type (even though images are often uint8 per channel, making it desirable to feed it directly).
Supporting this approach for dynamic inputs would require the casting logic to also be applied during dispatch which seems feasible but does add complexity to WebGPU interop because it would require adding an additional GPU shader to the pipeline. However I think @philloooo has already done this to make WebGPU interop work with Core ML so there's at least some precedent.
Reacted by Dwayne RobinsonDirectX ... specifies fp8_e4m3 and fp8_e5m2 in the new linalg extension, but doesn't tell us anything about the exact variant.
@mtavenrath: Yeah, adding more linalg float details is on Chris B's todo list, but linalg's "Fp8_E4M3" is really
float8m3e4s1fn_tlike above, and not an IEEE-likefloat8m3e4s1_tthat would have positive/negative infinity and multiple NaN's. See also: ONNX E4M3FN https://onnx.ai/onnx/technical/float8.html, https://www.opencompute.org/documents/ocp-microscaling-formats-mx-v1-0-spec-final-pdf, https://asawicki.info/articles/fp8_tables.php, https://arxiv.org/pdf/2209.05433).One convenient aspect of
float16m7e8s1_tis that it's trivial to upcast tofloat32m23e8s1_tfor backends that don't support it.I took a look at fp8 and bfloat16 in current graphics + compute APIs.
@mtavenrath: Updated the table with findings...
API Data type support (float8*, bfloat16) Apple CoreML x Apple BNNS BNNSDataTypeBFloat16 Apple MPS MPSDataType.bFloat16 Apple MLX x Google XNNPACK x Google ANN x Google TensorFlow tensorflow.bfloat16 Google TFLite/LiteRT kTfLiteBFloat16 Intel OpenVINO ov::element::Type_t::bf16
ov::element::Type_t::f8e4m3
ov::element::Type_t::f8e5m2
ov::element::Type_t::f4e2m1
ov::element::Type_t::f8e8m0
ov::element::Type_t::nv4Intel OneDNN dnnl_data_type_t::dnnl_f8_e4m3
dnnl_data_type_t::dnnl_f8_e5m2
dnnl_data_type_t::dnnl_bf16Microsoft CNTK x Microsoft DirectML x Microsoft Direct3D HLSL x Microsoft Direct3D LinAlg D3D12_LINEAR_ALGEBRA_DATATYPE_FLOAT8_E4M3FN
D3D12_LINEAR_ALGEBRA_DATATYPE_FLOAT8_E5M2NumPy x ONNX ONNX.TensorProto.DataType.FLOAT8E4M3FN
ONNX.TensorProto.DataType.FLOAT8E4M3FNUZ
ONNX.TensorProto.DataType.FLOAT8E5M2
ONNX.TensorProto.DataType.FLOAT8E5M2FNUZ
ONNX.TensorProto.DataType.BFLOAT16PyTorch torch.bfloat16
torch.float8_e4m3fn
torch.float8_e5m2
torch.float8_e4m3fnuz
torch.float8_e5m2fnuz
torch.float8_e8m0fnu
torch.float4_e2m1fn_x2StableHLO f8E4M3FN
f8E4M3FNUZ
f8E5M2
f8E5M2FNUZ
f8E4M3B11FNUZ
bf16Tencent NCNN bfloat16 TOSA tosa.bf16_t WebNN x Reacted by shiyi@fdwr , OpenVINO supports bf16, f8e4m3 and f8e5m2 element_type.hpp
Reacted by Dwayne RobinsonRESOLUTION: Draft a concrete proposal based on the survey results documented in the issue and update CONTRIBUTING.md with polyfill guidance. (issue #930)
@fdwr , OpenVINO supports bf16, f8e4m3 and f8e5m2 element_type.hpp
@huningxin: Thanks - updated table with newest enums.
I am adding float8 support to TFLite/LiteRT and XNNPACK. Then I found the hardware coverage is very limited today.
fp8: Now standardized by the Open Compute Project and natively supported by current-gen hardware (NVIDIA Hopper/Blackwell, Ryzen AI, RDNA4, Intel XE2 (Lunar Lake, ARC B-series)).
It's true for Nvidia/AMD GPUs.
But I wasn't able to find a way to enable it for Intel GPUs. Though OpenVINO supports bf16, f8e4m3 and f8e5m2, it does not seem to have fp8 kernels for the Arc GPU I have . Also, the GPU doc does not list fp8 as a supported data type: https://docs.openvino.ai/2026/openvino-workflow/running-inference/inference-devices-and-modes/gpu-device.html#supported-inference-data-types . The only supported data types are 'FP32', 'FP16' and 'INT8'. You can use fp8 with fake convert: https://docs.openvino.ai/2026/documentation/openvino-ir-format/operation-sets/operation-specs/quantization/fake-convert-13.html . But, I guess it would not bring any perf benefit.
I also checked a lot of different kinds of mobile NPUs. The situation is similar. They mainly only do integer maths.
Some Intel and ARM CPUs have FP8 support, but as of today they are all server CPUs. I believe eventually they will be available on consumer devices.
I didn't try Intel/AMD NPUs.
Reacted by shiyiReacted by Dwayne Robinson@snnn, we should distinct the hardware capabilities (native support in matrix multiplication units) and the software. For example, it is absolutely fine to keep the data in fp8 and then unpack it on the registers at runtime. That scenario keeps the memory footprint lower, as well as benefits memory-bound kernels due to reduced amount of data to transfer.
To illustrate this: oneDNN (OpenVINO's kernel provider) lists fp8 as supported data types for GEMMs and convolutions. There could be some limitations throughout the software stack, but looking at the list of PRs for OpenVINO, I'd say that they're working on it. And fp8 hardware support was added to NPU in Panther Lake.
Overall, even if the hardware doesn't support a specific data type in their matrix multiplication units, that doesn't necessarily rule out the benefit of using lower-precision data types for inference.
Then we lose the ground that these data types should be added because hardware already support them. Because, otherwise the list would be much broader. Almost any quantization data type can be unpacked to float. For example, all the data types in https://github.com/jax-ml/ml_dtypes .
I am adding float8 support to TFLite/LiteRT and XNNPACK.
@snnn Greetings Changming. I see you're adding them here. I wonder which ones specifically they correspond to? ⭐?
- ⭐
float8m3e4s1_t- mantissa: 3 bits, exponent: 4 bits, sign: 1 bit, infinities, NaNs float8m3e4s1fn_t- mantissa: 3 bits, exponent: 4 bits, sign: 1 bit, infinities, NaNsfloat8m3e4s1fnuz_t- mantissa: 3 bits, exponent: 4 bits, sign: 1 bit, no infinities, NaN as -0float8m3e4b11s1fnuz_t- mantissa: 3 bits, exponent: 4 bits with bias 11, sign: 1 bit, no infinities, NaN as -0- ⭐
float8m2e5s1_t- mantissa: 2 bits, exponent: 5 bits, sign: 1 bit, infinities, NaNs float8m2e5s1fnuz_t- mantissa: 2 bits, exponent: 5 bits, sign: 1 bit, no infinities, NaN as -0
I found the hardware coverage is very limited today. ... I also checked a lot of different kinds of mobile NPUs. The situation is similar. They mainly only do integer maths. ... Some Intel and ARM CPUs have FP8 support, but as of today they are all server CPUs. I believe eventually they will be available on consumer devices.
Thanks for the findings. Yes, more hardware is likely to come to the consumer.
- ⭐
@snnn, I think the baseline for this discussion is slightly different: it's not only what data types are natively supported by the hardware today, but rather what is supported by the vendor ecosystem and the corresponding software stack, or have a clear path to do so.
In this case, it is important that WebNN specification reflects real use cases and what major vendors are investing in: training approaches, quantization schemes, inference runtimes, silicon roadmaps, etc. It is true that many quantization schemes can be unpacked to
fp32, but with the swift development of AI we won't be able to keep pace with every new quantization format.Given that the silicon development lifecycle is long, it is natural for different data types to land with a software emulation first, get backported to previous generations, and then gain hardware acceleration as adoption and demand grow. In the case of
fp8, even the emulated path provides performance and memory footprint improvements overfp16, which is why it's getting more traction every day.See low-precision floating point data types explainer PR #938, for discussion on 2026-08-13.
Reacted by Dwayne Robinson
Problem Statement
Currently, the WebNN API supports standard data types like
float32,float16,int8, anduint8. However, the machine learning hardware ecosystem has rapidly migrated toward specialized, low-precision floating-point formats to handle the extreme memory bandwidth and compute demands of modern Large Language Models (LLMs) and diffusion models.Developers bringing state-of-the-art models to the web are currently forced into a strict dilemma:
bfloat16orfp8models tofloat32respectivefloat16preserves the model's accuracy and dynamic range, but immediately bloats the memory footprint and exacerbates memory bandwidth bottlenecks, which can be fatal for client-side inference.int8orint4achieves the desired performance and footprint, but sacrifices precision. Integers space values uniformly, which crushes the activation outliers prevalent in LLMs, leading to a degradation in reasoning quality and perplexity.Without native WebNN support for modern low-precision floating-point formats, the API forces developers to choose between performance and accuracy, losing out on the hardware-accelerated "best of both worlds" that these new formats were designed to provide.
The Case for Each Format
bfloat16(Brain Floating Point)float32, preventing gradient underflow/overflow.bfloat16-trained model to WebNN currently requires an offline conversion tofloat32(doubling the model size) orfloat16(which requires careful scaling to avoid out-of-bounds errors).fp8(OCP 8-bit Floating Point - E4M3 & E5M2)fp8acts as the perfect middle ground between the two extremes mentioned above. It provides the memory footprint ofint8while maintaining the dynamic range of a floating-point exponent, allowing it to handle activation outliers without precision loss.nvfp4(4-bit Microscaled Floating Point)Proposed Solution
Update the
MLOperandDataTypeenum and underlying buffer compatibility tables to include:bfloat16fp8(potentially differentiating between E4M3 and E5M2 variants)nvfp4(or a generic hardware-agnostic equivalent for 4-bit block-scaled floating point)Use Cases & Benefits
fp8,nvfp4) retains the dynamic range of activations much better than standard integer quantization, reducing perplexity degradation.