Skip to content

RFE: Add support for more floating point low-precision ML data types (bfloat16, fp8, nvfp4) #930

Description

@mtavenrath

Problem Statement
Currently, the WebNN API supports standard data types like float32, float16, int8, and uint8. However, the machine learning hardware ecosystem has rapidly migrated toward specialized, low-precision floating-point formats to handle the extreme memory bandwidth and compute demands of modern Large Language Models (LLMs) and diffusion models.

Developers bringing state-of-the-art models to the web are currently forced into a strict dilemma:

  1. Upcast to keep precision: Converting native bfloat16 or fp8 models to float32 respective float16 preserves the model's accuracy and dynamic range, but immediately bloats the memory footprint and exacerbates memory bandwidth bottlenecks, which can be fatal for client-side inference.
  2. Quantize to integer to save memory: Converting to int8 or int4 achieves the desired performance and footprint, but sacrifices precision. Integers space values uniformly, which crushes the activation outliers prevalent in LLMs, leading to a degradation in reasoning quality and perplexity.

Without native WebNN support for modern low-precision floating-point formats, the API forces developers to choose between performance and accuracy, losing out on the hardware-accelerated "best of both worlds" that these new formats were designed to provide.

The Case for Each Format

  1. bfloat16 (Brain Floating Point)

    • Context: Widely used as the default format for training modern AI models because its 8-bit exponent matches float32, preventing gradient underflow/overflow.
    • Issue: Porting a natively bfloat16-trained model to WebNN currently requires an offline conversion to float32 (doubling the model size) or float16 (which requires careful scaling to avoid out-of-bounds errors).
  2. fp8 (OCP 8-bit Floating Point - E4M3 & E5M2)

    • Context: Now standardized by the Open Compute Project and natively supported by current-gen hardware (NVIDIA Hopper/Blackwell, Ryzen AI, RDNA4, Intel XE2 (Lunar Lake, ARC B-series)).
    • Issue: fp8 acts as the perfect middle ground between the two extremes mentioned above. It provides the memory footprint of int8 while maintaining the dynamic range of a floating-point exponent, allowing it to handle activation outliers without precision loss.
  3. nvfp4 (4-bit Microscaled Floating Point)

    • Context: Introduced with NVIDIA's Blackwell architecture, this format represents the bleeding edge of sub-byte quantization. It utilizes a two-level microscaling block approach (sharing an FP8 scale across 16 values) to drastically reduce quantization error compared to pure integer 4-bit.
    • Issue: Allowing hardware-accelerated 4-bit floating-point inference in the browser would be a massive leap for client-side LLM execution, cutting memory usage by ~3.5x compared to FP16 while maintaining near-baseline accuracy.

Proposed Solution
Update the MLOperandDataType enum and underlying buffer compatibility tables to include:

  • bfloat16
  • fp8 (potentially differentiating between E4M3 and E5M2 variants)
  • nvfp4 (or a generic hardware-agnostic equivalent for 4-bit block-scaled floating point)

Use Cases & Benefits

  • Quicker Model Porting: Developers can load original weights directly into the graph without complex offline upcasting/re-quantization pipelines.
  • Higher Quantization Accuracy: Floating-point quantization (fp8, nvfp4) retains the dynamic range of activations much better than standard integer quantization, reducing perplexity degradation.
  • Reduced Memory Pressure: Keeping tensors in their ultra-low precision formats slashes both memory footprint and memory bandwidth bottlenecks during inference on edge devices.

Activity

  1. fdwr commented on Apr 28, 2026

    @fdwr
    Collaborator

    I support adding new data types to unlock improved performance and memory reduction (no doubt, goodness there), but we also want to judiciously weigh it with the fact that WebNN has many different backends of differing capabilities, and that these newer microscaled data types are a hot research area in flux, whereas Web standards should be durable and last for years ⚖️.

    So focusing on the broadly supported types improves compatibility and reduces fragmentation (granted, the queryable opSupportLimits does enable more opt-in with graceful fallback). Of those listed above, the float16m7e8s1_t type (brain float) is the most mature and longest-lived new kid on the block. I worry about these microscaled formats though with so many potential variants. I mean, there isn't simply one float8, as I've counted at least 8 in the wild! 🔍

    • float8m3e4s1_t - mantissa: 3 bits, exponent: 4 bits, sign: 1 bit, infinities, NaNs
    • float8m2e5s1_t - mantissa: 2 bits, exponent: 5 bits, sign: 1 bit, infinities, NaNs
    • float8m3e4s1fnuz_t - mantissa: 3 bits, exponent: 4 bits, sign: 1 bit, no infinities, NaN as -0
    • float8m2e5s1fnuz_t - mantissa: 2 bits, exponent: 5 bits, sign: 1 bit, no infinities, NaN as -0
    • float8m4e3s1_t - mantissa: 4 bits, exponent: 3 bits, sign: 1 bit, infinities, NaN's
    • float8m0e8s0fn_t - mantissa: 8 bits, exponent: 0 bits, sign: 0 bits, no infinities, NaN via all ones (purely range scaling)
    • float8m3e4b11s1fnuz_t - mantissa: 3 bits, exponent: 4 bits with bias 11, sign: 1 bit, no infinities, NaN as -0
    • float8m3e4s1fn_t - mantissa: 3 bits, exponent: 4 bits, sign: 1 bit, infinities, NaNs

    Additionally, many of these are not just the element by itself in a tensor but also entangle scaling factors and zero points wrapped into various block formats, which I have definitely lost count of now... 😵‍💫

    So, whatever we add, we would want to pick the ones:

    • supported widely by various hardware.
    • available in backends CoreML/TFLite/ORT... (granted there may be some chicken-or-egg-first considerations, because WebNN can't use a feature not found in the backends, but then backends might be more inclined to add support if WebNN had it...)
    • likely to stand the test of time (unlike say LSTM and GRU operators which had their highlight, but are now superseded by other operators, but add a lot of complexity to the API).
  2. mtavenrath commented on May 7, 2026

    @mtavenrath
    Author

    For fp8 I'd like follow the OCP standard (https://github.com/opencomputeproject/FP8) defined by NVIDIA, Intel, ARM, Google, AMD, and Meta, which defines float8m4e3 and float8m5e2.

    Further research told me that float8m5e2 is used for training due to its higher range and and float8me3 is used for inference due to its higher precision. Since WebNN targets inference only the logical choice would be to support only float8m4e3 as described in the OCP specification.

    Any specification for fp4 might be to early at the moment.

  3. reillyeon commented on May 7, 2026

    @reillyeon
    Contributor

    My suggestion from the working group meeting this morning was that implementations could support the specific constant -> cast subgraph where the constant type is not supported by the underlying framework as long as the output of the cast is. The value would be cast by the implementation before being provided to the underlying framework. This would maintain the benefit of reduced download size while keeping the site in control of the compute precision.

  4. anssiko commented on May 7, 2026

    @anssiko
    Member

    RESOLUTION: Survey the existing backends' support for low-precision floating-point data types

  5. fdwr commented on May 8, 2026

    @fdwr
    Collaborator

    For fp8 I'd like follow the OCP standard (https://github.com/opencomputeproject/FP8) ... Since WebNN targets inference only the logical choice would be to support only float8m4e3

    @mtavenrath: So concretely that would be this type, right?

    struct float8m3e4s1fn
    {
        uint8_t mantissa : 3;
        uint8_t exponent : 4; // Bias = 7 (pow(2, 4) - 1)
        uint8_t sign     : 1;
    };    
    
    • NaN = yes (all one's in mantissa and exponent)
    • Infinity = no
    • Signed and unsigned zero = yes
    • Subnormal/Denormal numbers = yes

    Let the backend survey commence! 👀🔍🙂

    ... implementations could support the specific constant -> cast subgraph where the constant type is not supported by the underlying framework as long as the output of the cast is. The value would be cast by the implementation before being provided to the underlying framework. This would maintain the benefit of reduced download size ...

    @reillyeon: 🤔 You know, that could be convenient for WebNN callers for dynamic inputs too, not just constants, for cases like CoreML graphs too where uint8 is supported internally as a data type, but not as a graph input type (even though images are often uint8 per channel, making it desirable to feed it directly).

  6. mtavenrath commented on May 8, 2026

    @mtavenrath
    Author

    I took a look at fp8 and bfloat16 in current graphics + compute APIs.

  7. reillyeon commented on May 8, 2026

    @reillyeon
    Contributor

    @reillyeon: 🤔 You know, that could be convenient for WebNN callers for dynamic inputs too, not just constants, for cases like CoreML graphs too where uint8 is supported internally as a data type, but not as a graph input type (even though images are often uint8 per channel, making it desirable to feed it directly).

    Supporting this approach for dynamic inputs would require the casting logic to also be applied during dispatch which seems feasible but does add complexity to WebGPU interop because it would require adding an additional GPU shader to the pipeline. However I think @philloooo has already done this to make WebGPU interop work with Core ML so there's at least some precedent.

  8. fdwr commented on May 8, 2026

    @fdwr
    Collaborator

    DirectX ... specifies fp8_e4m3 and fp8_e5m2 in the new linalg extension, but doesn't tell us anything about the exact variant.

    @mtavenrath: Yeah, adding more linalg float details is on Chris B's todo list, but linalg's "Fp8_E4M3" is really float8m3e4s1fn_t like above, and not an IEEE-like float8m3e4s1_t that would have positive/negative infinity and multiple NaN's. See also: ONNX E4M3FN https://onnx.ai/onnx/technical/float8.html, https://www.opencompute.org/documents/ocp-microscaling-formats-mx-v1-0-spec-final-pdf, https://asawicki.info/articles/fp8_tables.php, https://arxiv.org/pdf/2209.05433).

    One convenient aspect of float16m7e8s1_t is that it's trivial to upcast to float32m23e8s1_t for backends that don't support it.

    I took a look at fp8 and bfloat16 in current graphics + compute APIs.

    @mtavenrath: Updated the table with findings...

    API Data type support (float8*, bfloat16)
    Apple CoreML x
    Apple BNNS BNNSDataTypeBFloat16
    Apple MPS MPSDataType.bFloat16
    Apple MLX x
    Google XNNPACK x
    Google ANN x
    Google TensorFlow tensorflow.bfloat16
    Google TFLite/LiteRT kTfLiteBFloat16
    Intel OpenVINO ov::element::Type_t::bf16
    ov::element::Type_t::f8e4m3
    ov::element::Type_t::f8e5m2
    ov::element::Type_t::f4e2m1
    ov::element::Type_t::f8e8m0
    ov::element::Type_t::nv4
    Intel OneDNN dnnl_data_type_t::dnnl_f8_e4m3
    dnnl_data_type_t::dnnl_f8_e5m2
    dnnl_data_type_t::dnnl_bf16
    Microsoft CNTK x
    Microsoft DirectML x
    Microsoft Direct3D HLSL x
    Microsoft Direct3D LinAlg D3D12_LINEAR_ALGEBRA_DATATYPE_FLOAT8_E4M3FN
    D3D12_LINEAR_ALGEBRA_DATATYPE_FLOAT8_E5M2
    NumPy x
    ONNX ONNX.TensorProto.DataType.FLOAT8E4M3FN
    ONNX.TensorProto.DataType.FLOAT8E4M3FNUZ
    ONNX.TensorProto.DataType.FLOAT8E5M2
    ONNX.TensorProto.DataType.FLOAT8E5M2FNUZ
    ONNX.TensorProto.DataType.BFLOAT16
    PyTorch torch.bfloat16
    torch.float8_e4m3fn
    torch.float8_e5m2
    torch.float8_e4m3fnuz
    torch.float8_e5m2fnuz
    torch.float8_e8m0fnu
    torch.float4_e2m1fn_x2
    StableHLO f8E4M3FN
    f8E4M3FNUZ
    f8E5M2
    f8E5M2FNUZ
    f8E4M3B11FNUZ
    bf16
    Tencent NCNN bfloat16
    TOSA tosa.bf16_t
    WebNN x
  9. huningxin commented on May 21, 2026

    @huningxin
    Contributor

    @fdwr , OpenVINO supports bf16, f8e4m3 and f8e5m2 element_type.hpp

  10. anssiko commented on May 21, 2026

    @anssiko
    Member

    RESOLUTION: Draft a concrete proposal based on the survey results documented in the issue and update CONTRIBUTING.md with polyfill guidance. (issue #930)

  11. fdwr commented on May 21, 2026

    @fdwr
    Collaborator

    @fdwr , OpenVINO supports bf16, f8e4m3 and f8e5m2 element_type.hpp

    @huningxin: Thanks - updated table with newest enums.

  12. snnn commented on Jun 23, 2026

    @snnn

    I am adding float8 support to TFLite/LiteRT and XNNPACK. Then I found the hardware coverage is very limited today.

    fp8: Now standardized by the Open Compute Project and natively supported by current-gen hardware (NVIDIA Hopper/Blackwell, Ryzen AI, RDNA4, Intel XE2 (Lunar Lake, ARC B-series)).

    It's true for Nvidia/AMD GPUs.

    But I wasn't able to find a way to enable it for Intel GPUs. Though OpenVINO supports bf16, f8e4m3 and f8e5m2, it does not seem to have fp8 kernels for the Arc GPU I have . Also, the GPU doc does not list fp8 as a supported data type: https://docs.openvino.ai/2026/openvino-workflow/running-inference/inference-devices-and-modes/gpu-device.html#supported-inference-data-types . The only supported data types are 'FP32', 'FP16' and 'INT8'. You can use fp8 with fake convert: https://docs.openvino.ai/2026/documentation/openvino-ir-format/operation-sets/operation-specs/quantization/fake-convert-13.html . But, I guess it would not bring any perf benefit.

    I also checked a lot of different kinds of mobile NPUs. The situation is similar. They mainly only do integer maths.

    Some Intel and ARM CPUs have FP8 support, but as of today they are all server CPUs. I believe eventually they will be available on consumer devices.

    I didn't try Intel/AMD NPUs.

  13. mklimenko-nv commented on Jun 23, 2026

    @mklimenko-nv
    Contributor

    @snnn, we should distinct the hardware capabilities (native support in matrix multiplication units) and the software. For example, it is absolutely fine to keep the data in fp8 and then unpack it on the registers at runtime. That scenario keeps the memory footprint lower, as well as benefits memory-bound kernels due to reduced amount of data to transfer.

    To illustrate this: oneDNN (OpenVINO's kernel provider) lists fp8 as supported data types for GEMMs and convolutions. There could be some limitations throughout the software stack, but looking at the list of PRs for OpenVINO, I'd say that they're working on it. And fp8 hardware support was added to NPU in Panther Lake.

    Overall, even if the hardware doesn't support a specific data type in their matrix multiplication units, that doesn't necessarily rule out the benefit of using lower-precision data types for inference.

  14. snnn commented on Jun 23, 2026

    @snnn

    Then we lose the ground that these data types should be added because hardware already support them. Because, otherwise the list would be much broader. Almost any quantization data type can be unpacked to float. For example, all the data types in https://github.com/jax-ml/ml_dtypes .

  15. fdwr commented on Jun 23, 2026

    @fdwr
    Collaborator

    I am adding float8 support to TFLite/LiteRT and XNNPACK.

    @snnn Greetings Changming. I see you're adding them here. I wonder which ones specifically they correspond to? ⭐?

    • ⭐ float8m3e4s1_t - mantissa: 3 bits, exponent: 4 bits, sign: 1 bit, infinities, NaNs
    • float8m3e4s1fn_t - mantissa: 3 bits, exponent: 4 bits, sign: 1 bit, infinities, NaNs
    • float8m3e4s1fnuz_t - mantissa: 3 bits, exponent: 4 bits, sign: 1 bit, no infinities, NaN as -0
    • float8m3e4b11s1fnuz_t - mantissa: 3 bits, exponent: 4 bits with bias 11, sign: 1 bit, no infinities, NaN as -0
    • ⭐ float8m2e5s1_t - mantissa: 2 bits, exponent: 5 bits, sign: 1 bit, infinities, NaNs
    • float8m2e5s1fnuz_t - mantissa: 2 bits, exponent: 5 bits, sign: 1 bit, no infinities, NaN as -0

    I found the hardware coverage is very limited today. ... I also checked a lot of different kinds of mobile NPUs. The situation is similar. They mainly only do integer maths. ... Some Intel and ARM CPUs have FP8 support, but as of today they are all server CPUs. I believe eventually they will be available on consumer devices.

    Thanks for the findings. Yes, more hardware is likely to come to the consumer.

  16. mklimenko-nv commented on Jun 24, 2026

    @mklimenko-nv
    Contributor

    @snnn, I think the baseline for this discussion is slightly different: it's not only what data types are natively supported by the hardware today, but rather what is supported by the vendor ecosystem and the corresponding software stack, or have a clear path to do so.

    In this case, it is important that WebNN specification reflects real use cases and what major vendors are investing in: training approaches, quantization schemes, inference runtimes, silicon roadmaps, etc. It is true that many quantization schemes can be unpacked to fp32, but with the swift development of AI we won't be able to keep pace with every new quantization format.

    Given that the silicon development lifecycle is long, it is natural for different data types to land with a software emulation first, get backported to previous generations, and then gain hardware acceleration as adoption and demand grow. In the case of fp8, even the emulated path provides performance and memory footprint improvements over fp16, which is why it's getting more traction every day.

  17. anssiko commented on Jun 29, 2026

    @anssiko
    Member

    See low-precision floating point data types explainer PR #938, for discussion on 2026-08-13.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions