Skip to content

Core operator set #573

Description

@philloooo

In talking with internal Google ML frameworks teams, one theme has come up repeatedly when discussing ML execution: the need for a predictable, core set of operators with precisely defined behavior. Without this, frameworks can't provide predictable behavior, and can't reliably express higher level concepts if they are missing from a given execution runtime. We've seen work by internal and external ML frameworks towards defining these core operator sets, and believe this concept is important for WebNN to adopt, and ideally align with any emerging standards for core op sets.

We'd like to build consensus on the following:

  • WebNN should define a core op set, which focuses on the low level ops that are indecomposable and captures functional completeness of the API.
  • Implementations of WebNN MUST (in the RFC 2119 sense) implement this core op set.
  • The behavior of these core ops will be specified precisely with conformance tests in WPT.
  • We must validate with multiple ML frameworks that the identified core op set meets their needs.

Follow-up work:

  • Actually define the core op set - both the list of ops and their behavior
  • Have at least 2 implementations to make sure the interface including constraints specified can be supported by multiple platforms
  • Come up with a rubric for how rigorously the core op set is limited
    • E.g. Would we include both sin() and cos(), even though you can define one in terms of the other? Do we only need nand() because then you can make and/or/no/xor ?
  • Determine if a subset of a "standard" core op set is acceptable for v1 (i.e. do we need control flow Control flow operations: if, while #559 and bitwise operators Bitwise operators and logical operators naming (rename not to logicalNot) #496 ?)
  • Define core op set standardization / evolution over time (e.g. in conjunction with frameworks)


Related questions, but maybe out of scope for this issue:

  • What do we call non-core ops? (Composite? High-level? …)
  • Should all non-core ops be defined in terms of these core ops?
  • How should we structure the spec to make core vs non-core ops clear?
  • How precisely should the behavior of non-core ops be constrained?

There are some high level questions that need to be hashed out:

  • Similar to GPUSupportedLimits, whether/how to expose limits that are backend specific? This probably needs more implementation experience before answering.

See also:

Activity

  1. philloooo commented on Feb 21, 2024

    @philloooo
    ContributorAuthor

    One example we can compare with is the Pytorch Edge opset.

  2. wchao1115 commented on Feb 27, 2024

    @wchao1115
    Collaborator

    When we originally considered the design principles of the operator set for WebNN since even before the spec made it out of the community group, the topic of whether the core operator set should also include common high-level operations (commonly known as fusions) or just the rudimentary building block operations came up in many discussions.

    Without getting into a philosophical debate of what exactly, besides basic arithmetic operations, should be considered rudimentary operations, we set out to look at it objectively from the context of what already being developed in the industry at various software and hardware layers, and concluded that it is important both for practicality and performance reason to also include common fusions known to be implemented widely in the framework and the underlying platform (e.g. the operating system) layer, with an important caveat that for each fusion defined in the spec, every decomposed operations of its subgraph equivalence must also be defined.

    The main objective of this rule is to support the implementation that has yet to fully support a specific fusion to be able to carry on without failing. By pushing operator decomposition downward, we allow the implementation to catch up on it at a later time while simplifying the development of the framework's backend for WebNN. Also note that a keyword here is "common" fusions and not any fusion with unverifiable availability in the platform layers.

    For reference, we took the opportunity at that time to describe this rationale in this section of our explainer document.

    At the time of that writing, we used GRU and LSTM as de facto samples to describe such common fusions in the discussion. With the emergence of generative AI in the recent year, the better examples would be group-norm, layer-norm, and multi-headed attention -- operations that are widely used in both the diffusion and transformer models nowadays.

  3. anssiko commented on Feb 27, 2024

    @anssiko
    Member

    Thanks @philloooo for soliciting input from internal Google teams and @wchao1115 for your insights and various contributions in this space.

    The group has produced the following documentation on this topic:

    In addition, the group has received related position statements from Google in #453 and #573 (this issue), and input from an ONNX project participant on that project's approach. There may be more, but I was able to recall these.

    As @wchao1115 noted, this topic has been a long-term consideration and I note the topic re-emerges from time to time when new participants join. This suggests to me the group cares about this topic, and also that we could probably do better in documenting the group's current consensus position.

    To help transform this into action, I'd like to ask whether the group is happy with the current organization of related guidelines and documentation, or should we try to perhaps consolidate them somehow? Fold (more of) them into the specification? Is there some further investigation and research to be done to ensure we are well informed and data-driven? A review of widely deployed frameworks (expected users of this group's deliverables) with results shared in public?

    Regardless of where this content lives in our repo, I expect this documentation to evolve and be maintained. Everyone's feedback is welcome.

  4. philloooo commented on Mar 7, 2024

    @philloooo
    ContributorAuthor

    thanks @wchao1115 for elaborating the current design philosophy! I am overall agreed with current design. Thanks @anssiko for linking to the existing resources!

    This issue is not trying to exclude high-level operations from the spec. It's trying to bring up couple things related to this topic:

    • Currently comparing with an op set like pytorch edge set, webnn still has some core ops missing. We are also trying to work with internal teams to see if we can share something similar for StableHLO. So I created this issue to track the status of how "complete" is our core op set.
    • How we present the core op set with other high-level/composite ops in the spec is up to discussion. We can present as is now, so they are differentiated by whether there are decomposition defined. Or we put them to different sections, or we just add annotations for core ops.
    • We should ensure the core op set to behave the same across backends. The high level/composite ops are easier to diverge between backends though. Take lstml as example, the supported activations are different across backends. We need a solution to better support composite ops.
  5. anssiko commented on Nov 7, 2024

    @anssiko
    Member

    We'll revive this issue with a discussion on additional primitive ops informed by MLIR Linalg, PyTorch Prims IR, TOSA, others. Pointers to relevant background research can be shared here to inform the discussion.

  6. fdwr commented on Nov 12, 2024

    @fdwr
    Collaborator

    core set of operators with precisely defined behavior ... reliably express higher level concepts if they are missing from a given execution runtime

    I have a work-in-progress preliminary analysis of operator correspondence here Machine Learning Operator Mapping.xlsx which can help identify current primitive gaps (no firm recommendations yet), as the current WebNN standard has demonstrated viability of popular models, but it lacks breadth. However, implementing 800+ operators (PyTorch, TensorFlow, ...) would be untenable for a web specification that needs to be more rigorously documented and implemented across many user agents with many potential platform backends. Plus there will always be niche operators found in some ML libraries that aren't justifiably common enough to be part of WebNN itself but could still be constructed from WebNN operators, given WebNN has sufficient foundation. So the need is clear for compositional fundamentals (e.g. PyTorch prims, TOSA, StableHLO) which means potentially even adding operators to WebNN which on their own might not directly be that useful in neural networks (like sumPooling and bitwise operators) but are useful for composition.

    Current thinking:

  7. fdwr commented on Oct 9, 2025

    @fdwr
    Collaborator

    Related, we've had this concept of aggregate operators via subgraphs (Ningxin presented the idea at TPAC 2024), where you can define a composite operator and then execute it. Then if the backend has a compatible implementation of the subgraph as a builtin operator, passing that higher level of expression through can be simpler/more efficient than needing to recognize the patterns nested throughout the graph and refuse them. So long as we have a sufficiently mature list of core primitive operators, then we should be able to compose aggregates.

    #907

  8. anssiko commented on Oct 17, 2025

    @anssiko
    Member

    This issue was discussed at WebML WG Teleconference – 9 October 2025:

    Anssi: Fabio wanted to get back to the group after talking with the NVIDIA team

    Fabio: we're collecting all the ops that'd benefit from being in the set, one class is various attentions
    … also gathers, MoE, TopK
    … looking for other ops that'd benefit from not being composed

    Reilly: I'm curious about MoE and attentions, my concern with these high-level ops that are tied to particular model architectures, while they give performance boost, not necessarily long-lived
    … found out this by looking at e.g. LSTM where actual implementation details matter, and there were compatibility issues between implementations

    Also, this issue is slated for discussion at the upcoming F2F: webmachinelearning/meetings#35

    We can consider this to be a meta-issue and work on pieces in sub-issues.

  9. mtavenrath commented on Jul 20, 2026

    @mtavenrath

    What did we end up with MoE and TopK support in WebNN? For both operations I do not see a reasonable emulation path in WebNN and both are used by modern network.

    I am currently working myself through the transformers.js library networks and got blocked by the lack of MoE for in the HF openai\privacy-filter\onnx\model.onnx model.

  10. fdwr commented on Jul 20, 2026

    @fdwr
    Collaborator

    What did we end up with MoE and TopK support in WebNN?

    @mtavenrath, 👀 for topK (a component of an MoE graph), I see decent backend support that WebNN could use:

    Also, for an even more core operator, it seems a sort operator should exist, like StableHLO's sort (interestingly StableHLO has sort, but not topK). Since many ML libraries lack sort, sorting is evidently either unused by models directly, or models are achieving sort by abusing topK with full k. So adding sort would be more for foundational composition reasons. Depending on the backend, topK can be composed from sort and slice, whereas sort can be easily composed from topK and (k == full dimension size).

    For MoE itself, that feels more like a mini-graph rather than an operator (granted, we currently have gru/lstm which are mini-graphs), and I'm seeing different definitions of MoE across libraries, suggesting some instability still in the field 🤔. None of the primary Chromium ML backends have it. Though ONNX Runtime has a contributor operator. In any case, a particular variant of MoE should be composable, given topK and existing WebNN operators.

    Prototype

    Possible definition:

    dictionary MLTopKOptions : MLOperatorOptions
    {
        unsigned long axis;
        boolean ascending = false; // Default to descending order.
        boolean sorted = true;
    };
    
    dictionary MLSortOptions : MLOperatorOptions
    {
        unsigned long axis;
        boolean ascending = true; // Default to ascending order.
    };
    // Maybe these dictionaries could be combined cleverly?
    
    partial interface MLGraphBuilder
    {
        sequence<MLOperand> topK(MLOperand input, unsigned long k, optional MLTopKOptions options = {});
        // ⚠️ Would returning a named struct be clearer?
        //     Like: dictionary MLTopKResult {MLOperand values; MLOperand indices}?
        //     Or is returning a sequence simpler for existing logic given split/gru/lstm?
    
        sequence<MLOperand> sort(MLOperand input, optional MLSortOptions options = {});
    };
    
    partial dictionary MLOpSupportLimits
    {
        MLSingleInputSupportLimits topK;
        MLSingleInputSupportLimits sort;
    };

    Considerations

    • TFLite always sorts the topK output indices. So passing sorted as false would make no difference - would still be true.
    • TFLite always return descending order, and so for ascending order, you'd have to transpose the output, and then slice it to k.
    • TFLite is missing an axis parameter, always sorting the last dimension. So you'd have to transpose the input first.
    • I didn't include any "stable sorting" flag like StableHLO, since some API's break ties based on the element index location (which essentially means stable sorting is always enabled), and others don't guarantee it (meaning essentially false).
    • NaN's are an annoying conformance topic 🤔 (see here).

    Full Reference of other API's

    Apple MIL tensor_operation.topk

    classcoremltools.converters.mil.mil.ops.defs.iOS15.tensor_operation.topk(**kwargs)

    Returns a tensor containing top or bottom k values and the corresponding indices of the input tensor along a given axis.

    Parameters:

    • x: <*?, T> (Required) - Input tensor.
    • k: const (Optional) - Defaults to 1. Number of values/indices to be computed along each axis.
    • axis: const (Optional) - Defaults to -1 (last dimension). Axis to perform the operation.
    • ascending: const (Optional) - Defaults to False, sort in descending order. True to sort in ascending order.

    Attributes:

    • T: fp16, fp32, int32

    Returns:

    • tensor<*?, T>: Values of top/bottom k elements.
    • tensor<*?, int32>: Indices of the top/bottom k elements along axis.

    ONNX TopK

    Retrieve the top-K largest or smallest elements along a specified axis. Given an input tensor of shape [a_0, a_1, …, a_{n-1}] and integer argument k, return two outputs:

    • Value tensor of shape [a_0, a_1, …, a_{axis-1}, k, a_{axis+1}, … a_{n-1}] which contains the values of the top k elements along the specified axis
    • Index tensor of shape [a_0, a_1, …, a_{axis-1}, k, a_{axis+1}, … a_{n-1}] which contains the indices of the top k elements (original indices from the input tensor).
    • If “largest” is 1 (the default value) then the k largest elements are returned.
    • If “sorted” is 1 (the default value) then the resulting k elements will be sorted.
    • If “sorted” is 0, order of returned ‘Values’ and ‘Indices’ are undefined.

    Given two equivalent values, this operator uses the indices along the axis as a tiebreaker. That is, the element with the lower index will appear first.

    Attributes:

    • axis - INT (default is -1): Dimension on which to do the sort. Negative value means counting dimensions from the back. Accepted range is [-r, r-1] where r = rank(input).
    • largest - INT (default is 1): Whether to return the top-K largest or smallest elements.
    • sorted - INT (default is 1): Whether to return the elements in sorted order.

    Inputs:

    • X (heterogeneous) - T: Tensor of shape [a_0, a_1, …, a_{n-1}]
    • K (heterogeneous) - tensor(int64): A 1-D tensor containing a single positive value corresponding to the number of top elements to retrieve

    Outputs:

    Values (heterogeneous) - T: Tensor of shape [a_0, a_1, …, a_{axis-1}, k, a_{axis+1}, … a_{n-1}] containing top K values from the input tensor
    Indices (heterogeneous) - I: Tensor of shape [a_0, a_1, …, a_{axis-1}, k, a_{axis+1}, … a_{n-1}] containing the corresponding input tensor indices for the top K values.

    TFLite/LiteRT tfl.topk_v2 (TFL::TopKV2Op)

    Returns the top k largest element along each last dimensional slice of input and the indices of values within the last dimension of the input tensor. Results are always sorted in the descending order.

    Operands:

    • input: tensor of 32-bit float or 8-bit signless integer or 16-bit signless integer or 32-bit signless integer or 64-bit signless integer or 8-bit unsigned integer or QI8 type or QUI8 type values
    • k: tensor of 16-bit signless integer or 32-bit signless integer values
      Results:
    • values: tensor of 32-bit float or 8-bit signless integer or 16-bit signless integer or 32-bit signless integer or 64-bit signless integer or 8-bit unsigned integer or QI8 type or QUI8 type values
    • indices: tensor of 16-bit signless integer or 32-bit signless integer values

    PyTorch torch.topk

    torch.topk(input, k, dim=None, largest=True, sorted=True, *, out=None)

    Returns the k largest elements of the given input tensor along a given dimension.

    • If dim is not given, the last dimension of the input is chosen.
    • If largest is False then the k smallest elements are returned.
    • A namedtuple of (values, indices) is returned with the values and indices of the largest k elements of each row of the input tensor in the given dimension dim.
    • The boolean option sorted if True, will make sure that the returned k elements are themselves sorted

    When using torch.topk, the indices of tied elements are not guaranteed to be stable and may vary across different invocations.

    Parameters:

    • input (Tensor) – the input tensor.
    • k (int) – the k in “top-k”
    • dim (int, optional) – the dimension to sort along
    • largest (bool, optional) – controls whether to return largest or smallest elements
    • sorted (bool, optional) – controls whether to return the elements in sorted order

    Keyword Arguments:

    • out (tuple, optional) – the output tuple of (Tensor, LongTensor) that can be optionally given to be used as output buffers
  11. bhack commented on Jul 20, 2026

    @bhack
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions