Skip to content

Investigation: how WebNN / WebGPU interop could be happening #2500

Description

@huningxin

According to webmachinelearning/webnn#6, ML framework authors require a way to write custom ops (operations) that can interop with the built-in neural net ops of WebNN. That means having high-performance data exchange between custom ops and built-in ops.

In order to answer this requirement, in particular when the custom ops are written in WebGPU shaders, WebNN adds the device selection capability and the support for GPU resources from WebGPU and WebGL in webmachinelearning/webnn#162.

Specifically, developers are able to create an MLContext from a specific GPUDevice that is already in use by the application. In such case, developers could use the corresponding GPUBuffer resources as MLGraph constants, and use the GPUTexture and GPUBuffer as the inputs and outputs of MLGraph.compute. These GPU resources must be created from the same GPUDevice.

This interop capability of WebNN / WebGPU is also useful for integration with real-time video processing, where developers might want to establish a full-GPU-only pipeline by importing a video frame to GPUExternalTexture and feeding into WebNN graph compute.

Given those and thanks to @Kangz 's suggestion, this issue is opened for the coordination and investigation between WebGPU and WebNN groups on this topic.

/cc @anssiko @dontcallmedom @wchao @RafaelCintron @bbernhar

Activity

  1. Kangz commented on Jan 14, 2022

    @Kangz
    Contributor

    Thanks for opening this @huningxin, to try to better understand how WebGPU and WebNN would interoperate, I have bunch of questions:

    • When an MLContext is created from a GPUDevice, on the various target APIs of WebGPU and WebNN, how does that change how the MLContext is created? Is it completely switched to an implementation over WebGPU? Does it start using APIs like Metal's performance shaders and DirectML? Does it use the OS interop mechanisms between multiple APIs? Etc.
    • What if the WebGPU device is lost? What if it is destroyed?
    • Are there WebGPU devices that can't be used to create an MLContext? Do we need a way to request a WebGPU device that can? Should WebML hint at an adapter to do ML computations?
    • What are the synchronization guarantees between WebGPU's default queue and the MLGraph execution? Are the resources passed as MLGraph's input still usable by WebGPU?
    • If WebML is used an other accelerator in addition to WebGPU, how does the interaction happen? Is the rest of the WebGPU execution blocked by the accelerator?
    • What usages do the WebGPU objects require to be used with WebML? What other constraints are there on the resources? Why can't we specify ranges of resources like a GPUTextureView?
    • Does there need to be an explicit handoff of resources between WebGPU and WebML? Why?
    • What about GPUExternalTexture as an input?
  2. huningxin commented on Jan 17, 2022

    @huningxin
    Author

    Thanks for these questions, @Kangz . I am trying to go first round through them with my current knowledge. Probably more investigation and prototyping are required for more details. Others, please feel free to chime in and correct me if I made any mistakes.

    • When an MLContext is created from a GPUDevice, on the various target APIs of WebGPU and WebNN, how does that change how the MLContext is created? Is it completely switched to an implementation over WebGPU? Does it start using APIs like Metal's performance shaders and DirectML? Does it use the OS interop mechanisms between multiple APIs? Etc.

    In such case, I suppose the MLContext would try to use native API to create an ML execution context from the backend device of the GPUDevice, such as create a IDMLDevice by DMLCreateDevice from the ID3D12Device backing GPUDevice on Windows platform.

    • What if the WebGPU device is lost? What if it is destroyed?

    I suppose WebGPU device lost or destroyed would cause error (e.g. device lost) on the corresponding MLContext.

    • Are there WebGPU devices that can't be used to create an MLContext? Do we need a way to request a WebGPU device that can? Should WebML hint at an adapter to do ML computations?

    MLContext could check whether a GPUDevice supports ML features or not by native API.

    It would be good that if ML compute capability could be one GPU feature that developers can check and hint to request a WebGPU device.

    • What are the synchronization guarantees between WebGPU's default queue and the MLGraph execution? Are the resources passed as MLGraph's input still usable by WebGPU?

    On the high-level, I suppose an MLGraph execution would submit the command-list of the MLGraph to the queue of the GPUDevice. MLGraph execution might need to wait for the queue to complete. I guess the resources passed as MLGraph's input are still usable by WebGPU.

    • If WebML is used an other accelerator in addition to WebGPU, how does the interaction happen? Is the rest of the WebGPU execution blocked by the accelerator?

    If an MLContext created with an other accelerator, it would only accept ArrayBufferView and won't interop with WebGPU.

    • What usages do the WebGPU objects require to be used with WebML? What other constraints are there on the resources? Why can't we specify ranges of resources like a GPUTextureView?

    I guess the GPUBuffer as MLGraph compute input would have COPY_SRC usage, the output GPUBuffer would have COPY_DST usage. This may need more investigation.

    • Does there need to be an explicit handoff of resources between WebGPU and WebML? Why?

    I remember @RafaelCintron once mentioned an MLBuffer idea. It might be relevant to this question.

    • What about GPUExternalTexture as an input?

    This is required by a full-GPU-only video processing pipeline use case. It needs more investigation.

  3. chcunningham commented on Jan 21, 2022

    @chcunningham
  4. Kangz commented on Jan 24, 2022

    @Kangz
    Contributor

    Thanks for the answers. It seems there's significant assumptions baked in WebNN and more investigations needed in places.

    If an MLContext created with an other accelerator, it would only accept ArrayBufferView and won't interop with WebGPU.

    Are there possibilities where WebNN is running on a mix of multiple devices, or are will it guarantee that if created from a GPUDevice is will only use the physical device underlying that GPUDevice?

    It would be good that if ML compute capability could be one GPU feature that developers can check and hint to request a WebGPU device.

    That could be done similarly to the fallback adapter support in WebGPU.

    On the high-level, I suppose an MLGraph execution would submit the command-list of the MLGraph to the queue of the GPUDevice. MLGraph execution might need to wait for the queue to complete. I guess the resources passed as MLGraph's input are still usable by WebGPU.

    It is an issue if we both have the MLGraph wait for the commands to be executed on the GPU, and the resources still available to WebGPU: WebGPU could write to the resources while they are used by the MLGraph, causing synchronization issues. Or are we guaranteed that all operations using the GPU resource on the MLGraph are submitted to the queue before it waits for the completion. (for example during multi-stage GPU-CPU-GPU-CPU graphs?)

  5. added this to the post-V1 milestone on Jan 26, 2022
  6. huningxin commented on Jan 28, 2022

    @huningxin
    Author

    Are there possibilities where WebNN is running on a mix of multiple devices, or are will it guarantee that if created from a GPUDevice is will only use the physical device underlying that GPUDevice?

    If an MLContext is created from a GPUDevice, it would share the same physical device underlying the GPUDevice.

    Or are we guaranteed that all operations using the GPU resource on the MLGraph are submitted to the queue before it waits for the completion.

    I think the MLGraph::compute would submit all operations using the GPU resources to the queue. As webmachinelearning/webnn#230 (comment), we are discussing MLGraph::compute only submits the command lists without waiting for the queue to complete if the corresponding MLContext created from a GPUDevice. In this way, the caller of both WebGPU and WebNN could control the execution of the queue and order of submission. /cc @wchao1115

    WDYT?

  7. anssiko commented on Nov 20, 2024

    @anssiko

    MLTensor explainer now outlines the best-effort buffer-sharing proposal between WebNN and WebGPU APIs. Thanks to @Kangz and other WebGPU WG participants for your review and suggestions.

    Implementation work is ongoing and the first spec PR for this feature is currently under review at webmachinelearning/webnn#787 Review and comments welcome.

    @huningxin @a-sully please check the feedback in this issue is recorded in the webnn repo and suggest for closure as appropriate.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    apiWebGPU API

    Type

    No type

    Projects

    No projects

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions