Repository navigation
Investigation: how WebNN / WebGPU interop could be happening #2500
Description
Activity
Thanks for opening this @huningxin, to try to better understand how WebGPU and WebNN would interoperate, I have bunch of questions:
- When an
MLContextis created from aGPUDevice, on the various target APIs of WebGPU and WebNN, how does that change how theMLContextis created? Is it completely switched to an implementation over WebGPU? Does it start using APIs like Metal's performance shaders and DirectML? Does it use the OS interop mechanisms between multiple APIs? Etc. - What if the WebGPU device is lost? What if it is destroyed?
- Are there WebGPU devices that can't be used to create an
MLContext? Do we need a way to request a WebGPU device that can? Should WebML hint at an adapter to do ML computations? - What are the synchronization guarantees between WebGPU's default queue and the
MLGraphexecution? Are the resources passed as MLGraph's input still usable by WebGPU? - If WebML is used an other accelerator in addition to WebGPU, how does the interaction happen? Is the rest of the WebGPU execution blocked by the accelerator?
- What usages do the WebGPU objects require to be used with WebML? What other constraints are there on the resources? Why can't we specify ranges of resources like a
GPUTextureView? - Does there need to be an explicit handoff of resources between WebGPU and WebML? Why?
- What about
GPUExternalTextureas an input?
- When an
Thanks for these questions, @Kangz . I am trying to go first round through them with my current knowledge. Probably more investigation and prototyping are required for more details. Others, please feel free to chime in and correct me if I made any mistakes.
- When an
MLContextis created from aGPUDevice, on the various target APIs of WebGPU and WebNN, how does that change how theMLContextis created? Is it completely switched to an implementation over WebGPU? Does it start using APIs like Metal's performance shaders and DirectML? Does it use the OS interop mechanisms between multiple APIs? Etc.
In such case, I suppose the
MLContextwould try to use native API to create an ML execution context from the backend device of theGPUDevice, such as create aIDMLDevicebyDMLCreateDevicefrom theID3D12DevicebackingGPUDeviceon Windows platform.- What if the WebGPU device is lost? What if it is destroyed?
I suppose WebGPU device lost or destroyed would cause error (e.g.
device lost) on the correspondingMLContext.- Are there WebGPU devices that can't be used to create an
MLContext? Do we need a way to request a WebGPU device that can? Should WebML hint at an adapter to do ML computations?
MLContextcould check whether aGPUDevicesupports ML features or not by native API.It would be good that if ML compute capability could be one GPU feature that developers can check and hint to request a WebGPU device.
- What are the synchronization guarantees between WebGPU's default queue and the
MLGraphexecution? Are the resources passed as MLGraph's input still usable by WebGPU?
On the high-level, I suppose an
MLGraphexecution would submit the command-list of theMLGraphto the queue of theGPUDevice.MLGraphexecution might need to wait for the queue to complete. I guess the resources passed asMLGraph's input are still usable by WebGPU.- If WebML is used an other accelerator in addition to WebGPU, how does the interaction happen? Is the rest of the WebGPU execution blocked by the accelerator?
If an
MLContextcreated with an other accelerator, it would only acceptArrayBufferViewand won't interop with WebGPU.- What usages do the WebGPU objects require to be used with WebML? What other constraints are there on the resources? Why can't we specify ranges of resources like a
GPUTextureView?
I guess the
GPUBufferasMLGraphcompute input would haveCOPY_SRCusage, the outputGPUBufferwould haveCOPY_DSTusage. This may need more investigation.- Does there need to be an explicit handoff of resources between WebGPU and WebML? Why?
I remember @RafaelCintron once mentioned an
MLBufferidea. It might be relevant to this question.- What about
GPUExternalTextureas an input?
This is required by a full-GPU-only video processing pipeline use case. It needs more investigation.
- When an
cc @sandersdan
Thanks for the answers. It seems there's significant assumptions baked in WebNN and more investigations needed in places.
If an MLContext created with an other accelerator, it would only accept ArrayBufferView and won't interop with WebGPU.
Are there possibilities where WebNN is running on a mix of multiple devices, or are will it guarantee that if created from a
GPUDeviceis will only use the physical device underlying thatGPUDevice?It would be good that if ML compute capability could be one GPU feature that developers can check and hint to request a WebGPU device.
That could be done similarly to the fallback adapter support in WebGPU.
On the high-level, I suppose an MLGraph execution would submit the command-list of the MLGraph to the queue of the GPUDevice. MLGraph execution might need to wait for the queue to complete. I guess the resources passed as MLGraph's input are still usable by WebGPU.
It is an issue if we both have the
MLGraphwait for the commands to be executed on the GPU, and the resources still available to WebGPU: WebGPU could write to the resources while they are used by theMLGraph, causing synchronization issues. Or are we guaranteed that all operations using the GPU resource on theMLGraphare submitted to the queue before it waits for the completion. (for example during multi-stage GPU-CPU-GPU-CPU graphs?)Are there possibilities where WebNN is running on a mix of multiple devices, or are will it guarantee that if created from a
GPUDeviceis will only use the physical device underlying thatGPUDevice?If an
MLContextis created from aGPUDevice, it would share the same physical device underlying theGPUDevice.Or are we guaranteed that all operations using the GPU resource on the
MLGraphare submitted to the queue before it waits for the completion.I think the
MLGraph::computewould submit all operations using the GPU resources to the queue. As webmachinelearning/webnn#230 (comment), we are discussingMLGraph::computeonly submits the command lists without waiting for the queue to complete if the correspondingMLContextcreated from aGPUDevice. In this way, the caller of both WebGPU and WebNN could control the execution of the queue and order of submission. /cc @wchao1115WDYT?
MLTensorexplainer now outlines the best-effort buffer-sharing proposal between WebNN and WebGPU APIs. Thanks to @Kangz and other WebGPU WG participants for your review and suggestions.Implementation work is ongoing and the first spec PR for this feature is currently under review at webmachinelearning/webnn#787 Review and comments welcome.
@huningxin @a-sully please check the feedback in this issue is recorded in the webnn repo and suggest for closure as appropriate.
Reacted by mwyrzykowski and Eric Zhang
According to webmachinelearning/webnn#6, ML framework authors require a way to write custom ops (operations) that can interop with the built-in neural net ops of WebNN. That means having high-performance data exchange between custom ops and built-in ops.
In order to answer this requirement, in particular when the custom ops are written in WebGPU shaders, WebNN adds the device selection capability and the support for GPU resources from WebGPU and WebGL in webmachinelearning/webnn#162.
Specifically, developers are able to create an
MLContextfrom a specificGPUDevicethat is already in use by the application. In such case, developers could use the correspondingGPUBufferresources asMLGraphconstants, and use theGPUTextureandGPUBufferas the inputs and outputs ofMLGraph.compute. These GPU resources must be created from the sameGPUDevice.This interop capability of WebNN / WebGPU is also useful for integration with real-time video processing, where developers might want to establish a full-GPU-only pipeline by importing a video frame to
GPUExternalTextureand feeding into WebNN graph compute.Given those and thanks to @Kangz 's suggestion, this issue is opened for the coordination and investigation between WebGPU and WebNN groups on this topic.
/cc @anssiko @dontcallmedom @wchao @RafaelCintron @bbernhar