Repository navigation
Get devices used for a graph after graph compilation #836
Description
Activity
- changed the title
[-]Get devices used for a graph[/-][+]Get devices used for a graph after graph compilation[/+]on Apr 22, 2025 What is the use case for the website in providing the device information?
Given the discussion in #815 (comment, comment), the developer use case (for frameworks, not for websites) seems to be the following (please correct/update/complete this):
- Before downloading/loading a model, the developer wants to know if e.g. the GPU can be used for inference with WebNN.
- If no, then they might want to try other path than WebNN.
- If yes, then in some cases (e.g. CoreML) the model needs to be dispatched before knowing for sure whether it can be executed on GPU. For that we'd need this API discussed here and implemented in define graph.devices #854 .
Based on the answer, the developer can choose another option than WebNN.
Besides that, the feature permits gathering some data on typical graph accelerations (fingerprintable) which might help the spec work on device selection API.
I'm not sure if this provides the website with any useful or actionable information. The model could run on one device at one point in time but the same exact model could run on a different device 100ms later for instance.
So if we did provide this information, we could return 'cpu' when the developer queries due to system conditions despite if the model was run repeatedly, it would end up on the gpu or npu. I.e., the underlying frameworks restrict us from providing any useful information to the framework or website.
It would be great to have a more concrete use case prior to adding this to the WebNN API.
So if we did provide this information, we could return 'cpu' when the developer queries due to system conditions despite if the model was run repeatedly, it would end up on the gpu or npu. I.e., the underlying frameworks restrict us from providing any useful information to the framework or website.
Thanks for sharing that, I didn't know that it would return differently according to system resources. I think this can be addressed by making the API return async
graph.devices()instead and describe in the spec that this returns the devices that would get used at the current system situation.Re: concrete use case, we discussed a bit more and think it would be best to build a demo app that exercises the hypothetical use cases we have in mind and see if the graph.devices is necessary. We will report back on how that goes.
Reacted by Anssi Kostiainen, mwyrzykowski, Zoltan Kis and Ningxin HuRe: concrete use case, we discussed a bit more and think it would be best to build a demo app that exercises the hypothetical use cases we have in mind and see if the graph.devices is necessary. We will report back on how that goes.
Oh great, that would certainly be useful, thank you.
- added 3 commits that reference this issue
on Jun 17, 2025 - added a commit that references this issue
on Jun 27, 2025 - added a commit that references this issue
on Aug 8, 2025 he model could run on one device at one point in time but the same exact model could run on a different device 100ms later for instance.
I assume if a graph has been compiled, the system will try to execute it on the respective devices, and will try to get hold of the needed system conditions/resources for that, w.r.t. the cost of recompiling (reassigning parts of the graph is possible in some frameworks, but mostly not).
So I assume it should be still a valid thing to request such a device list.
(I am not addressing the utility of this for web apps, just the topic of validity/consistency of maintaining such a device list across implementations).@philloooo any new insights on this? It would also help #815 (and #884).
Reacted by mwyrzykowskiOne big use case for this is adaptation. Consider a web app with a bunch of ML workloads that execute on WebNN. Now consider the situation where the system is loaded enough that it can't handle all the load within responsiveness limits. In this case you'd like to stop some processing.
graph.devicescould help identify a) what resources a misbehaving model is using, and b) which models are candidates to stop that would help the situation.Reacted by Yajing TangOne big use case for this is adaptation. Consider a web app with a bunch of ML workloads that execute on WebNN. Now consider the situation where the system is loaded enough that it can't handle all the load within responsiveness limits. In this case you'd like to stop some processing.
graph.devicescould help identify a) what resources a misbehaving model is using, and b) which models are candidates to stop that would help the situation.If the website stops processing a certain model doesn't mean the UA or underlying OS won't move a completely unrelated task in a multi-process system to the resource the web app was attempting to free up.
Another way of achieving the same thing is the web app sorts its workloads in priority, terminating lower priority ones (1). Or some type of metric reporting that the model was stalled K ms waiting to run due to other work on the system and took S ms to complete (2).
But it would be great it show a PoC that (1) is insufficient and the UA can not do better without hints from the web app prior to adding (2) which may be a potential privacy concern.
As we've discussed in TPAC, we've agreed to merge npu and gpu to "accelerated".
To also signal a graph that gets run on a combination of cpu/gpu/npu, we can expose that as "hybrid".
So the revised plan is forgraph.deviceto return an enum:enum MLDeviceType { "accelerated", "not-accelerated", "hybrid" }Let me know what you think @handellm @mwyrzykowski @huningxin @reillyeon @fdwr . Once we reach a consensus I can work on implementation and spec changes.
How Is the
hybridoption useful @philloooo ? The percentage of what runs accelerated or not is not exposed.I would think it would be easier to start with:
[[isAccelerated]] of type [boolean](https://webidl.spec.whatwg.org/#idl-boolean)on the
MLGraphinterfaceSo the revised plan is for graph.device to return an enum:
...
"not_accelerated",
...(minor) Other enums use hyphen. We should be consistent:
enum MLDeviceType { "accelerated", "not-accelerated", "hybrid" } enum MLPowerPreference { "default", "high-performance", "low-power" }; enum MLInterpolationMode { "nearest-neighbor", "linear" };Reacted by Yajing Tang and Zoltan KisRight the percentage is not exposed, but I was thinking that knowing a graph gets partitioned and executed on multiple devices is a useful signal.
If developers' goal is to achieve full execution on accelerated device, it's a useful feedback that actually this model wasn't able to achieve that and ended up with graph partition. So they need to investigate into this.Alternatively if as long as some part of the model gets accelerated, we return
isAccelerated== true, that would hide such information.A "hybrid" return seems like useful information combined with deducing the model is fast enough or not. In that case we can avoid trying to run/download it the next time we try, an also also inform through the metrics pipeline whether it's a good idea.
To also signal a graph that gets run on a combination of cpu/gpu/npu, we can expose that as "hybrid".
To me "hybrid" seems too large a range (as NPU alone is) in sequential graph partitioning.
Some frameworks can use dynamic partitioning (CPU+NPU+GPU) with mostly GPU+NPU doing the heavy math and CPU the flow control and weird activation functions. This would be one kind of "hybrid".
Others may use NPU prefill/encode (compute bound) and GPU decode/generate (memory bound) - e.g. with generative LLMs. This would also be "hybrid" but no discernment from the previous case. Maybe none needed at the moment, but would be nice.
Some use cases use NPU+CPU for efficiency. Wouldn't this also be "hybrid"?
If we don't use "hybrid", but work only with the context option information (power pref, accelerated), we would still be able to express all information that we could with "hybrid":
- The combinations of "accelerated" and "low-power" maps well to NPU+CPU hybrid execution.
- The "accelerated" and "high-performance" maps well to GPU(+NPU) hybrid and non-hybrid scenarios.
A "hybrid" return seems like useful information combined with deducing the model is fast enough or not.
Since there is a large range for "hybrid", is there a certain generic threshold when this becomes useful (expressing "fast enough"), or is it graph dependent? To me it looks the use case here is to quantify "fast enough", not the hybrid-ness.
(I am not against using "hybrid", just ran an Occam challenge.)
FTR, this issue was discussed at Kobe F2F where we resolved the following:
RESOLUTION: Phillis to refine the proposal to reflect an accelerated status, with discussions on hybrid still TBD
Even if the group addressed @philloooo in this resolution, I appreciate that the group's other participants joined to co-design a solution. Spec design is not a solitary activity :-)
Let's step back and think about what the after graph compilation devices information is for.
I think it's a feedback mechanism to let developers understand whether the context option information (power pref, accelerated) was respected.To me it looks the use case here is to quantify "fast enough", not the hybrid-ness.
Yeah, I agree that the complex nature of how a graph can be "hybrid" makes it hard to correlate to whether it's "fast enough".
I think my original intent of including a "hybrid" was to give a clue when:
- Developers want the graph to be accelerated.
- User agent couldn't run the graph fully accelerated, because some ops are not supported on GPU/NPU, so it partitioned the graph to fallback the unsupported ops onto CPU.
If we don't have the "hybrid" enum, we would return the graph.accelerated == true, even though actually some part of the graph can't be accelerated.
If we do return the "hybrid" enum, developers would know that this model isn't fully compatible to be run on GPU/NPU yet, so they probably can do some adjustments to the model to achieve full execution on accelerated hardware.Reacted by Zoltan KisIf we don't have the "hybrid" enum
We can have it if we provide clarity on what it means. I wonder if we could (or should at all) distinguish between various levels of hybrid-ness? Or include (later) another parameter to quantify that? Strictly "CPU-free" execution is rare because the CPU almost always acts as the graph orchestrator or fallback for unsupported operators. Of course we could start with "hybrid" and later add more quantifiers. As it is, "hybrid" would be the equivalent to the dynamic partitioning scheduling policy (
allin CoreML andHETERO:GPU,NPU,CPUin OpenVINO).Apart from that, there is the "pipelined heterogeneity" use case for serving LLMs (prefill/encode + decode/generate), which mostly maps to "accelerated", but could also map to a variant with "hybrid". But I'd call all of these "accelerated". If later differentiation will be needed, we could add a new enum value like "pipelined" or similar.
By now, the naming
graph.devicelooks a bit misleading, as it's always a combination of devices, and devices may become a fluid notion in the future. If we wish to keep the word "device", wouldgraph.devicePolicybe a more exact name? Orgraph.deviceMapping? Or simplygraph.scheduling?@anssiko can we close this issue now based on the Kobe F2F resolution?
@mwyrzykowski thanks for the bump.
I believe the remaining open tasks are:
- document use case(s) for graph.devices, PR currently contains IDL changes only define graph.devices #854
- Adaptation use case from @handellm
- ... other use cases?
- check that this post-compile graph.devices proposal and any pre-compile hints and flags make a cohesive whole, proposal head for the latter at accelerated should be prior to powerPreference for device selection #911 (comment)
My understanding is ONNX Runtime is adding support for querying the device type to which the subgraph is assigned to. I expect that implementation experience to inform this work.
We will discuss this issue today in context of the accelerated and fallback topic if there is interest.
- document use case(s) for graph.devices, PR currently contains IDL changes only define graph.devices #854
We were discussing the topic of effective compute policy a few workgroup meetings ago and it didn't sound impossibly controversial (subjective). Could we re-purpose this issue to solve the problem that way?
RESOLUTION: Close issue #836 and open a new issue for Effective MLComputePolicy exposure discussion.
Hi! This is deriving from discussion here #815 (comment) .
Eventually which devices are used can only be known after a model is provided. It's a different problem from the context level device query mechanism so I decided to create a separate issue.
I've tested with CoreML's MLComputePlan which can give you op level device selections after a model is compiled. (Before calling
dispatch).Tflite also gives which delegate gets used for each op, after a graph is compiled.
So for WebNN, we can attach such information to MLGraph which represents a compiled graph.
The first high level information we can attach is a list of
devicesthat will be used to execute the graph, since sometimes a graph could be executed using multiple devices.We could further expose op level device selection with a map or list:
The only thing is we would need identifiers for each op, we could: 1. auto generate op identifiers using op name + auto incrementing index. 2. Use
labeland only return ops that have labels defined.Maybe just the high level
graph.devicesis good enough for app developers to make decisions. But I want to provide what's possible with current backends right now.