Repository navigation
Support multimodal tool input/output #41
Description
Activity
Providing some reference material for the tool result types in the MCP spec:
Included in the ToolResult object is an array of Content Blocks
export type ContentBlock = | TextContent | ImageContent | AudioContent | ResourceLink | EmbeddedResource; export interface CallToolResult extends Result { /** * A list of content objects that represent the unstructured result of the tool call. */ content: ContentBlock[]; ...
Images can be returned as an ImageContent or an EmbeddedResource. Both of which request the image be represented as a Base64-encoded string.
/** * An image provided to or from an LLM. */ export interface ImageContent { type: "image"; /** * The base64-encoded image data. * * @format byte */ data: string; /** * The MIME type of the image. Different providers may support different image types. */ mimeType: string; /** * Optional annotations for the client. */ annotations?: Annotations; /** * See [General fields: `_meta`](/specification/2025-06-18/basic/index#meta) for notes on `_meta` usage. */ _meta?: { [key: string]: unknown }; }
I believe most inference providers accept base64. It would nice if the browser did the conversion so the server code could be as simple as this
window.navigator.modelContext.registerTool({ name: "get_product_image", description: "Get the product image from the page", inputSchema: { type: "object", properties: { productId: { type: "string" } } }, handler: async ({ productId }) => { const img = document.querySelector(`#product-${productId} img`); return { content: [{ type: "image", element: img }] }; } });
- added a commit that references this issue
on Oct 24, 2025 - added a commit that references this issue
on Oct 24, 2025 Thanks for the pointer @MiguelsPizza! +1 to this: "It would nice if the browser did the conversion so the server code could be as simple as this". There are a few other data types to reference pixel content in Web APIs, canvas2d is a good reference for that. Documentation here.
Other than
dataandmimeTypewhich will come from the image source above, MCP data structure hasannotationsand_meta. Those are shared types that we should discuss separately. Doesn't need to block the addition of a return type for image resources, we can always define a base structure and add to it going forward.Related, there should likely be a way to reference the image content in a json response. If you're returning a list of products, you likely have a json with text info and placeholders for the image of each product. Does MCP already support that concept? Some overlap with #9 too.
We should likely do what the Prompt API does here:
- For image input, the Web platform has the concept of
ImageBitmapSourceandBufferSource. - For audio input, there's
AudioBuffer, againBufferSource, or plainBlob.
- For image input, the Web platform has the concept of
quick revival :)
polyfill prototype exists #260 supporting both shapes; spec should bless native types per tomayac's direction, with {type:"image", data, mimeType} as the serialized fallback.- marked Proposal: Tool result content types beyond text #86 as a duplicate of this issue
on Sep 28, 2026 - changed the title
[-]Support image input/output in tools[/-][+]Support multimodal tool input/output[/+]on Sep 28, 2026 +1 to the direction here: the browser converts native image types, with {type: "image", data, mimeType} as the serialized form.
One use case this doesn't cover yet. I'm building an app where images live in server storage, not in the page, and access can be revoked mid-task. Returning inline base64 means the data can't be taken back, so we return a short-lived reference instead and check permissions on every call.
Would the group consider allowing a reference form (URL or handle, MIME type, alt text) next to inline data? It might also answer @khushalsagar's question about pointing to images from a JSON result.
Also: should untrustedContentHint cover text derived from images, like captions or OCR? That text can carry prompt injection.
The API proposal doesn't include support for image input to the tool and for the tool output to include images. This seems like a basic use-case we should support. Worth looking at how MCP is doing it but we'll need something web specific, allowing devs to use image elements for this as an example.