Skip to content

Support multimodal tool input/output #41

Description

@khushalsagar

The API proposal doesn't include support for image input to the tool and for the tool output to include images. This seems like a basic use-case we should support. Worth looking at how MCP is doing it but we'll need something web specific, allowing devs to use image elements for this as an example.

Activity

  1. MiguelsPizza commented on Oct 23, 2025

    @MiguelsPizza
    Contributor

    Providing some reference material for the tool result types in the MCP spec:

    Included in the ToolResult object is an array of Content Blocks

    export type ContentBlock =
      | TextContent
      | ImageContent
      | AudioContent
      | ResourceLink
      | EmbeddedResource;
    
    export interface CallToolResult extends Result {
      /**
       * A list of content objects that represent the unstructured result of the tool call.
       */
      content: ContentBlock[];
    ...

    Images can be returned as an ImageContent or an EmbeddedResource. Both of which request the image be represented as a Base64-encoded string.

    /**
     * An image provided to or from an LLM.
     */
    export interface ImageContent {
      type: "image";
    
      /**
       * The base64-encoded image data.
       *
       * @format byte
       */
      data: string;
    
      /**
       * The MIME type of the image. Different providers may support different image types.
       */
      mimeType: string;
    
      /**
       * Optional annotations for the client.
       */
      annotations?: Annotations;
    
      /**
       * See [General fields: `_meta`](/specification/2025-06-18/basic/index#meta) for notes on `_meta` usage.
       */
      _meta?: { [key: string]: unknown };
    }

    I believe most inference providers accept base64. It would nice if the browser did the conversion so the server code could be as simple as this

    window.navigator.modelContext.registerTool({
        name: "get_product_image",
        description: "Get the product image from the page",
        inputSchema: {
          type: "object",
          properties: {
            productId: { type: "string" }
          }
        },
        handler: async ({ productId }) => {
          const img = document.querySelector(`#product-${productId} img`);
          
          return { content: [{ type: "image", element: img }] };
        }
    });
  2. added a commit that references this issue on Oct 24, 2025
  3. added a commit that references this issue on Oct 24, 2025
  4. khushalsagar commented on Oct 24, 2025

    @khushalsagar
    CollaboratorAuthor

    Thanks for the pointer @MiguelsPizza! +1 to this: "It would nice if the browser did the conversion so the server code could be as simple as this". There are a few other data types to reference pixel content in Web APIs, canvas2d is a good reference for that. Documentation here.

    Other than data and mimeType which will come from the image source above, MCP data structure has annotations and _meta. Those are shared types that we should discuss separately. Doesn't need to block the addition of a return type for image resources, we can always define a base structure and add to it going forward.

    Related, there should likely be a way to reference the image content in a json response. If you're returning a list of products, you likely have a json with text info and placeholders for the image of each product. Does MCP already support that concept? Some overlap with #9 too.

  5. tomayac commented on Oct 27, 2025

    @tomayac

    We should likely do what the Prompt API does here:

  6. arnabwithab commented on Aug 31, 2026

    @arnabwithab

    quick revival :)
    polyfill prototype exists #260 supporting both shapes; spec should bless native types per tomayac's direction, with {type:"image", data, mimeType} as the serialized fallback.

  7. changed the title [-]Support image input/output in tools[/-] [+]Support multimodal tool input/output[/+] on Sep 28, 2026
  8. srirammk0 commented on Oct 5, 2026

    @srirammk0

    +1 to the direction here: the browser converts native image types, with {type: "image", data, mimeType} as the serialized form.

    One use case this doesn't cover yet. I'm building an app where images live in server storage, not in the page, and access can be revoked mid-task. Returning inline base64 means the data can't be taken back, so we return a short-lived reference instead and check permissions on every call.

    Would the group consider allowing a reference form (URL or handle, MIME type, alt text) next to inline data? It might also answer @khushalsagar's question about pointing to images from a JSON result.

    Also: should untrustedContentHint cover text derived from images, like captions or OCR? That text can carry prompt injection.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    enhancementNew feature or request

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions