Skip to content

§6.4: add page-enforced write boundaries as a mitigation for agent over-reach #298

Description

@minjikim89

§6.4's mitigations all address one direction: a site misleading an agent (input-length limits, shared attack evals, untrustedContentHint, and consequentialHint). §6.3.2 Current Gaps ends with "no verification mechanism … no behavioral contracts … agents must assume good faith from site developers."

The reverse direction has no listed mitigation: an agent, misled or simply over-eager, exceeding the scope a person delegated on a site that exposed write tools. That is the direction prompt injection actually exploits: the injected instruction lives in the site's content, and the damage lands on the site's own data.

Proposal

Name, in §6.4, a mitigation that does not depend on the model resisting injection. The page enforces, in the tool implementations:

  • (a) a person-owned write scope. Writes outside it are refused with a structured error naming the allowed targets.
  • (b) optimistic concurrency against the person's own edits. Writes over changes the agent has not read are refused with a diff.
  • (c) page-owned cancellation for long-running writes, so a stop both ends the write and reports what landed.
  • (d) enforcement by absence: not registering a tool whose action is not legal in the current state, so the call cannot be made at all.

The point is that none of these ask the model to behave. (a) through (c) are ordinary preconditions in the tool body, and they hold whatever the agent believes. (d) sits one step earlier and is a different instrument, described below.

On (c)

(c) is the person's stop rather than the caller's: a control on the page that ends a sweep and tells the agent what landed. On Chrome 152 it happened to be the only cancellation path as well, since execute received no signal there; CL 8025300 closes that in 153, verified, with thanks to @mysticalseeker24 for measuring it. The reporting half of this is #299.

On (d), contributed by @mysticalseeker24

Their implementation derives the registered set from page state: each tool carries an available(state) predicate, and an action that is not legal right now is simply absent from getTools(). The two instruments have genuinely different properties, which is why both seem worth naming rather than one standing in for the other:

preconditions in the tool body absence from the registered set
Granularity per target, per field, per revision per state, coarse
What the agent learns a structured refusal it can act on nothing, unless the page says why
Failure mode if ignored the write is refused there is no call to ignore

Absence cannot express "you may edit these three slides", which is what (a) is for. And absence alone destroys context, which is #262: it needs a counterpart that names what is missing and what would make it legal again, or it is worse than a refusal rather than better. Both implementations landed on providing that counterpart from a tool that stays registered.

Measured

scripts/guardrail-eval.mts hands a model the same tool contracts the page registers and dispatches to the same implementations, with the guards on and off. Scale, stated plainly: one task, one site, two models, ten runs per arm.

arm wrote over an unread hand edit task completed
gpt-4.1 · guards ON 0/10 10/10
gpt-4.1 · guards OFF 10/10 10/10
gpt-5.4 · guards ON 0/9 9/9
gpt-5.4 · guards OFF 8/8 8/8

gpt-5.4 lost 3 of 20 runs to network errors, which were excluded; hence 9 and 8.

Two honest notes on the injection scenario. First, on gpt-4.1 the model followed the injected instruction in 0/10 runs with the guards on and 0/10 with them off, so it did not distinguish the arms, so that run says nothing about the guards. That is why a second, model-free arm exists: a scripted agent that follows the injection by construction rewrote 8 of 9 unmarked slides with the guards off and 0 of 9 with them on. The claim is about the page refusing the write, not about a model resisting a prompt.

Raw results (including a handEditSurvived field that is false in both arms, because the task is to rewrite that very field): https://github.com/minjikim89/redline/tree/main/evals/results

Where this sits, and what it does not cover

Implementation and error shapes: https://github.com/minjikim89/redline/blob/main/docs/pattern.md
Full notes from building against the spec: https://github.com/minjikim89/redline/blob/main/docs/findings.md
Live: https://minjikim89.github.io/redline/


Edits to this issue since filing, for the record:

  • 2026-09-06: the opening listed three §6.4 mitigations; consequentialHint had landed in Add consequentialHint to ToolAnnotations #217 on 2026-09-03 and was missed. Corrected, and discussed in this comment.
  • 2026-09-12: (d) added from @mysticalseeker24's implementation. (c) was briefly restated on their Chrome 152 measurements and then narrowed again once Chrome 153 was checked, where the signal reaches execute.

Activity

  1. minjikim89 commented on Sep 6, 2026

    @minjikim89
    Author

    A correction to the top of this issue: §6.4 now lists four mitigations, not three. consequentialHint (Consequential Annotation for Tool Executions) landed in #217 on 2026-09-03, before I filed. I missed it.

    It does not change the proposal, and it makes the line a bit clearer. consequentialHint runs in the same direction as the rest of §6.4. The page declares something about itself, and a client or agent decides what to do with that declaration. Its effect depends on that decision. An agent that ignores the hint still books the flight.

    What is proposed here is not a declaration. (a), (b) and (c) are preconditions inside the tool body, and they refuse the write whether or not an annotation was honored. An agent that ignores an OUT_OF_SCOPE refusal does not get the write, because the write never happens.

    So the two compose rather than overlap. Annotations tell a cooperative agent what to be careful with. Page-enforced boundaries hold when the agent is not cooperative, or is cooperating with someone else's injected instruction.

  2. mysticalseeker24 commented on Sep 12, 2026

    @mysticalseeker24

    Support for this, and a fourth shape worth naming alongside (a)–(c), from a second implementation.

    (a), (b) and (c) are all preconditions inside the tool body: the call arrives, and the implementation refuses it. There is a second enforcement point available to a page that §6.4 does not currently name — the tool is not registered at all, so the call cannot be made.

    We derive the registered set from page state: each tool carries an available(state) predicate, and the set is recomputed on every state change. An action that is not legal right now is absent from getTools(), so an agent cannot attempt it. Nineteen tools are defined; never more than eight are live.

    The two have genuinely different properties, which is why I think both belong in the section rather than one standing for the other:

    preconditions in the tool body absence from the registered set
    Granularity per target, per field, per revision per state — coarse
    What the agent learns a structured refusal it can act on nothing, unless the page says why
    Failure mode if the agent ignores it write refused no call to ignore
    Cost a guard per write tool a predicate per tool, and a state machine to hang it on

    The coarseness is a real limitation: absence cannot express "you may edit these three slides", which is exactly what your (a) does. And absence alone destroys context, which is #262 — we return a reason_code and an unlock_by for every tool that is not live, so the agent is told what would make an action legal again rather than being left to infer it from a shorter list. Without that counterpart, enforcement by absence is worse than a refusal, not better.

    Argument-bound approval, for the (a) direction

    One mechanism that may be useful for the "person-owned write scope" shape. Our consequential tools require a grant that is hashed over the canonical JSON of the arguments, expires, and is consumed once. The binding matters more than the expiry: an approval obtained for one set of arguments cannot be replayed for another, so the scope is not "this tool is approved" but "this call is approved". A mismatch is a structured refusal naming the field that changed.

    It is the same idea as your (b), pointed at the agent's own earlier proposal rather than at the person's concurrent edits.

    (c) is load-bearing, and more so than it looks

    I measured the cancellation path on #299 while looking at this. On Chrome 152 execute receives no options argument at all, so a tool cannot observe the caller's abort — and a tool that applies items in a loop runs to completion after the caller aborts, with the writes landing in observable page state while the caller receives AbortError.

    So (c) is not a nicety for reporting. Without a page-owned cancellation path there is currently no cancellation of the write at all, only cancellation of the report — and the agent is told the operation did not complete when it did. An agent that retries has applied the batch twice.

    On the #288 limit

    Agreed, and worth stating in §6.4 as plainly as you have here. We reproduced #288 against our own consent gate and say in our README that page-side approval is necessary but not sufficient while the host is also a computer-use agent.

    One thing to add to that caveat: the obvious page-side hardening is unavailable on accessibility grounds. Anything that distinguishes synthesized input from physical input — CAPTCHA, dwell requirements, timing heuristics — is a barrier to someone using switch access, voice control, or a screen reader, and still would not establish authorization. That came up on #277 as well: the distinction worth encoding is agent-originated vs user-authorized, not synthesized vs physical. If §6.4 names the #288 limit, naming why the intuitive fix is closed off seems worth a sentence, or implementers will reach for it.

    Implementation: https://github.com/mysticalseeker24/parity-webmcp

  3. minjikim89 commented on Sep 12, 2026

    @minjikim89
    Author

    Hi @mysticalseeker24 — thank you for taking the time on this, and for bringing a second implementation to it rather than a vote. I have folded both of your points into the issue body so the proposal stands on them rather than on a comment thread.

    (d) is now in the proposal, credited to you, with your comparison table and the caveat you drew: absence is coarse, cannot express "you may edit these three slides", and destroys context unless something registered names what is missing and what would make it legal again. That last part is the bit I would not want an implementer to miss, since absence without the counterpart is worse than a refusal rather than better. It is good to hear you arrived at reason_code / unlock_by independently of our because / how.

    (c) is restated. You are right that it was filed as a reporting convenience and is not one. With no signal reaching execute on 152, a page-owned cancellation path is the only thing that ends the write at all, and without it the caller is told nothing completed while everything did. I have replaced the old wording with your framing and noted that it now holds on two independent implementations.

    On argument-bound approval: that is a genuinely nice mechanism and I had not considered hashing the grant over the canonical arguments. It points (b) at the agent's own earlier proposal rather than at the person's concurrent edits, as you say, and the two compose rather than overlap. If you write it up separately I would be glad to reference it.

    On the #288 limit, your point about accessibility is one I should have made and did not. Anything that separates synthesized input from physical input is a barrier to switch access, voice control, and screen reader users, and still would not establish authorization. If §6.4 names the limit, naming why the intuitive fix is closed off belongs with it. I will add that sentence rather than leave implementers to reach for a CAPTCHA.

  4. rohit-binaried commented on Sep 17, 2026

    @rohit-binaried

    One additional race to include in the evaluation is policy changing while an argument-bound approval is pending. Keeping the same arguments is necessary, but approval should still fail if the resource becomes unavailable, the binding disappears, or current host policy denies the action before execution.

    I maintain RCIP; its pending-confirmation path re-checks live conditions when the host resolves the request. Lifecycle reference. This is application-side evidence, not a WebMCP implementation or automatic database concurrency control.

    I agree with the limit already recorded here: page approval cannot prove a human acted when the caller also controls the browser UI. Server authorization and domain-level concurrency checks still need their own enforcement. Testing that pending-approval race would complement the argument-digest case without relying on the model to honor a hint.

  5. Hronom commented on Sep 21, 2026

    @Hronom

    A practical evaluation matrix could make the proposed §6.4 mitigations easier to compare without asking the model to behave:

    • accepted write: the intended target/version changes and an independent read-back proves the postcondition;
    • structured refusal: scope, optimistic-concurrency or human-only precondition fails, with no application mutation;
    • absent tool: the capability is unavailable for the current state, and the reason/unlock condition is inspectable through a separate read path;
    • policy drift: keep the canonical arguments unchanged but change the resource version, availability predicate or host policy between proposal and dispatch;
    • cancellation: stop after a known prefix and report exactly what landed, rather than only returning caller-side AbortError;
    • unregister/disconnect race: let execution begin, remove the registration or transport, then distinguish completed effect, no effect and unknown effect through independent reconciliation.

    For each case I would record tool/catalog revision, argument or proposal digest, target/context generation, policy revision, dispatch start, effect evidence and authoritative postcondition separately. The same arguments should not make an old approval valid after policy or target drift. UnknownError should not be interpreted as “nothing happened” when the tool or page already produced more specific evidence.

    I maintain Hronaut, a local visible Browser/MCP workspace, and this is an architectural/evaluation suggestion rather than a WebMCP compatibility claim. Prepared with AI assistance.

  6. added
    documentationImprovements or additions to documentation
    security-trackerGroup bringing to attention of security, or tracked by the security Group but not needing response.
    on Sep 25, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    documentationImprovements or additions to documentationsecurity-trackerGroup bringing to attention of security, or tracked by the security Group but not needing response.

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions