Skip to content

Encourage always-escaping ampersand character. - #11988

Open
dmsnell wants to merge 1 commit into
whatwg:mainfrom
dmsnell:syntax-errors/always-escape-amp
Open

dmsnell wants to merge 1 commit into
whatwg:mainfrom
dmsnell:syntax-errors/always-escape-amp

Conversation

@dmsnell

@dmsnell dmsnell commented Dec 4, 2025

Copy link
Copy Markdown
Contributor

In the example highlighting ambiguities from missing semicolons on named character references, a "correct" encoding is provided, but that example makes no mention of the fact that the fragment was ambiguous precisely because the ampersand wasn't escaped.

This patch adds a clarifying note explaining how this situation is avoided by always escaping the ampersand.

  • At least two implementers are interested (and none opposed):
    • …
    • …
  • Tests are written and can be reviewed and commented upon at:
    • …
  • Implementation bugs are filed:
    • Chromium: …
    • Gecko: …
    • WebKit: …
    • Deno (only for timers, structured clone, base64 utils, channel messaging, module resolution, web workers, and web storage): …
    • Node.js (only for timers, structured clone, base64 utils, channel messaging, and module resolution): …
  • Corresponding HTML AAM & ARIA in HTML issues & PRs:
  • MDN issue is filed: …
  • The top of this comment includes a clear commit message to use.

(See WHATWG Working Mode: Changes for more details.)

In the example highlighting ambiguities from missing semicolons on named
character references, a "correct" encoding is provided, but that example
makes no mention of the fact that the fragment was ambiguous precisely
because the ampersand wasn't escaped.

This patch adds a clarifying note explaining how this situation is
avoided by always escaping the ampersand.

Co-authored-by: Jon Surrell <jon.surrell@automattic.com>
GitHub-PR: 11988
GitHub-PR-URL: whatwg#11988
@dmsnell
dmsnell force-pushed the syntax-errors/always-escape-amp branch from d1fb385 to 9753779 Compare December 4, 2025 19:50
@dmsnell

dmsnell commented Dec 4, 2025

Copy link
Copy Markdown
Contributor Author

As a side note, I overlooked adding my name to the list of contributors in my first submission.

@sirreal

sirreal commented Dec 5, 2025 •

Copy link
Copy Markdown
Contributor

I was surprised to find no recommendation about escaping & with character references anywhere in the HTML standard. The section this PR touches seems to encourage not escaping & if it is not ambiguous (bold mine):

Thus, the correct way to express the above cases is as follows:

<a href="?bill&ted">Bill and Ted</a> <!-- &ted is ok, since it's not a named character reference -->
<a href="?art&amp;copy">Art and Copy</a> <!-- the & has to be escaped, since &copy is a named character reference -->

I read this as if &amp;ted would be wrong in some way, since it isn't the correct way. However, it seems much simpler to me to escape the ampersand here as &amp;.

I would change this section to something like the following:

-<!-- &ted is ok, since it's not a named character reference -->
+<!-- "&ted" is ok because "ted" is not a named character reference. 
+<!-- "&amp;ted" is equivalent and less error-prone because "&amp;" explicitly decodes to "&". -->

There is precedent for such a recommendation. Section 4.12.1.3 Restrictions for contents of script elements has a prominent note with an encoding recommendation:

The easiest and safest way to avoid the rather strange restrictions described in this section is to always escape an ASCII case-insensitive match for "<!--" as "\x3C!--", "<script" as "\x3Cscript", and "</script" as "\x3C/script" when these sequences appear in literals in scripts (e.g. in strings, regular expressions, or comments), and to avoid writing code that uses such constructs in expressions. Doing so avoids the pitfalls that the restrictions in this section are prone to triggering: namely, that, for historical reasons, parsing of script blocks in HTML is a strange and exotic practice that acts unintuitively in the face of these sequences.


Section 13.1.4 Character references seems like a good place to add a similar note. For example

Note

Where character references are allowed, it's a good idea to always encode & with its character reference &amp;. This prevents any ambiguity as to whether the & is part of a character reference or a literal &.

I would consider mention the most common characters that are useful to escape in different contexts, but the note about & seems particularly helpful.

@annevk

annevk commented Dec 5, 2025

Copy link
Copy Markdown
Member

https://html.spec.whatwg.org/multipage/syntax.html#character-references already requires this so I'm not sure we need to state it again in the parser section. Is the problem that the parser doesn't flag it?

@dmsnell

dmsnell commented Dec 5, 2025

Copy link
Copy Markdown
Contributor Author

Is the problem that the parser doesn't flag it?

I believe the problem here is that the illustrative example in the syntax-error section explicitly states that the correct way to produce HTML text containing & is to not escape it if what follows is not a legitimately-parsed character reference.

The example illustrates that a parser will correctly identify &ted as that raw string, but suggests that &ted is more appropriate than &amp;ted.

So basically this is just a confusing aspect for implementers and it seems like we could tweak the wording to maintain the demonstration of how these errors are handled without encouraging people to lean on syntax errors in cases where they produce the right output.

@annevk

annevk commented Dec 5, 2025

Copy link
Copy Markdown
Member

I see, this is part of https://html.spec.whatwg.org/multipage/introduction.html#syntax-errors.

We don't disallow &ted currently so unless we also change the HTML Writing requirements in some way I'd be a bit hesitant to change it in this one place.

@dmsnell

dmsnell commented Dec 5, 2025

Copy link
Copy Markdown
Contributor Author

@annevk thanks. I’m very open to trying out different ideas, but I think the spec is actually a bit vague on this.

already requires this

Unless I’m wrong, the spec does not require that & be escaped as &amp;, only that when mixing character references with text that they must begin with & and be followed by the correct syntax.

However, if someone is authoring HTML and not intending to produce a character reference, a stray & is both properly decoded by the parser and not forbidden.

I think we all agree that the intention is to always escape & as &amp;, but in the nitty gritty, unless it’s hidden in some other section none of us have scoured up yet, it’s not explicitly normalized as such. The only reference we’ve been able to find that isn’t implied is the one in this PR, where the spec assertively states that it’s correct to omit the escaping.

@dmsnell

dmsnell commented Dec 5, 2025

Copy link
Copy Markdown
Contributor Author

I apologize for omitting the before/after screenshots, but I took a before shot and was waiting to add it to the description until I had the parser previews generated but then they never appeared and I forgot to upload the before-shot anyway. Here is the relevant context from the modified section.

Screenshot 2025-12-04 at 12 51 32 PM

@annevk

annevk commented Dec 5, 2025

Copy link
Copy Markdown
Member

That's what I'm saying as well though in my latest comment. The Writing section explicitly allows you to do this. So I don't want to accept this PR as-is, as it'll contradict the Writing section.

@zcorpan was involved in some of the details here and should probably weigh in.

@dmsnell

dmsnell commented Dec 5, 2025

Copy link
Copy Markdown
Contributor Author

sounds great, and I have no wish that this be as-is. in fact, I was hoping for further input because I myself struggled to figure out how best to represent it. @sirreal is the author of the original suggestion.

interestingly enough, the HTML 3 spec was clearer on this point, but that entire document comprises only a handful of ill-defined paragraphs 🙃

Because certain characters will be interpreted as markup, they should be represented by markup…for instance the character "&" must be represented by the entity &amp;.

@zcorpan

zcorpan commented Dec 9, 2025

Copy link
Copy Markdown
Member

I think it's worth considering switching to require escaped ampersands. The rules for when it's allowed are non-trivial and it's surprising that &ted is OK but &copy is not OK, or that the behavior is different between in data and in attribute values.

Always escape & is clear and easy to understand.

This was my position in 2007 also: https://lists.whatwg.org/pipermail/whatwg-whatwg.org/2007-September/012457.html

cc @hsivonen @sideshowbarker

@sirreal

sirreal commented Dec 9, 2025 •

Copy link
Copy Markdown
Contributor

Always escape & is clear and easy to understand.

This is what I'd really like to address with at least a recommendation in the HTML standard that & is best escaped where applicable.

@dmsnell linked to the HTML3 spec. HTML4 also makes a recommendation:

Authors should use "&amp;" (ASCII decimal 38) instead of "&" to avoid confusion with the beginning of a character reference (entity reference open delimiter).

Escaping & is something we understand implicitly and it's apparent in functions like PHP's htmlspecialchars or Python's html.escape.

An explicit recommendation in the standard about & escaping would be a service to web developers.

@zcorpan

zcorpan commented Dec 11, 2025

Copy link
Copy Markdown
Member

I think we should make it a parse error if we change this.

@sideshowbarker

Copy link
Copy Markdown
Member

I think we should make it a parse error if we change this.

I would rather we don't. I say that because, I don't actually want to implement an error or warning for this in the checker — despite whatever the spec may end up being changed to say here. I don't think it will actually be good for users to be getting new errors or warnings from the checker about this.

But if it's made an actual parse error in the spec, I would somewhat be forced into it, regardless — because for errors from the HTML parser, the checker basically just bubbles all those up as-is.

That said, I would also not personally implement a parse error for it in the HTML parser sources. But there's nothing that would prevent any other contributor (or code owner) for the parser code from implementing it.

@zcorpan

zcorpan commented Dec 12, 2025

Copy link
Copy Markdown
Member

Thanks @sideshowbarker .

I think unescaped ampersand falls into at least:
https://html.spec.whatwg.org/multipage/introduction.html#syntax-errors

  • Unintuitive error-handling behavior (different parsing in data vs attribute values is unintuitive)
  • Errors involving fragile syntax constructs (there are 2000+ named charrefs, knowing when & followed by text is ok is hard)

It's true that a new check means people will be presented with errors that were previously ok, which is a cost. But we improve the learnability of HTML and could avoid errors where entities are replaced but they were intended to be text.

@tabatkins

Copy link
Copy Markdown
Contributor

The problem is that virtually every <a> will trigger this error, if it contains any query parameters. It's uncommon for people to escape the & separating params.

Yes, it's confusing that in <a href="foo?bar&copy=bar">foo?bar&copy=bar</a> the attribute works correctly but the text shows a copyright symbol, but forcing checkers to flag all such links as invalid would be a huge issue, I think.

@zcorpan

zcorpan commented Feb 27, 2026

Copy link
Copy Markdown
Member

I ran a quick query on 0.1% of httparchive to find how common unescaped & vs escaped &amp; is in src or href attributes.

pages_total pages_with_unescaped pages_with_amp pages_with_both unescaped_attrs amp_attrs pct_pages_with_unescaped pct_pages_with_amp ratio_pages_unescaped_to_amp ratio_unescaped_attrs_of_total ratio_amp_attrs_of_total ratio_unescaped_attrs_to_amp
17152 7045 6116 2868 56239 54723 0.4107392724 0.3565764925 1.151896664 0.5068311674 0.4931688326 1.02770316

~41% of pages use &, and ~36% of pages use &amp;. When counting the number of elements, it's ~50/50.

So yes, it is indeed common. It could be compared to ~73% of pages that have syntax errors about mismatched tags.

Things to consider:

  • If conformance checkers report this as an error, and people fix it, was their time well spent? Probably no.
  • If we change this, does it make it easier to learn the rules of the HTML syntax? IMO yes.

Maybe we could instead disallow unescaped & when it appears as text, i.e. not in attribute values? That would avoid the issue of wasting time "fixing" harmless & in URLs in attribute values, but make the rules simpler and avoid surprises for & as text.

query
WITH params AS (
  SELECT
    r'(?i)\s(?:src|href)\s*=\s*["\'][^"\']*&[a-z0-9_-]+[&=][^"\']*["\']' AS re_unescaped,
    r'(?i)\s(?:src|href)\s*=\s*["\'][^"\']*&amp;[^"\']*["\']'            AS re_amp
),
per_page AS (
  SELECT
    page,
    rank,
    ARRAY_LENGTH(REGEXP_EXTRACT_ALL(response_body, p.re_unescaped)) AS n_unescaped_attrs,
    ARRAY_LENGTH(REGEXP_EXTRACT_ALL(response_body, p.re_amp))       AS n_amp_attrs
  FROM `httparchive.crawl.requests` TABLESAMPLE SYSTEM(0.1 PERCENT)
  CROSS JOIN params p
  WHERE date = '2026-02-01'
    AND client = 'desktop'
    AND page = url
),
agg AS (
  SELECT
    COUNT(*) AS pages_total,

    -- page-level
    COUNTIF(n_unescaped_attrs > 0) AS pages_with_unescaped,
    COUNTIF(n_amp_attrs > 0)       AS pages_with_amp,
    COUNTIF(n_unescaped_attrs > 0 AND n_amp_attrs > 0) AS pages_with_both,

    -- element-level (attribute matches)
    SUM(n_unescaped_attrs) AS unescaped_attrs,
    SUM(n_amp_attrs)       AS amp_attrs
  FROM per_page
)
SELECT
  pages_total,
  pages_with_unescaped,
  pages_with_amp,
  pages_with_both,

  unescaped_attrs,
  amp_attrs,

  -- ratios: pages
  SAFE_DIVIDE(pages_with_unescaped, pages_total) AS pct_pages_with_unescaped,
  SAFE_DIVIDE(pages_with_amp,       pages_total) AS pct_pages_with_amp,
  SAFE_DIVIDE(pages_with_unescaped, pages_with_amp) AS ratio_pages_unescaped_to_amp,

  -- ratios: elements/attributes
  SAFE_DIVIDE(unescaped_attrs, unescaped_attrs + amp_attrs) AS ratio_unescaped_attrs_of_total,
  SAFE_DIVIDE(amp_attrs,       unescaped_attrs + amp_attrs) AS ratio_amp_attrs_of_total,
  SAFE_DIVIDE(unescaped_attrs, amp_attrs)                   AS ratio_unescaped_attrs_to_amp
FROM agg;

@dmsnell

dmsnell commented Feb 27, 2026

Copy link
Copy Markdown
Contributor Author

There seems to be a similarity to this and how angle brackets were changed to always-escape in serialization, and I don’t think that added any parser error when reading unescaped angle brackets.

Are we potentially overthinking a clarification to what is already a non-normative statement describing how the parser works but which is unintentionally suggesting writing in an ambiguous style?

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Development

Successfully merging this pull request may close these issues.

6 participants