Repository navigation
Escape "<" and ">" in attributes when serializing HTML #6235
Description
Activity
- addedneeds implementer interestMoving the issue forward requires implementers to express interestMoving the issue forward requires implementers to express interest
on Dec 17, 2020 cc @whatwg/html-parser
- addedsecurity-trackerGroup bringing to attention of security, or tracked by the security Group but not needing response.Group bringing to attention of security, or tracked by the security Group but not needing response.security/privacyThere are security or privacy implicationsThere are security or privacy implications
on Dec 18, 2020 Example like, DOMPurify bypass is website/email old virus when you click on advertisement image/button/hyperlink/text.
It was resolved by checking malicious code during HTML DOM Parsing.SOLUTION
< and > to < and >
Change at code level as per your requirement like *lt".IMO, the true reason of the described behavior is that
<style>is a raw text element so there is no "elements" and "attributes" inside it from the HTML DOM perspective, only text content (as correctly shown in the 3rd DOM example, with the#text: <a title="child node of thestyleelement). Changing this behavior doesn't seem to be web compatible since at least the>character is widely used in in-page styles as a CSS child combinator.Could you please explain why do you expect<svg></p><style><a title="</style><img src onerror=alert(1)>">and<svg><p></p><style><a title="</style><img src onerror=alert(1)>"></style></svg>to be parsed differently?@SelenIT because there it's an SVG
styleelement and those parse differently. Foreign content is tricky.I understand the difference between HTML
styleand SVGstyle. I didn't get why the same HTMLpelement breaks the foreign element parsing in the second case and doesn't break in the first. Is it the difference introduced by the "fragment case" (https://html.spec.whatwg.org/#parsing-main-inforeign)?Parsing of
<svg></p>is actually a spec bug that still works in Safari. Check: #5113[edit]: But the spec bug isn't the only case in which escaping of
<and>in attributes would help. Check this famous XSS in Google Search video.It had the following payload:
<noscript><p title="</noscript><img src onerror=alert(1)>">
It abused the difference in NOSCRIPT parsing when scripting is enabled and disabled. If the code would be serialized to:
<noscript><p title="</noscript><img src onerror=alert(1)>"></p></noscript>
then the XSS would also not have been possible.
Reacted by Ilya StreltsynThanks a lot for the explanation!
I agree that escaping these characters will prevent the XSS in these cases. But isn't the existence of parsing modes/cases where attribute-like sequences aren't actually parsed as attributes the significant part of the problem here?
@SelenIT are there known mutation XSS exploits with parsing modes that don't also use "<" or ">" in an attribute value?
I think this change is a good idea and is probably doable.
The main unknown is the web compat risk. In 2008, all browsers except IE escaped "<" and ">" in attribute values, and the spec was changed based on feedback from myself, citing a single page that was broken in Opera but worked in IE. I don't recall any fallout of newly broken pages from making that change back then. However, it's been over a decade, and new content might have come to rely on the current behavior.
Any breakage here seems hard to find through static analysis, since it requires
innerHTMLor something to serialize some HTML, and then something else to expect unescaped "<"s or ">"s. The example from 2008 was something like this:<a href="javascript: if (foo < 10) doSomething()" id="theLink">link</a> ... <script> theLink.onclick = function() { eval(this.outerHTML.match(/href="([^"]+)"/)[1]); return false; } </script>Yes, extremely silly code, but it's stuff like this that can break.
Unless someone comes up with an idea of how to identify and measure regressions beforehand, I think I would suggest experimenting with making this change in a browser and see what if anything breaks during the dev/beta period.
40 remaining items
@zcorpan This means that it will be available on the stable channel for everyone in Firefox 140, is that correct? (I'm currently preparing a blog post about this change so this information will be useful there!)
Reacted by Frederik BraunYes.
Reacted by Michał Bentkowski and Frederik Braunfwiw @securityMB if your blog post is still in progress, the fix is in ladybird as of a few weeks ago as well. LadybirdBrowser/ladybird#4839.
@ADKaster the blog post is actually already published:
Reacted by Andrew Kaster and Frederik BraunReacted by Frederik BraunReacted by Frederik BraunReacted by Frederik BraunLooks like we have the first bug report caused by this change: https://issues.chromium.org/issues/425325063 (Figma affected). This should be easy to fix though.
Is there a fix you envision other than asking the website to make it work on their end?
I think it's one of these cases where the website itself has to be fixed. The only way to "fix" it on our end would be to revert the change.
Reacted by Frederik Braun- added a commit that references this issue
on Aug 2, 2025 securityMB commented
on Aug 3, 2025 on Aug 3, 2025 · Hidden as off-topicAuthorshow commentMore actions- added a commit that references this issue
on Aug 7, 2025 - added a commit that references this issue
on Aug 8, 2025 - added a commit that references this issue
on Aug 14, 2025 - added a commit that references this issue
on Aug 14, 2025 MasterInQuestion commented
on Aug 31, 2025 on Aug 31, 2025 · Hidden as off-topicshow commentMore actions
I'm submitting this issue after a short discussion on Twitter with @zcorpan today.
I think we should change the rules of escaping a string in attribute mode, and also escape
<and>to<and>respectively.The fact that these characters are not escaped led to some security issues in HTML parsers and sanitizers.
As an example, see this DOMPurify bypass. The bug was that the following markup
was parsed into the following DOM tree in Chromium and Safari:
Now because the typical usage of sanitizers is as follows:
it means that the markup is serialized and then reparsed.
Now consider the following markup:
which is parsed into the following DOM tree:
It doesn't contain any harmful markup, so it is serialized to:
However, after reparsing a different DOM tree is created:
Leading to cross-site scripting. The reason for that is in fact that
pbreaks out foreign content.Please note that if
<and>were escaped, then the markup would be serialized to:Making this particular bypass (and many similar ones) impossible.