Repository navigation
Unmatched </p> or </br> inside foreign context needs a special parser rule #5113
Description
Activity
/cc @whatwg/html-parser
(Additional impl data:) parse5 is currently consistent with Blink/Webkit
I should also have referenced the corresponding Chromium bug.
- added a commit that references this issue
on Nov 27, 2019 html5lib leaves the <p> or <br> inside <svg>:
$ echo "<svg></p></svg>" | python -c "from sys import stdin; \ import html5lib; from lxml import html; \ doc = html5lib.parse(stdin, treebuilder='lxml', namespaceHTMLElements=False); \ print html.tostring(doc)" <html><head></head><body><ns0:svg xmlns:ns0="http://www.w3.org/2000/svg"><p></p></ns0:svg> </body></html> $ echo "<svg></br></svg>" | python -c "from sys import stdin; \ import html5lib; from lxml import html; \ doc = html5lib.parse(stdin, treebuilder='lxml', namespaceHTMLElements=False); \ print html.tostring(doc)" <html><head></head><body><ns0:svg xmlns:ns0="http://www.w3.org/2000/svg"><br></ns0:svg> </body></html>- added a commit that references this issue
on Nov 30, 2019 - addedsecurity/privacyThere are security or privacy implicationsThere are security or privacy implications
on Dec 2, 2019 No problem - I just marked it all public. I had essentially been treating it as public given the blog post.
Reacted by Anne van Kesteren- added a commit that references this issue
on Dec 5, 2019 - added a commit that references this issue
on Dec 5, 2019 - added a commit that references this issue
on Mar 10, 2020 1 remaining item
While in general I'm pretty reluctant to change the parser, this seems like an obvious oversight in the current parser (given the odd
</p>/</br>parsing), and hence I'd support changing the behaviour to match Gecko here.Here's another sanitizer bypass that appears to be caused by the same issue: GHSA-vv2x-vrpj-qqpq
When writing tests for this, here's one for html5lib
foreign-fragment.dat, but should also test this in regular parsing mode (without#document-fragment)#data <svg></p><foo> #errors 9: HTML end tag “p” in a foreign namespace context. #document-fragment div #document | <svg svg> | <p> | <foo>- added a commit that references this issue
on Jun 4, 2021 - added a commit that references this issue
on Jun 23, 2021 - added a commit that references this issue
on Jul 26, 2021 - added a commit that references this issue
on Oct 25, 2021 - added a commit that references this issue
on Jun 3, 2022 - added a commit that references this issue
on Jul 29, 2022 - added a commit that references this issue
on May 1, 2025 - added a commit that references this issue
on May 28, 2025 - added a commit that references this issue
on Sep 3, 2025
As the current parser spec is written, <svg></p></svg> and <svg></br></svg> both result in <p> and <br> DOM nodes as children of the <svg>. As mentioned in this Chromium bug and this blog post, this can be exploited as a sanitizer bypass. Here is an example DOM Viewer link showing the behavior.
By my reading of the spec:
Current implementations:
I believe the spec should follow current Gecko behavior. I think the easiest way to change the spec would be to add a special case within the foreign context section for end tags whose tag name is "p" or "br", which closes the foreign context and then processes the </p> or </br> as normal for a non-foreign context.