Skip to content

Encoding detection unpredictable #105

Description

@marob

When validating multiple files in concurrency, the encoding detection is not always predictable.
The easiest way to explain is probably to reproduce my test case:

  • launch the validation server (15.6.29)
java -cp vnu.jar nu.validator.servlet.Main 8888
ab -T"text/html; encoding=utf-8" -p testCase.txt -c 100 -n 10000 -v 4 http://localhost:8888/?out=json | grep Unmappable

You should obtain some results of the form (if not, try to restart the server and re-validate):

{"messages":[{"type":"info","message":"The Content-Type was “text/html”. Using the HTML parser."},{"type":"error","message":"The character encoding was not declared. Proceeding using “big5”."},{"type":"error","lastLine":7,"lastColumn":47,"message":"Unmappable byte sequence: “c2”, “a0”."},{"type":"info","message":"Using the schema for HTML5 + SVG 1.1 + MathML 3.0 + RDFa Lite 1.1."}]}

As you can see, the detected encoding is sometimes "big5", but only once or twice among thousands... hence the unpredictability.
When this encoding is detected, the unbreakable space triggers the "Unmappable byte sequence" error.

I think the charset detection should be predictable.
My test demonstrate it is not.

Regards

Activity

  1. added a commit that references this issue on Nov 3, 2015
  2. sideshowbarker commented on Nov 4, 2015

    @sideshowbarker
    Member

    I’ve been able to reproduce this to the degree that I can get some Proceeding using “big5” messages to appear—but I can only do it just right after I’ve re-started the validator. But after getting the messages to appear that first time, I can’t get them to re-appear again no matter how many times I re-run ab with the testCase.txt file and the supplied parameters.

    So to me it’s not clear that this is a bug that’s going to affect the running service in practice, nor is it clear that it’s caused by the validator code and not a bug in ab. But it would be good to hear if @hsivonen has any insights on this.

  3. marob commented on Nov 4, 2015

    @marob
    Author

    This is definitely not a bug in ab as I've only used it to create an easily reproducible test case. The charset detection bug does occur in real use cases that doesn't involve ab.

    Let me explain how I've come to identify this bug.

    I first encountered the bug on a real-life project by using https://github.com/nikestep/grunt-html-angular-validate on my angular HTML templates (approximately 190 files).
    As it occurred randomly, I tried to determine if the bug came from the tool (grunt-html-angular-validate) or the validator. I then monitored the network with tcpdump and achieved to compare the 2 cases (working one and failing one). In those 2 cases, the HTTP requests were identical (byte to byte) while the responses were distinct (one without the "Unmappable byte sequence" and one with the "Unmappable byte sequence" and the "big5" charset detected).
    That demonstrates the problem is server-side (in the validator) as it should be idempotent.

    I then tried to create an easily reproducible test case which led to the one I've exposed in my initial comment.

    About the fact it occurs only after a server restart, I've got the same behavior with my test case, but in my real-life project, the bug continues to occur randomly without restarting the server (I'm using a locally deployed service in a long-running tomcat).

    FYI, I've found a workaround for my project by configuring grunt-html-angular-validate to provide a charset in an HTML meta markup (nikestep/html-angular-validate#8), but it doesn't solve the underlying bug in the validator as it only bypasses the automatic charset detection.

  4. hsivonen commented on Nov 4, 2015

    @hsivonen
    Member

    Very weird. If you specify charset=UTF-8 in the Content-Type of the input, heuristic detection should never run. In any case, the autodetection really needs to go away from the parser to match browsers better.

  5. sideshowbarker commented on Nov 5, 2015

    @sideshowbarker
    Member

    @marob Thanks for the detailed explanation. Based on that it’s clear this is a bug the validator code or in one of the libraries that are among its dependencies, if not strictly the validator code itself. The actual charset detection is actually done in some of that library code, and I kind of wonder if the case is that that code may be working as expected and it’s just that the input to that code is getting subtly corrupted somehow before it gets there—by, I dunno, the Jetty code. Anyway, I will try to make time at some point to troubleshoot this but I am very unlikely will not be able to put any find time to look into it much further myself within the next several weeks.

  6. marob commented on Nov 5, 2015

    @marob
    Author

    @hsivonen About the (unexpected) triggering of the heuristic detection, I've found why.
    When monitoring the network while using grunt-html-angular-validate, I obtained a Content-Type header with the value text/html; encoding=utf-8 which I copy-pasted in my ab command line to reproduce the bug.
    But the header should have been "text/html; charset=utf-8" (charset instead of encoding). I've created a pull request (thomasdavis/w3cjs#23) to solve this issue.

    Still, if heuristic detection is triggered (no charset is defined, neither in the http headers nor in an html meta tag), the behavior of the validator is unpredictable.

    @sideshowbarker I completely understand it can't be solved in the minute as it involves concurrency and/or randomness. Good luck solving this bug! :)

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions