Repository navigation
Encoding detection unpredictable #105
Description
Activity
I’ve been able to reproduce this to the degree that I can get some
Proceeding using “big5”messages to appear—but I can only do it just right after I’ve re-started the validator. But after getting the messages to appear that first time, I can’t get them to re-appear again no matter how many times I re-runabwith thetestCase.txtfile and the supplied parameters.So to me it’s not clear that this is a bug that’s going to affect the running service in practice, nor is it clear that it’s caused by the validator code and not a bug in
ab. But it would be good to hear if @hsivonen has any insights on this.This is definitely not a bug in
abas I've only used it to create an easily reproducible test case. The charset detection bug does occur in real use cases that doesn't involveab.Let me explain how I've come to identify this bug.
I first encountered the bug on a real-life project by using https://github.com/nikestep/grunt-html-angular-validate on my angular HTML templates (approximately 190 files).
As it occurred randomly, I tried to determine if the bug came from the tool (grunt-html-angular-validate) or the validator. I then monitored the network with tcpdump and achieved to compare the 2 cases (working one and failing one). In those 2 cases, the HTTP requests were identical (byte to byte) while the responses were distinct (one without the "Unmappable byte sequence" and one with the "Unmappable byte sequence" and the "big5" charset detected).
That demonstrates the problem is server-side (in the validator) as it should be idempotent.I then tried to create an easily reproducible test case which led to the one I've exposed in my initial comment.
About the fact it occurs only after a server restart, I've got the same behavior with my test case, but in my real-life project, the bug continues to occur randomly without restarting the server (I'm using a locally deployed service in a long-running tomcat).
FYI, I've found a workaround for my project by configuring grunt-html-angular-validate to provide a charset in an HTML meta markup (nikestep/html-angular-validate#8), but it doesn't solve the underlying bug in the validator as it only bypasses the automatic charset detection.
Very weird. If you specify
charset=UTF-8in theContent-Typeof the input, heuristic detection should never run. In any case, the autodetection really needs to go away from the parser to match browsers better.@marob Thanks for the detailed explanation. Based on that it’s clear this is a bug the validator code or in one of the libraries that are among its dependencies, if not strictly the validator code itself. The actual charset detection is actually done in some of that library code, and I kind of wonder if the case is that that code may be working as expected and it’s just that the input to that code is getting subtly corrupted somehow before it gets there—by, I dunno, the Jetty code. Anyway, I will try to make time at some point to troubleshoot this but I am very unlikely will not be able to put any find time to look into it much further myself within the next several weeks.
@hsivonen About the (unexpected) triggering of the heuristic detection, I've found why.
When monitoring the network while using grunt-html-angular-validate, I obtained aContent-Typeheader with the valuetext/html; encoding=utf-8which I copy-pasted in myabcommand line to reproduce the bug.
But the header should have been"text/html; charset=utf-8"(charset instead of encoding). I've created a pull request (thomasdavis/w3cjs#23) to solve this issue.Still, if heuristic detection is triggered (no charset is defined, neither in the http headers nor in an html meta tag), the behavior of the validator is unpredictable.
@sideshowbarker I completely understand it can't be solved in the minute as it involves concurrency and/or randomness. Good luck solving this bug! :)
When validating multiple files in concurrency, the encoding detection is not always predictable.
The easiest way to explain is probably to reproduce my test case:
You should obtain some results of the form (if not, try to restart the server and re-validate):
As you can see, the detected encoding is sometimes "big5", but only once or twice among thousands... hence the unpredictability.
When this encoding is detected, the unbreakable space triggers the "Unmappable byte sequence" error.
I think the charset detection should be predictable.
My test demonstrate it is not.
Regards