Skip to content

iso8859-1 vs windows-1252 #25851

Description

@hashseed

This is somewhat related to #13722, but not quite.

Wikipedia contains the gist:

It is very common to mislabel Windows-1252 text with the charset label ISO-8859-1. [...] Most modern web browsers and e-mail clients treat the media type charset ISO-8859-1 as Windows-1252 to accommodate such mislabeling. This is now standard behavior in the HTML5 specification, which requires that documents advertised as ISO-8859-1 actually be parsed with the Windows-1252 encoding.

Chromium's ICU interprets "iso8859-1" to mean Windows-1252. Node.js does not. The WHATWG spec suggests Chromium's behavior to be correct.

The consequence of all of this is that when I build Node.js with Chromium's ICU, test/parallel/test-icu-transcode.js fails due to the character "€", which Windows-1252 includes, but ISO-8859-1 does not.

I propose:

  1. Change the test to pass for both interpretations.
  2. Conform to Chromium's behavior.

Activity

  1. hashseed commented on Jan 31, 2019

    @hashseed
    MemberAuthor
  2. added
    i18n-apiIssues and PRs related to Node.js internationalization support.
    on Jan 31, 2019
  3. added
    bufferIssues and PRs related to the buffer subsystem.
    string_decoderIssues and PRs related to the string_decoder subsystem.
    on Feb 1, 2019
  4. joyeecheung commented on Feb 1, 2019

    @joyeecheung
    Member

    The consequence of all of this is that when I build Node.js with Chromium's ICU, test/parallel/test-icu-transcode.js fails due to the character "€"

    From a glance of the test I couldn't see how the encoding differences would come into play if the test itself was decoded properly. Was the test file stored in Windows-1252? And then got interpreted as utf8 by the CJS loader?

  5. hashseed commented on Feb 1, 2019

    @hashseed
    MemberAuthor

    The test loads the euro sign as utf8 and converts to latin1. With iso8859-1 it doesn't encode so it gets replaced by "?".

  6. richardlau commented on Feb 1, 2019

    @richardlau
    Member

    Is this affected by full-icu?

  7. richardlau commented on Feb 1, 2019

    @richardlau
    Member

    Also #13722.

  8. hashseed commented on Feb 1, 2019

    @hashseed
    MemberAuthor

    I can check again, but iirc full icu doesn't fix this issue. If it did, the full icu build would fail the above mentioned test.

  9. hashseed commented on Feb 1, 2019

    @hashseed
    MemberAuthor

    Same actually also applies to "ascii". According to WHATWG it is to be interpreted as Windows-1252 as well.

  10. mhdawson commented on Feb 1, 2019

    @mhdawson
    Member

    @srl295 as our ICU expert.

  11. TimothyGu commented on Feb 2, 2019

    @TimothyGu
    Member

    We have a dedicated WHATWG text encoder/decoder util.TextDecoder and util.TextEncoder, which should strictly follow the semantics of the Encoding Standard. buffer.transcode() should indeed just fall back on whatever the linked ICU does. I believe #25866 is a good change because of this.

  12. hashseed commented on Feb 5, 2019

    @hashseed
    MemberAuthor

    Test fix landed.

  13. srl295 commented on Feb 14, 2019

    @srl295
    Member

    sorry to miss this, but the fix looks good.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    bufferIssues and PRs related to the buffer subsystem.i18n-apiIssues and PRs related to Node.js internationalization support.string_decoderIssues and PRs related to the string_decoder subsystem.

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions