Skip to content

Mismatch between glibc and X11 locale.alias #64286

Description

@serhiy-storchaka
BPO 20087
Nosy @malemburg, @loewis, @benjaminp, @serhiy-storchaka, @Licht-T
PRs
  • update locale aliases for glibc 2.24 #422
  • bpo-20087: Revert "make the glibc alias table take precedence over thee X11 one (#422)" #713
  • bpo-20087: Update locale alias mapping with glibc 2.27 supported locales. #6708
  • [3.7] bpo-20087: Update locale alias mapping with glibc 2.27 supported locales. (ПР-6708) #6713
  • [3.6] bpo-20087: Update locale alias mapping with glibc 2.27 supported locales. (ПР-6708) #6714
  • [2.7] bpo-20087: Update locale alias mapping with glibc 2.27 supported locales. (GH-6708). #6717
  • Note: these values reflect the state of the issue at the time it was migrated and might not reflect the current state.

    Show more details

    GitHub fields:

    assignee = None
    closed_at = <Date 2021-02-17.12:12:03.181>
    created_at = <Date 2013-12-28.09:29:49.864>
    labels = ['3.7', 'type-bug', 'library']
    title = 'Mismatch between glibc and X11 locale.alias'
    updated_at = <Date 2021-02-17.12:12:03.180>
    user = 'https://github.com/serhiy-storchaka'

    bugs.python.org fields:

    activity = <Date 2021-02-17.12:12:03.180>
    actor = 'lemburg'
    assignee = 'none'
    closed = True
    closed_date = <Date 2021-02-17.12:12:03.181>
    closer = 'lemburg'
    components = ['Library (Lib)']
    creation = <Date 2013-12-28.09:29:49.864>
    creator = 'serhiy.storchaka'
    dependencies = []
    files = []
    hgrepos = []
    issue_num = 20087
    keywords = ['patch']
    message_count = 33.0
    messages = ['207025', '288871', '289174', '289176', '289179', '289205', '289210', '289222', '289223', '289231', '289232', '289242', '289277', '289282', '289283', '289284', '289286', '289290', '289340', '289377', '289386', '289439', '289787', '290131', '290273', '316214', '316216', '316224', '316226', '316227', '316228', '316234', '387148']
    nosy_count = 6.0
    nosy_names = ['lemburg', 'loewis', 'benjamin.peterson', 'Arfrever', 'serhiy.storchaka', 'licht-t']
    pr_nums = ['422', '713', '6708', '6713', '6714', '6717']
    priority = 'normal'
    resolution = 'fixed'
    stage = 'resolved'
    status = 'closed'
    superseder = None
    type = 'behavior'
    url = 'https://bugs.python.org/issue20087'
    versions = ['Python 2.7', 'Python 3.5', 'Python 3.6', 'Python 3.7']

    Activity

    1. serhiy-storchaka commented on Dec 28, 2013

      @serhiy-storchaka
      MemberAuthor

      The locale module uses locale alias table derived from X11 locale.alias file for mapping bare locale names without encodings to locale names with encodings. However sometimes glibc default encoding for a locale differs from that used in X11 locale.alias.

      Here is full differences table:

                   GLibc                 X11 locale.alias
      

      az_az az_AZ.UTF-8 az_AZ.ISO8859-9E
      ca_ad ca_AD.ISO8859-15 ca_AD.ISO8859-1
      ca_fr ca_FR.ISO8859-15 ca_FR.ISO8859-1
      ca_it ca_IT.ISO8859-15 ca_IT.ISO8859-1
      cy_gb cy_GB.ISO8859-14 cy_GB.ISO8859-1
      en_in en_IN.UTF-8 en_IN.ISO8859-1
      et_ee et_EE.ISO8859-1 et_EE.ISO8859-15
      fi_fi fi_FI.ISO8859-1 fi_FI.ISO8859-15
      gd_gb gd_GB.ISO8859-15 gd_GB.ISO8859-1
      hi_in hi_IN.UTF-8 hi_IN.ISCII-DEV
      iu_ca iu_CA.UTF-8 iu_CA.NUNACOM-8
      iw_il iw_IL.ISO8859-8 he_IL.ISO8859-8
      ka_ge ka_GE.GEORGIAN_PS ka_GE.GEORGIAN-ACADEMY
      lo_la lo_LA.UTF-8 lo_LA.MULELAO-1
      mi_nz mi_NZ.ISO8859-13 mi_NZ.ISO8859-1
      nr_za nr_ZA.UTF-8 nr_ZA.ISO8859-1
      nso_za nso_ZA.UTF-8 nso_ZA.ISO8859-15
      ru_ru ru_RU.ISO8859-5 ru_RU.UTF-8
      rw_rw rw_RW.UTF-8 rw_RW.ISO8859-1
      sq_al sq_AL.ISO8859-1 sq_AL.ISO8859-2
      ss_za ss_ZA.UTF-8 ss_ZA.ISO8859-1
      ta_in ta_IN.UTF-8 ta_IN.TSCII-0
      tg_tj tg_TJ.KOI8_T tg_TJ.KOI8-C
      th_th th_TH.TIS_620 th_TH.ISO8859-11
      tn_za tn_ZA.UTF-8 tn_ZA.ISO8859-15
      ts_za ts_ZA.UTF-8 ts_ZA.ISO8859-1
      tt_ru tt_RU.UTF-8 tt_RU.TATAR-CYR
      ur_pk ur_PK.UTF-8 ur_PK.CP1256
      uz_uz uz_UZ.ISO8859-1 uz_UZ.UTF-8
      uz_uz@cyrillic uz_UZ.UTF-8@cyrillic uz_UZ.UTF-8
      vi_vn vi_VN.UTF-8 vi_VN.TCVN
      zh_cn zh_CN.GB2312 zh_CN.gb2312
      zh_tw zh_TW.BIG5 zh_TW.big5
      zh_tw.euctw zh_TW.EUC_TW zh_TW.eucTW

      For example with the en_IN encoding:

      >>> import locale, _locale
      >>> _locale.setlocale(locale.LC_CTYPE)
      'en_IN'
      >>> locale.getlocale()
      ('en_IN', 'ISO8859-1')
      >>> locale.nl_langinfo(locale.CODESET)
      'UTF-8'
      >>> locale.setlocale(locale.LC_CTYPE, locale.getlocale())
      Traceback (most recent call last):
        File "<stdin>", line 1, in <module>
        File "/home/serhiy/py/cpython/Lib/locale.py", line 592, in setlocale
          return _setlocale(category, locale)
      locale.Error: unsupported locale setting
    2. added
      stdlibStandard Library Python modules in the Lib/ directory
      type-bugAn unexpected behavior, bug, or error
      on Dec 28, 2013
    3. serhiy-storchaka commented on Mar 3, 2017

      @serhiy-storchaka
      MemberAuthor

      Needed a test for few common locales (en_IN, ru_RU) and maybe for unusual locales (uz_uz, uz_uz@cyrillic).

      I would prefer to have a separate issue that updates the aliases table to glibc 2.24.

    4. malemburg commented on Mar 7, 2017

      @malemburg
      Member

      I agree that it's reasonable to have glibc's aliases override
      the X.org ones, but this patch makes some pretty significant changes to Python's default assumptions with respect to default encodings for several locales.

      While some changes obviously make sense (e.g. 'ca_AD.ISO8859-1' to 'ca_AD.ISO8859-15'), others are less clear (e.g. 'cy_GB.ISO8859-1' to 'cy_GB.ISO8859-14' or 'tg_TJ.KOI8-C' to 'tg_TJ.KOI8-T' or several of the moves from ISO encodings to UTF-8). Is there some reference for why glibc chose different values than X.org for these ?

      I also don't understand why some "xx.utf-8" locale mappings were removed - I don't think we should remove those, unless they are no lot needed due to some other logic implying these mappings.

      Since these are major changes, we need an appropriate warning in the NEWS file (and the "What's New" document), an update of the top comment (under "### Database") to mention that the glibc database takes precedence and where to find it,

    5. serhiy-storchaka commented on Mar 7, 2017

      @serhiy-storchaka
      MemberAuthor

      'cy_GB.ISO8859-1' to 'cy_GB.ISO8859-14'

      Looks as just fixing an error. The default West-European ISO8859-1 is changed to Celtic cy_GB.ISO8859-14. This looks better option for Welsh.

      'tg_TJ.KOI8-C' to 'tg_TJ.KOI8-T'

      KOI8-C is not supported by Python, but KOI8-T is supported. I don't know what KOI8-C means, there are several rarely used incompatible encodings with this name.

      I also don't understand why some "xx.utf-8" locale mappings were removed - I don't think we should remove those, unless they are no lot needed due to some other logic implying these mappings.

      The aliases table is a table of exceptions. Removed entries no longer are exceptional.

    6. malemburg commented on Mar 7, 2017

      @malemburg
      Member

      On 07.03.2017 18:23, Serhiy Storchaka wrote:

      Serhiy Storchaka added the comment:

      > 'cy_GB.ISO8859-1' to 'cy_GB.ISO8859-14'

      Looks as just fixing an error. The default West-European ISO8859-1 is changed to Celtic cy_GB.ISO8859-14. This looks better option for Welsh.

      > 'tg_TJ.KOI8-C' to 'tg_TJ.KOI8-T'

      KOI8-C is not supported by Python, but KOI8-T is supported. I don't know what KOI8-C means, there are several rarely used incompatible encodings with this name.

      While all this may make sense, I'm missing some more reasoning
      behind the differences between X.org and glibc.

      This change also looks strange:

      -    'ka_ge':                                'ka_GE.GEORGIAN-ACADEMY',
      +    'ka_ge':                                'ka_GE.GEORGIAN_PS',
           'ka_ge.georgianacademy':                'ka_GE.GEORGIAN-ACADEMY',
           'ka_ge.georgianps':                     'ka_GE.GEORGIAN-PS',
           'ka_ge.georgianrs':                     'ka_GE.GEORGIAN-ACADEMY',

      Why is GEORGIAN_PS written with an underscore whereas the other
      mappings use dashes ?

      Or this one:

      • 'fi_fi': 'fi_FI.ISO8859-15',
        + 'fi_fi': 'fi_FI.ISO8859-1',

      Why would a locale switch away from an encoding having
      the Euro sign to one without it ?

      Or why is this latin variant removed:

      • 'nan_tw@latin': 'nan_TW.UTF-8@latin',

      Why should Russians switch back to ISO ?

      • 'ru_ru': 'ru_RU.UTF-8',
        + 'ru_ru': 'ru_RU.ISO8859-5',

      or from ISO to KOI ?

      • 'russian': 'ru_RU.ISO8859-5',
        + 'russian': 'ru_RU.KOI8-R',

      The more I look at these changes, the more I believe we
      should not simply take everything we find in the files
      for granted. They obviously both have bugs.

      > I also don't understand why some "xx.utf-8" locale mappings were removed - I don't think we should remove those, unless they are no longer needed due to some other logic implying these mappings.

      The aliases table is a table of exceptions. Removed entries no longer are exceptional.

      It's not a table of exceptions, it's a table mapping commonly
      used locale settings to ones which the lib C understands :-)

      But regardless, I checked the code and it is already
      smart enough to convert lib C incompatible spellings such
      as "utf8" to "UTF-8", so these entries can indeed be
      removed, but only if the locale is otherwise listed.

      In some cases, it's probably better to drop the ".utf8"
      to have more generic mappings, e.g.

      + 'bhb_in.utf8': 'bhb_IN.UTF-8',

      or

       'de_li.utf8':                           'de_LI.UTF-8',
      

      though I'd expect that mapping to be:

       'de_li':                           'de_LI.ISO8859-1',
      

      as for all other "de" entries.

    7. benjaminp commented on Mar 8, 2017

      @benjaminp
      Contributor

      Why is the X11 locale alias map used at all? It seems like it can only create confusion with libc.

    8. serhiy-storchaka commented on Mar 8, 2017

      @serhiy-storchaka
      MemberAuthor

      Not all platforms use glibc 2.24 as libc.

      Ideally most of entries should even not exist. We should ask libc for the default encoding if it is not included in the locale name. The aliases table should be used only for mapping commonly used but unsupported by libc locales to supported by libc locales.

    9. malemburg commented on Mar 8, 2017

      @malemburg
      Member

      On 08.03.2017 08:20, Serhiy Storchaka wrote:

      Serhiy Storchaka added the comment:

      Not all platforms use glibc 2.24 as libc.

      True. Many don't even use glibc.

      Ideally most of entries should even not exist. We should ask libc for the default encoding if it is not included in the locale name. The aliases table should be used only for mapping commonly used but unsupported by libc locales to supported by libc locales.

      I think you have a wrong understanding of what this alias table
      is used for: we need it to determine the lib C compatible locale
      name without using lib C APIs such as setlocale(), since these are
      not thread safe and have side-effects for the whole process.

      The alias table is there to avoid having to go to the lib C
      to ask it indirectly for more details. Unfortunately, there are
      no cross-platform lib C APIs which would allow querying these
      details without also changing the local settings of the process.

      I know that Python still plays the usual "save current locale,
      run setlocale(), revert to previous locale" trick in a couple
      of places and this works if Python is the only thread running,
      but it doesn't when embedded into other applications.

      Regarding the patch: we cannot simply use the output from the
      script to set new values. The changes have to be manually
      reviewed as well.

      E.g. this entry in the table is clearly a typo:

      'en_zw.utf8':                           'en_ZS.UTF-8',
      

      (it should read en_ZW.UTF-8)

      This entry appears wrong as well:

      'eo':                                   'eo_XX.ISO8859-3',
      

      (XX is not a valid country ISO code)

      How should we go about this ? Mark all the problems in the PR ?

    10. serhiy-storchaka commented on Mar 8, 2017

      @serhiy-storchaka
      MemberAuthor

      The problem is that that table can get incorrect result for non-Linux platforms (or for Linux with old glibc).

    11. malemburg commented on Mar 8, 2017

      @malemburg
      Member

      On 08.03.2017 07:27, Benjamin Peterson wrote:

      Why is the X11 locale alias map used at all? It seems like it can only create confusion with libc.

      Because it was the only such maintained mapping available at the
      time. It's also used for the X.org system, which has a rather strong
      focus on user interfaces where locale matter a lot, unlike
      the lib C :-)

    12. malemburg commented on Mar 8, 2017

      @malemburg
      Member

      On 08.03.2017 10:37, Serhiy Storchaka wrote:

      The problem is that that table can get incorrect result for non-Linux platforms (or for Linux with old glibc).

      Sure, it's a best effort approach.

      Also note that on today's systems you often don't have the full set of
      locales available anymore - instead these have to either be installed
      separately or generated on the target system.

      Our locale database works on all these system, regardless of
      what's installed or not.

    13. malemburg commented on Mar 8, 2017

      @malemburg
      Member

      Why was the PR merged while we were still discussing it ?

    14. 8 remaining items

    15. serhiy-storchaka commented on Mar 10, 2017

      @serhiy-storchaka
      MemberAuthor

      I'm feeling there is something wrong with the current locale design. See issues bpo-504219, bpo-10466, bpo-20088, bpo-25191, bpo-29571.

    16. benjaminp commented on Mar 11, 2017

      @benjaminp
      Contributor

      I'm still confused about what getlocale() is supposed to do. Why do we attempt to return an encoding anyway if the underlying setlocale call doesn't return one? Is getlocale() not supposed to a simple wrapper over the C locale? If not, how is one supposed to get the encoding associated with the C locale?

      The old alias table code meant that the encoding returned from getlocale() could be related to or completely unrelated to the actual C locale. Misunderstanding this results in issues like bpo-29571.

    17. malemburg commented on Mar 17, 2017

      @malemburg
      Member

      The main purpose of the alias table is to support normalization and this is used for getdefaultencoding() which was created to be able to determine the default encoding based on what X.org uses as default without doing temporary setlocale() tricks.

      Now, normalization also happens when passing a locale value to the underlying setlocale(), mainly to avoid many common bugs due to setlocale() being extremely picky about the locale value. A side effect of this is that normalization will also kick in to add the encoding in case no encoding is given in the parameter.

      Note that no normalization is necessary to simply set the configured default locale configured on the system. In such a case, you'd run setlocale('LC_ALL') and get what's configured.

      If you run the lib C setlocale() with a locale without encoding, the encoding used by the system entirely on what's configured on the system. The SUPPORTED file only gives a hint at what glibc think it should install per default, but any admin or distributor could change these settings simply by running localedef with some other encoding (charmap in locale speak).

      I suppose that we could resolve some of the confusion by adding a parameter to disable this normalization in setlocale().

    18. benjaminp commented on Mar 24, 2017

      @benjaminp
      Contributor

      New changeset df82808 by Benjamin Peterson in branch 'master':
      bpo-20087: Revert "make the glibc alias table take precedence over the X11 one (#422)" (#713)
      df82808

    19. benjaminp commented on Mar 24, 2017

      @benjaminp
      Contributor

      New changeset 02371e0 by Benjamin Peterson in branch 'master':
      make the glibc alias table take precedence over the X11 one (#422)
      02371e0

    20. Licht-T commented on May 5, 2018

      Licht-Tmannequin
      Mannequin

      Hi all,

      The locale in the latest Ubuntu 18.04 contains en_IL as valid locale, but Python cannot resolve this.
      This makes test failure in pandas.
      pandas-dev/pandas#20957

      en_IL has significant impact because this is English locale and now supported in the latest Ubuntu. Is there any plan to add only en_IL?

      (Note that I've already created the PR. ( #6707 ))

      (pandas-dev) [pandas] locale -a
      C
      C.UTF-8
      en_AG
      en_AG.utf8
      en_AU.utf8
      en_BW.utf8
      en_CA.utf8
      en_DK.utf8
      en_GB.utf8
      en_HK.utf8
      en_IE.utf8
      en_IL
      en_IL.utf8
      en_IN
      en_IN.utf8
      en_NG
      en_NG.utf8
      en_NZ.utf8
      en_PH.utf8
      en_SG.utf8
      en_US.utf8
      en_ZA.utf8
      en_ZM
      en_ZM.utf8
      en_ZW.utf8
      ja_JP.utf8
      POSIX
      
    21. serhiy-storchaka commented on May 5, 2018

      @serhiy-storchaka
      MemberAuthor

      Benjamin's patch did two things: 1) made the glibc alias table taking precedence over the X11 one; 2) updated the alias mapping with new glibc. The first part is controversial, but updating the alias mapping with new glibc is made regularly. PR 6708 updates it with glibc 2.27. This adds 39 new aliases and fixes bpo-32781 and bpo-33432.

    22. serhiy-storchaka commented on May 6, 2018

      @serhiy-storchaka
      MemberAuthor

      New changeset cedc9b7 by Serhiy Storchaka in branch 'master':
      bpo-20087: Update locale alias mapping with glibc 2.27 supported locales. (ПР-6708)
      cedc9b7

    23. serhiy-storchaka commented on May 6, 2018

      @serhiy-storchaka
      MemberAuthor

      New changeset 6049bda by Serhiy Storchaka (Miss Islington (bot)) in branch '3.7':
      [3.7] bpo-20087: Update locale alias mapping with glibc 2.27 supported locales. (GH-6708) (GH-6713)
      6049bda

    24. serhiy-storchaka commented on May 6, 2018

      @serhiy-storchaka
      MemberAuthor

      New changeset b1c70d0 by Serhiy Storchaka (Miss Islington (bot)) in branch '3.6':
      [3.6] bpo-20087: Update locale alias mapping with glibc 2.27 supported locales. (GH-6708) (GH-6714)
      b1c70d0

    25. serhiy-storchaka commented on May 6, 2018

      @serhiy-storchaka
      MemberAuthor

      New changeset a55ac80 by Serhiy Storchaka in branch '2.7':
      [2.7] bpo-20087: Update locale alias mapping with glibc 2.27 supported locales. (GH-6708). (GH-6717)
      a55ac80

    26. malemburg commented on May 6, 2018

      @malemburg
      Member

      Thanks, Serhiy.

    27. malemburg commented on Feb 17, 2021

      @malemburg
      Member

      I believe we can close this old issue.

      The discussion was certainly a useful one. I guess we should stop updating the alias table automatically and instead add new aliases or change existing ones based on more research and using the X11 files as well as glibc and other resources to help.

    28. transferred this issue fromon Apr 10, 2022
    Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

    Metadata

    Metadata

    Assignees

    No one assigned

      Labels

      3.7 (EOL)end of lifestdlibStandard Library Python modules in the Lib/ directorytype-bugAn unexpected behavior, bug, or error

      Projects

      No projects

        Milestone

        No milestone

        Relationships

        None yet

        Development

        No branches or pull requests

        Issue actions