Repository navigation
Mismatch between glibc and X11 locale.alias #64286
Description
Activity
The locale module uses locale alias table derived from X11 locale.alias file for mapping bare locale names without encodings to locale names with encodings. However sometimes glibc default encoding for a locale differs from that used in X11 locale.alias.
Here is full differences table:
GLibc X11 locale.aliasaz_az az_AZ.UTF-8 az_AZ.ISO8859-9E
ca_ad ca_AD.ISO8859-15 ca_AD.ISO8859-1
ca_fr ca_FR.ISO8859-15 ca_FR.ISO8859-1
ca_it ca_IT.ISO8859-15 ca_IT.ISO8859-1
cy_gb cy_GB.ISO8859-14 cy_GB.ISO8859-1
en_in en_IN.UTF-8 en_IN.ISO8859-1
et_ee et_EE.ISO8859-1 et_EE.ISO8859-15
fi_fi fi_FI.ISO8859-1 fi_FI.ISO8859-15
gd_gb gd_GB.ISO8859-15 gd_GB.ISO8859-1
hi_in hi_IN.UTF-8 hi_IN.ISCII-DEV
iu_ca iu_CA.UTF-8 iu_CA.NUNACOM-8
iw_il iw_IL.ISO8859-8 he_IL.ISO8859-8
ka_ge ka_GE.GEORGIAN_PS ka_GE.GEORGIAN-ACADEMY
lo_la lo_LA.UTF-8 lo_LA.MULELAO-1
mi_nz mi_NZ.ISO8859-13 mi_NZ.ISO8859-1
nr_za nr_ZA.UTF-8 nr_ZA.ISO8859-1
nso_za nso_ZA.UTF-8 nso_ZA.ISO8859-15
ru_ru ru_RU.ISO8859-5 ru_RU.UTF-8
rw_rw rw_RW.UTF-8 rw_RW.ISO8859-1
sq_al sq_AL.ISO8859-1 sq_AL.ISO8859-2
ss_za ss_ZA.UTF-8 ss_ZA.ISO8859-1
ta_in ta_IN.UTF-8 ta_IN.TSCII-0
tg_tj tg_TJ.KOI8_T tg_TJ.KOI8-C
th_th th_TH.TIS_620 th_TH.ISO8859-11
tn_za tn_ZA.UTF-8 tn_ZA.ISO8859-15
ts_za ts_ZA.UTF-8 ts_ZA.ISO8859-1
tt_ru tt_RU.UTF-8 tt_RU.TATAR-CYR
ur_pk ur_PK.UTF-8 ur_PK.CP1256
uz_uz uz_UZ.ISO8859-1 uz_UZ.UTF-8
uz_uz@cyrillic uz_UZ.UTF-8@cyrillic uz_UZ.UTF-8
vi_vn vi_VN.UTF-8 vi_VN.TCVN
zh_cn zh_CN.GB2312 zh_CN.gb2312
zh_tw zh_TW.BIG5 zh_TW.big5
zh_tw.euctw zh_TW.EUC_TW zh_TW.eucTWFor example with the en_IN encoding:
>>> import locale, _locale >>> _locale.setlocale(locale.LC_CTYPE) 'en_IN' >>> locale.getlocale() ('en_IN', 'ISO8859-1') >>> locale.nl_langinfo(locale.CODESET) 'UTF-8' >>> locale.setlocale(locale.LC_CTYPE, locale.getlocale()) Traceback (most recent call last): File "<stdin>", line 1, in <module> File "/home/serhiy/py/cpython/Lib/locale.py", line 592, in setlocale return _setlocale(category, locale) locale.Error: unsupported locale setting
- addedstdlibStandard Library Python modules in the Lib/ directoryStandard Library Python modules in the Lib/ directorytype-bugAn unexpected behavior, bug, or errorAn unexpected behavior, bug, or error
on Dec 28, 2013 Needed a test for few common locales (en_IN, ru_RU) and maybe for unusual locales (uz_uz, uz_uz@cyrillic).
I would prefer to have a separate issue that updates the aliases table to glibc 2.24.
I agree that it's reasonable to have glibc's aliases override
the X.org ones, but this patch makes some pretty significant changes to Python's default assumptions with respect to default encodings for several locales.While some changes obviously make sense (e.g. 'ca_AD.ISO8859-1' to 'ca_AD.ISO8859-15'), others are less clear (e.g. 'cy_GB.ISO8859-1' to 'cy_GB.ISO8859-14' or 'tg_TJ.KOI8-C' to 'tg_TJ.KOI8-T' or several of the moves from ISO encodings to UTF-8). Is there some reference for why glibc chose different values than X.org for these ?
I also don't understand why some "xx.utf-8" locale mappings were removed - I don't think we should remove those, unless they are no lot needed due to some other logic implying these mappings.
Since these are major changes, we need an appropriate warning in the NEWS file (and the "What's New" document), an update of the top comment (under "### Database") to mention that the glibc database takes precedence and where to find it,
'cy_GB.ISO8859-1' to 'cy_GB.ISO8859-14'
Looks as just fixing an error. The default West-European ISO8859-1 is changed to Celtic cy_GB.ISO8859-14. This looks better option for Welsh.
'tg_TJ.KOI8-C' to 'tg_TJ.KOI8-T'
KOI8-C is not supported by Python, but KOI8-T is supported. I don't know what KOI8-C means, there are several rarely used incompatible encodings with this name.
I also don't understand why some "xx.utf-8" locale mappings were removed - I don't think we should remove those, unless they are no lot needed due to some other logic implying these mappings.
The aliases table is a table of exceptions. Removed entries no longer are exceptional.
On 07.03.2017 18:23, Serhiy Storchaka wrote:
Serhiy Storchaka added the comment:
> 'cy_GB.ISO8859-1' to 'cy_GB.ISO8859-14'
Looks as just fixing an error. The default West-European ISO8859-1 is changed to Celtic cy_GB.ISO8859-14. This looks better option for Welsh.
> 'tg_TJ.KOI8-C' to 'tg_TJ.KOI8-T'
KOI8-C is not supported by Python, but KOI8-T is supported. I don't know what KOI8-C means, there are several rarely used incompatible encodings with this name.
While all this may make sense, I'm missing some more reasoning
behind the differences between X.org and glibc.This change also looks strange:
- 'ka_ge': 'ka_GE.GEORGIAN-ACADEMY', + 'ka_ge': 'ka_GE.GEORGIAN_PS', 'ka_ge.georgianacademy': 'ka_GE.GEORGIAN-ACADEMY', 'ka_ge.georgianps': 'ka_GE.GEORGIAN-PS', 'ka_ge.georgianrs': 'ka_GE.GEORGIAN-ACADEMY',
Why is GEORGIAN_PS written with an underscore whereas the other
mappings use dashes ?Or this one:
- 'fi_fi': 'fi_FI.ISO8859-15',
+ 'fi_fi': 'fi_FI.ISO8859-1',
Why would a locale switch away from an encoding having
the Euro sign to one without it ?Or why is this latin variant removed:
- 'nan_tw@latin': 'nan_TW.UTF-8@latin',
Why should Russians switch back to ISO ?
- 'ru_ru': 'ru_RU.UTF-8',
+ 'ru_ru': 'ru_RU.ISO8859-5',
or from ISO to KOI ?
- 'russian': 'ru_RU.ISO8859-5',
+ 'russian': 'ru_RU.KOI8-R',
The more I look at these changes, the more I believe we
should not simply take everything we find in the files
for granted. They obviously both have bugs.> I also don't understand why some "xx.utf-8" locale mappings were removed - I don't think we should remove those, unless they are no longer needed due to some other logic implying these mappings.
The aliases table is a table of exceptions. Removed entries no longer are exceptional.
It's not a table of exceptions, it's a table mapping commonly
used locale settings to ones which the lib C understands :-)But regardless, I checked the code and it is already
smart enough to convert lib C incompatible spellings such
as "utf8" to "UTF-8", so these entries can indeed be
removed, but only if the locale is otherwise listed.In some cases, it's probably better to drop the ".utf8"
to have more generic mappings, e.g.+ 'bhb_in.utf8': 'bhb_IN.UTF-8',
or
'de_li.utf8': 'de_LI.UTF-8',though I'd expect that mapping to be:
'de_li': 'de_LI.ISO8859-1',as for all other "de" entries.
- 'fi_fi': 'fi_FI.ISO8859-15',
Why is the X11 locale alias map used at all? It seems like it can only create confusion with libc.
Not all platforms use glibc 2.24 as libc.
Ideally most of entries should even not exist. We should ask libc for the default encoding if it is not included in the locale name. The aliases table should be used only for mapping commonly used but unsupported by libc locales to supported by libc locales.
On 08.03.2017 08:20, Serhiy Storchaka wrote:
Serhiy Storchaka added the comment:
Not all platforms use glibc 2.24 as libc.
True. Many don't even use glibc.
Ideally most of entries should even not exist. We should ask libc for the default encoding if it is not included in the locale name. The aliases table should be used only for mapping commonly used but unsupported by libc locales to supported by libc locales.
I think you have a wrong understanding of what this alias table
is used for: we need it to determine the lib C compatible locale
name without using lib C APIs such as setlocale(), since these are
not thread safe and have side-effects for the whole process.The alias table is there to avoid having to go to the lib C
to ask it indirectly for more details. Unfortunately, there are
no cross-platform lib C APIs which would allow querying these
details without also changing the local settings of the process.I know that Python still plays the usual "save current locale,
run setlocale(), revert to previous locale" trick in a couple
of places and this works if Python is the only thread running,
but it doesn't when embedded into other applications.Regarding the patch: we cannot simply use the output from the
script to set new values. The changes have to be manually
reviewed as well.E.g. this entry in the table is clearly a typo:
'en_zw.utf8': 'en_ZS.UTF-8',(it should read en_ZW.UTF-8)
This entry appears wrong as well:
'eo': 'eo_XX.ISO8859-3',(XX is not a valid country ISO code)
How should we go about this ? Mark all the problems in the PR ?
The problem is that that table can get incorrect result for non-Linux platforms (or for Linux with old glibc).
On 08.03.2017 07:27, Benjamin Peterson wrote:
Why is the X11 locale alias map used at all? It seems like it can only create confusion with libc.
Because it was the only such maintained mapping available at the
time. It's also used for the X.org system, which has a rather strong
focus on user interfaces where locale matter a lot, unlike
the lib C :-)On 08.03.2017 10:37, Serhiy Storchaka wrote:
The problem is that that table can get incorrect result for non-Linux platforms (or for Linux with old glibc).
Sure, it's a best effort approach.
Also note that on today's systems you often don't have the full set of
locales available anymore - instead these have to either be installed
separately or generated on the target system.Our locale database works on all these system, regardless of
what's installed or not.Why was the PR merged while we were still discussing it ?
8 remaining items
I'm feeling there is something wrong with the current locale design. See issues bpo-504219, bpo-10466, bpo-20088, bpo-25191, bpo-29571.
I'm still confused about what getlocale() is supposed to do. Why do we attempt to return an encoding anyway if the underlying setlocale call doesn't return one? Is getlocale() not supposed to a simple wrapper over the C locale? If not, how is one supposed to get the encoding associated with the C locale?
The old alias table code meant that the encoding returned from getlocale() could be related to or completely unrelated to the actual C locale. Misunderstanding this results in issues like bpo-29571.
The main purpose of the alias table is to support normalization and this is used for getdefaultencoding() which was created to be able to determine the default encoding based on what X.org uses as default without doing temporary setlocale() tricks.
Now, normalization also happens when passing a locale value to the underlying setlocale(), mainly to avoid many common bugs due to setlocale() being extremely picky about the locale value. A side effect of this is that normalization will also kick in to add the encoding in case no encoding is given in the parameter.
Note that no normalization is necessary to simply set the configured default locale configured on the system. In such a case, you'd run setlocale('LC_ALL') and get what's configured.
If you run the lib C setlocale() with a locale without encoding, the encoding used by the system entirely on what's configured on the system. The SUPPORTED file only gives a hint at what glibc think it should install per default, but any admin or distributor could change these settings simply by running localedef with some other encoding (charmap in locale speak).
I suppose that we could resolve some of the confusion by adding a parameter to disable this normalization in setlocale().
Hi all,
The locale in the latest Ubuntu 18.04 contains en_IL as valid locale, but Python cannot resolve this.
This makes test failure in pandas.
pandas-dev/pandas#20957en_IL has significant impact because this is English locale and now supported in the latest Ubuntu. Is there any plan to add only en_IL?
(Note that I've already created the PR. ( #6707 ))
(pandas-dev) [pandas] locale -a C C.UTF-8 en_AG en_AG.utf8 en_AU.utf8 en_BW.utf8 en_CA.utf8 en_DK.utf8 en_GB.utf8 en_HK.utf8 en_IE.utf8 en_IL en_IL.utf8 en_IN en_IN.utf8 en_NG en_NG.utf8 en_NZ.utf8 en_PH.utf8 en_SG.utf8 en_US.utf8 en_ZA.utf8 en_ZM en_ZM.utf8 en_ZW.utf8 ja_JP.utf8 POSIXBenjamin's patch did two things: 1) made the glibc alias table taking precedence over the X11 one; 2) updated the alias mapping with new glibc. The first part is controversial, but updating the alias mapping with new glibc is made regularly. PR 6708 updates it with glibc 2.27. This adds 39 new aliases and fixes bpo-32781 and bpo-33432.
Thanks, Serhiy.
I believe we can close this old issue.
The discussion was certainly a useful one. I guess we should stop updating the alias table automatically and instead add new aliases or change existing ones based on more research and using the X11 files as well as glibc and other resources to help.
Note: these values reflect the state of the issue at the time it was migrated and might not reflect the current state.
Show more details
GitHub fields:
bugs.python.org fields: