Repository navigation
Option to skip padding for base64 urlsafe encoding/decoding #73613
Description
Activity
Suggest changing base64 module to better handle encoding schemes that don't use padding.
Because RFC4648 [1] allows other RFCs that implement RFC4648-compliant base64url encoding to explicitly stipulate that there is no padding. Dropping the padding is lossless when we know the length [2].
Various standard specifications require this - often crypto related (e.g., JWS [3] or named hashes [4]).RFC4648 specifically makes an exemption for this and it should be better supported in Python's standard library. There is a related closed issue [5] asking for the padding to be removed or altered which wouldn't comply with the spec. This request is different with a view to better support the wider specification.
Proposed behaviour adapted from resolution that ruby discussion on same topic [6]:
- base64.urlsafe_b64encode(s) should continue to produce padded output, but have an additional argument, padding, which defaults to True.
- base64.urlsafe_b64decode(s) should accept both padded and unpadded inputs. It can still reject incorrectly-padded input.
If that sounds sensible I'd like to put a patch/PR together.
From wikipedia [7]:
Some variants allow or require omitting the padding '=' signs to avoid them being confused with field separators, or require that any such padding be percent-encoded. Some libraries will encode '=' to '.'.
- [1] https://tools.ietf.org/html/rfc4648#page-4
- [2] http://stackoverflow.com/questions/4080988/why-does-base64-encoding-requires-padding-if-the-input-length-is-not-divisible-b
- [3] https://tools.ietf.org/html/rfc7515
- [4] https://tools.ietf.org/html/rfc6920#section-3
- [5] http://bugs.python.org/issue1661108
- [6] https://bugs.ruby-lang.org/issues/10740
- [7] https://en.wikipedia.org/wiki/Base64#Output_Padding
Reacted by codedust and Peter Kong- added3.7 (EOL)end of lifeend of lifestdlibStandard Library Python modules in the Lib/ directoryStandard Library Python modules in the Lib/ directorytype-bugAn unexpected behavior, bug, or errorAn unexpected behavior, bug, or error
on Feb 3, 2017 This sounds reasonable. I ran into a similar issue today trying to decode a JSON Web Key. Although I don't have any real say, I'd say that if you put together a patch it may have a higher chance to get reviewed.
I wonder about the following:
- What about adding a new kwarg to b64decode, passed through by urlsafe_b64decode, called "checkpad=True" which validates padding? Then we can just set that False when we need.
- At the same time it might be nice to pass "validate=False" through from urlsafe_b64decode and friends, so we can have some nicer validation of data.
- I like adding the "padding=True" arg to encode, but it may not be necessary given the ease of ".rstrip('=')" as an alternative. Anyway, if you will add it to encode, please add it to b64encode and pass through from the variant encoders to unify the API somewhat.
If you are still interested in putting together a patch, post a comment. Otherwise I may work on a patch for this.
Hi Robert, It would be at least a week or two before I could take another look at this so please feel free to work on it. Not sure why I didn't write a patch at the time!
Hi,
What is the status of this issue? there are many application that require url safe base64 encoding such as JOSE (JWT) which require ugly workarounds because of this bug or actually fail because of the bug.
From the (RFC-4648)[https://www.rfc-editor.org/rfc/rfc4648#section-5]
"""
The pad character "=" is typically percent-encoded when used in an
URI [9], but if the data length is known implicitly, this can be
avoided by skipping the padding; see section 3.2.
"""The python bytes/str length falls into the definition of implicitly.
Please support this behavior per RFC, there is no need for an additional parameter, just to read up to the end of the buffer.
Regards,
The main problem is in naming. Base 85 encoders support the
padargument to pad input to a multiple of 4 before encoding. Base 64 and Base 32 codecs use the pad character, and we want an option to omit pad characters in the encoding output and accept input without pad characters. What name should we use to avoid confusion with thepadargument in Base 85 encoders?Other problem is that while in encoders we only need a boolean value (add padding/do not add padding), in decoders we may need more options:
- require padding, error on excessive padding
- require padding, ignore excessive padding
- ignore any padding
- accept valid padding and no padding, error on invalid padding
- error on any padding
(and maybe more variants).
We now have the
ignorecharsoption, including=in it will switch between "error on excessive padding" and "ignore excessive padding". This could also applied for other cases, but we still can need more than a boolean value. Although I am not sure that all such variants are useful, so 4 variants (the boolean x=inignorechars) might be enough. Anyway we should have this in mind when choosing the parameter name.- added a commit that references this issue
on Apr 1, 2026 #147974 adds the
paddedparameter to all functions related to Base32 and Base64 codecs, exceptstandard_b64encode()andstandard_b64decode()(if this is standard, then padding is required, althoughstandard_b64decode()does not implement the current standard strictly).In the encoding functions it controls whether the pad character can be added in the output, in the decoding functions it controls whether padding is required in input. Inclusion of the pad character in
ignorechars(and also the value ofvalidate/strict_mode) controls whether the padding is ignored or treated as error whenpaddedis false. I hope this will be enough. I preferred simplicity and efficiency of implementation.The default value is True, except in
base64.urlsafe_b64decode(). So, padding of input no longer required inbase64.urlsafe_b64decode()by default.Reacted by Roy Hyunjin Han
Note: these values reflect the state of the issue at the time it was migrated and might not reflect the current state.
Show more details
GitHub fields:
bugs.python.org fields:
Linked PRs