Repository navigation
Excess Base64 data ignored after padding by default #145264
Description
Activity
- addedtype-bugAn unexpected behavior, bug, or errorAn unexpected behavior, bug, or error3.13only security fixesonly security fixes3.14bugs and security fixesbugs and security fixes3.15pre-release feature fixes, bugs and security fixespre-release feature fixes, bugs and security fixes
on Feb 26, 2026 For example, using standard CLI utility
base64on Linux:$ echo YW==Jj | base64 -di abc
$ echo -n YWJ=j | base64 -di abc
Using Python:
$ echo -n YW==Jj | python3 -m base64 -d a
$ echo -n YWJ=j | python3 -m base64 -d ab
- added a commit that references this issue
on Mar 23, 2026 - added a commit that references this issue
on Mar 23, 2026 - added a commit that references this issue
on Mar 23, 2026 - marked "Short circuiting" in base64's b64decode, decode, decodebytes #79013 as a duplicate of this issue
on Apr 6, 2026 11 remaining items
- added a commit that references this issue
on Apr 19, 2026 - added a commit that references this issue
on Apr 21, 2026 @serhiy-storchaka I tried the
base64example in various containers and can't reproduce your findings.debian:latestaka. trixie, which comes with GNU coreutils:# dpkg -S bin/base64 coreutils: /usr/bin/base64 # dpkg -l | grep coreutils | awk '{print $2 " " $3}' coreutils 9.7-3 # echo YW==Jj | base64 -di abase64: invalid input # echo -n YWJ=j | base64 -di abbase64: invalid inputNote the leading
aorabin the output, it means that some characters are decoded before the command fails.ubuntu:latestaka. 26.04, which comes with Rust' uutils:# dpkg -S bin/base64 coreutils-from-uutils: /usr/bin/base64 # dpkg -l | grep uutils | awk '{print $2 " " $3}' coreutils-from-uutils 0.0.0~ubuntu25 # base64 --version base64 (uutils coreutils) 0.8.0 # echo YW==Jj | base64 -di base64: error: invalid input # echo -n YWJ=j | base64 -di base64: error: invalid inputAnd Alpine 3.23.4, which comes with busybox:
# realpath /bin/base64 /bin/busybox # busybox | head -1 BusyBox v1.37.0 (2025-12-16 14:19:28 UTC) multi-call binary. # echo YW==Jj | base64 -di abase64: truncated input # echo -n YWJ=j | base64 -di abbase64: truncated inputWhat distro did you use for your tests?
In my tests both coreutils and busybox decoded
aandabwhich is exactly what Python did. Am I missing something?@shadchin Interesting. It seems that all these examples worked when the length of the encoded data was not the multiple of 3. They worked for tokens which have fixed length.
@elboulangero It was on Ubuntu 25.10.
I believe that the base64 implementation in Ubuntu 26.04 does not comply with RFC 4648, because it a) does not error '=' not at the end even without the
-ioption, b) interprets '=' not at the end as the padding instead of ignoring it.$ echo -n QQ==QQ== | base64 -d | hd 00000000 41 41 |AA| 00000002 $ echo -n QQQQ | base64 -d | hd 00000000 41 04 10 |A..| 00000003
But this is not related to Python implementation.
@elboulangero It was on Ubuntu 25.10.
Trying in Ubuntu 25.10 container then!
# dpkg -S bin/base64 coreutils-from-uutils: /usr/bin/base64 # dpkg -l | grep uutils | awk '{print $2 " " $3}' coreutils-from-uutils 0.0.0~ubuntu24 # base64 --version base64 (uutils coreutils) 0.2.2 # echo YW==Jj | base64 -di; echo abc # echo -n YWJ=j | base64 -di; echo abcSo I can confirm your findings now, as you can see! And it also means that Ubuntu 25.10 was an outlier, it was the only one to behave this way.
Ubuntu 25.10 came with
uutils 0.2.2, and Ubuntu 26.04 comes withuutils 0.8.0, so it looks like the behavior changed between those 2 versions.I believe that the base64 implementation in Ubuntu 26.04 does not comply with RFC 464
FWIW I get the same output on a Debian trixie system (GNU coreutils 9.7-3):
# echo -n QQ==QQ== | base64 -d | hd 00000000 41 41 |AA| 00000002 # echo -n QQQQ | base64 -d | hd 00000000 41 04 10 |A..| 00000003I suppose uutils has to balance between RFC compliance and bug-for-bug compatibility with GNU coreutils.
(note: uutils is the Rust re-implementation of GNU coreutils that is now default in Ubuntu)
@serhiy-storchaka I have one last question. Should this change be backported to older Python releases? I ask as I'm doing CVE backport for Debian. From the outside, it's not clear why it was backported to 3.13 and no further.
The issue looked like potential security issue, but we did not have real examples of possible attacks, so it was not classified as security issue. This change can potentially break the user code (and it turned out that it really breaks the user code), so it was safer to only apply it to maintained versions. If something goes wrong, it will be easier to revert or alter it in maintained versions.
I reported the bug in uutils coreutils: uutils/coreutils#12204 .
@serhiy-storchaka thanks for clarifying!
- added a commit that references this issue
on May 27, 2026 - added 2 commits that reference this issue
on Jul 22, 2026 - added a commit that references this issue
on Aug 21, 2026
Bug report
After adding the
ignorecharsparameter for the Base64 decoder (see #144001), decoding in non-strict mode is almost equivalent to decoding withignorecharsincluding all characters. Except for one detail -- in non-strict mode the first valid padding stops decoding. Any following data is silently ignored. This leads to issues like #137687.This contradicts RFC 4648, section 3.3 which only allows to ignore the pad character if it is present before the end of the encoded data.
b'YW==Jj'andb'YWJ=j'should be decoded tob'abc', not tob'a'orb'ab'.So, how are we going to fix this issue? We can simply change the behavior by default -- this may be a breaking change, but it is a bugfix, it breaks incorrect behavior. We can start long process of emitting a FutureWarning, and then changing the behavior few releases later. We can add a new option to alter the behavior and start emitting a FutureWarning by default.
In 3.15+ we can simply pass the
ignorecharsargument to enable RFC 4648 complaining lenient behavior. The question is about the default behavior.Linked PRs