Skip to content

Excess Base64 data ignored after padding by default #145264

Description

@serhiy-storchaka

Bug report

After adding the ignorechars parameter for the Base64 decoder (see #144001), decoding in non-strict mode is almost equivalent to decoding with ignorechars including all characters. Except for one detail -- in non-strict mode the first valid padding stops decoding. Any following data is silently ignored. This leads to issues like #137687.

This contradicts RFC 4648, section 3.3 which only allows to ignore the pad character if it is present before the end of the encoded data.

Furthermore, such specifications MAY ignore the pad
character, "=", treating it as non-alphabet data, if it is present
before the end of the encoded data. If more than the allowed number
of pad characters is found at the end of the string (e.g., a base 64
string terminated with "==="), the excess pad characters MAY also be
ignored.

b'YW==Jj' and b'YWJ=j' should be decoded to b'abc', not to b'a' or b'ab'.

So, how are we going to fix this issue? We can simply change the behavior by default -- this may be a breaking change, but it is a bugfix, it breaks incorrect behavior. We can start long process of emitting a FutureWarning, and then changing the behavior few releases later. We can add a new option to alter the behavior and start emitting a FutureWarning by default.

In 3.15+ we can simply pass the ignorechars argument to enable RFC 4648 complaining lenient behavior. The question is about the default behavior.

Linked PRs

Activity

  1. added
    type-bugAn unexpected behavior, bug, or error
    3.13only security fixes
    3.14bugs and security fixes
    3.15pre-release feature fixes, bugs and security fixes
    on Feb 26, 2026
  2. added 2 commits that reference this issue on Feb 26, 2026
  3. serhiy-storchaka commented on Feb 27, 2026

    @serhiy-storchaka
    MemberAuthor

    For example, using standard CLI utility base64 on Linux:

    $ echo YW==Jj | base64 -di
    abc
    $ echo -n YWJ=j | base64 -di
    abc

    Using Python:

    $ echo -n YW==Jj | python3 -m base64 -d
    a
    $ echo -n YWJ=j | python3 -m base64 -d
    ab
  4. added 2 commits that reference this issue on Mar 22, 2026
  5. added a commit that references this issue on Mar 23, 2026
  6. added a commit that references this issue on Mar 23, 2026
  7. added a commit that references this issue on Mar 23, 2026
  8. 11 remaining items

  9. added a commit that references this issue on Apr 25, 2026
  10. elboulangero commented on May 7, 2026

    @elboulangero

    @serhiy-storchaka I tried the base64 example in various containers and can't reproduce your findings.

    debian:latest aka. trixie, which comes with GNU coreutils:

    # dpkg -S bin/base64
    coreutils: /usr/bin/base64
    # dpkg -l | grep coreutils | awk '{print $2 " " $3}'
    coreutils 9.7-3
    
    # echo YW==Jj | base64 -di
    abase64: invalid input
    # echo -n YWJ=j | base64 -di
    abbase64: invalid input
    

    Note the leading a or ab in the output, it means that some characters are decoded before the command fails.

    ubuntu:latest aka. 26.04, which comes with Rust' uutils:

    # dpkg -S bin/base64
    coreutils-from-uutils: /usr/bin/base64
    # dpkg -l | grep uutils | awk '{print $2 " " $3}'
    coreutils-from-uutils 0.0.0~ubuntu25
    # base64 --version
    base64 (uutils coreutils) 0.8.0
    
    # echo YW==Jj | base64 -di
    base64: error: invalid input
    # echo -n YWJ=j | base64 -di
    base64: error: invalid input
    

    And Alpine 3.23.4, which comes with busybox:

    # realpath /bin/base64
    /bin/busybox
    # busybox | head -1
    BusyBox v1.37.0 (2025-12-16 14:19:28 UTC) multi-call binary.
    
    # echo YW==Jj | base64 -di
    abase64: truncated input
    # echo -n YWJ=j | base64 -di
    abbase64: truncated input
    

    What distro did you use for your tests?

    In my tests both coreutils and busybox decoded a and ab which is exactly what Python did. Am I missing something?

  11. serhiy-storchaka commented on May 7, 2026

    @serhiy-storchaka
    MemberAuthor

    @shadchin Interesting. It seems that all these examples worked when the length of the encoded data was not the multiple of 3. They worked for tokens which have fixed length.

    @elboulangero It was on Ubuntu 25.10.

    I believe that the base64 implementation in Ubuntu 26.04 does not comply with RFC 4648, because it a) does not error '=' not at the end even without the -i option, b) interprets '=' not at the end as the padding instead of ignoring it.

    $ echo -n QQ==QQ== | base64 -d | hd
    00000000  41 41                                             |AA|
    00000002
    $ echo -n QQQQ | base64 -d | hd
    00000000  41 04 10                                          |A..|
    00000003

    But this is not related to Python implementation.

  12. elboulangero commented on May 8, 2026

    @elboulangero

    @elboulangero It was on Ubuntu 25.10.

    Trying in Ubuntu 25.10 container then!

    # dpkg -S bin/base64
    coreutils-from-uutils: /usr/bin/base64
    # dpkg -l | grep uutils | awk '{print $2 " " $3}'
    coreutils-from-uutils 0.0.0~ubuntu24
    # base64 --version
    base64 (uutils coreutils) 0.2.2
    
    # echo YW==Jj | base64 -di; echo
    abc
    # echo -n YWJ=j | base64 -di; echo
    abc
    

    So I can confirm your findings now, as you can see! And it also means that Ubuntu 25.10 was an outlier, it was the only one to behave this way.

    Ubuntu 25.10 came with uutils 0.2.2, and Ubuntu 26.04 comes with uutils 0.8.0, so it looks like the behavior changed between those 2 versions.

  13. elboulangero commented on May 8, 2026

    @elboulangero

    I believe that the base64 implementation in Ubuntu 26.04 does not comply with RFC 464

    FWIW I get the same output on a Debian trixie system (GNU coreutils 9.7-3):

    # echo -n QQ==QQ== | base64 -d | hd
    00000000  41 41                                             |AA|
    00000002
    
    # echo -n QQQQ | base64 -d | hd
    00000000  41 04 10                                          |A..|
    00000003
    

    I suppose uutils has to balance between RFC compliance and bug-for-bug compatibility with GNU coreutils.

    (note: uutils is the Rust re-implementation of GNU coreutils that is now default in Ubuntu)

  14. elboulangero commented on May 8, 2026

    @elboulangero

    @serhiy-storchaka I have one last question. Should this change be backported to older Python releases? I ask as I'm doing CVE backport for Debian. From the outside, it's not clear why it was backported to 3.13 and no further.

  15. serhiy-storchaka commented on May 8, 2026

    @serhiy-storchaka
    MemberAuthor

    The issue looked like potential security issue, but we did not have real examples of possible attacks, so it was not classified as security issue. This change can potentially break the user code (and it turned out that it really breaks the user code), so it was safer to only apply it to maintained versions. If something goes wrong, it will be easier to revert or alter it in maintained versions.

  16. serhiy-storchaka commented on May 9, 2026

    @serhiy-storchaka
    MemberAuthor

    I reported the bug in uutils coreutils: uutils/coreutils#12204 .

  17. elboulangero commented on May 11, 2026

    @elboulangero

    @serhiy-storchaka thanks for clarifying!

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    3.13only security fixes3.14bugs and security fixes3.15pre-release feature fixes, bugs and security fixestype-bugAn unexpected behavior, bug, or error

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions