Repository navigation
Add bytes_per_line parameter to binascii.b2a_base64 #141966
Description
Activity
- addedtype-featureA feature request or enhancementA feature request or enhancement
on Nov 26, 2025 Hi team,
I would like to work on improving
binascii.b2a_base64by adding an optionalbytes_per_lineparameter. Currently, modules like base64, email, imaplib, and plistlib split input bytes in Python before callingb2a_base64, which creates multiple Python objects and increases memory copies.My proposal:
- Add
bytes_per_line: Optional[int] = Nonetob2a_base64. - If set, the output will wrap encoded bytes at that line length, inserting newlines as needed.
- Default behavior remains unchanged for backward compatibility.
This change will reduce overhead, simplify code, and improve efficiency for large inputs. I’m happy to provide a patch with tests and benchmarks.
Please assign this issue to me if it’s acceptable.
Thank you!
- Add
Hi @AnshArya927 great enthusiasm! In the CPython repository issues aren't typically assigned to individuals. In this case there's already an existing patch linked in the issue / code ready, the issue is here for discussion around what the API should look like and if CPython wants this change in general.
- addedperformancePerformance or resource usagePerformance or resource usagestdlibStandard Library Python modules in the Lib/ directoryStandard Library Python modules in the Lib/ directory
on Nov 26, 2025 FTR,
b2a_base64must conform to RFC 3548.email& co follow MIME RFC which has the notion of line feeds and cutoff. Here is the relevant paragraph for RFC 3548:MIME [3] is often used as a reference for base 64 encoding. However,
MIME does not define "base 64" per se, but rather a "base 64
Content-Transfer-Encoding" for use within MIME. As such, MIME
enforces a limit on line length of base 64 encoded data to 76
characters. MIME inherits the encoding from PEM [2] stating it is
"virtually identical", however PEM uses a line length of 64
characters. The MIME and PEM limits are both due to limits within
SMTP.Implementations MUST NOT not add line feeds to base encoded data
unless the specification referring to this document explicitly
directs base encoders to add line feeds after a specific number of
characters.AFAIU, this means that we could have an interface that automatically adds whatever needs to be added (newlines and splits). So I would probably be in favor of this change. Some observations and questions:
- Are performances affected? how much do we gain with the for-loop version vs this one? how much do we lose for the paths that don't use this feature?
newlineis meant to be used exactly once, namely add a trailing newline. If I usenewline=Falseandbytes_per_line=100, it doesn't always make sense. Sometimes it fits 100 chars, sometimes not. In particular, one parameter affects the other which is usually not a good design choice for Python itself (it's usually a pattern we want to avoid). Now, in this case, it looks like worth the change (it's not unheard of for people to wrap b64 data).bytes_per_lineseems overly verbose but I don't think ew have a better name. Are there precedents elsewhere (other languages)?
Re: Performance, it should be still working on gathering numbers (the list of buffers + join -> just build a bytes mirrors the optimization in gh-139877 and I suspect will be about the same magnitude)
re:
bytes_per_line, existing comments around these pieces suggestedmax_line_lengthbut the code functions on number of input bytes not number of output bytes. Definitely open to suggestions, I tend to bias towards verbose names.Oh, I missed this issue and opened almost identical one: #143214. Details are slightly different:
- A newline at the end and newlines between lines are controlled by orthogonal options, because we often need multiline output without a trailing newline.
- The option specifies the length of the output line, not the number of input bytes per line.
bytes_per_linelooks similar tobytes_per_sepinbytes.hex()which has a counterintuitive behavior -- you usually need to specify a negative value for it.Reacted by Cody Maloney
Feature or enhancement
Proposal:
Currently
base64,email.contentmanager,imaplib, andplistliball have code which takes a contiguous bytes, splits it into at mostbytes_per_line(ormaxbinsize) length chunks. These chunks are passed tobinascii.b2a_base64, the results collected, and then joined back together to create a final result. For example:cpython/Lib/base64.py
Lines 565 to 573 in 33efd71
Internally
b2a_base64is usingPyBytesWriterto manage the buffer and that could hold the final joined together bytes. To do that, need to teach it to handlebytes_per_lineterminating lines after that many bytes and inserting a newline if required.That reduces the amount of code to call as well as increasing efficiency of these cases by reducing the number of Python objects involved as well as the number of times data is copied.
Proposed new signature:
Sample implementation: cmaloney@705bd9b
Has this already been discussed elsewhere?
This is a minor feature, which does not need previous discussion elsewhere
Links to previous discussion of this feature:
No response