Skip to content

Size the initial decompress buffer from the input - #1

Open
bdraco wants to merge 6 commits into
developfrom
preallocate-max-length-output-buffer
Open

bdraco wants to merge 6 commits into
developfrom
preallocate-max-length-output-buffer

Conversation

@bdraco

@bdraco bdraco commented Sep 14, 2026 •

Copy link
Copy Markdown
Owner

Experiment for pycompression#256, on this fork to test before proposing anything upstream.

Rule

Decompress.decompress sizes its initial output buffer from the input:

  • Start at 8 times the size of the input.
  • Never start below 16 KiB (DEF_BUF_SIZE, the current starting size).
  • Never start above 1 MiB.
  • Never start above max_length when one is given.
  • If the output turns out larger, the buffer grows by doubling exactly as it does today.

The compress side uses a separate hint:

  • compress(), igzip_lib.compress() and compressobj().compress() start at the input size plus 25% at level 0 or plus 6.25% at levels 1 to 3, plus 32 KiB.
  • Inputs under 16 KiB keep the 16 KiB start, and the hint is capped at 16 MiB.

IgzipDecompressor, Decompress.flush, and the one-shot decompress() functions are unchanged.

Why these numbers

  • 8x comes from the orjson sample payloads. Typical JSON compresses 3 to 8 to 1, so 8 times the input covers most real payloads without growing.
  • The 16 KiB floor matters for small websocket messages. With context takeover, small messages with repeated structure compress 10 to 47 to 1, far beyond 8x. The floor keeps every 4 KiB message growth-free, which is also today's behaviour.
  • The 1 MiB cap bounds the transient over-allocation for incompressible input, which never compresses. Among the orjson samples only canada.json has a guess above 1 MiB, and it pays one extra grow for it.
  • This replaces an earlier version on this branch that preallocated max_length, bounded by the maximum deflate expansion of 1032:1. As the maintainer pointed out in the issue, that sizes every buffer for a decompression bomb: a 20 KiB JSON message with a 4 MiB limit allocated 4 MiB.

Benchmarks

Everything below compares this PR against develop, run back to back on the same machine: macOS 26.6 on arm64 with Python 3.14.5, and Linux aarch64 in Docker with glibc 2.36 and Python 3.14.
In the comparison columns, "faster" and "slower" compare speed: "25% faster" means this PR does the same work 1.25 times as fast. Differences under 2% are shown as "same", and differences of 5% or more are in bold.

Real traffic: Home Assistant websocket frames

Receiving two frames captured from a live Home Assistant instance, the way aiohttp does it: one decompressor per connection, decompress(unconsumed_tail + payload + b"\x00\x00\xff\xff", max_msg_size + 1). The varied streams change timestamps, ids and states in every message. Time per message, median of 5 interleaved runs. Lower is better.

These frames already fit the 16 KiB starting buffer on develop, so no change is expected here. The table confirms the new sizing adds no cost to the common small-message path.

macOS:

frame size develop this PR this PR vs develop
state change 183 B 985 ns 978 ns same
two-event batch 423 B 1339 ns 1320 ns same
state change, identical repeats 182 B 4352 ns 4312 ns same
two-event batch, identical repeats 444 B 4435 ns 4368 ns 2% faster

Linux:

frame size develop this PR this PR vs develop
state change 183 B 932 ns 924 ns same
two-event batch 423 B 1157 ns 1149 ns same
state change, identical repeats 182 B 5742 ns 5798 ns same
two-event batch, identical repeats 444 B 5651 ns 5763 ns 2% slower

orjson payloads received as one websocket message

The four orjson sample files, each received as a single message, compressed either by aiohttp (isal level 0) or by a zlib level 6 peer. The last row sends every JSON object from three of the files as its own message. Time per message, median of 6 interleaved runs. Lower is better.

macOS:

payload sender uncompressed size develop this PR this PR vs develop
github.json aiohttp (isal level 0) 55 KiB 68 µs 66 µs 3% faster
twitter.json aiohttp (isal level 0) 617 KiB 483 µs 476 µs 2% faster
citm_catalog.json aiohttp (isal level 0) 1.65 MiB 398 µs 400 µs same
canada.json aiohttp (isal level 0) 2.15 MiB 4.91 ms 4.93 ms same
github.json zlib level 6 peer 55 KiB 47 µs 45 µs 5% faster
twitter.json zlib level 6 peer 617 KiB 282 µs 271 µs 4% faster
citm_catalog.json zlib level 6 peer 1.65 MiB 214 µs 232 µs 8% slower
canada.json zlib level 6 peer 2.15 MiB 3.54 ms 3.54 ms same
one message per JSON object (average) aiohttp (isal level 0) 3 KiB 2558 ns 2542 ns same
one message per JSON object (average) zlib level 6 peer 3 KiB 2392 ns 2367 ns same

Linux:

payload sender uncompressed size develop this PR this PR vs develop
github.json aiohttp (isal level 0) 55 KiB 47 µs 41 µs 17% faster
twitter.json aiohttp (isal level 0) 617 KiB 325 µs 317 µs 2% faster
citm_catalog.json aiohttp (isal level 0) 1.65 MiB 309 µs 297 µs 4% faster
canada.json aiohttp (isal level 0) 2.15 MiB 4.04 ms 3.98 ms same
github.json zlib level 6 peer 55 KiB 32 µs 32 µs 2% faster
twitter.json zlib level 6 peer 617 KiB 208 µs 194 µs 7% faster
citm_catalog.json zlib level 6 peer 1.65 MiB 209 µs 203 µs 3% faster
canada.json zlib level 6 peer 2.15 MiB 3.01 ms 3.02 ms same
one message per JSON object (average) aiohttp (isal level 0) 3 KiB 2150 ns 2134 ns same
one message per JSON object (average) zlib level 6 peer 3 KiB 2415 ns 2476 ns 2% slower

The one consistent slowdown on macOS, citm_catalog from a zlib level 6 peer, is where each build's doubling lands, not extra work. That message compresses 45 to 1. develop grows 7 times from 16 KiB and copies about 1.95 MiB. This PR starts at about 300 KiB, grows 3 times and copies about 2.16 MiB, because its last doubling overshoots the 1.65 MiB output further. Payloads that land the other way, such as github.json from the same sender, are faster. On macOS every one of those grows moved the buffer, as did the final shrink, in both builds. On Linux the same citm_catalog message is 3% faster with this PR, and no orjson payload is more than 2% slower.

Allocation behaviour

Counted with a realloc interposer, one warm run per case. "grows" is the number of realloc calls to a larger size, the ones that may copy depending on heap layout.

orjson samples received the way aiohttp does for websockets:

macOS:

stream sender messages develop grows this PR grows messages with no grow, this PR
the four whole files, 55 KiB to 2.15 MiB aiohttp (isal level 0) 4 23 4 50%
the four whole files, 55 KiB to 2.15 MiB zlib level 6 peer 4 23 6 25%
one message per JSON object aiohttp (isal level 0) 373 0 0 100%
one message per JSON object zlib level 6 peer 373 0 0 100%
about 4 KiB per message aiohttp (isal level 0) 192 0 0 100%
about 4 KiB per message zlib level 6 peer 192 0 0 100%

Linux:

stream sender messages develop grows this PR grows messages with no grow, this PR
the four whole files, 55 KiB to 2.15 MiB aiohttp (isal level 0) 4 23 4 50%
the four whole files, 55 KiB to 2.15 MiB zlib level 6 peer 4 23 6 25%
one message per JSON object aiohttp (isal level 0) 373 0 0 100%
one message per JSON object zlib level 6 peer 373 0 0 100%
about 4 KiB per message aiohttp (isal level 0) 192 0 0 100%
about 4 KiB per message zlib level 6 peer 192 0 0 100%

benchmark_scripts/benchmark_output_buffer.py cases, --size-mib 64:

macOS:

case develop grows this PR grows
stream-128K-decompressobj 1506 0
stream-128K-igzipdecompressor 0 0
aiohttp-ws-recv-512K-x 1200 800
aiohttp-ws-recv-512K-mixed 1200 0
aiohttp-ws-recv-100B 0 0
aiohttp-ws-send-512K-x 0 0
aiohttp-ws-send-512K-mixed 1000 0
aiohttp-ws-send-100B 0 0
aiohttp-http-body-64M 1004 0
oneshot-64M 13 13
compressobj-128M-incompressible 15 5
compress-128K-chunks 400 0

Linux:

case develop grows this PR grows
stream-128K-decompressobj 1506 0
stream-128K-igzipdecompressor 0 0
aiohttp-ws-recv-512K-x 1200 800
aiohttp-ws-recv-512K-mixed 1200 0
aiohttp-ws-recv-100B 0 0
aiohttp-ws-send-512K-x 0 0
aiohttp-ws-send-512K-mixed 1000 0
aiohttp-ws-send-100B 0 0
aiohttp-http-body-64M 1004 0
oneshot-64M 13 13
compressobj-128M-incompressible 15 5
compress-128K-chunks 400 0

aiohttp-ws-recv-512K-x is aiohttp's own benchmark payload, b"x" * 512 KiB, which compresses 100 to 250 to 1. No input-proportional guess reaches that without over-allocating for everything else, so it still grows 4 times per message instead of 6.

Synthetic cases

python benchmark_scripts/benchmark_output_buffer.py --rounds 7, best of 3 interleaved runs. Total time per case. Lower is better.

macOS:

case develop this PR this PR vs develop
stream-128K-decompressobj 789.78 ms 773.52 ms 2% faster
stream-128K-igzipdecompressor 783.84 ms 783.66 ms same
aiohttp-ws-recv-512K-x 5.76 ms 5.08 ms 13% faster
aiohttp-ws-recv-512K-mixed 454.86 ms 445.57 ms 2% faster
aiohttp-ws-recv-100B 88.78 ms 86.90 ms 2% faster
aiohttp-ws-send-512K-x 4.55 ms 4.50 ms same
aiohttp-ws-send-512K-mixed 215.76 ms 208.79 ms 3% faster
aiohttp-ws-send-100B 9.99 ms 9.94 ms same
aiohttp-http-body-64M 293.52 ms 296.39 ms same
oneshot-256M 819.43 ms 812.30 ms same
compressobj-128M-incompressible 543.35 ms 534.26 ms 2% faster
compress-128K-chunks 76.53 ms 73.69 ms 4% faster

Linux:

case develop this PR this PR vs develop
stream-128K-decompressobj 545.07 ms 540.81 ms same
stream-128K-igzipdecompressor 536.96 ms 544.64 ms same
aiohttp-ws-recv-512K-x 51.22 ms 50.35 ms 2% faster
aiohttp-ws-recv-512K-mixed 258.48 ms 252.58 ms 2% faster
aiohttp-ws-recv-100B 118.36 ms 117.69 ms same
aiohttp-ws-send-512K-x 24.32 ms 24.91 ms 2% slower
aiohttp-ws-send-512K-mixed 183.59 ms 191.46 ms 4% slower
aiohttp-ws-send-100B 14.96 ms 14.73 ms 2% faster
aiohttp-http-body-64M 137.91 ms 135.82 ms 2% faster
oneshot-256M 541.16 ms 537.34 ms same
compressobj-128M-incompressible 647.48 ms 645.73 ms same
compress-128K-chunks 93.97 ms 93.66 ms same

Several threads decompressing at once

Each thread has its own decompressor. Three workloads run through it: 26 MiB of JSON in 128 KiB blocks, 4 KiB websocket messages, and 180 B state change frames. Numbers are MiB of output per second across all threads, best of 3 within a run and median of 3 interleaved runs. Higher is better. This checks that the larger initial buffers do not add contention.

macOS:

workload develop, 1 thread develop, 4 threads this PR, 1 thread this PR, 4 threads 4 threads, this PR vs develop
stream 128 KiB blocks 855 3046 862 3327 9% faster
receive 4 KiB messages 2377 834 2381 836 same
receive 180 B state changes 165 43 167 42 2% slower

Linux:

workload develop, 1 thread develop, 4 threads this PR, 1 thread this PR, 4 threads 4 threads, this PR vs develop
stream 128 KiB blocks 1282 4452 1328 5034 13% faster
receive 4 KiB messages 2418 625 2433 645 3% faster
receive 180 B state changes 174 36 173 36 same

For both builds, four threads receiving small messages are slower than one. Decompress.decompress releases the GIL on every call, and for messages this small the hand-off costs more than the work. That behaviour predates this PR, and #2 addresses it on the compression side.

Instruction counts

Executed instructions counted with callgrind on Linux, including interpreter start-up. These counts do not depend on heap layout or CPU contention. The small-message receive does the same work on both builds, and the block-streaming case, where develop reallocated 1506 times, executes 0.7% fewer instructions.

workload develop, instructions executed this PR, instructions executed difference
receive 5000 state change frames 892,088,557 892,012,349 -0.01%
receive the four orjson files, aiohttp sender 349,208,859 349,031,696 -0.05%
decompress 16 MiB of JSON in 128 KiB blocks 660,419,254 655,676,952 -0.72%

Tests

  • tests/test_compat.py: exact-size chunks for max_length from 1 to 512 KiB + 1, max_length larger than the output, exact fill, repeated empty input, a websocket-like sync-flush stream, and compress roundtrips of zeros and random data at all levels for sizes around 16 KiB, 128 KiB, 1 MiB and 17 MiB.
  • tests/test_igzip_lib.py: an exact-size chunk loop for IgzipDecompressor, as a regression test for the unchanged code.
  • The full suite passes locally on 3.14 on macOS, including the CPython-derived compliance suites.

🤖 Generated with Claude Code

https://claude.ai/code/session_01R5pkuwR7q3BJ3gj5hAwGRt

bdraco and others added 2 commits September 13, 2026 20:45
Decompress.decompress started its output buffer at 16 KiB, clamped that
down to max_length, and then doubled it with _PyBytes_Resize until the
output fit or max_length was reached. Every doubling is a realloc that may
or may not copy the whole buffer depending on heap layout, which makes the
run time depend on what else the process allocated (pycompression#256).

IgzipDecompressor.decompress already allocated min(max_length, 16 MiB) up
front (aa79253). Move that policy into a shared helper in isal_shared.h and
use it from both decompressors. Callers that pass max_length almost always
fill it (fixed-size block reads, websocket message size limits), so this is
one allocation that is shrunk at most once instead of a realloc chain:

- decompress(chunk, 128 KiB): one 128 KiB allocation, no realloc at all.
- decompress(msg, 4 MiB + 1) with a 512 KiB message (aiohttp): one
  allocation and one in-place shrink instead of six doublings.
- no max_length, or sys.maxsize: unchanged, so large one-shot calls keep
  the current growth behaviour.

The concatenated output is byte-identical. Because ISA-L reads ahead into
its internal output buffer each time output space runs out, the split
between the returned bytes and unconsumed_tail for a single capped call can
differ from before.

Add tests that lock in exact-size chunks for a range of max_length values,
and benchmark_scripts/benchmark_output_buffer.py covering the cases from the
issue and the call shapes aiohttp uses.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01R5pkuwR7q3BJ3gj5hAwGRt
Compress.compress, isal_zlib.compress and igzip_lib.compress started
their output buffer at 16 KiB regardless of the input size, so compressing
a 512 KiB message of mostly incompressible data grew the buffer five times
by doubling, and a 128 KiB chunk twice.

Start from a level-aware estimate of the output instead: level 0 uses
static Huffman tables and expands incompressible input by about 23%,
levels 1-3 fall back to stored blocks so the output stays close to the
input size. A 32 KiB slack covers the wrapper and output that ISA-L held
back from an earlier call (about 22 KiB for levels 1-2). Inputs smaller
than 16 KiB keep the 16 KiB buffer, and the estimate is capped at 16 MiB,
above which the buffer grows as before. This is only a hint: the growth
path is unchanged for outputs that exceed it.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01R5pkuwR7q3BJ3gj5hAwGRt
@bdraco
bdraco force-pushed the preallocate-max-length-output-buffer branch from da0c162 to 971f329 Compare September 14, 2026 01:45
bdraco and others added 2 commits September 13, 2026 20:58
Preallocating max_length is wasteful when the input is tiny: a 100 byte
websocket message with aiohttp's default 4 MiB limit allocated 4 MiB and
shrank it to 100 bytes on every message. Deflate expands by at most
1032:1, so cap the preallocation at pending output ISA-L still holds plus
(input + 8 bytes of bit buffer) * 1032 plus 16 KiB of slack. A 100 byte
message now gets a ~132 KiB buffer, a 128 KiB input slice or a 300 KiB
websocket payload still gets the full max_length, and unbounded calls are
unchanged. The bound is an upper bound, so the buffer still never has to
grow on these paths.

Also add a tiny-message websocket send case to the benchmark.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01R5pkuwR7q3BJ3gj5hAwGRt
- Merge initial_output_buffer_size and inflate_output_bound into one
  inflate_initial_buffer_size used by both decompressors. The input bound
  only applies when the input cannot fill the capped size, which also
  removes the overflow guard. Use ISAL_DEF_MAX_MATCH and the bit buffer
  size instead of magic numbers.
- Compute the Decompress.decompress buffer size after taking the lock,
  since it now reads the inflate state.
- Collapse IgzipDecompressor's hard_limit if/else and simplify
  compress_initial_buffer_size.
- Tests: size data by max_length and compress it once, drop cases that
  existing tests already cover. The new tests took about 140 s, mostly
  max_length=1 over 3.5 MB, and now take under 2 s.
- Benchmark: build each case's data only when it runs, so --cases skips
  unrelated setup and one case's data does not stay alive during the
  next; use timeit.repeat.
- Shorten the changelog entries and the isal_shared.h header note.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01R5pkuwR7q3BJ3gj5hAwGRt
The previous approach preallocated max_length, bounded by the maximum
deflate expansion of 1032:1. As pointed out in pycompression#256, that sizes every
buffer for a decompression bomb: a 20 KiB JSON websocket message with a
4 MiB limit allocated 4 MiB.

Start from a guess instead: 8 times the input, at least 16 KiB and at most
1 MiB, never above max_length. On the orjson sample payloads 8 times
covers typical JSON without growing. Small messages often compress far
better because of earlier context, so the 16 KiB floor keeps them
growth-free. Larger or more compressible outputs grow as before. The 1 MiB
cap bounds the transient over-allocation for incompressible input.

IgzipDecompressor is left as upstream has it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01R5pkuwR7q3BJ3gj5hAwGRt
@bdraco bdraco changed the title Preallocate the output buffer when max_length is given Size the initial decompress buffer from the input Sep 14, 2026
@aiolibsbot

aiolibsbot commented Sep 14, 2026 •

Copy link
Copy Markdown

Previous review — superseded by a newer review below.

@aiolibsbot aiolibsbot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Tip

No blocking issues found — ready to merge.

- Restore the original isal_shared.h header rationale for the single
  output buffer instead of claiming it grows in place without copies.
- Rename the decompress sizing constants to DEF_DECOMP_GUESS_RATIO and
  DEF_MAX_DECOMP_GUESS_SIZE and define them next to DEF_BUF_SIZE, so they
  are not mistaken for the exported DECOMP_* flags.
- Fix the MAX_LENGTHS comment, which still mentioned the 4 MiB limit.
- Benchmark: every case now returns the uncompressed bytes it processed
  and is checked against the expected size before timing, and the stream
  and HTTP body cases fail if decompression does not reach the end of the
  stream, so truncated output cannot look like a speed-up.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01R5pkuwR7q3BJ3gj5hAwGRt
@aiolibsbot

aiolibsbot commented Sep 14, 2026 •

Copy link
Copy Markdown

PR Review — Size the initial decompress buffer from the input

Merge-ready: all five findings from the last review are fixed at fb6919e, and this pass found nothing new.

What's done well:

  • decompress_initial_buffer_size (src/isal/isal_shared.h:279-286) cannot overflow. It only multiplies when input_len < 131072, and it always returns at most hard_limit, so assert(length <= max_length) in arrange_output_buffer_with_maximum still holds. With empty input it starts at min(16 KiB, max_length), the same as before.
  • Sizing from the input avoids allocating for a decompression bomb. A payload with a 4 MiB limit now starts at no more than 1 MiB, and only if its input is at least 128 KiB.
  • compress_initial_buffer_size caps input_len before adding to it. Both callers set zst.level before asking for the hint (isal_shared.h:351 then 359, isal_zlibmodule.c:748 then 832). One-shot isal_zlib.compress reaches the hint through igzip_lib_compress_impl (isal_zlibmodule.c:511).
  • Paths the description calls unchanged are unchanged: IgzipDecompressor (igzip_libmodule.c:95-110), Compress.flush (isal_zlibmodule.c:1016) and one-shot decompress.
  • The tests check behaviour at the edges that matter: max_length=1, exact fill, repeated empty input, a sync-flush websocket stream, and level-0 random input past the 16 MiB hint cap.

Prior findings, now fixed:

  • Header comment: no longer changed; the Python 3.9 _BlocksOutputBuffer explanation is intact.
  • Constants: renamed to DEF_DECOMP_GUESS_RATIO / DEF_MAX_DECOMP_GUESS_SIZE and moved next to DEF_BUF_SIZE (isal_shared.h:42-45).
  • MAX_LENGTHS comment in tests/test_compat.py now matches the list.
  • Benchmark: each case runs once and is checked against nbytes before it is timed.
  • Benchmark: the stream and HTTP-body loops now raise if the stream did not reach eof, or if unused_data is left over.

Not verified: Python is not available in this review shell, so I did not build the extension or run the tests or benchmarks. The realloc counts and timings are the author's. I also could not confirm that the uncompressed test.fastq.gz is over 3 MiB, which is the largest DATA[:size] slice the tests use. It is 1.66 MiB compressed, so it almost certainly is.


✅ Resolved since last review (3)

Previously-flagged issues verified fixed
  • src/isal/isal_shared.h:9 Header comment swaps a correct reason for an inaccurate one
  • src/isal/isal_shared.h:276 DECOMP_* prefix is already used for public flag constants
  • tests/test_compat.py:185 Stale comment: MAX_LENGTHS does not include the 4 MiB limit


Checklist

  • Logic correct at edge cases (empty input, max_length=1, max_length=0, overflow)
  • Initial buffer never exceeds the hard limit (no decompression-bomb preallocation)
  • Compress hint uses zst.level after it is set, in every caller
  • Unchanged paths (IgzipDecompressor, flush, one-shot decompress) left intact
  • No hardcoded secrets or unsafe operations
  • New branches covered by behavioural tests
  • Comments and docs accurate (prior findings 1 and 3 fixed)
  • Naming and placement consistent with DEF_* sizing constants (prior finding 2 fixed)
  • Benchmark checks output size and eof before reporting timings (prior findings 4 and 5 fixed)
  • Python 3.9+ compatibility of new test/benchmark code (randbytes, removesuffix)
  • PR description matches diff

Silent Failure Analysis

🟡 **1. MEDIUM** — discarded result / vacuous validation
benchmark_scripts/benchmark_output_buffer.py:131-137

Risk: The benchmark's docstring says truncated output fails rather than looking like a speed-up, but the compress cases (aiohttp-ws-send-*, compressobj-128M-incompressible, compress-128K-chunks) throw away the compressed output and return the input length, so a broken or empty compressed result would pass the size check and look faster.

for message in messages:
    compressor.compress(message) + compressor.flush(isal_zlib.Z_SYNC_FLUSH)
    total += len(message)
return total

Fix: In the untimed check run, decompress what these cases produce (for example with zlib.decompressobj(wbits=-15) for the websocket stream) and compare it to the input, or return the decompressed length so the existing nbytes check means something for compression.


Automated review by Kōan (Claude) HEAD=fb6919e 1 min 53s

@aiolibsbot aiolibsbot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Tip

No blocking issues found — ready to merge.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants