Repository navigation
Memory corruption and crash when streaming to zlib #45268
Description
Activity
- addedzlibIssues and PRs related to the zlib module and its compression dependencies.Issues and PRs related to the zlib module and its compression dependencies.
on Nov 1, 2022 I can reproduce on Node.js main but it seems to be fixed by #44412.
I can no longer reproduce the issue with Node.js 19.1.0. Can you please confirm?
We also landed #45387 (not included in a release yet) and I also cannot reproduce the issue with Node.js main.
I can no longer reproduce the issue with Node.js 19.1.0. Can you please confirm?
We also landed #45387 (not included in a release yet) and I also cannot reproduce the issue with Node.js main.
Scratch that. I can reproduce with both Node.js 19.1.0 and Node.js main on Linux.
This patch seems to fix the issue
diff --git a/deps/zlib/zlib.gyp b/deps/zlib/zlib.gyp index 547143e19d..50c8ae8be4 100644 --- a/deps/zlib/zlib.gyp +++ b/deps/zlib/zlib.gyp @@ -82,7 +82,6 @@ 'defines': [ 'ADLER32_SIMD_SSSE3', 'INFLATE_CHUNK_SIMD_SSE2', - 'CRC32_SIMD_SSE42_PCLMUL', 'DEFLATE_SLIDE_HASH_SSE2' ], 'sources': [
but I' not sure if the issue is Node.js or in the zlib fork we use.
I can not reproduce it in
v20.0.0-preandv16.14.1. Are you still suffering from this issue? @MaddieLowe@ywave620 I can still reproduce with Node.js main.
Then it's weird🤔️, I ran the script with the main version of Node in
Linux VM-79-215-centos 3.10.107-1-tlinux2_kvm_guest-0049 #1 SMP Tue Jul 30 23:46:29 CST 2019 x86_64 x86_64 x86_64 GNU/Linuxmultiple times, but I didn't find any crash.I can reproduce on every run on WSL and sometimes on macOS (13.1). I've also received this from a failed assertion in a recent run
/Users/luigi/code/node/node[1577]: ../src/node_zlib.cc:270:virtual node::(anonymous namespace)::CompressionStream<node::(anonymous namespace)::ZlibContext>::~CompressionStream() [CompressionContext = node::(anonymous namespace)::ZlibContext]: Assertion `(zlib_memory_) == (0)' failed. 1: 0x10c6b9645 node::Abort() [/Users/luigi/code/node/out/Release/node] 2: 0x10c6b9451 node::Assert(node::AssertionInfo const&) [/Users/luigi/code/node/out/Release/node] 3: 0x10c77cdd4 node::(anonymous namespace)::ZlibStream::~ZlibStream() [/Users/luigi/code/node/out/Release/node] 4: 0x10c9e50b5 unsigned long v8::internal::GlobalHandles::InvokeFirstPassWeakCallbacks<v8::internal::GlobalHandles::Node>(std::__1::vector<std::__1::pair<v8::internal::GlobalHandles::Node*, v8::internal::GlobalHandles::PendingPhantomCallback>, std::__1::allocator<std::__1::pair<v8::internal::GlobalHandles::Node*, v8::internal::GlobalHandles::PendingPhantomCallback> > >*) [/Users/luigi/code/node/out/Release/node] 5: 0x10ca49f73 v8::internal::Heap::PerformGarbageCollection(v8::internal::GarbageCollector, v8::internal::GarbageCollectionReason, char const*) [/Users/luigi/code/node/out/Release/node] 6: 0x10ca47239 v8::internal::Heap::CollectGarbage(v8::internal::AllocationSpace, v8::internal::GarbageCollectionReason, v8::GCCallbackFlags) [/Users/luigi/code/node/out/Release/node] 7: 0x10cae6461 v8::internal::ScavengeJob::Task::RunInternal() [/Users/luigi/code/node/out/Release/node] 8: 0x10c723f2e node::PerIsolatePlatformData::RunForegroundTask(std::__1::unique_ptr<v8::Task, std::__1::default_delete<v8::Task> >) [/Users/luigi/code/node/out/Release/node] 9: 0x10c722857 node::PerIsolatePlatformData::FlushForegroundTasksInternal() [/Users/luigi/code/node/out/Release/node] 10: 0x10d22c79b uv__async_io [/Users/luigi/code/node/out/Release/node] 11: 0x10d2408e8 uv__io_poll [/Users/luigi/code/node/out/Release/node] 12: 0x10d22cc78 uv_run [/Users/luigi/code/node/out/Release/node] 13: 0x10c5ed713 node::SpinEventLoopInternal(node::Environment*) [/Users/luigi/code/node/out/Release/node] 14: 0x10c6fcde6 node::NodeMainInstance::Run() [/Users/luigi/code/node/out/Release/node] 15: 0x10c680f64 node::LoadSnapshotDataAndRun(node::SnapshotData const**, node::InitializationResultImpl const*) [/Users/luigi/code/node/out/Release/node] 16: 0x10c6811cd node::Start(int, char**) [/Users/luigi/code/node/out/Release/node] 17: 0x7ff801a73310 start [/usr/lib/dyld] Abort trap: 6I'm not sure if it is related.
After a stress test in
Darwin MacBook-Pro.local 21.6.0 Darwin Kernel Version 21.6.0: Wed Aug 10 14:25:27 PDT 2022; root:xnu-8020.141.5~2/RELEASE_X86_64 x86_64, I got a different assertion error... sending 759 bytes, buf size is 204747767 parse_chunk, length 759, total bytes received 1943006 sending 4381 bytes, buf size is 204747008 parse_chunk, length 4381, total bytes received 1947387 sending 3557 bytes, buf size is 204742627 parse_chunk, length 3557, total bytes received 1950944 sending 4974 bytes, buf size is 204739070 parse_chunk, length 4974, total bytes received 1955918 /Users/roger/Work/node-contribute/out/Debug/node[81933]: ../src/node_zlib.cc:752:void node::(anonymous namespace)::ZlibContext::Close(): Assertion `status == 0 || status == (-3)' failed. 1: 0x1015e411d node::DumpBacktrace(__sFILE*) [/Users/roger/Work/node-contribute/out/Debug/node] 2: 0x10176223b node::Abort() [/Users/roger/Work/node-contribute/out/Debug/node] 3: 0x101761d95 node::Assert(node::AssertionInfo const&) [/Users/roger/Work/node-contribute/out/Debug/node] 4: 0x101989385 node::(anonymous namespace)::ZlibContext::Close() [/Users/roger/Work/node-contribute/out/Debug/node] 5: 0x10198917e node::(anonymous namespace)::CompressionStream<node::(anonymous namespace)::ZlibContext>::Close() [/Users/roger/Work/node-contribute/out/Debug/node] 6: 0x101985141 node::(anonymous namespace)::CompressionStream<node::(anonymous namespace)::ZlibContext>::Close(v8::FunctionCallbackInfo<v8::Value> const&) [/Users/roger/Work/node-contribute/out/Debug/node] 7: 0x101d04e12 v8::internal::FunctionCallbackArguments::Call(v8::internal::CallHandlerInfo) [/Users/roger/Work/node-contribute/out/Debug/node] 8: 0x101d03b3f v8::internal::MaybeHandle<v8::internal::Object> v8::internal::(anonymous namespace)::HandleApiCallHelper<false>(v8::internal::Isolate*, v8::internal::Handle<v8::internal::HeapObject>, v8::internal::Handle<v8::internal::FunctionTemplateInfo>, v8::internal::Handle<v8::internal::Object>, unsigned long*, int) [/Users/roger/Work/node-contribute/out/Debug/node] 9: 0x101d027db v8::internal::Builtin_HandleApiCall(int, unsigned long*, v8::internal::Isolate*) [/Users/roger/Work/node-contribute/out/Debug/node] 10: 0x102d562f9 Builtins_CEntry_Return1_DontSaveFPRegs_ArgvOnStack_BuiltinExit [/Users/roger/Work/node-contribute/out/Debug/node]Given that the original crash message provided by MaddieLowe, I suspect that bad memory allocation/deallocation is to blame. Since node performs (de)compression in bg thread, I wonder are
mallocandfreethread safe? Folks in https://stackoverflow.com/questions/855763/is-malloc-thread-safe confirm that the one in glibc is thread safe, how about the one in mac/windows?Given that the original crash message provided by MaddieLowe, I suspect that bad memory allocation/deallocation is to blame. Since node performs (de)compression in bg thread, I wonder are
mallocandfreethread safe? Folks in https://stackoverflow.com/questions/855763/is-malloc-thread-safe confirm that the one in glibc is thread safe, how about the one in mac/windows?Yes, they are thread safe in mac(See https://opensource.apple.com/source/libmalloc/libmalloc-53.1.1/src/malloc.c.auto.html)
@lpinca reported removing
CRC32_SIMD_SSE42_PCLMULfixed it for him and that sounds plausible.Just a hunch, but the xmm0-xmm15 registers are caller-saved. The SIMD version of zlib's crc32 definitely clobbers them. Maybe something somewhere is not saving the xmm registers when it should?
An easy way to test is to run under qemu twice, once with CPU = Intel Nehalem, then CPU = Intel Westmere.
If the first one works okay and the second one crashes, then that's a pretty good indicator; Westmere was when PCLMULQDQ was introduced.
@MaddieLowe Could you share some thought on why you write this project in the first place? what problem you want to conquer? how you come up with the
anonymized-data, because I can not identify any straightforward purpose by reading the code. It may be helpful for debugging :)@ywave620 from the issue description
The application this is currently affecting is used for streaming medical imaging data to a viewer.
14 remaining items
I got a few questions that I hope could better insulate what the problem is, please check below:
a) The report of reproducing on Mac, is it a M1 (Apple Silicon) macintosh or an Intel based?
b) Is there a standalone repro case (i.e. relies only on zlib proper) that I could have access?
The macro CRC32_SIMD_SSE42_PCLMUL guards the x86 vectorized implementation of CRC-32 in Chromium zlib.
I would be quite surprised if that is the culprit for the observed misbehavior, since we have shipped it since 2018.
There is one macro that guards an optimization that can be impacted by misbehaved client code though i.e. code that relies on undefined behavior present in canonical zlib that has changed thanks to this specific optimization.
The macro is INFLATE_CHUNK_SIMD_SSE2, and I even mentioned it in the source code (i.e. https://github.com/nodejs/node/blob/main/deps/zlib/contrib/optimizations/inflate.c#L1256).
What happens if you disable the above macro?
a) The report of reproducing on Mac, is it a M1 (Apple Silicon) macintosh or an Intel based?
Intel based Mac (3,2 GHz 8-Core Intel Xeon W). I have just tried again compiling the main branch, and it crashed after 2/3 runs
node(54825,0x7ff8568e6700) malloc: Incorrect checksum for freed object 0x7f881d00c200: probably modified after being freed. Corrupt value: 0x0 node(54825,0x7ff8568e6700) malloc: *** set a breakpoint in malloc_error_break to debug Abort trap: 6b) Is there a standalone repro case (i.e. relies only on zlib proper) that I could have access?
I don't have one.
The macro is INFLATE_CHUNK_SIMD_SSE2, and I even mentioned it in the source code (i.e. https://github.com/nodejs/node/blob/main/deps/zlib/contrib/optimizations/inflate.c#L1256).
What happens if you disable the above macro?
I'm currently compiling Node.js on WSL where the issue is easier to reproduce. I will let you know.
I have to sleep, I'll update tomorrow.
With #49511 reverted I can still reproduce the issue after applying this patch
diff --git a/deps/zlib/zlib.gyp b/deps/zlib/zlib.gyp index 49de2a6c6d..70a05e40db 100644 --- a/deps/zlib/zlib.gyp +++ b/deps/zlib/zlib.gyp @@ -134,43 +134,6 @@ '<!@pymod_do_main(GN-scraper "<(ZLIB_ROOT)/BUILD.gn" "\\"zlib_crc32_simd\\".*?sources = ")', ], }, # zlib_crc32_simd - { - 'target_name': 'zlib_inflate_chunk_simd', - 'type': 'static_library', - 'conditions': [ - ['target_arch in "ia32 x64" and OS!="ios"', { - 'defines': [ 'INFLATE_CHUNK_SIMD_SSE2' ], - 'conditions': [ - ['target_arch=="x64"', { - 'defines': [ 'INFLATE_CHUNK_READ_64LE' ], - }], - ], - }], - ['arm_fpu=="neon"', { - 'defines': [ 'INFLATE_CHUNK_SIMD_NEON' ], - 'conditions': [ - ['target_arch=="arm64"', { - 'defines': [ 'INFLATE_CHUNK_READ_64LE' ], - }], - ], - }], - ], - 'include_dirs': [ '<(ZLIB_ROOT)' ], - 'direct_dependent_settings': { - 'conditions': [ - ['target_arch in "ia32 x64" and OS!="ios"', { - 'defines': [ 'INFLATE_CHUNK_SIMD_SSE2' ], - }], - ['arm_fpu=="neon"', { - 'defines': [ 'INFLATE_CHUNK_SIMD_NEON' ], - }], - ], - 'include_dirs': [ '<(ZLIB_ROOT)' ], - }, - 'sources': [ - '<!@pymod_do_main(GN-scraper "<(ZLIB_ROOT)/BUILD.gn" "\\"zlib_inflate_chunk_simd\\".*?sources = ")', - ], - }, # zlib_inflate_chunk_simd { 'target_name': 'zlib', 'type': 'static_library', @@ -199,8 +162,11 @@ }], # Incorporate optimizations where possible. ['(target_arch in "ia32 x64" and OS!="ios") or arm_fpu=="neon"', { - 'dependencies': [ 'zlib_inflate_chunk_simd' ], - 'sources': [ '<(ZLIB_ROOT)/slide_hash_simd.h' ] + # 'dependencies': [ 'zlib_inflate_chunk_simd' ], + 'sources': [ + '<(ZLIB_ROOT)/inflate.c', + '<(ZLIB_ROOT)/slide_hash_simd.h', + ] }, { 'defines': [ 'CPU_NO_SIMD' ], 'sources': [ '<(ZLIB_ROOT)/inflate.c' ],
Also, if the issue was with
INFLATE_CHUNK_SIMD_SSE2it would still be reproducible in #49511, no?, but it is not.The macro CRC32_SIMD_SSE42_PCLMUL guards the x86 vectorized implementation of CRC-32 in Chromium zlib. I would be quite surprised if that is the culprit for the observed misbehavior, since we have shipped it since 2018.
It's quite possible chromium's zlib is blameless. It's been a while since I last looked at this issue but I remember seeing hints of xmm registers not being saved by callers. I have some jotted-down notes where I was apparently suspicious of xmm7 getting clobbered across calls.
Reacted by Luigi Pinca and Adenilson Cavalcanti- added a commit that references this issue
on Sep 10, 2023 @bnoordhuis It seems that disabling CRC32_PCLMUL will be a regression of around -20% in decompression speed and -15% in compression speed.
Could you (or someone from the node.js team) elaborate further on how the calls are performed between node.sj <--> zlib?
The relevant code is in src/node_zlib.cc. We're not doing anything out of the ordinary, our C++ code just call zlib's public API.
I entertained the idea that we may be miscompiling node or zlib somehow (missing or wrong compiler flags) but nothing Obviously Wrong stood out to me at the time.
I think there is something wrong in our benchmarks or I ran the job incorrectly. See #49511 (comment).
@lpinca I measured using zlib_bench (https://source.chromium.org/chromium/chromium/src/+/main:third_party/zlib/contrib/bench/zlib_bench.cc), where the data is all in memory and it calculates raw compression/decompression speeds of zlib.
If values are smaller than what I've reported, that probably means that there are other bottlenecks probably in moving data between node.js <--> zlib, such that data integrity checks (i.e. CRC32) is not the limiting factor.
@Adenilson yes I understand. What I find surprising is that our benchmarks show improvements without the CRC32 SIMD optimization. They just compare before and after.
@lpinca what would be distribution of lengths of data vectors passed to zlib (i.e. the 'len' in crc32_z, https://source.chromium.org/chromium/chromium/src/+/main:third_party/zlib/crc32.c;l=702) during the benchmarks?
If lengths are short (i.e. smaller than the threshold to enter vectorized code), we fallback to the portable code.
This is really hot code and by disabling CRC32_SIMD_SSE42_PCLMUL quite a few branches are removed from the the hot path.
- added a commit that references this issue
on Sep 28, 2023
Version
v19.0.0
Platform
Linux 5.4.0-131-generic #147-Ubuntu SMP Fri Oct 14 17:07:22 UTC 2022 x86_64 x86_64 x86_64 GNU/Linux
Subsystem
No response
What steps will reproduce the bug?
The bug happens when I send certain data in random-sized chunks over a socket and then pipe it to gzip. I tried to reproduce it with a smaller data sample, but unfortunately it seems that very specific data is needed to reproduce this bug. The script and data to reproduce the issue are here: https://github.com/morpheus-med/node-defect. Repro steps are:
unzip anonymized-data.zipnode index.jsHow often does it reproduce? Is there a required condition?
It should crash in 1-2 runs. It crashes every time for me.
What is the expected behavior?
The expected behaviour is for the program to read and pipe the data to gzip without crashing.
What do you see instead?
The program crashes with:
double free or corruption (!prev)Aborted (core dumped)or
node: malloc.c:4036: _int_malloc: Assertion (unsigned long) (size) >= (unsigned long) (nb)' failed.Aborted (core dumped)or
corrupted size vs. prev_sizeAborted (core dumped)Additional information
The application this is currently affecting is used for streaming medical imaging data to a viewer. We've currently got around the problem by downgrading node, but we would like to be able to upgrade node in the future to keep up with security updates.
The problem isn't reproducible with node 12.16.3, but is reproducible with versions 12.17.0 and later. I've done a git bisect and it seems to have been introduced with this commit: 9e33f97
Thanks for your help