Repository navigation
multiprocessing forkserver does not flush output before fork (was: preloading '__main__' with forkserver has been broken for a long time) #98552
Description
Activity
- addedtype-bugAn unexpected behavior, bug, or errorAn unexpected behavior, bug, or error
on Oct 22, 2022 - added3.11only security fixesonly security fixes3.12only security fixesonly security fixes
on Oct 23, 2022 A PR fixing this and adding unittest coverage would be great.
BTW as you are using forkserver, be aware of the recently disclosed #97514 if you ever run on 3.9 or later. (that issue can be worked around)
Thanks for the pointer on the CVE.
The simple fix for this is pretty straightforward and I'll try to get it together quickly (first cpython PR so may take a little longer). It is just to deal with the renamed dict entry.
However, I've discovered some other "quirks" with forkserver, and was wondering if they are known or being worked on, if they could use improvement, or are expected and it's just my newness in looking at it. (I'm on linux but trying to use multiprocessing in programs that also use threading, and have gotten burned a few times by deadlocks because of it, leading me to move to using forkserver. But maybe I should be considering another way forward?)
Anyways, with the break fixed by using the right dict key to get
main_pathinto main, preloading__main__works, but imports in the main module can fail, becausesys.pathis not set in forkserver's main before importing main. (The correctsys_pathis actually fed intoforkserver.main, but not used?! It's been that way since forkserver was added 9 years ago...)Not having
sys.pathset leads to other weirdness even when not importing main. The modules specified to preload will succeed/fail depending on the cwd of the process, sopython a/foo.pybehaves different thencd a; python foo.py. And I'm pretty sure I've seen the same module imported twice, once at the name given as preload and once as the name used by another module (e.g.,a.my_moduleandmy_module).The
try/catcharound each preloaded module hides all the failures too. In my case, I'd much prefer to get the exception, because if the preload fails, then every process created by the forkserver is going to have to do expensive module loading itself, and I don't realize it until I notice the performance problem.When
spawnis used, we set/clear_inheritingand callpreparewith the result ofget_preparation_datainspawn._main. My completely uninformed inclination is to do the same inforkserver.main, instead of what's currently happening. Is that a reasonable thing to look into?I did some experiments. Using
spawn.prepareinforkserver.pyto consume the dict fromget_prepare_dataseems to work well. It fixessys.pathto be correct so imports work regardless of cwd, fixes preloading__main__in a less brittle way, and does a bunch of other stuff that I'm guessing should be happening too. But I don't know enough about forkserver and spawn to be certain.I'll put together a PR for this approach too, to solicit input.
forkserver.maingets its parameters from the string created inforkserver.ensure_running, so the dict fromget_prepare_datahas to be turned into a string. A quick way to do this and not have to worry about escaping quotes/etc is to remove theauthkeyfrom the dict, thenjson.dumpstoutf-8tob64encodeit. Thenmaindoes the reverse to get it back. (Thought aboutpicklebut don't know enough to know the security implications here.)The downside (for me) of this more general solution is that it alters code on the new process side, versus the not as complete fix that is on the original process side. This means the latter can be done with monkey-patching, but the former cannot, as far as I understand.
Updating the tests for these changes led to some comical confusion. Turns out there was another bug waiting... forkserver does not flush its stdout/stderr before each fork, and that makes things really confusing.
I was using print statements in files to make sure preload was working and they were only imported in the forkserver and not in the new forked processes. But the output said otherwise, even though everything was actually working...
- addedstdlibStandard Library Python modules in the Lib/ directoryStandard Library Python modules in the Lib/ directory
on Nov 28, 2023 I think this one is fixed now: #99515 (comment).
I think this one is fixed now: #99515 (comment).
Ah, apologies @aggieNick02 for not finding this issue and your fix when I looked at #126631. You even discovered and fixed the same stdout buffering issue that so confused me!
preloading main got fixed as noted above, while this one was overlooked as having had a fix for it already in the shuffle. thanks for tying things together. does the output buffering issue still remain?
Reacted by Sam James- changed the title
[-]preloading '__main__' with forkserver has been broken for a long time[/-][+]multiprocessing forkserver does not flush output before fork (was: preloading '__main__' with forkserver has been broken for a long time)[/+]on Nov 22, 2025 3 remaining items
the remaining purpose i had for keeping this open wound up a dupe of #135335 which was already done elsewhere in the code that I hadn't noticed. reverting those. closing this out. a new issue #141860 tracks the continuation of your feature to be able to see the errors from preload module imports in the forkserver process.
- added a commit that references this issue
on Nov 23, 2025 preloading main got fixed as noted above, while this one was overlooked as having had a fix for it already in the shuffle. thanks for tying things together. does the output buffering issue still remain?
As you noted, this was fixed under #135335. There is still an open PR, #138686, which simplifies the unit test added for #126631. With the buffering bug fixed we can print from
__main__, which makes things much easier. It would be lovely if that could be merged, assuming it looks good ofc 🙂Reacted by Gregory P. Smith
Bug report
The
forkserverstart method provides the ability to callset_forkserver_preloadon the multiprocessing context to load modules into and configure the forkserver process. By doing this carefully, you can avoid having to do module loading and other work each time the forkserver process is forked to create a new process. Without doing such work,forkservercan be way slower than the traditionalforkstart methodYou can specify the module
'__main__'in theset_forkserver_preloadlist, and the forkserver source has special code when you do this. It ensures that the main file path does not have to be configured/loaded after each fork. To do this, inmultiprocessing.forkserver.ensure_running, it callsmultiprocessing.spawn.get_preparation_dataand then uses themain_pathentry that may be in the returned dict.Unfortunately, 3 months after it was introduced, this functionality was broken in commit 9a76735. That commit renamed the
main_pathdictionary entry returned inget_preparation_datatoinit_main_from_path, but didn't update the use inmultiprocessing.fork_serverNot having the ability to load and configure main on the forkserver ends up being unusually painful for my recent scenario, which led to tracking this down. I have a python program on a share that spawns short-lived processes at a high rate. Then multiple machines run this program from the share. Huge slowdown ensues as smbd processes on the server go crazy responding to every new process on every client reading the file and stat-ing the directory the file is contained in.
A simple fix in
multiprocessing.forkserveraccounting for the changed name rectifies the problem. I'll work on putting that PR together. Any thoughts on a workaround that doesn't require modifying the python source are welcome, as I imagine it will be a while until I'm on a python with the fix.Here's a simple repro.
Your environment
Ubuntu 20.04.4 LTS, CPython 3.8
Python source code examination indicates this bug is still present in the current version of CPython.
multiprocessing.set_forkserver_preload#141859