Repository navigation
[Windows] test_wmi: test_wmi_query_error test is flaky #130727
Description
Activity
- addedtestsTests in the Lib/test dirTests in the Lib/test dirtype-bugAn unexpected behavior, bug, or errorAn unexpected behavior, bug, or error
on Mar 1, 2025 Yeah, unfortunately we're stuck between the OS being unreliable on timings here, and people complaining about waiting for the OS to be ready. It used to have a long enough timeout to be reliable on all but the slowest machines.
Possibly what we need is to do some warmup calls to WMI before actually testing it, so we can be pretty sure that it's been loaded. It's worth trying, at least.
- changed the title
[-]`test_wmi_query_error` test is flaky[/-][+][Windows] test_wmi: `test_wmi_query_error` test is flaky[/+]on Mar 4, 2025 - added a commit that references this issue
on Mar 4, 2025 I can reproduce some errors on Windows by stressing my virtual machine using the command
python -m test -j12 -r.I wrote #130832 to make the test more reliable. Using my PR, I cannot reproduce these errors anymore even if the system load is high (34.17). The change also makes the test faster since it removes the LOOPBACK_TIMEOUT sleep (10 seconds!) before a retry.
With my PR, I still get errors when the system load is high:
Invalid descriptor(error -2147024890).ERROR: test_wmi_query_repeated (test.test_wmi.WmiTests.test_wmi_query_repeated) ---------------------------------------------------------------------- Traceback (most recent call last): File "C:\victor\python\main\Lib\test\test_wmi.py", line 42, in test_wmi_query_repeated self.test_wmi_query_os_version() ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~^^ File "C:\victor\python\main\Lib\test\test_wmi.py", line 30, in test_wmi_query_os_version r = wmi_exec_query("SELECT Version FROM Win32_OperatingSystem").split("\0") ~~~~~~~~~~~~~~^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ File "C:\victor\python\main\Lib\test\test_wmi.py", line 18, in wmi_exec_query return _wmi.exec_query(query) ~~~~~~~~~~~~~~~^^^^^^^ OSError: [WinError -2147024890] Descripteur non valideMaybe we should retry on this error as well.
Another random error:
FAIL: test_wmi_query_repeated (test.test_wmi.WmiTests.test_wmi_query_repeated) ---------------------------------------------------------------------- Traceback (most recent call last): File "C:\victor\python\main\Lib\test\test_wmi.py", line 42, in test_wmi_query_repeated self.test_wmi_query_os_version() ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~^^ File "C:\victor\python\main\Lib\test\test_wmi.py", line 31, in test_wmi_query_os_version self.assertEqual(1, len(r)) ~~~~~~~~~~~~~~~~^^^^^^^^^^^ AssertionError: 1 != 2- added a commit that references this issue
on Mar 4, 2025 The initial issue should be fixed by the change f67ff9e. Remaining issues seem more complex to fix and might require further work.
I think the invalid handle error (0x80070006, which is GetLastError()==6) is going to be a race condition with the timeout, probably on one of the events. Deciding which thread is responsible to close the events is tricky, though in theory we should be able close them immediately after setting them as the threads have already been awoken (though haven't yet run any more code).
In any case, the place where we use it (in
platform) handles allOSErrorand falls back onto other code, so it isn't going to affect users apart from those who use the internal API.I can only assume the second one is a partially-written buffer? It's splitting on a null separator and checking the total length of the result, but I bet if it pre-filtered to only non-empty entries then it'd be fine. Again, the real code will only return from valid entries, so users won't notice this in practice (unless it's truncated actual data, but then, it'll be a one-off in their logs, assuming they're following the
platformmodule instructions).Reacted by Victor StinnerIn any case, the place where we use it (in platform) handles all OSError and falls back onto other code
I don't understand this. If the fallback is good enough and WMI is flaky, why don't we always use the fallback code?
The fallbacks can be inaccurate (due to OS compatibility shims or incremental builds) or even slower (launches a separate process).
WMI is the only correct way, and it's only flaky due to multithreading so that we don't have to "poison" the main thread with COM. Though I'm ~90% sure we could probably get away with declaring the main CPython thread as always being multi-threaded under COM, that's not enough for me to want to break users who rely on it being in a single-threaded apartment (who I'm only assuming exist... they may not?). The Store install (for now) is always a multi-threaded apartment and I haven't heard any issues relating to that, hence 90%, but if we were going to change it on purpose then I'd want to do it at startup, and not rely on whether you'd invoked a particular function or not.
Also, I'm not sure we'd get any better than a 5 second timeout running the API directly. And that was too long for some users, which is another reason it's on its own thread.
- added a commit that references this issue
on May 20, 2025
Bug report
Seen in https://github.com/python/cpython/actions/runs/13606431591/job/38038447105?pr=130724 on both the default and free threading builds:
In these cases the tests succeeded when retried.
Linked PRs