Skip to content

Handle a non-importable __main__ in multiprocessing #64145

Description

@ogrisel
mannequin
BPO 19946
Nosy @brettcannon, @ncoghlan, @pitrou, @larryhastings, @tiran, @ericsnowcurrently, @zware, @ogrisel
Files
  • issue19946.diff
  • issue19946_pep_451_multiprocessing_v2.diff: Work in progress patch (Fork+, ForkServer-, Spawn--)
  • test_multiprocessing_main_handling.py: Final test case
  • issue19946.diff
  • skip_forkserver.patch
  • Note: these values reflect the state of the issue at the time it was migrated and might not reflect the current state.

    Show more details

    GitHub fields:

    assignee = 'https://github.com/ncoghlan'
    closed_at = <Date 2013-12-20.14:33:20.194>
    created_at = <Date 2013-12-10.13:08:59.356>
    labels = ['type-bug', 'library', 'release-blocker']
    title = 'Handle a non-importable __main__ in multiprocessing'
    updated_at = <Date 2015-02-24.09:59:58.484>
    user = 'https://github.com/ogrisel'

    bugs.python.org fields:

    activity = <Date 2015-02-24.09:59:58.484>
    actor = 'ncoghlan'
    assignee = 'ncoghlan'
    closed = True
    closed_date = <Date 2013-12-20.14:33:20.194>
    closer = 'ncoghlan'
    components = ['Library (Lib)']
    creation = <Date 2013-12-10.13:08:59.356>
    creator = 'Olivier.Grisel'
    dependencies = []
    files = ['33091', '33162', '33175', '33201', '33222']
    hgrepos = []
    issue_num = 19946
    keywords = ['patch']
    message_count = 47.0
    messages = ['205810', '205831', '205832', '205833', '205847', '205854', '205857', '205885', '205893', '205894', '205895', '205898', '205904', '206115', '206118', '206120', '206126', '206129', '206133', '206134', '206137', '206138', '206140', '206180', '206224', '206228', '206229', '206230', '206252', '206260', '206296', '206297', '206298', '206301', '206421', '206424', '206478', '206559', '206575', '206607', '206650', '206653', '206655', '206659', '206677', '206678', '206683']
    nosy_count = 10.0
    nosy_names = ['brett.cannon', 'ncoghlan', 'pitrou', 'larry', 'christian.heimes', 'python-dev', 'sbt', 'eric.snow', 'zach.ware', 'Olivier.Grisel']
    pr_nums = []
    priority = 'release blocker'
    resolution = 'fixed'
    stage = 'resolved'
    status = 'closed'
    superseder = None
    type = 'behavior'
    url = 'https://bugs.python.org/issue19946'
    versions = ['Python 3.4']

    Activity

    1. ogrisel commented on Dec 10, 2013

      ogriselmannequin
      MannequinAuthor

      Here is a simple python program that uses the new forkserver feature introduced in 3.4b1:

      name: checkforkserver.py
      """
      import multiprocessing
      import os

      def do(i):
          print(i, os.getpid())
      
      
      def test_forkserver():
          mp = multiprocessing.get_context('forkserver')
          mp.Pool(2).map(do, range(3))
      
      
      if __name__ == "__main__":
          test_forkserver()
      """

      When running this using the "python check_forkserver.py" command everything works as expected.

      When running this using the nosetests launcher ("nosetests -s check_forkserver.py"), I get:

      """
      Traceback (most recent call last):
        File "<string>", line 1, in <module>
        File "/opt/Python-HEAD/lib/python3.4/multiprocessing/forkserver.py", line 141, in main
          spawn.import_main_path(main_path)
        File "/opt/Python-HEAD/lib/python3.4/multiprocessing/spawn.py", line 252, in import_main_path
          methods.init_module_attrs(main_module)
        File "<frozen importlib._bootstrap>", line 1051, in init_module_attrs
      AttributeError: 'NoneType' object has no attribute 'loader'
      """

      Indeed, the spec variable in multiprocessing/spawn.py's import_main_path
      function is None as the nosetests script is not a regular python module: in particular is does not have a ".py" extension.

      If I copy or symlink or renamed the "nosetests" script as "nosetests.py" in the same folder, this works as expected. I am not familiar enough with the importlib machinery to suggest a fix for this bug.

      Also there is a typo in the comment: "causing a psuedo fork bomb" => "causing a pseudo fork bomb".

      Note: I am running CPython head updated today.

    2. added
      stdlibStandard Library Python modules in the Lib/ directory
      type-crashA hard crash of the interpreter, possibly with a core dump
      on Dec 10, 2013
    3. pitrou commented on Dec 10, 2013

      @pitrou
      Member

      This sounds related to the ModuleSpec changes.

    4. brettcannon commented on Dec 10, 2013

      @brettcannon
      Member

      The returning of None means that importlib.find_spec() didn't find the spec/loader for the specified module. So the question is exactly what module is being passed to importlib.find_spec() and why isn't it finding a spec/loader for that module.

      Did this code work in Python 3.3?

    5. brettcannon commented on Dec 10, 2013

      @brettcannon
      Member

      I see Oliver says he is testing a new forkserver feature from 3.4b1, so it might not necessarily be importlib's fault then. Does using the old importlib.find_loader() approach work?

    6. ogrisel commented on Dec 10, 2013

      ogriselmannequin
      MannequinAuthor

      So the question is exactly what module is being passed to importlib.find_spec() and why isn't it finding a spec/loader for that module.

      The module is the nosetests python script. module_name == 'nosetests' in this case. However, nosetests is not considered an importable module because of the missing '.py' extension in the filename.

      Did this code work in Python 3.3?

      This code did not exist in Python 3.3.

    7. brettcannon commented on Dec 10, 2013

      @brettcannon
      Member

      So at the bare minimum, the multiprocessing code should raise an ImportError when it can't find the spec for the module to help debug this kind of thing. Also that typo should get fixed.

      Second, there is no way that 'nosetests' will ever succeed as an import since, as Oliver pointed out, it doesn't end in '.py' or any other identifiable way for a finder to know it can handle the file. So this is not a bug and simply a side-effect of how import works. The only way around it would be to symlink nosetests to nosetests.py or to somehow pre-emptively set up 'nosetests' for supported importing.

    8. self-assigned this
      on Dec 10, 2013
    9. changed the title [-]multiprocessing crash with forkserver or spawn when run from a non ".py" ending script[/-] [+]Have multiprocessing raise ImportError when spawning a process that can't find the "main" module[/+] on Dec 10, 2013
    10. added
      type-bugAn unexpected behavior, bug, or error
      and removed
      type-crashA hard crash of the interpreter, possibly with a core dump
      on Dec 10, 2013
    11. sbt commented on Dec 10, 2013

      sbtmannequin
      Mannequin

      I guess this is a case where we should not be trying to import the main module. The code for determining the path of the main module (if any) is rather crufty.

      What is sys.modules['__main__'] and sys.modules['__main__'].__file__ if you run under nose?

    12. ericsnowcurrently commented on Dec 11, 2013

      @ericsnowcurrently
      Member

      There is definitely room for improvement relative to module specs and __main__ (that's the topic of issue bpo-19701). That issue is waiting for __main__ to get a proper spec (see issues bpo-19700 and bpo-19697).

    13. ogrisel commented on Dec 11, 2013

      ogriselmannequin
      MannequinAuthor

      I agree that a failure to lookup the module should raise an explicit exception.

      Second, there is no way that 'nosetests' will ever succeed as an import since, as Oliver pointed out, it doesn't end in '.py' or any other identifiable way for a finder to know it can handle the file. So this is not a bug and simply a side-effect of how import works. The only way around it would be to symlink nosetests to nosetests.py or to somehow pre-emptively set up 'nosetests' for supported importing.

      I don't agree that (unix) Python programs that don't end with ".py" should be modified to have multiprocessing work correctly. I think it should be the multiprocessing responsibility to transparently find out how to spawn the new process independently of the fact that the program ends in '.py' or not.

      Note: the fork mode works always under unix (with or without the ".py" extension). The spawn mode always work under windows as AFAIK there is no way to have Python programs that don't end in .py under windows and furthermore I think multiprocessing does execute the __main__ under windows (but I haven't tested if it's still the case in Python HEAD).

    14. 33 remaining items

    15. sbt commented on Dec 17, 2013

      sbtmannequin
      Mannequin

      Thanks for your hard work Nick!

    16. tiran commented on Dec 18, 2013

      @tiran
      Member

      The commit broken a couple of buildbots like all Windows bots and OpenIndiana.

    17. reopened this on Dec 18, 2013
    18. zware commented on Dec 19, 2013

      @zware
      Member

      The problem on Windows at least is that the skips for the 'fork' and 'forkserver' start methods aren't firing due to setUpClass being improperly set up in MultiProcessingCmdLineMixin: it's not decorated as a classmethod and the 'u' is lower-case instead of upper. Just fixing that makes for some unusual output ("skipped '"fork" start method not available'" with no indication of which test was skipped) and a variable number of tests depending on available start methods, so a better fix is to just do the check in setUp.

      Unrelated to the failure, but we're also in the process of moving away from using test_main(), preferring unittest.main().

      The attached patch addresses both, passes on Windows and Linux, and I suspect should help on OpenIndiana as well judging by the tracebacks it's giving.

    19. python-dev commented on Dec 19, 2013

      python-devmannequin
      Mannequin

      New changeset 460961e80e31 by Nick Coghlan in branch 'default':
      Issue bpo-19946: appropriately skip new multiprocessing tests
      http://hg.python.org/cpython/rev/460961e80e31

    20. tiran commented on Dec 19, 2013

      @tiran
      Member

      The OpenIndiana tests are still failing. OpenIndiana doesn't support forkserver because it doesn't implement the send handle feature. The patch skips the forkserver tests if HAVE_SEND_HANDLE is false.

    21. ncoghlan commented on Dec 19, 2013

      @ncoghlan
      Contributor

      I think that needs to be fixed on the multiprocessing side rather than just
      in the tests - we shouldn't create a concrete context for a start method
      that isn't going to work on that platform. Finding that kind of discrepancy
      was part of my rationale for basing the skips on the available contexts
      (although my main motivation was simplicity).

      There may also be docs implications in describing which methods are
      supported on different platforms (although I haven't looked at how that is
      currently documented).

    22. sbt commented on Dec 20, 2013

      sbtmannequin
      Mannequin

      On 19/12/2013 10:00 pm, Nick Coghlan wrote:

      I think that needs to be fixed on the multiprocessing side rather than just
      in the tests - we shouldn't create a concrete context for a start method
      that isn't going to work on that platform. Finding that kind of discrepancy
      was part of my rationale for basing the skips on the available contexts
      (although my main motivation was simplicity).

      There may also be docs implications in describing which methods are
      supported on different platforms (although I haven't looked at how that is
      currently documented).

      If by "concrete context" you mean _concrete_contexts['forkserver'], then
      that is supposed to be private. If you write

           ctx = multiprocessing.get_context('forkserver')

      then this will raise ValueError if the forkserver method is not
      available. You can also use

      'forkserver' in multiprocessing.get_all_start_methods()
      

      to check if it is available.

    23. ncoghlan commented on Dec 20, 2013

      @ncoghlan
      Contributor

      Ah, I should have looked more closely at the docs to see if there was a public API for that before poking around in the package internals.

      In that case, I suggest we change this bit in the test:

          # We look inside the context module to find out which
          # start methods we can check
          from multiprocessing.context import _concrete_contexts

      to use the appropriate public API:

          # Need to know which start methods we should test
          import multiprocessing
          AVAILABLE_START_METHODS = set(multiprocessing.get_all_start_methods())

      And then adjust the skip check to look in AVAILABLE_START_METHODS rather than _concrete_contexts.

      I'll make that change tonight if nobody beats me to it.

    24. python-dev commented on Dec 20, 2013

      python-devmannequin
      Mannequin

      New changeset 00d09afb57ca by Nick Coghlan in branch 'default':
      Issue bpo-19946: use public API for multiprocessing start methods
      http://hg.python.org/cpython/rev/00d09afb57ca

    25. ncoghlan commented on Dec 20, 2013

      @ncoghlan
      Contributor

      Pending a clean bill of health from the stable buildbots :)

    26. ncoghlan commented on Dec 20, 2013

      @ncoghlan
      Contributor

      Now passing on all the stable buildbots (the two red Windows bots are for other issues, such as bpo-15599 for the threaded import test failure)

    27. transferred this issue fromon Apr 10, 2022
    Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

    Metadata

    Metadata

    Assignees

    Labels

    release-blockerstdlibStandard Library Python modules in the Lib/ directorytype-bugAn unexpected behavior, bug, or error

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions