Skip to content

Offer suggestions on AttributeError and NameError #82711

Description

@pablogsal
BPO 38530
Nosy @vstinner, @aroberge, @serhiy-storchaka, @1st1, @pablogsal, @tirkarthi, @isidentical, @sweeneyde
PRs
  • bpo-38530: Offer suggestions on AttributeError #16850
  • bpo-38530: Offer suggestions on AttributeError #16856
  • bpo-38530: Offer suggestions on NameError #25397
  • bpo-38530: Clean exceptions if dir() fails when making suggestions for AttributeError #25408
  • bpo-38530: Optimize the calculation of string sizes when offering suggestions #25412
  • bpo-38530: Match exactly AttributeError and NameError when offering suggestions #25443
  • bpo-38530: Properly extend UnboundLocalError from NameError #25444
  • bpo-38530: Include builtins in NameError suggestions #25460
  • bpo-38530: Cover more error paths in error suggestion functions #25462
  • bpo-38530: Surround suggestions by quotes #25473
  • bpo-38530: Require 50% similarity in NameError suggestions #25584
  • bpo-38530: Refactor AttributeError suggestions #25776
  • Note: these values reflect the state of the issue at the time it was migrated and might not reflect the current state.

    Show more details

    GitHub fields:

    assignee = None
    closed_at = <Date 2021-04-14.23:32:49.056>
    created_at = <Date 2019-10-19.18:57:31.623>
    labels = ['interpreter-core', 'type-feature', '3.9']
    title = 'Offer suggestions on AttributeError and NameError'
    updated_at = <Date 2021-05-03.15:47:35.739>
    user = 'https://github.com/pablogsal'

    bugs.python.org fields:

    activity = <Date 2021-05-03.15:47:35.739>
    actor = 'pablogsal'
    assignee = 'none'
    closed = True
    closed_date = <Date 2021-04-14.23:32:49.056>
    closer = 'pablogsal'
    components = ['Interpreter Core']
    creation = <Date 2019-10-19.18:57:31.623>
    creator = 'pablogsal'
    dependencies = []
    files = []
    hgrepos = []
    issue_num = 38530
    keywords = ['patch']
    message_count = 40.0
    messages = ['354954', '354955', '354956', '354957', '354958', '354959', '354960', '354964', '354965', '354967', '354969', '354970', '354971', '354972', '354973', '354975', '354979', '354981', '355499', '355503', '355527', '355530', '355531', '355533', '359977', '391026', '391080', '391090', '391108', '391123', '391219', '391224', '391312', '391316', '391834', '392007', '392513', '392543', '392581', '392817']
    nosy_count = 9.0
    nosy_names = ['vstinner', 'aroberge', 'serhiy.storchaka', 'yselivanov', 'james', 'pablogsal', 'xtreak', 'BTaskaya', 'Dennis Sweeney']
    pr_nums = ['16850', '16856', '25397', '25408', '25412', '25443', '25444', '25460', '25462', '25473', '25584', '25776']
    priority = 'normal'
    resolution = 'fixed'
    stage = 'resolved'
    status = 'closed'
    superseder = None
    type = 'enhancement'
    url = 'https://bugs.python.org/issue38530'
    versions = ['Python 3.9']

    Activity

    1. pablogsal commented on Oct 19, 2019

      @pablogsal
      MemberAuthor

      To improve the debugging experience in both interactive and non-interactive code, I propose to offer suggestions when attribute access fails. For example:

      >>> class A: foo = None
      ... 
      >>> A.fou
      Traceback (most recent call last):
        File "<stdin>", line 1, in <module>
      AttributeError: type object 'A' has no attribute 'fou'

      Did you mean: foo?

      This also applies to imports from modules and other situations:

      >>> import collections
      >>> collections.NamedTuple
      Traceback (most recent call last):
        File "<stdin>", line 1, in <module>
      AttributeError: module 'collections' has no attribute 'NamedTuple'

      Did you mean: namedtuple?

    2. pablogsal commented on Oct 19, 2019

      @pablogsal
      MemberAuthor

      PR 16850 shows an initial prototype for the idea

    3. isidentical commented on Oct 19, 2019

      @isidentical
      SponsorMember

      It already exists as a 3rd party module and it would be really cool to have this in core level.

      https://github.com/dutc/didyoumean (by James Powell)

    4. pablogsal commented on Oct 19, 2019

      @pablogsal
      MemberAuthor

      I am not super convinced that this is a great idea because it has some performance cost (although somewhat controlled) but I want to open a discussion.

    5. tirkarthi commented on Oct 19, 2019

      @tirkarthi
      Member

      Ruby has it integrated into the core : https://bugs.ruby-lang.org/issues/11252 . It was initially a gem that got merged into core.

      methosd
      undefined local variable or method methosd' for main:Object Did you mean? methods method (repl):1:in

      '

    6. pablogsal commented on Oct 19, 2019

      @pablogsal
      MemberAuthor

      Idea: we could only do this in interactive mode if we consider that is expensive enough.

    7. pablogsal commented on Oct 19, 2019

      @pablogsal
      MemberAuthor

      I am running pyperformance to check the performance cost of this.

    8. pablogsal commented on Oct 19, 2019

      @pablogsal
      MemberAuthor

      Slower (27):

      • pathlib: 25.2 ms +- 0.4 ms -> 105 ms +- 2 ms: 4.18x slower (+318%)
      • sympy_str: 315 ms +- 3 ms -> 500 ms +- 3 ms: 1.59x slower (+59%)
      • sympy_sum: 203 ms +- 2 ms -> 286 ms +- 2 ms: 1.41x slower (+41%)
      • sqlalchemy_imperative: 36.5 ms +- 0.9 ms -> 41.8 ms +- 0.8 ms: 1.14x slower (+14%)
      • python_startup: 12.4 ms +- 0.1 ms -> 13.9 ms +- 0.0 ms: 1.12x slower (+12%)
      • python_startup_no_site: 9.19 ms +- 0.04 ms -> 10.3 ms +- 0.0 ms: 1.12x slower (+12%)
      • sympy_integrate: 24.9 ms +- 0.1 ms -> 27.6 ms +- 0.3 ms: 1.11x slower (+11%)
      • scimark_sparse_mat_mult: 5.21 ms +- 0.05 ms -> 5.49 ms +- 0.06 ms: 1.05x slower (+5%)
      • unpickle_list: 5.78 us +- 0.07 us -> 6.08 us +- 0.07 us: 1.05x slower (+5%)
      • 2to3: 392 ms +- 2 ms -> 411 ms +- 10 ms: 1.05x slower (+5%)
      • nbody: 155 ms +- 2 ms -> 160 ms +- 2 ms: 1.03x slower (+3%)
      • scimark_fft: 423 ms +- 2 ms -> 432 ms +- 3 ms: 1.02x slower (+2%)
      • float: 142 ms +- 2 ms -> 145 ms +- 2 ms: 1.02x slower (+2%)
      • unpack_sequence: 68.6 ns +- 2.5 ns -> 69.9 ns +- 2.1 ns: 1.02x slower (+2%)
      • pickle_list: 5.65 us +- 0.05 us -> 5.76 us +- 0.07 us: 1.02x slower (+2%)
      • fannkuch: 569 ms +- 3 ms -> 579 ms +- 6 ms: 1.02x slower (+2%)
      • xml_etree_parse: 196 ms +- 2 ms -> 199 ms +- 2 ms: 1.02x slower (+2%)
      • sqlalchemy_declarative: 210 ms +- 2 ms -> 213 ms +- 3 ms: 1.01x slower (+1%)
      • unpickle: 16.9 us +- 0.1 us -> 17.2 us +- 0.7 us: 1.01x slower (+1%)
      • regex_effbot: 3.93 ms +- 0.12 ms -> 3.97 ms +- 0.04 ms: 1.01x slower (+1%)
      • scimark_monte_carlo: 128 ms +- 2 ms -> 129 ms +- 2 ms: 1.01x slower (+1%)
      • scimark_lu: 186 ms +- 5 ms -> 188 ms +- 6 ms: 1.01x slower (+1%)
      • regex_dna: 272 ms +- 1 ms -> 274 ms +- 1 ms: 1.01x slower (+1%)
      • xml_etree_iterparse: 130 ms +- 2 ms -> 131 ms +- 3 ms: 1.01x slower (+1%)
      • genshi_xml: 74.3 ms +- 0.8 ms -> 74.8 ms +- 0.8 ms: 1.01x slower (+1%)
      • regex_v8: 29.7 ms +- 0.3 ms -> 29.9 ms +- 0.2 ms: 1.01x slower (+1%)
      • mako: 19.3 ms +- 0.3 ms -> 19.4 ms +- 0.2 ms: 1.01x slower (+1%)

      The current approach is too expensive, so I'm closing PR 16850.

    9. serhiy-storchaka commented on Oct 19, 2019

      @serhiy-storchaka
      Member

      AFAIK there is existing issue for this idea.

      I have doubts about performance. I added _PyObject_LookupAttr in particularly to avoid an overhead of raising and silencing an AttributeError. I believe most performance sensitive code in the core now uses it and will not be affected, but there is other code which silences it (_PyObject_LookupAttr itself silences an AttributeError raised in called functions), third-party code which uses PyObject_HasAttr or PyObject_GetAttr can be affected.

      It might be simpler to implement it in sys.excepthook or sys.displayhook, but at that point we do not have attribute name and a reference to the object. There is an issues and a PEP about adding references to AttributeError. It could help to implement this feature.

    10. serhiy-storchaka commented on Oct 19, 2019

      @serhiy-storchaka
      Member

      I am surprised that it was SO expensive.

      Pathlib would largely benefit from cached_property if it be compatible with slots.

    11. pablogsal commented on Oct 19, 2019

      @pablogsal
      MemberAuthor

      Serhiy, do you think we could attach the object and the name to some private fields of the AttributeError and check that in sys.excepthook if they are present?

    12. pablogsal commented on Oct 19, 2019

      @pablogsal
      MemberAuthor

      I will also repeat the pyperformance results locally just in case something was off on the speed.python.org server.

    13. 17 remaining items

    14. pablogsal commented on Apr 14, 2021

      @pablogsal
      MemberAuthor

      New changeset 3fc65b9 by Pablo Galindo in branch 'master':
      bpo-38530: Optimize the calculation of string sizes when offering suggestions (GH-25412)
      3fc65b9

    15. vstinner commented on Apr 15, 2021

      @vstinner
      Member
    16. pablogsal commented on Apr 16, 2021

      @pablogsal
      MemberAuthor

      New changeset 3b82cae by Pablo Galindo in branch 'master':
      bpo-38530: Properly extend UnboundLocalError from NameError (GH-25444)
      3b82cae

    17. pablogsal commented on Apr 16, 2021

      @pablogsal
      MemberAuthor

      New changeset 0ad81d4 by Pablo Galindo in branch 'master':
      bpo-38530: Match exactly AttributeError and NameError when offering suggestions (GH-25443)
      0ad81d4

    18. pablogsal commented on Apr 17, 2021

      @pablogsal
      MemberAuthor

      New changeset 3ab4bea by Pablo Galindo in branch 'master':
      bpo-38530: Include builtins in NameError suggestions (GH-25460)
      3ab4bea

    19. pablogsal commented on Apr 17, 2021

      @pablogsal
      MemberAuthor

      New changeset 0b1c169 by Pablo Galindo in branch 'master':
      bpo-38530: Cover more error paths in error suggestion functions (GH-25462)
      0b1c169

    20. sweeneyde commented on Apr 25, 2021

      @sweeneyde
      Member

      I opened PR 25584 to fix this current behavior:

      >>> v
      Traceback (most recent call last):
        File "<stdin>", line 1, in <module>
      NameError: name 'v' is not defined. Did you mean: 'id'?
      >>> vv
      Traceback (most recent call last):
        File "<stdin>", line 1, in <module>
      NameError: name 'vv' is not defined. Did you mean: 'id'?
      >>> vvv
      Traceback (most recent call last):
        File "<stdin>", line 1, in <module>
      NameError: name 'vvv' is not defined. Did you mean: 'abs'?
    21. pablogsal commented on Apr 27, 2021

      @pablogsal
      MemberAuthor

      New changeset 284c52d by Dennis Sweeney in branch 'master':
      bpo-38530: Require 50% similarity in NameError and AttributeError suggestions (GH-25584)
      284c52d

    22. sweeneyde commented on Apr 30, 2021

      @sweeneyde
      Member

      Some research of other projects:

      LLVM [1][2]
      -----------

      • Compute Levenshtein
        • Using O(n) memory rather than O(n^2)
      • Uses UpperBound = (len(typo) + 2) // 3

      GCC [3]
      -------

      • Uses Damerau-Levenshtein distance
        • Counts transpositions like "abcd" <-> "bacd" as one move
      • Swapping Case as in "a" <-> "A" counts as half a move
      • cutoff = (longer + 2) // 3 if longer - shorter >= 2 else max(longer // 3, 1)

      Rust [4]
      --------

      • "maximum allowable edit distance defaults to one-third of the given word."
      • First checks for exact case-insensitive match, then check for Levenshtein distance small enough, then check if sorted(a.split("_")) == sorted(b.split("_"))

      Ruby [5]
      --------

      • Quickly filter out words with bad Jaro–Winkler distance
        • threshold = input.length > 3 ? 0.834 : 0.77
      • Only compute Levenshtein for words that remain
        • threshold = (input.length * 0.25).ceil
        • Output all good enough words
      • If no word was good enough then output the closest match.

      I think there are some good ideas here.

      [1] https://github.com/llvm/llvm-project/blob/d480f968ad8b56d3ee4a6b6df5532d485b0ad01e/llvm/include/llvm/ADT/edit_distance.h#L42
      [2] https://github.com/llvm/llvm-project/blob/e2b3b89bf1ce74bf889897e0353a3e3fa93e4452/clang/lib/Sema/SemaLookup.cpp#L4263
      [3] https://github.com/gcc-mirror/gcc/blob/16e2427f50c208dfe07d07f18009969502c25dc8/gcc/spellcheck.c
      [4] https://github.com/rust-lang/rust/blob/673d0db5e393e9c64897005b470bfeb6d5aec61b/compiler/rustc_span/src/lev_distance.rs#L44
      [5] https://github.com/ruby/ruby/blob/48b94b791997881929c739c64f95ac30f3fd0bb9/lib/did_you_mean/spell_checker.rb

    23. pablogsal commented on Apr 30, 2021

      @pablogsal
      MemberAuthor

      Hi Dennis, this is a fantastic investigation!

      I think I really like GCC approach here. We may want to invest into porting some of their ideas into our solution.

    24. sweeneyde commented on May 1, 2021

      @sweeneyde
      Member

      PR 25776 is a work in progress for what it might look like to do a few things:

      • Make case-swaps half the cost of any other edit
      • Refactor Levenshtein code to not use memory allocator, and to bail early on no match.
      • Add comments to Levenshtein distance code
      • Add test cases for Levenshtein distance behind a debug macro
      • Set threshold to (name_size + item_size + 3) * MOVE_COST / 6.
        • Reasoning: similar to difflib.SequenceMatcher.ratio() >= 2/3:
          "Multiset Jaccard similarity" >= 2/3
          matching letters / total letters >= 2/3
          (name_size - distance + item_size - distance) / (name_size + item_size) >= 2/3
          1 - (2*distance) / (name_size + item_size) >= 2/3
          1/3 >= (2*distance) / (name_size + item_size)
          (name_size + item_size) / 6 >= distance
          With rounding:
          (name_size + item_size + 3) // 6 >= distance

      Re: Damerau-Levenshtein (transpositions as single edits), if that were to get implemented, I don't see a way to do that without using a buffer of at least 3x the size, storing the most recent 3 rows of the matrix.

    25. pablogsal commented on May 3, 2021

      @pablogsal
      MemberAuthor

      New changeset 80a2a4e by Dennis Sweeney in branch 'master':
      bpo-38530: Refactor and improve AttributeError suggestions (GH-25776)
      80a2a4e

    26. transferred this issue fromon Apr 10, 2022
    Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

    Metadata

    Metadata

    Assignees

    No one assigned

      Labels

      3.9 (EOL)end of lifeinterpreter-core(Objects, Python, Grammar, and Parser dirs)type-featureA feature request or enhancement

      Projects

      No projects

        Milestone

        No milestone

        Relationships

        None yet

        Development

        No branches or pull requests

        Issue actions