Repository navigation
_PyPegen_is_memoized() has a complexity of O(n) #93289
Description
Activity
- addedtype-bugAn unexpected behavior, bug, or errorAn unexpected behavior, bug, or errorperformancePerformance or resource usagePerformance or resource usage
on May 27, 2022 @pablogsal @lysnikolaou @isidentical: Would it be possible to maintain a hashtable to make this lookup function faster?
If you consider that it's not worth it, just you can close my issue ;-)
The benchmark comes from: #93106
@pablogsal @lysnikolaou @isidentical: Would it be possible to maintain a hashtable to make this lookup function faster?
If you consider that it's not worth it, just you can close my issue ;-)
We already tried to use a hash table, a btree, a static array, a dynamic array and a skip list and all of them were slower. If you manage to get a hash table to perform faster, I'm happy to merge it :)
I tried yet again to use the hash table from
pycore_hashtable.hand is a performance disaster. I took also some statistics and the linked lists average number of elements is between 1 and 10 values so a hash table is an overkill, either per token or one in the parser. Also I checked and the average percentage of times the function returns true (so a value is found) is around 75%, which means that adding a bloom filter or similar in front is also going to slow down the thing quite a lot.I'm closing the issue, but if you manage to get a hash table or other structure running and is faster, please feel free to reopen.
@isidentical @lysnikolaou If you want to give it a go, please feel free to try :)
I tried yet again to use the hash table from pycore_hashtable.h and is a performance disaster.
That's what I would try. I trust you that it's worse ;-) Ok, good that you already knew about it!
Using Linux perf, I noticed that the Python parser spends a significant time in the _PyPegen_is_memoized() function which iterates on a linked list to find a value:
script.py:
Linux perf says that overall, Python spent 7% of its runtime in this function: