Repository navigation
Conversation
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Thanks for contributing!
Before submitting your pull request please have a look at the
following checklist:
pytest)ruff)What this PR does
Two micro-optimizations in
TokenListthat together yield a measurablethroughput improvement on parse-heavy workloads.
1.
token_next_by: fast-path for single-criterion lookupstoken_next_byis one of the hottest functions in sqlparse — profilingshows 69,800 calls for a single complex
format(reindent=True)run.The current implementation always creates a lambda closure and routes
through
_token_matching→imt(), which checks all three branches(
i,m,t) on every token. However, ~95% of call sites pass exactlyone criterion.
This PR adds an inline fast-path that, when only one of
t,m, oriis provided (and it's not a list), scans
self.tokensdirectly — skippinglambda creation, tuple wrapping, and multi-branch
imt()dispatch.When multiple criteria are passed (rare), the original
_token_matchingfallback is used unchanged.
2.
token_index: uselist.index(token, start)instead of sliceThe original
start + self.tokens[start:].index(token)creates atemporary list slice on every call. Passing
startdirectly tolist.index()is semantically identical but avoids the copy.Benchmark results
Python 3.14, 4 realistic SQL queries, 300 iterations × 2 ops (parse +
format), 5 runs, median:
parse()onlyparse()+format(reindent=True)All 506 tests pass,
ruff checkclean.Optimization by ARCHON EVO engine — autonomous performance analysis.