fix(usage): correct Claude cost accounting - #9019
Conversation
85de09c to
cc39ac0
Compare
There was a problem hiding this comment.
Four findings, all in the new usage UI: two shared-primitive recreations (date inputs, micro icon action), one new disclosure control that is mouse-only, and one same-PR duplication of a helper/formatter. Details inline.
Posted via Macroscope — UI Consistency
cc39ac0 to
e8abc6a
Compare
e8abc6a to
d57d8d4
Compare
d57d8d4 to
4c0bd05
Compare
The usage window was locked to four presets, so a spike on the chart could not be inspected without eyeballing dates. Dragging across any daily chart now commits the selection as the date window, double-click returns to the preset, and date fields beside the presets accept any custom range directly. The server already accepted arbitrary day windows; this is web-only. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Cap custom windows before enumeration, keep date controls reachable in compact layouts, and complete brush gestures synchronously through pointer capture.
d25635c to
615c119
Compare
615c119 to
703b4fd
Compare
703b4fd to
2913a86
Compare
|
Macroscope skipped reviewing this pull request. Per-review cost limit exceeded (workspace setting). This review would cost an estimated $10.45, which exceeds your per-review limit of $10.00. The top 3 files driving up this estimate:
Tip To get this pull request reviewed, you can:
|
3a2c112 to
5d53442
Compare
# Conflicts: # apps/server/src/usage/UsageService.test.ts # apps/server/src/usage/UsageService.ts
There was a problem hiding this comment.
Cursor Bugbot has reviewed your changes using default effort and found 1 potential issue.
❌ Bugbot Autofix is OFF. To automatically fix reported issues with cloud agents, enable autofix in the Cursor dashboard.
Reviewed by Cursor Bugbot for commit a31345a. Configure here.
Bugbot is paused — on-demand spend limit reachedBugbot uses usage-based billing for this team and has hit its on-demand spend limit. A team admin can raise the spend limit in the Cursor dashboard, or wait for the next billing cycle to continue. |

What Changed
Claude usage accounting now treats progressive events as snapshots, attributes fallback usage and provider-reported cost to the serving iteration once, and preserves five-minute, one-hour, and unclassified cache creation separately. Thinking remains a subset of output, and long-context and partial-TTL pricing flow through summaries and thread rows without inflating token totals.
Why
Claude transcripts can repeat a request as it progresses and can include several fallback attempts before one serves the response. Counting each event or assigning all top-level cost to every attempt overstates spend; discarding partial TTL details makes the component estimate inaccurate even when the total happens to match.
Stack and Review Scope
This PR is cumulative and lands after #9018. Review the Claude-accounting increment; #9136 and #9308 include it.
Verification
Current candidate
9226b2c32905dc29c07328e6e463e1aff0589baeincludes upstream main39802c06117fae0b3da43624b0d54309c5437c72. Focused UsageService and chart-gesture suites passed (38 tests), and the server package typecheck passed. Repository test doubles now implement the imported-transcript interface. Earlier evidence below retains its original scope and revision; integrated web evidence and its attribution limits are described below.A direct Claude Opus 5 high review attempt for this integration batch exited 1 because OAuth expired; no Claude review occurred.
Candidate:
fe497f655c8aa8b3a0dcfb9a025a3d99b564d1e4, including upstream main82c2b7ffb4d450572f5baf9a70693f1e25cfdfed.ba873b8passed 138 tests across 7 files, plus the affected package typechecks.82c2b7fdelta removes an unused pricing helper: 11 pricing tests, server typecheck, Knip andgit diff --checkpassed on the pre-repair main composition. The chart/runtime contribution was unchanged by that final delta.Focused commands
Refresh correctness checks
The refresh repair passed the focused UsageService suite (17 tests), 3 chart-gesture regression tests, and 8 web usage-state tests. After the final thread-cache guard, 3 affected service tests were rerun and passed; the earlier full-suite count is not presented as a new full-suite run. The optional thread-refresh contract passed 2 schema tests and the contracts package typecheck. Server and web package typechecks, changed-file formatting/lint and diff checks passed. Independent source review found no actionable issue in this frozen repair.
These repairs cancel a chart gesture if its date window changes, let each environment progress after its own pricing request, and coalesce concurrent tokened summary scans. Thread-enabled branches await their environment summary before refreshing thread rows. The source-scan revision guard prevents older aggregation from pruning files discovered by a newer scan. Integrated web evidence and remaining client limits are described below.
UI Changes
The fixture reconciles the visible provider total ($57.30) with the model breakdown and displays estimated cache-write cost. This is a representative synthetic client sanity pass, complementing the per-head Claude parser/accounting tests; it is not an external billing reconciliation.
Before: upstream main
39802c06117fae0b3da43624b0d54309c5437c72. After: cumulative #9308 integration7e248026469ec7be6d5c2bdaccf86f380e0c2476, captured in an isolated web preview with two synthetic projects, three conversations and generated transcripts. This reuses the retained feature in the integrated stack; it is not a separate capture of this PR’s individual head. Focused tests above were run on the individual published head. The 30-day preset ends at the last complete day on the integration; main includes the current day. Both captures contain the same fourteen days of synthetic usage.Clean breakdown recording · Annotated breakdown recording. This shows project → thread → expand → collapse; idle time is trimmed from 109.7 s to 12.57 s, without a performance claim.
Limits: Real billing, every progressive-message variant, and platform-specific transcript paths are covered by the stated focused tests or remain outside this web capture. That recording is web-only; the separate Android proof below covers this increment’s retained cache-write estimate and wording. No iOS or SwiftUI capture is claimed. Clean/annotated packet validation passed; uploaded derivatives were retrieved and hash-verified.
React Native Android integration proof
A disposable Pixel 8 / Android API 36 emulator ran the development APK from clean composition
52ba1db0923b537fcb1ecdb5450a83aa2b4db86a: #93087e248026469ec7be6d5c2bdaccf86f380e0c2476, #10055637051a5e08ee346537eade291f2cc422adf1d4b, and #8727e1338a57e64c30b554b3df1635d292c636f367f0. The additional Usage-screen change only selects the initial tab from widget navigation parameters; accounting, charts, coverage, and refresh logic retain the #9308 implementation. This is cumulative Android proof, not an individual-head or iOS recording.The app paired to an isolated synthetic environment with three threads. Usage showed $62.00 for 30 days, $32.73 for 7 days, then $62.00 on return. Tokens showed 25.3M across three sessions; returning to Cost restored $62.00. Coverage displayed data through September 4. The APK build, signature check, mobile typecheck, and 35 focused Usage/widget/navigation tests passed.
Clean Android window/metric recording · Annotated Android window/metric recording. The matched 20.8-second viewing cuts preserve four real taps; idle intervals and later unlogged scrolling are omitted. The sparse raw Android recording was normalized to 30 fps before editing; no timing or performance claim. The floating gear is the development-client control. Packet validation passed and published derivatives were retrieved and hash-verified.
The retained React Native cache-write metric renders $20.03 for 5.34M tokens, with “not an expiry measure”; the chart states that costs are local public-list estimates. This directly exercises the RN wording/metric increment, while progressive-event and TTL edge cases remain covered by the focused accounting tests.
Checklist
Implemented and reviewed with GPT-6 and GPT-5.6 Sol in the Codex harness. A direct Claude Opus 5 high availability attempt exited 1 because OAuth had expired; no Opus review occurred.
Note
Fix Claude cost accounting with multi-record parsing and TTL-aware cache pricing
readThreadBreakdownin UsageService.ts, theserverGetUsageThreadBreakdownRPC in rpc.ts, and rendering in UsageThreadTable.tsxcwdor TTL fields are discarded bydecodeScanCache. Existing rate tables lacking one-hour cache or long-context fields fall back to base rates, andsameRatenow treats differing one-hour or long-context pricing as different ratesMacroscope summarized 9226b2c.