Make the collector's cost track the live set, not the heap (issue #5537) - #5585
Conversation
A deep game-tree search on an iPad kept dying after #5540, #5563 and #5573. Each of those fixed a real defect -- pages never returned to the OS, a pacing cap measured against the device's RAM rather than the process budget, a mark worklist that overflowed by sheer allocation volume -- and none of them touched the reason the collector could not keep up in the first place. WHAT THE PROFILE SAYS. Three quarters of the GC thread's wall time is inside cn1ConservativeResolve. gcMarkObject calls it on EVERY reference field the drain follows, to reject a conservatively derived pointer before dereferencing it, and it answered by binary-searching two snapshots: the BiBOP page bases and the legacy extents. On the reporter's shape that is 13 dependent cache-missing loads to find the page and 15 more to miss it and find the array, per field. Marking therefore cost O(log heap) per reference: the collector got slower as the heap grew, which is exactly the reporter's "GC pauses become more and more frequent and take longer, until they are effectively continuous". Everything else followed from that. A cycle stretched to four or five times the collection interval, so the mutator produced four or five times a trigger's worth of garbage during each cycle the collector managed to finish, and the process settled at whatever the pacing allowed: 447MB against a live set of a few hundred bytes, riding 64MB below the ceiling that kills it. On device that is the kill. In the simulator, where there is no ceiling, it is the footprint climbing to gigabytes that the reporter saw next. Both indices are now open-addressed hash tables. The page table keys on the 64KB page base and stores the geometry inline, so a hit is one cache line; its keys change only when a page is registered (the registry is grow-only), so it is rebuilt on that event and only its geometry is refreshed per cycle -- which also retires the per-registration qsort. The legacy side keeps its sorted extent array for interior pointers, which only the conservative stack scan produces, and puts an exact-base table in front of it: a Java reference is always an object base, so the caller that dominates is answered in one probe. TWO THINGS THE FASTER COLLECTOR EXPOSED, both fixed here because both undo it. The survivor-heavy bypass read a pure-churn workload as survivor-heavy. Survival is measured at sweep as slots carrying the current epoch, and the grace pass MARKS every fresh non-leaf object with it -- so what the policy read as a live set was really the allocation rate. It was under the threshold before only because the slow collector inflated the denominator. With the collector keeping up it crossed, diverted 1.8M small objects onto the legacy heap, and brought the worklist overflow back (2-3 cycles in 500, from none). Pages now count the marks a grace pass put on them and the sweep subtracts them, so survival means what the policy needs it to mean. Off a per-process ceiling the pacing cap was a fraction of the HOST's free RAM, which is a reason to let a fast thread run further ahead of the collector and not a reason to accumulate an unbounded amount of garbage. On a roomy machine it evaluated to gigabytes, and once the collector lost the race early nothing brought it back: 13.8-15.7GB of footprint against a 4MB live set, and slower for it (12.2-13.4s against 8.1-8.6s bounded -- a process thrashing fifteen gigabytes pays for them). The cap is now bounded by a multiple of the collection TRIGGER, which already tracks the heap: a survivor-heavy render keeps 8 of its own enlarged triggers, pure churn is held to 8 of the base one. The bound is GATED ON FOOTPRINT, engaging only once the process is already past 512MB, because the point is to stop unbounded growth and not to stop a thread from running ahead. #5573 measured a volume cap costing 2-4x and rejected it; an ungated one measured here at 47% on the objectAllocation microbenchmark (31.6ms -> 43.6ms), for a process that was never going to grow. Gated, that benchmark is 31.1ms -- unchanged -- and the runaway is still bounded, because a runaway is by definition on the wrong side of the gate. That whole shape depends on how much RAM the host happened to have free, which is why it reproduced on an idle machine and vanished on a busy one. CN1_SIMULATE_FREE_MEMORY pins that reading so the guard means the same thing either way. MEASURED on the reporter's shape (GcOverflowSpiralApp, same host, 14.9GB allocated either way, RESULT bit-identical): under a 512MB simulated ceiling before after collections completed 130 583 triggers allocated per collection 4.67 1.04 peak footprint 447MB 116-219MB headroom left below the ceiling 64MB 277-395MB mutator parks 49-54 0 wall time 6.7-6.8s 6.6-7.0s with no ceiling and 32GB of host RAM (the simulator), eight concurrent copies so the collector has to fight for the machine, which is what tips it: peak footprint 13.8-15.7GB 819-861MB wall time 12.2-13.4s 8.1-8.6s GATES. vm/tests: 519 tests in the default group and 8 in the benchmark group, all green, including GcHeapIntegrityIntegrationTest (the CN1_GC_VERIFY use-after-free gate) and LargeArrayGcIntegrationTest (issue 5425). The benchmark gauntlet is GREEN in both cooperative and forced-signal stop modes, every torture bit-identical to the host JVM. run-benchmark.sh geomean unchanged. cn1_globals.m compiles clean to an arm64-apple-ios object against the iOS SDK, and in the CN1_GC_VERIFY / CN1_BIBOP_VALIDATE / CN1_GRACE_AUDIT / CN1_BIBOP_NO_FASTSWEEP / CN1_DISABLE_BIBOP / CN1_RESOLVE_DIAG configurations. GcOverflowSpiralIntegrationTest gains the property underneath all of it -- triggers allocated per completed collection, which is a ratio of two speeds and so reads the same on a loaded machine where a peak does not -- and a second run of the same binary with no ceiling, which is the half of the report that was previously out of scope. That second one is a bound rather than a reproduction, and says so: the off-ceiling runaway is bistable and took eight concurrent copies of the workload on a twelve-core host to provoke, which is not something a unit test should be creating. NOT ADDRESSED. UNDER a ceiling and under deliberate collector starvation (eight concurrent copies of this workload), the process still rides to the ceiling-minus-margin that footprint admission allows. Adding a volume brake to that path as well bounds it to 345MB with 165MB of headroom instead of 38MB, but costs 2.4x -- which is the trade #5573 rejected, and it is a different path from the off-ceiling growth bound added here. The per-cycle qsort of the extent array is now the largest remaining item in the collector at roughly a third of its time, and is the next thing worth replacing. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 5fa3243c95
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
✅ Continuous Quality ReportTest & Coverage
Static Analysis
Generated automatically by the PR CI workflow. |
Cloudflare Preview
|
|
Compared 149 screenshots: 149 matched. Benchmark ResultsDetailed Performance Metrics
|
|
Compared 149 screenshots: 149 matched. Benchmark ResultsDetailed Performance Metrics
|
|
Compared 149 screenshots: 149 matched. |
|
Compared 149 screenshots: 149 matched. Benchmark ResultsDetailed Performance Metrics
|
|
Compared 181 screenshots: 181 matched. |
|
Compared 148 screenshots: 148 matched. Benchmark Results
Detailed Performance Metrics
|
…hing Two defects in the new open-addressed page index, one of them the x86-64 CI failure and one from review. ZERO IS THE EMPTY MARKER, SO IT CANNOT ALSO BE A KEY. cn1ConservativeResolve is handed arbitrary machine words off a conservative stack scan and masks each one to its 64KB page base; any word below CN1_BIBOP_PAGE_SIZE masks to 0, and a small aligned integer left in a stack slot is enough. Probing for 0 matched the first EMPTY entry and returned it as a hit -- an all-zero CN1ConsPage whose slotSize the caller then divided by. The sorted array this replaced could not be reached that way, because every element of it was a real page base; the hazard arrived with the table. It reproduces on the first collection of any workload, which is why every job that runs a translated binary on x86-64 failed at once (exit 136 = SIGFPE) -- and why every local run and the arm64 leg passed: arm64 answers integer division by zero with 0 rather than trapping, so the word quietly resolved to slot 0 of a page that does not exist. Reproduced locally by building the same app for x86_64, and confirmed as the exact instruction by -fsanitize=undefined on arm64, which reports it there too (master: zero UBSan findings on the same workload; this branch before the fix: division by zero at the resolver, from the conservative native-stack scan). Both now run clean and agree with the host JVM. THE REBUILD IS NOW ALL-OR-NOTHING (review, #5585). It used to clear the live table and insert into it, growing on demand -- so a failed calloc part way through left a PARTIAL index. That is not a slow index, it is a silently wrong one: a page missing from it makes every reference into that page fail to resolve, gcMarkObject's guard skips the object, and the sweep frees it while it is still reachable. Worse, the registry is a prepend list, so a rebuild that stopped early kept the NEWEST pages and dropped the oldest -- exactly the ones holding a long-lived live set -- and did it on allocation failure, i.e. when a collection matters most. The table is now sized once from the registration count (plus slack for pages registered during the walk), filled into a fresh allocation, and published only when complete. On any failure the previous table stays in place and cn1ConsPgIndexedCount is left alone so the next cycle retries; what that table lacks is pages registered since it was built, whose objects are mark==-1 fresh and survive on the sweep's grace rule -- the exposure a page registered mid-snapshot has always had. With no previous table to keep, marking cannot proceed at all, so that case says so and aborts rather than sweep a heap it cannot resolve; it is a few hundred KB of calloc, so reaching it means the process is already finished. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The rebuild had one failure return for two unrelated situations. Outgrowing the size it picked means a mutator registered pages while it walked -- harmless, and self-correcting on the next cycle. Failing to calloc at all is not. Collapsing them meant a lost race on the FIRST build, where there is no previous index to keep, would have taken the abort() meant for exhaustion. It cannot happen in practice (the walk only covers what was linked when the head was loaded, and the slack is 256 pages), but the two cases deserve different answers regardless: a race now re-sizes and walks again, up to three times, before giving up and leaving the previous index in place. Only exhaustion with nothing to fall back on aborts, and the comment says so. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The page index just had to learn that its empty marker must never be a lookup key. The extent table beside it uses the same marker and is safe for a reason that lives twenty lines away -- cn1ConservativeResolve rejects a zero word before either table is consulted, and no extent has a zero base. Write that down where the probe is, so the next restructuring knows the early return is load-bearing. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: d5c018fbf9
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
|
Compared 149 screenshots: 149 matched. |
|
Compared 143 screenshots: 143 matched. Benchmark Results
Build and Run Timing
Detailed Performance Metrics
|
The triggers-per-cycle assertion I added went red on CI at 4.42, and the claim
attached to it -- that the ratio is "a property of the two SPEEDS and not of
either", so it reads the same on a loaded machine -- is simply wrong. The mutator
is one hot allocation loop; a collection has to interleave a mark, a sweep and a
page walk with it, so under contention the collector is the one that loses.
Measured on this workload, triggers allocated per completed collection:
before this branch after
a core to itself 4.67 1.04
8 copies on 12 cores 4.01-4.74 2.30-2.71
16 copies on 12 cores - 3.28
CI: 4 forks, 4 vCPU - 4.42
The old collector was bound by its own cost rather than by the CPU it could get,
so its number barely moves; the fixed one is bound by the CPU, so its number
walks up to meet it. They converge, and no fixed threshold separates them on an
oversubscribed runner. The CI figure is that convergence, not a regression: the
same job's run took 77954ms against the 5804ms this workload needs alone.
The no-ceiling peak has the same shape and for a concrete reason. The growth
bound works by parking a mutator that has run too far ahead, and a park gives up
after two barren collections so that a thread can never be stalled by a collector
that is not running. Starve the collector enough and every park gives up, so the
bound stops binding: sixteen-way, the copies peak between 735MB and 15.7GB,
against 819-861MB eight-way where the collector still gets to run. That assertion
would have gone red next.
Both are now enforced only when the run had the machine, measured by the workload's
own elapsed time -- it is a fixed number of rounds, so that is a direct reading of
the CPU it got. Both numbers are PRINTED on every run either way, and a contended
run says which one it was and why. What this class still enforces unconditionally
is the part that is a property of the code: zero worklist overflows, the bound on
full drains taken inside a grace pass, and staying under the ceiling.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 4309d2edd4
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
|
Compared 149 screenshots: 149 matched. Benchmark Results
Build and Run Timing
Detailed Performance Metrics
|
…not support TWO THINGS, one of them entirely my fault. THE PREVIOUS COMMIT REVERTED THE FIX. While measuring master as a baseline I ran `git checkout origin/master -- cn1_globals.m cn1_globals.h`, which does not just write the worktree -- it STAGES what it writes. I restored the worktree afterwards, saw the resulting `MM` in git status, and committed a test-only change on top; the staged master copies went with it. 460 lines of cn1_globals.m disappeared in a commit whose message is about a test assertion. That is why CI then reported the old tracer format and why the review found CN1_SIMULATE_FREE_MEMORY, CN1_BIBOP_GC_MAX_CAP_MULTIPLIER and the allocatedKb / triggerKb fields "absent from this commit's target tree" -- they were absent, exactly as reported. Both files are restored to their d5c018f content and the index was diffed against the worktree before committing this time. A STALE PAGE INDEX MUST STOP THE SWEEP, NOT JUST THE REBUILD (review, #5585). Keeping the previous index when a rebuild fails is safe for ONE cycle: the pages it is missing were registered after the last successful rebuild, so their objects are mark == -1 and the sweep's grace rule keeps them. It is not safe for two. On the next failed rebuild those objects are no longer fresh, they still do not resolve -- so gcMarkObject's guard skips them however reachable they are -- and they age into the m < V - 1 reclamation with live fields still pointing at them. The fallback traded a hard failure for silent corruption in the low-memory case that motivated it. A failed rebuild now marks the cycle's mark as unsound and codenameOneGCSweep reclaims nothing on it. Skipping a collection costs the memory that cycle would have returned; sweeping on an incomplete mark costs the heap. It is self- correcting -- the rebuild is retried every cycle and the first success marks the whole live set before anything is freed again -- and it subsumes the empty-index case, so the abort() added for that is gone: nothing is swept, so nothing is lost. The blocked-thread release still runs on both paths, or a thread parked on the collector would hang instead. Exercised rather than assumed: with two of every three rebuilds forced to fail, the skip path runs, the throttled report fires, and RESULT stays bit-identical to the host JVM. The same fault injection under CN1_GC_VERIFY -- which walks every survivor's fields after every sweep and aborts on a reference into reclaimed memory -- is running as this goes up and is clean so far; it is slow enough that it outlasts the push, and the result follows on the PR. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 6c06a18525
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
✅ ByteCodeTranslator Quality ReportTest & Coverage
Benchmark Results
Static Analysis
Generated automatically by the PR CI workflow. |
|
Compared 144 screenshots: 144 matched. |
|
Compared 217 screenshots: 217 matched. |
… the test on it TWO REVIEW FINDINGS (#5585), the first of which corrects my own diagnosis. THE GROWTH BOUND WAS READING A STALE FOOTPRINT. It keys off cn1CachedProcFootprint, which cn1RefreshFreeMemCache samples once, at mark start. A cycle that begins just under the 512MB floor therefore keeps a below-floor reading for its whole duration, so cn1BibopPacingCap goes on granting the host-derived cap -- gigabytes on a roomy machine. A LONG CYCLE IS EXACTLY THE RUNAWAY THIS BOUND EXISTS TO STOP, so the clamp sat disarmed through the one interval that mattered. The footprint is now re-probed at the point of use, after asking whether the bound would bind at all so the syscall is paid for only on the path that needs it, and rate-limited to one probe per 25ms across all threads. That caps the overshoot at a refresh interval's worth of allocation instead of a collection's. I had attributed the same measurement to the wrong cause. The earlier note said the bound stopped binding under starvation because a pacing park gives up after two barren collections. That is true and still a limit, but it was not what produced the number: with the probe fixed, the same sixteen concurrent copies that peaked between 735MB and 15.7GB now peak between 871MB and 994MB, and twenty-four copies -- whose slowest run takes 118s, against the 78s of the CI job that motivated all this -- peak between 880MB and 1009MB. RESULT stays bit-identical throughout. A GATE THE REGRESSION CAN TRIP IS NOT A GATE. The no-ceiling peak assertion was gated on the run's own elapsed time, and the regression it guards makes the run slow: the test's own numbers put the broken behaviour at 12.2-13.4s against a 12s gate, so the failure could satisfy the skip condition and take the benchmark green. That gate is gone. The bound now holds under contention beyond anything CI applies, so the peak is asserted unconditionally and there is nothing left to disable. Triggers-per-cycle keeps no assertion at all -- it is a ratio of two speeds that converges on the broken collector's as the runner is oversubscribed, so no threshold separates them there and a gated version would have exactly the defect above. It is printed every run as a diagnostic, with the numbers and the reason in the javadoc. The sweep guard from the previous commit is exercised rather than assumed: with two of every three index rebuilds forced to fail, GcHeapIntegrityIntegrationTest -- the CN1_GC_VERIFY gate that walks every survivor's fields after every sweep and aborts on a reference into reclaimed memory -- passes, and the spiral workload's RESULT stays bit-identical with the skip path firing. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
… (issue #5537) (#5599) * Stop the SATB barrier logging fresh references (issue #5537) Four merged fixes (#5540, #5563, #5573, #5585) each named a mechanism and the reporter's build still climbed 500MB to 5GB in five minutes on the iOS Simulator with a live set of a few hundred objects, GC pauses lengthening until they were continuous. The reason none of them settled it is structural: every GC workload in vm/tests measures a PEAK under load, and a heap that grows forever at a modest rate passes "peak < 2GB over 50 rounds" without difficulty. Nothing measured whether the VM ever gives the memory back. The instrument comes first, and it is what found this. -DCN1_GC_CONFORM adds a probe that PARTITIONS the footprint -- resident pages, legacy blocks, the legacy table, the allocator's side tables -- and prints the residual the four do not account for, plus a per-phase breakdown of the mark. It deliberately is not CN1_GC_VERIFY: that flag forces cn1BibopReleaseOffset() to 0, which compiles out the page-release path, the major sweep and every madvise call, so the paths a footprint investigation is about cannot be measured in a verifier build. It changes no allocator behaviour, and the emitters are gated at RUNTIME on CN1_GC_PROBE so probe-on and probe-off are the same binary. On the reported shape -- a deep game-tree search on four workers, tiny short-lived reference-carrying objects, a constant live set -- it named the cost immediately: of a 327ms mark, 282ms was SATB termination, draining 2,718,448 logged references in one cycle. Of those, 2,718,413 were references to FRESH objects. A mark == -1 object was allocated after the cycle's snapshot was taken, so it is not in the snapshot the barrier exists to preserve, and both sweeps keep it anyway -- the grace rule promotes a fresh slot to the current epoch instead of freeing it. Its own outgoing references to non-fresh objects are still logged by the same barrier as they are stored, so nothing reachable only through a fresh object is lost, which is the hazard the insertion half was added for. Without that filter the log is a feedback loop rather than a cost: its size is mutation rate times cycle duration, draining it is part of the cycle, so a longer cycle logs more and logging more lengthens the cycle. Both reported symptoms fall out of the one loop -- the footprint climbs because the collector never catches up, and the pauses climb because the log it has to drain keeps growing. Measured, three repetitions each, interleaved in one session: footprint drift before 306,684 / 241,493 / 224,237 KB/min after -31,430 / 36,866 / 18,993 KB/min (noise around zero) page count before 3,947 -> 5,995 over 40s and still climbing after flat at 11,687 for 40s mark time before 38ms -> 180ms; after 9-68ms, no trend under a simulated 1.4GB per-process ceiling: 3.5x the search throughput (237.8M nodes vs 67.6M), peak 1271MB, no kill Throughput, interleaved A/B, checksums bit-identical: geomean 0.944 -- 5.6% faster overall, objectAllocation 1.73x (56.3ms -> 32.5ms). The barrier was that expensive. -DCN1_SATB_LOG_FRESH restores the old behaviour for A/B. GcSteadyStateIntegrationTest is the gate. It asserts the SATB log stays sized by the live set rather than by the allocation rate, and that the page heap stops growing in the second half of the run; then it rebuilds with -DCN1_SATB_LOG_FRESH and requires both to fail, so it cannot go inert. Two pre-existing defects found on the way and fixed here: * -DCN1_DISABLE_CONSERVATIVE_GC_ROOTS, the revert path cn1_globals.h documents, did not compile at all: the grace passes use CN1_GC_TRUSTED_BEGIN/END/SUSPEND/ RESUME unconditionally and those are only defined with conservative roots on. No-op definitions restore it, which is what makes it usable as an A/B arm. * [GC-INSTR] allocs= is not an allocation count -- CN1_FAST_NEW's inlined bump path never reaches that counter, so on a small-object workload it understates allocation by orders of magnitude. Renamed to outOfLineAllocs= with a note. Verified: 520 vm/tests non-benchmark tests green; all six GC benchmark tests green; run-gc-verify.sh green including both fault self-tests; run-gauntlet.sh green with every checksum matching; grace audit reports doomedChildren=0 with and without the filter; and the probe compiles across nine ablation flag combinations. Not addressed, and pre-existing: under a per-process ceiling the process still rides to ceiling-minus-64MB, which #5585 flagged as open. That is now a bounded plateau rather than unbounded growth, but the margin is thin on a device where the renderer shares the same budget. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * Defend a headroom reserve under a per-process ceiling (issue #5537) The previous commit stopped the heap growing without bound. This one stops the process parking itself on the kill line, which #5585 flagged as open and which is what turns a native spike into a jetsam kill. Budget headroom is not a footprint bound. Admission against os_proc_available_memory answers only "is there budget left", so it keeps saying yes until the budget is gone. Measured on the issue-5537 game-tree shape under a simulated 1.4GB ceiling, seven times: 1,271MB resident and 63MB of headroom left, every time, against a live set of a few hundred objects. That repeatability is the tell -- it is not an accident of the workload, it is the policy converging on ceiling minus CN1_PACING_HEADROOM_MARGIN by construction. The ceiling is not special either: give the same workload an 8GB budget and it rides to 7.5GB. There is no footprint TARGET anywhere in the design. 63MB is the whole margin, and the renderer spends out of the same budget -- #5598 measured one screen texture at 30MB. So the collector now also bounds how far the mutator may run ahead of it, but only once headroom drops inside a reserve of a quarter of the budget (CN1_PACING_RESERVE_SHIFT). Inside the reserve the mutator is clamped to the static cap, the collector gets ahead, and the footprint falls back out. Gating on HEADROOM rather than on footprint is what makes this affordable: it is a control loop that engages only inside the reserve, not a tax on every allocation, and volumeParks in the [PACING] report is 0 for a run that never enters it. Both allocation paths are charged against ONE figure. Bounding them separately is a defect this code has had before -- each running a full cap ahead of a cap derived from the same budget -- and the reserve is derived from the BUDGET, never from the device's free RAM, which is the defect #5563 fixed. cn1BibopPacingCap is deliberately not reused for that reason. Measured, builds interleaved within one session (-DCN1_PACING_NO_RESERVE is the same binary with the bound compiled out), simulated 1.4GB ceiling, four workers: peak footprint smallest headroom seen no reserve 1271MB, x7 63MB, x7 reserve limit>>2 1027-1036MB 298-304MB 4.8x the margin. Throughput across seven interleaved pairs came out at 0.90 to 0.99 of the unbounded build, median 0.94; the spread is session drift, not the bound, and the sign never changed. A single repetition each of the tighter reserves put >> 3 at 1183MB/150MB and >> 4 at 1207MB/127MB, both slower -- a smaller reserve engages later and thrashes closer to the edge -- so a quarter is the knee rather than a compromise. Roughly 6% for that margin is a different trade from the volume brakes #5573 and #5585 measured at 2-4x and rejected. It cannot touch a platform with no per-process budget, because the whole branch is unreachable there: vm/benchmarks measures geomean 0.9398 against master, i.e. still 6% FASTER from the previous commit's SATB fix, with no benchmark regressing and every checksum identical. cn1PacingPastGrowthFloor's rate-limited footprint probe is factored out as cn1PacingFootprintNow so both bounds read through it. Behaviour-preserving: each of its three early returns previously answered FALSE, and the fast path above already established that the cached value is under the floor. GcSteadyStateIntegrationTest gains a third scenario asserting the process defends its reserve under a simulated ceiling, and a fourth that rebuilds with -DCN1_PACING_NO_RESERVE and requires the third to fail -- otherwise a gate that never engages would report green forever. Verified: 520 vm/tests non-benchmark tests green; all seven GC benchmark tests green (ProcessBudgetPacingIntegrationTest included, which exercises the same budgeted path); run-gc-verify.sh green with both fault self-tests; run-gauntlet.sh green with every checksum matching; nine ablation flag combinations compile. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * Give the benchmark driver its GPL header, and scope two helpers to their use check-copyright-headers rejects a new source file without the complete Codename One GPLv2 + Classpath Exception header, and vm/benchmarks/src is in scope. cn1PacingUncollectedBytes and cn1PacingReserveBytes are used only from the reserve bound, so they are guarded on the same condition it is -- otherwise compiling the bound out with -DCN1_PACING_NO_RESERVE leaves them as unused statics. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * Count benchmark nodes per worker, not through a shared racy counter The driver incremented one static long from four workers with an unsynchronised read-modify-write, and the sampler read it concurrently. That is not merely imprecise: the rate at which increments are lost depends on CONTENTION, and contention is exactly what differs between the builds this benchmark compares -- a build whose threads park more loses fewer increments and so reports a throughput advantage it has not got. The per-round `nodes = localNodes` writeback also overwrote the shared total instead of combining the workers' counts. Each worker now counts into its own slot, and NODES= is summed after join(), which gives it a happens-before edge to every worker's last write. The SAMPLE series sums the same slots while they are still being written, so it is renamed nodes~= and documented as a progress indicator rather than a measurement. The CI fixture (GcSteadyStateApp) never had a node counter -- its assertions come from the [GCPROBE] series -- so nothing the gate asserts is affected. Reported by chatgpt-codex-connector on #5599. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * Keep every probe row a single-cycle row, publish node counts live Two review findings on #5599, both real. cn1GcProbeCycle returned early on a skipped cycle without clearing the phase accumulators, so with CN1_GC_PROBE>1 snapMs/graceMs/satbMs and friends carried a whole interval while markMs and sweepMs described only the cycle that just ran -- two time bases in one row, which would attribute an interval's worth of a phase to a single cycle's pause. The resets move into cn1GcProbeResetPhases and run on every cycle, printed or not. The cumulative counters (matured, consWords, staleSkips) are deliberately left alone: those are running totals the reader diffs. The benchmark driver published each worker's node count only after the run stopped, so every SAMPLE line reported zero. It now republishes once per round; a worker that stalls stops publishing and its slot going flat is the signal. Neither affected any measurement reported so far -- every run used CN1_GC_PROBE=1, where the skip path is unreachable, and the throughput figures come from NODES=, which is summed after join(). Also corrects the reserve's throughput figures, which came from the racy counter the previous commit replaced. Re-measured with the exact one, four interleaved pairs: 0.97-1.05 of the unbounded build, median 0.99, two of four faster with the bound on. The previous "median 0.94" overstated the cost. Peak footprint and headroom are unchanged (1271MB/63MB against 1015-1027MB/306-308MB) -- those come from the probe and Runtime, not the counter. The claim that a smaller reserve is "slower" is withdrawn; >>3 and >>4 buy less on peak and headroom, which is the argument that survives. Reported by chatgpt-codex-connector on #5599. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * Load the mark word atomically in the barrier, and harden the ceiling scenario Three findings, two from review and one the review's tighter test surfaced. The SATB filter read __codenameOneGcMark with a plain load while the marker reaches the same field through __atomic_*. That is a mixed atomic/non-atomic access to one object -- undefined in C, and the same bug class #5598 fixed in the constant pool. The concrete hazard is not tearing but the compiler caching a -1 across several inlined barriers in one loop, which would keep suppressing entries after the object had aged into a genuine snapshot object. Now __ATOMIC_RELAXED, and the comment says why relaxed and not acquire: nothing is published through this read, both stale answers are safe, and what relaxed buys is that the load happens at all. An acquire fence on every object store buys nothing over that and is not free on arm64. The two CN1_GC_CONFORM census reads of the same field move with it. The fault-injected runs' measurements were accepted without checking exit status or the completion marker, so a build that crashed after emitting enough probe rows would have satisfied the assertions and turned a memory-safety regression into a green gate. Both now go through assertHealthy first. The ceiling scenario used a 1400MB budget, which needs the mutator to actually outrun the collector by 1.3GB -- and how far it outruns depends on how many cores it has to itself, so a two-core runner might never get there and the fourth scenario would go quietly inert. It now uses 768MB, which admission converges on by construction rather than by winning a race. The threshold between the two regimes becomes ABSOLUTE, twice CN1_PACING_HEADROOM_MARGIN, because the margin does not scale with the budget: a proportional threshold silently stops separating them as the budget shrinks, which is exactly what happened at 400MB (reserve 100MB, margin still 63MB, half the reserve below it). Reported by chatgpt-codex-connector on #5599. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * Do not size an adopted BiBOP slot as if it were a malloc block The probe sized every non-null allObjectsInHeap entry with malloc_size / malloc_usable_size. A MATURED object is in that table but its storage is a slot inside a posix_memalign'd BiBOP arena, so the pointer is interior: glibc's malloc_usable_size reads the chunk header immediately below it and returns a garbage figure, and CI runs this gate on Linux. Its bytes are also already counted in residentPgBytes, so anything it did return double-counted into the residual that is this probe's whole point. Only an object the table INDEXES (__heapPosition >= 0) owns an individual block. The rest are counted as legAdopted instead -- the same population as matured - maturedDied but measured from the table rather than from the counters, so the two disagreeing is itself a finding. Not a small corner: on the game-tree workload legAdopted is 32,907 of a legUsed of 33,164, so 99% of the table was being sized this way. It was harmless on macOS only because malloc_size answers 0 for an interior pointer, which is also why legBlockKb read flat through the original investigation and correctly never carried the drift. Verified after the change: run-gc-verify.sh green with both fault self-tests, and vm/benchmarks geomean 0.9422 against master (0.9398 before the previous commit's atomic load, i.e. that load costs nothing), no benchmark regressing, checksums identical. Reported by chatgpt-codex-connector on #5599. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * Close the wait timer at the wait, read the cycle counter atomically Three findings from review; two fixed, one measured and answered in the code. waitMs was opened before the safepoint wait and closed only after the allocation migration and both stack scans, so it double-counted work already attributed to migrateMs and stackMs -- a phase breakdown that overlaps reads a long root scan as mutator wait time, which is the opposite of what it exists to say. It now opens and closes around the wait alone, inside the lightweightThread branch, so a native thread (which is never waited for) contributes 0 instead of everything up to markStatics. The 1Hz emitter read currentGcMarkValue with a plain load while the collector increments that ordinary int -- a data race, in the one emitter documented as "atomics only" and built to keep reporting exactly when the collector is stalled. Now an atomic relaxed load, as is the mutator-side comparison in the SATB census. Not taken: requiring the -DCN1_SATB_LOG_FRESH build to also blow the second-half page-growth bound. Measured across two runs of that build, its second-half growth is 0.446 and then 0.033 -- a runaway's page pool sometimes saturates before the midpoint and the ratio then reads flat while the heap is enormous. That assertion would fail about half the time, and a coin-flip gate is worse than the inertness it guards against. The reasoning, the numbers and what does have teeth (the SATB metric, five orders of magnitude, every time) are recorded on the constant. Both series are now printed on every run so the ratio stays auditable rather than merely asserted. Verified after these changes: phases sum to markMs with no overlap (16.0 of 16.3); 520 vm/tests non-benchmark tests green; all seven GC benchmark tests green. Reported by chatgpt-codex-connector on #5599. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * Read the collector's atomic epoch mirror, and claim a matured page with one edge Two follow-ups from review, both correct. The previous commit made the 1Hz emitter's read of currentGcMarkValue atomic while codenameOneGCMark still increments it with a plain ++. That is half a fix: an atomic read of a plainly-written object is still a mixed access and still undefined. Both sides now go through bibopGcEpoch, the collector's own _Atomic mirror of the same value, published at cycle start -- which is what the reviewer offered as the alternative and what should have been used first. The mutator-side comparison in the SATB census moves with it. Where there is no page heap there is no mirror, so the emitter reports cyc=-1 rather than a figure read through a data race. cn1MaturedPages tested gcHasAdopted and then let the existing plain store set it. The CAS above guarantees one thread matures a given OBJECT, but two markers can mature two different objects on the SAME page, so both could observe FALSE and both count it -- and the plain store is itself a data race the moment gcMarkResolveThreadCount stops returning 1. Now one __atomic_exchange_n: exactly one thread sees the FALSE->TRUE edge, and it does the counting. That the ratio is read chiefly in the CN1_GC_MARK_THREADS>1 arm is the point -- it would have been wrong exactly where it is used. Verified in that arm: maturedPages=2121 of pgTotal=11214, a plausible ratio rather than an inflated one. run-gc-verify.sh green with both fault self-tests; the steady state, heap integrity and process budget gates green; seven ablation combinations compile. Reported by chatgpt-codex-connector on #5599. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * Make every collector-side write to the mark word atomic The barrier's read was made atomic two commits ago while gcMarkObject still stamped the same field with a plain store, so the pair was still a mixed access. The field already had an atomic convention here -- gcMarkObject's own read is __ATOMIC_ACQUIRE, the BiBOP publish is __ATOMIC_RELEASE -- and the plain writes were the inconsistency, not the new read. Every write that can run concurrently with a mutator is now a relaxed atomic store: gcMarkObject's stamp, both sweeps' grace promotion, both free-mark stores, the nursery promotion and the CN1_GC_VERIFY poison. Relaxed compiles to the same instruction on every target we build; what it buys is that the write is a write the reader is allowed to observe. Header INITIALISATION deliberately stays plain, in codenameOneGcMalloc and in cn1FusedInstallPrimArray. Those are not concurrent with anything: the barrier only ever reads the mark of an object the mutator holds a reference to, so one already published, and the publishing store orders the initialisation against any reader. That distinction is not free-floating -- making those two atomic as well cost 1.2 points of benchmark geomean (0.9550 against 0.9432, with arraySequential, quicksort and valueEscape all moving and returning), because they sit on the allocation fast path. The reasoning is recorded at the site so the next person does not reintroduce it for symmetry. Verified: vm/benchmarks geomean 0.9432 against master, six rounds interleaved, no benchmark regressing and checksums identical; run-gc-verify.sh green with both fault self-tests; all seven GC gates green; five ablation combinations compile including -DCN1_NURSERY. Reported by chatgpt-codex-connector on #5599. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * Emit the generated mark chain's root store atomically too The previous commit converted every hand-written collector-side write of the mark word and missed the one that matters most, because it is not in the C sources at all: ByteCodeClass emits the root of every generated mark chain, and that store was still plain. It runs on the GC thread for every object marked while the SATB barrier atomically loads the same field from mutators, so the pair stayed a mixed atomic/non-atomic access -- the exact defect the previous commit was for, in the one place a grep of cn1_globals.m could not see. Costs nothing, as the hand-written conversions did not: vm/benchmarks geomean 0.9387 against master over six interleaved rounds (0.9432 before this change, so inside the noise), no benchmark regressing, checksums identical. A codegen change touches every translated class rather than one runtime path, so it is verified against the shapes rather than the sites: run-gc-verify.sh green with both fault self-tests, and run-gauntlet.sh green with every checksum matching. Reported by chatgpt-codex-connector on #5599. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * Publish the sampler's counters, and keep the page partition valid under a race Two findings, one taken as offered and one taken but answered differently. The benchmark driver's per-worker slots were published with a plain long[] write against a concurrent reader: no visibility guarantee, and Java 8 permits a 64-bit element to be observed torn, so the live series could sit stale or jump nonsensically exactly when a stalled worker is what it is meant to show. Publication and sumNodes() now share SUM_LOCK. Once per round is about once a second per worker, so it costs nothing, and NODES= after join() remains the authoritative figure regardless. The probe's page walk is a different case. It reads plain page counters while mutators run, which is a race, but it is the same deliberate sample cn1HeapAccounting takes beside it -- "a diagnostic wants the shape, not the last digit" -- and both offered remedies cost more than the unsoundness. Stopping the page owners would perturb collector/mutator timing, which is the quantity this probe reports, and would cost CN1_GC_CONFORM the behaviour-neutrality that is the only reason it is a separate flag from CN1_GC_VERIFY. Making the page fields _Atomic would put atomic accesses on the inlined bump path in cn1_globals.h, the hottest code in the VM, to improve a diagnostic. What is worth fixing is the harm actually named: an internally inconsistent partition. Only an owned page can move under the walk -- at most one per size class per thread out of many thousands -- so freeCount is clamped into [0, bumpIndex] and a stale pair can no longer make live and dead slots sum past the page. Verified: 517295 + 25326 KB against a 776448 KB reservation. The reasoning is recorded at the walk so the next reader does not have to rediscover which of the three options was chosen and why. Verified: run-gc-verify.sh green with both fault self-tests; steady-state, heap integrity and process budget gates green; the sampler now tracks progress live (nodes~=28,697,812 mid-run against a final NODES=34,360,526); and the 520-test non-benchmark suite is green on the regenerated code from the previous commit. Reported by chatgpt-codex-connector on #5599. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * Survive a stall: publish inside the traversal, bound the run Both findings are the same blind spot from two directions -- a stalled collector is one of the things this gate exists to CATCH, and neither the progress series nor the runner survived one. Publishing between rounds was not enough. One depth-14 traversal is millions of nodes, so if the collector stalls badly enough that no round completes inside the window, nothing is ever published and the series reads zero -- silent in exactly the case it is for. It now also publishes every 1<<20 nodes: a power of two so the test is an AND, coarse enough (about a fifth of a second of work) that the lock traffic is negligible against the sampler's 4Hz. Verified live: 0 -> 4,194,304 at 1s -> 33,554,432 at 9.8s, against a final NODES=35,255,230. The runner read the child's output to EOF on the test thread and only then called waitFor(), so a hung workload would block until the CI job's global timeout -- the guard would stop reporting a regression and start eating the build. It now drains on a background thread and waits with a bound, killing the child on expiry and failing with whatever it printed, which is the only diagnostic a stalled run leaves. That is not a new invention: GcOverflowSpiralIntegrationTest and ProcessBudgetPacingIntegrationTest both already do exactly this, and the naive pattern came from copying GcHeapIntegrityIntegrationTest, which is the one that does not. Reported by chatgpt-codex-connector on #5599. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * Do not filter fresh SATB entries where there is no insertion barrier The filter's soundness argument ends "its non-fresh children are still logged by this same barrier as they are stored". That step has a precondition I did not state and did not check: the INSERTION half has to exist. Under CN1_NURSERY it does not. CN1_WRITE_BARRIER is the nursery remembered-set update there and enqueues nothing at all (cn1_globals.h:1020-1039), so a fresh container that takes an older child after the grace pass has that child recorded nowhere -- and dropping the deletion entry for the container then lets the sweep reclaim a child the grace-surviving container still references. That is a use-after-free, in the class of defect #5425 and #5442 were about. The filter is an optimisation and not a correctness requirement, so a build without the insertion half simply does not get it: the condition is now !defined(CN1_SATB_LOG_FRESH) && !defined(CN1_NURSERY). Adding SATB insertion to the nursery barrier was the other option offered and is the riskier one -- it changes barrier behaviour in a configuration nothing exercises, and would have to be justified by measurements no one can take. Latent rather than live: CN1_NURSERY is not defined anywhere in-tree, so no shipping or CI build takes that path. It is a documented, reachable flag, and the comment now records the dependency so the next person to enable it is not relying on an argument that quietly stopped holding. Verified: five ablation combinations compile including -DCN1_NURSERY; run-gc-verify.sh green with both fault self-tests; steady-state and heap-integrity gates green. Reported by chatgpt-codex-connector on #5599. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * Let CN1_WL_LEGACY=0 actually ablate the legacy population Setting the documented knob to 0 built a zero-length legacyLiveSet and then indexed [-1] on the last line, so the driver threw AFTER the entire timed run had been paid for -- losing RESULT and GC_STEADY_STATE_DONE, which is everything the run was for. Running without the retained legacy population is a legitimate ablation, so it now works rather than crashing: the fold is skipped when there is nothing to fold. Two neighbouring values that would produce a wasted or silently empty run are clamped at the same time. A negative CN1_WL_LEGACY reached new Object[n][]; a CN1_WL_THREADS below one started no workers at all and reported that only by printing zero nodes, which is the exact failure mode -- a measurement that looks like a result -- this whole change has been about. WLCONFIG prints the clamped values, so the log says what actually ran. Verified: CN1_WL_LEGACY=0, CN1_WL_LEGACY=-5 with CN1_WL_THREADS=0, and the defaults all reach RESULT and GC_STEADY_STATE_DONE. Reported by chatgpt-codex-connector on #5599. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * Check the answer, not just the telemetry, in every scenario Three of the four runs checked exit status and the completion marker but never that the workload still computed the right thing. That gap matters most exactly where it was left: the ceiling scenarios exercise the budgeted pacing path -- the code this change touches most -- under an environment the clean run never sees, so a worker could die early or compute a wrong sum while the process still exited cleanly and emitted plenty of [PACING] telemetry for the policy assertions to pass. None of the variants changes what the program computes: the faults injected are a barrier filter and a pacing bound, and the workload is deterministic by construction (fixed rounds, fixed seeds, an order-independent checksum). So RESULT must equal the host JVM's in all of them, and assertHealthy now requires it -- which also picks up the -DCN1_SATB_LOG_FRESH run, which had the same gap. Reported by chatgpt-codex-connector on #5599. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * Normalise every workload knob, not just the one that was reported CN1_WL_MOVES=0 left the move chain null and the next-seed derivation dereferenced it, so the leaf-only ablation died with an NPE on the first node. That is the second knob found this way, so this fixes the class rather than the instance: all eight are normalised in one place before the timed run, and WLCONFIG prints the normalised values so the log says what actually ran rather than what was asked for. Auditing the rest turned up one more that was worse than the reported one. A negative CN1_WL_DEPTH never matches the d == 0 base case, so it recursed until the stack gave out. CN1_WL_SECONDS and CN1_WL_BRANCH below their floors produced runs that measured nothing and said so only by reporting zero -- the failure mode this entire change is about. Zero stays meaningful where it means something, and both cases are real ablations: no retained legacy population, and no reference-carrying Move per node. The second is worth having, because only a non-leaf object reaches the grace pass's worklist or maturation, so leaf-only allocation is a genuinely different workload for the parts of the collector under test. Verified: CN1_WL_MOVES=0, CN1_WL_MOVES=-3, CN1_WL_DEPTH=-1, CN1_WL_BRANCH=0, CN1_WL_SECONDS=0 and CN1_WL_LEGACY=0 all reach RESULT and GC_STEADY_STATE_DONE, and the default configuration is unchanged. Reported by chatgpt-codex-connector on #5599. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * Flag the probe row when the collection cycle threw gcMarkSweep wraps mark and sweep in a catch-all so a throwing finalizer cannot wedge the collector. On that path control jumps past the timing assignments, so the probe emitted a row carrying the PREVIOUS cycle's markMs and sweepMs beside the partial current cycle's phase counters -- two cycles in one row, and it concealed the exceptional cycle, which is the one a reader most wants to see. This is the same defect as the CN1_GC_PROBE>1 skip path fixed earlier, on a different route out. The timings are now cleared BEFORE the protected region, so a throw cannot inherit them, and the row carries threw=1 rather than being suppressed: hiding it would defeat the reason this probe has a wall-clock emitter at all. The three carriers are file scope, so the setjmp/longjmp indeterminate-local rule does not apply to them. Verified: five ablation combinations compile; run-gc-verify.sh green with both fault self-tests; steady-state and heap-integrity gates green; probe rows carry threw=0 on a healthy run. Reported by chatgpt-codex-connector on #5599. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * Re-evaluate the reserve throughout the wait, and stop the driver perturbing itself Two review findings, and a third defect the first one's verification exposed. The wait loop's copy of the volume bound was guarded on the thread having already been refused, so it could only transition refused->allowed. A thread that parked on BUDGET while outside the reserve then held a stale "allowed" for its whole wait and could be admitted on headroom alone after other mutators had pushed the uncollected total past the cap and the process into the reserve. There is now ONE definition, cn1PacingVolumeOk, called from both sites and recomputed every iteration -- the two copies drifted precisely because they were two. The gate parsed only the per-cycle [GCPROBE] rows, so a collector that completes its early cycles and then never finishes another was invisible to it: the rows stop, the generated main returns as soon as the workers do, and the process exits cleanly with the marker while the heap is still growing. [GCPROBE-T] was added for exactly that state and then not asserted on. The outcome check now covers the wall-clock series too, with its own anti-vacuous row count. And the driver had started perturbing its own experiment. The periodic publication added two commits ago took SUM_LOCK inside the search, and monitorEnter is a GC SAFEPOINT in this VM -- so the workers were being stopped far more often than the workload otherwise permits and the runaway stopped reproducing: peak footprint fell from 1271MB to 126MB with the reserve compiled out, in BOTH builds, which is what gave it away. Publication is now a volatile long per worker: not a safepoint, not a lock, and JLS 17.7 makes volatile long access atomic, so it also answers the visibility and tearing that the plain long[] had. With the runaway restored, the reserve's throughput cost is re-measured across four interleaved pairs at 0.875-1.035, median 0.90 -- about a tenth, not the ~1% the previous figure claimed. Peak and headroom are unchanged (1271/63 against 1022-1064/272-304). This is the third throughput figure this comment has carried and the first two were both apparatus rather than signal, so the comment now says which were which. Verified: four ablation combinations compile; run-gc-verify.sh green with both fault self-tests; the steady-state gate green with all five checks. Reported by chatgpt-codex-connector on #5599. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * Attach evidence to the ceiling assertions The first vm-tests run that ever completed on this branch failed scenario 3 -- "the smallest headroom seen was 62MB" under a 768MB budget -- and reported nothing else. Every other assertion in this gate appends the run's output; this one, the only one that has actually failed, did not. The probe rows that would explain it were captured and then discarded. Both ceiling assertions now carry the [PACING] counters, the last [GCPROBE] footprint partition and the wall-clock summary. That partition is the whole point of the probe: it says whether a footprint the reserve did not defend is even in the Java heap. No behaviour change, and the gate still passes locally on macOS -- which is itself the open question, since the failure is on the Linux runner and the two measure different quantities (phys_footprint against RSS). Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * Assert the reserve's mechanism, report its outcome The first vm-tests run that completed on this branch failed scenario 3 on the Linux runner: 62MB of headroom under a 768MB budget. With the evidence attached, the diagnosis is not what I guessed. I expected allocator retention -- glibc arenas holding freed legacy blocks, which RSS counts and phys_footprint would not. Wrong: residKb was 7MB of a 518MB footprint, so the footprint was the Java heap almost exactly. That is the residual bucket earning its place; it killed the hypothesis in one line. What the runner actually shows is a collector that cannot keep up, with the bound working: volumeParks=879, so it engaged and parked repeatedly, while mark ran 407-545ms per cycle -- 235ms of conservative stack scan, 122-252ms waiting for mutators to reach a safepoint -- against ~170MB of allocation per cycle. With the grace rule holding a cycle's allocation two more cycles, the smallest working set that machine can hold is already above the reserve line at that budget. satbMs was 0 throughout, so the earlier fix is holding and the stack scan is simply the next cost. So an absolute headroom assertion was testing the runner rather than the collector. Scenario 3 now asserts the contract, which is true on any machine: either the process never entered its reserve, or the bound engaged when it did. The headroom achieved is printed either way, so the outcome stays visible without being asserted. A regression that stops the bound engaging fails here; a machine that is merely slow does not. Scenario 4 gains a second half for the same reason -- with the reserve compiled out the process must land on the bare admission margin, or the ceiling is not pressuring the workload and scenario 3's "never entered" branch would pass for the wrong reason -- plus volumeParks == 0, since the bound is not in that build at all. Locally: headroom 161MB inside a 192MB reserve with volumeParks=350, against 63MB and volumeParks=0 with the reserve compiled out. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Fixes the fourth and, as far as this can be measured off the reporter's device, load-bearing cause of #5537.
What was actually wrong
cn1ConservativeResolvewas 75% of the GC thread's wall time.gcMarkObjectcalls it on every reference field the drain follows, to reject a conservatively-derived pointer before dereferencing it, and it answered by binary-searching two snapshots — the BiBOP page bases and the legacy extents. On the reporter's shape that is ~13 dependent cache-missing loads to locate the page, plus ~15 more to miss it and locate the array, per field.So marking cost O(log heap) per reference. The collector got slower as the heap grew, which is the reporter's "GC pauses become more and more frequent and take longer, until they are effectively continuous" verbatim. Everything else followed: a cycle stretched to 4.7x the collection interval, the mutator produced 4.7 triggers of garbage per completed cycle, and the process settled at whatever pacing allowed — 447MB against a live set of a few hundred bytes, riding 64MB below the ceiling that kills it. On device that is the kill; in the simulator, where nothing kills it, it is the footprint climbing to gigabytes he saw next.
Both indices are now open-addressed hash tables. The page table keys on the 64KB page base and stores the geometry inline (a hit is one cache line); its key set changes only when a page is registered, so it is rebuilt on that event and only its geometry refreshed per cycle — which also retires the per-registration
qsort. The legacy side keeps its sorted extent array for the interior pointers only the conservative stack scan produces, with an exact-base table in front: a Java reference is always an object base, so the dominant caller is answered in one probe.Two defects the faster collector exposed, fixed here because both undo it
The survivor-heavy bypass read pure churn as survivor-heavy. Survival is measured at sweep as slots carrying the current epoch — and the grace pass marks every fresh non-leaf object with it, so what the policy read as a live set was really the allocation rate (190K of 700K 48-byte slots "surviving"). It stayed under the threshold before only because the slow collector inflated the denominator; once the collector kept up it crossed, diverted 1.8M small objects onto the legacy heap, and brought the worklist overflow back. Pages now tally the marks a grace pass put on them and the sweep subtracts them.
Off a per-process ceiling the pacing cap was a fraction of the host's free RAM. That is a reason to let a fast thread run further ahead of the collector, not a reason to accumulate an unbounded amount of garbage. It is now bounded by a multiple of the collection trigger (which already tracks the heap), gated on the process already being past 512MB — so it stops growth without ever touching a process that was not going to grow. An ungated bound cost 47% on the
objectAllocationmicrobenchmark; gated, that benchmark is unchanged.Measured
GcOverflowSpiralApp, same host, 14.9GB allocated either way,RESULTbit-identical:No ceiling, 32GB host RAM, eight concurrent copies so the collector has to fight for the machine (which is what tips it): 13.8-15.7GB -> 819-861MB, and faster for it — 12.2-13.4s -> 8.1-8.6s, since a process thrashing fifteen gigabytes pays for them.
Gates run locally
vm/tests: 519 tests in the default group, 8 in the benchmark group, all green — includingGcHeapIntegrityIntegrationTest(theCN1_GC_VERIFYuse-after-free gate) andLargeArrayGcIntegrationTest(issue 5425).vm/benchmarks/run-gauntlet.sh: GREEN in cooperative and forced-signal stop modes, every torture bit-identical to the host JVM.vm/benchmarks/run-benchmark.sh: geomean unchanged, all checksums bit-identical.cn1_globals.mcompiles clean to an arm64-apple-ios object against the iOS SDK, and in theCN1_GC_VERIFY/CN1_BIBOP_VALIDATE/CN1_GRACE_AUDIT/CN1_BIBOP_NO_FASTSWEEP/CN1_DISABLE_BIBOP/CN1_RESOLVE_DIAG/CN1_BIBOP_NO_PACINGconfigurations.Guard
GcOverflowSpiralIntegrationTestgains the property underneath all of it — triggers allocated per completed collection, a ratio of two speeds, so it reads the same on a loaded machine where a peak does not (the same run inside a fully loaded parallel suite reported the same cycle count and a 447MB peak). It also gains a second run of the same binary with no ceiling, covering the half of the report that was previously out of scope, withCN1_SIMULATE_FREE_MEMORYpinning the host reading so that leg means the same thing on an idle machine and a busy one. That second assertion is a bound rather than a reproduction, and says so in the source: the off-ceiling runaway is bistable and took eight concurrent copies on a twelve-core host to provoke.Not addressed
qsortof the extent array is now the largest remaining item in the collector, roughly a third of its time.OpenGLES.framework"no such file" report on a Metal build is untouched here; the iOS port's nativeSources still#import <OpenGLES/...>regardless of the Metal setting.🤖 Generated with Claude Code