feat(perf): #263-A — the measurement instrument (CALIBRATION_ONLY; no admissible baseline yet) - #351
Open
PhysShell wants to merge 42 commits into
Open
feat(perf): #263-A — the measurement instrument (CALIBRATION_ONLY; no admissible baseline yet)#351PhysShell wants to merge 42 commits into
PhysShell wants to merge 42 commits into
Conversation
…alled The instrument for #263's P-022 performance baselines: how a number is produced, never what the number should be. It permits calibration and refuses the decisive measurement, which is #263-B behind the D7 freeze. No thresholds, no repetition count N, no engine-comparison statistic — those are D7's, and a harness that offered any of them would be choosing them. Preflight, against primary source. Live #263 and #262's Performance-gates reconcile as: the brief's phase list is #262's seven gate phases plus frontend extraction, recorded and explicitly NOT D7-gated. Both populations are frozen — A, the G3 cutover subset, feeds D7; B, the IDE-foundation remainder, does not unless separately ratified. What this instrument does NOT measure is written down rather than left implied, so completing #263-A cannot narrow #263's own acceptance by omission: the .own workloads, OwnIR serialization, no-op and single-file-edit recheck, the latency proxy, allocation profiles and IDE budgets are owed on B's track. The honest limit, recorded rather than routed around. Neither engine exposes per-phase timing through its production surface and #263-A may not add production instrumentation, so phases are measured by a LADDER of real production invocations, each recorded as the composed interval it actually is with the phases inside it named. bridge-lowering and analysis are marked NOT SEPARATELY OBSERVABLE. Derived views — parse by subtraction, core work as end-to-end minus extraction — are labelled derived and never presented as measurements. #263 asks for unavailable stages to be marked, not imputed. The firewall is structural, not clerical. Exactly one function starts a clock and it refuses a decisive workload unless the identity gate is armed, which needs a D7 attestation that does not exist. Deliberately stricter than "no paired runs": a single-engine timed run of a decisive workload is refused too, because a stopwatch appears nowhere in the permitted list of enumerate / hash / fetch / availability-check / non-timed smoke. "We did not save the delta" is accounting-clean and epistemically worthless — whoever watched the two numbers has already peeked. The four carried obligations live in the harness. The identity gate exists NOW and dormant, because D7's C1 freezes the harness digest and a gate added after that freeze makes this a different instrument. Arming is a JSON file: a control proves the digest does not move when the gate arms and that the firewall then opens. Verification completes before any clock starts, and a control counts digest calls INSIDE a measured interval and requires zero — a gate that pays for itself out of the startup benchmark is a defect wearing a safeguard's coat. The candidate binary is a separate domain: frozen at session start, its drift refuses the remainder rather than measuring half the cells on another binary. Peak RSS comes from os.wait4's per-child ru_maxrss rather than /usr/bin/time, which is a package that can simply be absent — it is absent here, and "the tool was missing" is not a memory measurement. Windows uses a Job Object. Where nothing is available the value is null WITH A REASON, never a confident zero. Two defects of my own, found while building this and worth recording. The comparison guard first grepped the source for "ratio" and refused its own report — the sentence explaining the prohibition matched it, and then so did "calibration" and "iterations". It now reads the report's structure and matches whole tokens. And an aggregate of bytes was being emitted under a median_ns key: a unit lie in a report whose entire purpose is measurement. Calibrated: 44 cells, both engines, all six rungs, run twice on one machine with every cell's median reproducing within the noise floor that run itself recorded. CI calibrates on Linux and Windows and asserts the same, plus that no decisive workload was timed and the gate stayed dormant. scripts/perf_baseline.py is a new wrapper that reaches the launcher, so it is added to the Stage-2 census entry points and the job classified — the census enumerates wrappers, which makes that list a maintenance obligation, now stated in the ledger. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01CYLNQy6tLXqsV1CbuNqsSb
Recorded on a clean tree at f09fa00, dirty:false, with the D7 gate dormant and no decisive workload timed. 44 cells both engines, all six rungs, process-cold and warm reproduced 44/44 medians within the noise floor the run itself recorded RSS os.wait4 ru_maxrss, per child, on every cell noise not invalidated; opening and closing probes agree decisive timed none The report validates the INSTRUMENT and nothing else. It is not evidence about either engine, it cannot choose a threshold, a repetition count or a comparison statistic, and it carries no engine comparison — the harness refuses to emit one, checked on the way out as well as on the way in. The Python reference commit is recorded beside the interpreter version so a later artifact cannot silently combine comparisons bound to two different reference states. The environment is a development container, recorded as such in the fingerprint (runner_class: local); CI calibrates on ubuntu and windows runners and asserts the same properties there. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01CYLNQy6tLXqsV1CbuNqsSb
|
Important
This repository does not receive automatic reviews because it has fewer than 10 stars. ⚙️ Run configurationConfiguration used: Path: .coderabbit.yaml Review profile: CHILL Plan: Advanced Run ID: Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
The Windows leg refused its own calibration on the first CI run: the opening noise probe measured a relative IQR of 1.807 against the 0.35 floor, with 0.870 drift between the opening and closing probes. That is the instrument doing exactly what §4 and §7 ask — refusing to report numbers taken on a machine that was not holding still — and the job read it as a failure. So CI now asks two questions instead of one. Did the HARNESS stand up and report honestly: tagged CALIBRATION_ONLY, gate dormant, zero decisive workloads timed, cells produced, raw per-iteration data retained, a named RSS mechanism, complete provenance. That is the gate, on both platforms. And separately: is this ENVIRONMENT measurement-grade — recorded per run, and allowed to be "no", with the reasons required when it is. The noise floor was deliberately NOT raised. An instrument-validity constant chosen after seeing which runs it rejects is a threshold fitted to a result, and fitting thresholds to results is the entire failure mode #263-A exists to prevent. The number that would have made this green is the one number I am least entitled to pick. The finding itself matters more than the fix: a GitHub-hosted Windows runner is not measurement-grade, so the decisive Windows measurement will need a single-tenant machine. #263-B needed to know that before it started, and now it does — from the instrument's first run rather than from a surprise halfway through a decisive session. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01CYLNQy6tLXqsV1CbuNqsSb
On ONE commit, two Windows runs minutes apart disagreed about their own environment. The first refused it — relative IQR 1.807 against the 0.35 floor, 0.870 drift between probes. The second accepted it and completed a full 44-cell calibration that reproduced within 0.35. Same commit, same runner class, opposite validity verdicts. The risk is therefore not that a hosted-Windows measurement would be noisy. It is that whether it is ADMISSIBLE turns on scheduling luck, so a decisive session started on the lucky run and continued into the unlucky one would be half a measurement. D7 should not preregister a Windows protocol that assumes a hosted runner. Both legs are otherwise green and complete: 44 cells each, 44 reproduced, peak RSS through os.wait4 on Linux and the Job Object on Windows, gate dormant, zero decisive workloads timed. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01CYLNQy6tLXqsV1CbuNqsSb
…t back Two findings from review, both confirmed directly in the code. P0 — the C1/C2 verifier read a proxy, not the thing. IdentityGate.load() observed harness_digest, workload_manifest_sha256 and python_reference_commit, and armed on any JSON carrying those three values. They are values anyone holding this repository can compute in one line, so the gate proved the instrument was the instrument and called that a freeze: it answered "is this the harness?" when the question is "did D7 happen?". There was no check for a C1 section, no payload commit, no git blob identity, no payload hash and no ratification. The control was no better — it built a temp JSON out of exactly those three fields and required the firewall to open. The freeze is now two objects, because a payload cannot name the commit that contains it: that sha would have to be inside the bytes hashed into it, which is the reason detached signatures exist. A payload carries kind/schema, a C1 section and a C2 section, and hashes over its own canonical body. A separately committed ratification names that body hash and the commit the payload is frozen at. Arming requires all of: the discriminator, both sections complete, the recomputed body hash, C1 matching observation, a ratification of that same hash, the payload's working-tree bytes being the blob at the named commit, and the ratification itself committed and unmodified. C2 is checked for presence and never read — its values are thresholds, and an artifact from this stage that quoted one would leak the number the module exists to keep out. Arming still costs no source patch, so the harness digest D7 freezes does not move when the gate arms. It now costs two reviewed commits instead of one text editor. What it still is not is a signature: anyone with write access could author both objects, so a declared key is verified against the ratification commit and a declared "none" is recorded as unsigned. Both stated in the note's limitations rather than implied away. P1 — the ladder floor claimed to be a phase it merely contains. core-usage was recorded as observability "direct" with the note that the interval "IS core startup", while the invocation it times starts the process, parses argv, writes a usage refusal and exits. It is now composed over process-startup-core plus a named cli-argv-refusal phase, and described as a lower bound. The consequence is corrected too: subtracting it from core-parse-refused does not leave parse, it leaves read+parse under an assumption this instrument never measures, so the difference is a derived bound. Controls: perf-gate-payload-identity damages one property of a real committed freeze at a time, thirteen cases, each required to be refused by the check that owns it. One case reformats the committed payload without changing its canonical hash, so only git object identity can catch it; if that case passes, the blob check is decorative. perf-gate-arms-by-data now builds a real freeze in a throwaway repository and additionally asserts no C2 value is recorded. perf-phase-attribution checks the rung table against itself and against the shipped report, which turns "re-record after changing a rung" into a gate. perf-notary-outside now counts git invocations as well as digests: the gate forks, and obligation (c) is violated by forking inside an interval just as much as by hashing inside one. The harness digest changes, so the calibration report no longer certifies the shipped instrument and is re-recorded in the next commit. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01CYLNQy6tLXqsV1CbuNqsSb
…ment The gate repair and the startup-attribution fix both change perf_baseline.py, so the harness digest moves from 34fa2717f24d to 5c1cdf0c69a9. The previous report certified an instrument that no longer exists; a report whose provenance names a tree that has been superseded is not evidence, it is a receipt. Re-recorded on a clean tree at 5d6dab4: 44 cells, gate dormant, no decisive workload timed, environment valid, 44/44 cells reproduced within the 0.35 noise floor the earlier run recorded. The report now also carries gate_requirements (the eleven properties a D7 freeze must satisfy) and gate_pinned, so the evidence describes its own verifier instead of asking a reader to trust the harness source. perf-phase-attribution checks the shipped report against the shipped rung table, so this re-recording is enforced rather than remembered: the control failed on the stale report and passes on this one. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01CYLNQy6tLXqsV1CbuNqsSb
…ong things CI was green on the previous head and proved exactly what it had been asked to prove. Four blockers, found by reading rather than by running. 1. The production D7 freeze could never have armed. The verifier took the payload's own directory as the repository, so it asked git for `<commit>:p022-263a-d7-ratification.json` for a file that lives at `docs/evidence/...`. `<rev>:<path>` resolves from the TREE ROOT; only a path starting `./` or `../` is read relative to the current directory. Every control passed because the throwaway fixtures put both objects at the repository root, where the wrong model is accidentally right — a control proving the verifier on a layout production never has. The root now comes from `git rev-parse --show-toplevel`, every object path is resolved against it, and the fixtures use the nested production layout. The ratification must also name the exact repository-relative path this instrument reads its payload from, rather than being free to point at any blob with the same bytes. 2. The advertised signed-ratification path was written wrong and could not have worked. It read `git verify-commit --raw` from stdout; git writes that raw status to stderr, and the helper discards stderr. A correctly signed commit would have verified cryptographically and then been refused for not naming its own key. No control caught it because no environment here holds a signing key. Rather than fix a path whose accept direction still could not be exercised, signed mode is withdrawn from the accepted contract: `"signature"` must be exactly `"none"` and the freeze is recorded as unsigned. An advertised path that is observably wrong is worse than an absent one. 3. The calibration timed failure and called it reproducible. With no .NET on PATH the launcher rungs exited 127 before doing anything, and nothing looked at exit codes: 12 of 44 committed cells reproducibly measured a command-not-found path, and the CI gate did not check either, since it asked for cells, raw timings, RSS, provenance and reproducibility but never whether a cell had done its work. Each rung now declares the exit codes that mean it did its job and a post-condition proving it, verified once, untimed, with output captured. A cell that fails is not timed at all, the calibration is refused, the CI gate reads the verdict, and reproducibility compares outcome identity as well as medians — two runs agreeing about the wrong path is the worst possible reassurance. The exit codes and evidence were probed against both engines before being written down. The first version of the floor check asserted the banner contains the word "usage"; it does not, and asserting instead of measuring is the exact habit this contract exists to break. 4. P1 was not closed. `launcher-extract` invoked the launcher with `--emit-facts` and was documented as "no core runs, so this isolates the frontend stage". The launcher copies the intermediate facts and then runs Stage 2 anyway, so the interval contained the whole pipeline while recording itself as launcher startup plus extraction — the same defect as the old `core-usage` label, one level up, hidden locally because rc 127 never reached Stage 2. Changing the production launcher is out of scope and a direct extractor process is not the launcher's extraction stage, so the rung is withdrawn and launcher-scoped extraction isolation is recorded unavailable. `frontend-extraction` stays measured as a member of `launcher-e2e`. Also: `perf-provenance-complete` now requires the committed report's harness and manifest digests to EQUAL the shipped ones, not merely to be present. Presence made a stale report tick every box. Controls: 16 damaged freezes, each matched to the check that owns it, including a basename `payload_path`, a ratification naming another file, and a signed ratification. New `perf-rung-outcome` proves 127 is never success, a silent zero-exit is refused, and the real production floor invocation is accepted. The harness digest moves again, so the calibration is re-recorded next commit — this time with a toolchain on PATH and every cell doing its rung's work. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01CYLNQy6tLXqsV1CbuNqsSb
…ally there The previous report was recorded with .NET absent from PATH — installed all along at ~/.dotnet, never on PATH, which is also why the environment fingerprint recorded an empty dotnet_sdk and nobody looked. Twelve of its forty-four cells were command-not-found exits, timed, summarised, and counted as reproduced. Re-recorded at 55aad67 on a clean tree with the toolchain present: 40 cells, every one timed and every outcome valid launcher-e2e now exits 0 having actually run the pipeline, not 127 dotnet_sdk 8.0.425 recorded rather than empty 40/40 reproduced within the 0.35 floor the run itself recorded no cell whose outcome changed between runs gate dormant, no decisive workload timed Forty rather than forty-four because launcher-extract is withdrawn: its four cells measured an interval that contained the whole pipeline while claiming to isolate extraction. Two controls now make this re-recording enforced rather than remembered. perf-provenance-complete requires the report's harness and manifest digests to equal the shipped ones, and perf-phase-attribution requires the report's rung table to equal the instrument's. Both failed on the stale report and pass on this one. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01CYLNQy6tLXqsV1CbuNqsSb
Caught by the control added in the previous commit, on the Windows CI leg: the same tree hashed af32f04ddbad on Linux and 51ef2ca2422a on windows-latest, where core.autocrlf rewrites source on checkout. One instrument, two identities. That matters more than a mismatched string. D7's C1 freezes the harness digest and #263-B has to run on both platforms, so a freeze would have armed on one and been refused on the other — and the refusal would have looked like tampering. This repository already knew the defect class. .gitattributes pins docs/evidence/*.json to LF because the mutation campaigns' definition hashes hit it first, which is exactly why the workload manifest digest matched across platforms while the harness digest did not. An attribute only governs files git checks out under it, though: a tree that predates the rule, a zip download or a different local config all still differ. An identity D7 will freeze should depend on content, so it is normalized in the digest rather than delegated to a checkout setting. sha256_file stays raw. The candidate binary's identity is its actual bytes, and a hasher that normalized them would be a different kind of wrong; the control asserts both directions. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01CYLNQy6tLXqsV1CbuNqsSb
…ed digest The digest fix changes the harness identity a third time, from af32f04ddbad to 69a3a3f8b074 — this time to a value that is the same on both platforms, which is the point. Re-recorded at 2feb496 on a clean tree with the toolchain present: 40 cells, every outcome valid and every cell timed, launcher-e2e exiting 0 having run the pipeline, dotnet_sdk 8.0.425, 40/40 reproduced with no cell whose outcome changed, environment not invalidated, gate dormant, no decisive workload timed. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01CYLNQy6tLXqsV1CbuNqsSb
One phase name covered two different actions. `cli-argv-refusal` bundled argv parsing with the usage-refusal message, so it appeared on core-parse-refused, core-full-human and core-full-sarif — none of which write a usage refusal. The instrument's own outcome contract already distinguishes them: usage-help versus door-refusal. The committed calibration carried the same wrong taxonomy, and perf-phase-attribution compared the report against the rung table and agreed they matched, which is how a schema stays perfectly self-consistent while saying something false. Two identical tables are one claim written twice, not evidence. The methodology note was already more accurate than the code: its table never put a usage refusal in the successful rungs. The vocabulary now separates three things. cli-argv-parse is in every real core invocation. cli-usage-refusal is only the floor rung's driver banner. ownir-door-refusal is only the strict door's version refusal. A refusal phase may appear exactly on the rung whose outcome evidence proves that refusal happened: cli-usage-refusal iff usage-help, ownir-door-refusal iff door-refusal, and neither on any rung that reaches a verdict. Stated as an equivalence in both directions, so a rung can neither claim a refusal it did not perform nor omit one it did. Applying the same rule the other way, launcher-e2e now also names process-startup-core and cli-argv-parse. The launcher spawns the core, so those actions are inside that interval and a list that means "this interval did these things" has to include them. Over-claiming was the reported defect; under-claiming is the same rule read in the other direction. The rules were verified against the defect itself: re-adding cli-usage-refusal to core-full-human makes perf-phase-attribution fail on both the equivalence and the reaches-a-verdict rule. Harness digest moves again, so the calibration is re-recorded next commit. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01CYLNQy6tLXqsV1CbuNqsSb
… taxonomy Splitting cli-argv-refusal into cli-argv-parse, cli-usage-refusal and ownir-door-refusal changes perf_baseline.py, so the harness digest moves to 84b3e98fc4d9 and the previous report certifies an instrument that no longer exists. Re-recorded on a clean tree at 889c285: 40 cells, all outcomes valid and all timed, 40/40 reproduced within the noise floor the run itself recorded, no cell whose outcome changed, environment valid, gate dormant, no decisive workload timed. Every cell's phase list now names only what its interval actually did: core-full-human no longer claims a usage refusal it never performs, and launcher-e2e names the core startup and argv parse that happen inside it. perf-phase-attribution fails on a stale report by design, so this re-recording is enforced rather than remembered. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01CYLNQy6tLXqsV1CbuNqsSb
The Windows leg failed on dfdce55 with two cells outside tolerance: core-full-human|rust|cal-facts-medium|warm moved 0.355 and core-usage|python|cal-facts-tiny|process-cold moved 0.362, against 0.35. Every outcome was identical between the two runs, every exit code matched, and the noise probe passed at both ends. That is not a defect in the harness. It is the harness correctly detecting that a hosted Windows runner does not hold still — the exact risk this PR already records as its most important finding for #263-B — and the gate converting that detection into "the instrument is broken" threw the signal away. Reproducibility now answers the two questions separately. The INSTRUMENT reproduces when both runs did the same work: same cells, identical outcomes. A disagreement there is a defect and still fails. The ENVIRONMENT reproduces when the timings agree within the policy; a disagreement there is a property of the machine, recorded with its numbers and allowed to be "no". The tolerance did not move. 0.35 is still 0.35, still derived from the noise floor the run itself recorded rather than chosen, and the failing cells are still named with their measured drift. Raising it to make a red run green would be a threshold fitted to a result, and it is the one number I am least entitled to pick. The shipped evidence gets stricter, not looser. A CI leg may now record "this environment is not measurement-grade" and pass; the COMMITTED calibration may not. perf-provenance-complete now requires the report of record to have reproduced on both counts and to carry a run that was never invalidated. Only the diagnosis of a hosted runner became more honest; the artifact that feeds D7 became harder to produce. Stated plainly because the timing invites suspicion: this changes a check that had just gone red. What it changes is which of two questions a disagreement answers, not the threshold that decides it. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01CYLNQy6tLXqsV1CbuNqsSb
…nstant Owner review of the proposed two-axis split. All three points were right. 1. Two axes were not enough, and the second was named for a conclusion rather than an observation. Calling a timing disagreement "the environment" asserts one of three possible causes without evidence: a contended runner, a genuinely variable workload, and a median estimate that is uncertain at this repetition count are indistinguishable from where the harness stands. The contract is now outcomes_reproduced, timings_reproduced and environment_valid, with an admissibility block requiring all three for a report of record. A CI leg exists to prove the instrument stands up and may pass while recording timings_reproduced false as a diagnostic; the committed calibration may not, and perf-provenance-complete checks all three on the shipped artifact. reproducibility.reproduced keeps the meaning it has always had — everything agreed — and remains the conjunction rather than a renamed subset. A red result must not disappear because a word changed profession. 2. The report's claim that the tolerance was "the noise floor recorded by the earlier run, not a chosen number" was false twice over. The recorded limit IS the hard-coded constant, so reading it back was reading the constant through a JSON detour, and nothing about it tightened on a quiet machine. Worse, one constant was bounding three different statistical quantities: dispersion within a probe, drift between the opening and closing probes, and a cell's median change between runs. They are now three named policies. The VALUE is 0.35 for all three and is deliberately unchanged — it is chosen, and what makes it admissible is that it was chosen before the runs it judges and is not adjusted after seeing which of them it rejects. 3. The sample-size argument was a hypothesis wearing a proof's clothes. More samples reduce the uncertainty of the median ESTIMATE; they do not narrow the intrinsic spread of the workload's distribution. "Observed IQR 0.404 exceeds the tolerance, therefore too few samples" does not follow. The escalation is therefore bounded in advance and recorded in the note: n=15 is the only permitted sizing step from the observed n=5, the tolerance does not move, the n=5 evidence is preserved as exploratory and non-admissible, and if n=15 also fails timings_reproduced the answer is to stop rather than to keep climbing until CI turns green. Also recorded, because it is a hole in the provenance rather than a tidy story: the specific n=5 report that failed was reverted from the working tree before the preservation rule existed, since a control forbids committing a non-reproducing calibration and reverting was the reflex. Its numbers survive in the note; the file does not. The exploratory evidence is a fresh n=5 pair taken under the shipped instrument, not the run that prompted the change. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01CYLNQy6tLXqsV1CbuNqsSb
… is intermittent The owner's rule was that failed n=5 evidence must survive the sizing decision it influenced. It nearly did not: the specific report that failed was reverted from the working tree before the rule existed, because a control forbids committing a non-reproducing calibration and reverting was the reflex. That hole is recorded in the note rather than papered over. What is preserved here is a fresh n=5 pair taken under the shipped instrument — and it changes the conclusion. It reproduced: admissible on all three axes, nothing outside tolerance, noise not invalidated. So the original n=5 failure was intermittent rather than systematic, and "n=5 is too few samples" is no longer merely unproven, it is weakly contradicted. That makes the sizing choice precautionary rather than demonstrated, and says so. The report of record will use n=15 because a report of record should carry the better-estimated median, which was true before anything went red and does not depend on n=15 having been the only green option. It was not: both counts reproduced. What n=5 showed is that its admissibility is luck-dependent on this machine — a reason to prefer more samples for the artifact, and not evidence about the workload's intrinsic spread. The file is named p022-263a-sizing-n5.linux.json so that it cannot match the glob identifying a report of record. Exploratory evidence that could be mistaken for canonical evidence is worse than none. CI keeps n=5: its question is whether the instrument stands up, it may record timings_reproduced false as a diagnostic, and tripling every leg's runtime would buy an artifact property that job does not produce. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01CYLNQy6tLXqsV1CbuNqsSb
Recorded at ad4c302 with the tree clean, harness digest 2b36de7a7bc9 matching the shipped instrument: 40 cells, all timed, all outcomes valid, admissible on all three axes — outcomes_reproduced, timings_reproduced and environment_valid. n=15 rather than n=5, and not because n=5 was red. A fresh n=5 pair reproduced cleanly and is preserved alongside as sizing evidence, so both counts are admissible on this machine. The record uses more samples because a report of record should carry the better-estimated median, which was true before anything went red. What the intermittent n=5 failure showed is that its admissibility is luck-dependent here, not that the workload's spread is wide. The first attempt at this report was discarded rather than committed: it carried tree_dirty true, because writing the sizing evidence had modified the tree before the run. A report of record with dirty provenance is not a report of record. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01CYLNQy6tLXqsV1CbuNqsSb
Four findings, all provenance rather than architecture. 1. A second knob moved. The authorised escalation was exactly repetitions 5 -> 15 with every other rule unchanged. The canonical report was recorded with --warmup 3 while the shipped default and the n=5 sizing run both use 2, so two parameters were turned after seeing which runs went red, not one. Warmup is now a named policy default, DEFAULT_WARMUP_DISCARDS, and a control refuses a committed calibration that does not use it. Repetitions may differ, because that is the sanctioned escalation; warmup may not, because changing it should be a deliberate policy edit in one place rather than a flag someone passes. 2. Only half of each pair was committed. "This pair reproduced" could not be recomputed from evidence: run B's raw samples were committed, run A's existed only inside the program that had already reached the verdict. reproduce() now records the earlier run's path, sha256, byte length, tree, harness digest and both calibration knobs; a control requires that file to be committed beside its mate, to hash to the recorded value, and to agree with it on digest, repetitions and warmup. A pair measured under different settings is not a reproducibility check. 3. The corrected truth about the tolerance was shipping next to the old lie. The implementation was right and the emitted report was right, but reproduce()'s own docstring still said the tolerance was "the noise floor the run itself recorded" and "tightens on a quiet machine", and the methodology note said the same. Both are false: the recorded limit is the constant, read back out of JSON and returning disguised as a measurement, and it never tightened. Source and note now say it is chosen, and that what makes it admissible is having been chosen before the runs it judges. 4. The rationale for 15 was the one the owner had already rejected. "A report of record should carry the better-estimated median" does not pick 15 — it picks 25, then 100. n=15 is used because a review given before any result was seen permitted exactly one bounded escalation with policies unchanged and a mandatory stop on failure. That it came back green is a result, not the reason. The note now says that; the historical commit messages are left alone, since they are the record of how the reasoning changed. The note also stops softening the destroyed evidence: the original failing n=5 artifact is lost, what is preserved is a fresh pair, and the original observations survive only as narrative, which is not evidence. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01CYLNQy6tLXqsV1CbuNqsSb
…tree Run A of each pair, recorded at 6c33e2a with tree_dirty false. Written outside the repository and copied in afterwards, because writing run A into the working tree is precisely what gave the previous attempt dirty provenance: run B then measured a tree whose sha no longer described what was on disk. An earlier attempt produced both pairs admissible, with matching hashes, digests and knobs, and was discarded anyway — both reports carried tree_dirty true, and a report of record whose tree_sha does not identify what was measured is not a report of record. That standard was set two rounds ago and it costs nothing to keep except a re-run. Neither filename matches the report-of-record glob p022-263a-calibration.*.json, so a run-A half can never be mistaken for a canonical artifact. Run B of each pair follows in the next commit, measured against these committed files on a clean tree. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01CYLNQy6tLXqsV1CbuNqsSb
…s withdrawn The preregistered escalation is spent and it did not work. n=5 1 of 40 cells outside REPRODUCIBILITY_MAX_MEDIAN_CHANGE, worst 0.478 n=15 3 of 40 cells outside it, worst 0.516 Both pairs reproduced their outcomes exactly and both environments were valid. Only the timings disagreed, and tripling the repetition count produced more disagreement rather than less. That is the predicted result, predicted before the run: more samples tighten the uncertainty of the median estimate, they do not narrow the intrinsic spread of the distribution being sampled. Reading a single n=5 failure as "too few samples" was a hypothesis, and this is the evidence against it. Across four observed n=5 pairs it has now failed twice and passed twice, so n=5 is intermittent rather than systematically insufficient. So there is no admissible calibration of record at this head, and none is manufactured. The permitted escalation was exactly one step with a mandatory stop on failure. There is no n=25, no n=45, and 0.35 does not move: a tolerance adjusted after seeing which runs it rejects is a threshold fitted to a result. p022-263a-calibration.linux.json is DELETED rather than left certifying a superseded instrument. What ships instead is all four halves of both failed pairs as exploratory evidence, each recorded on a clean tree, each run B naming its run A by path and sha256 so the verdicts can be recomputed rather than trusted. None of the four matches the report-of-record glob. Two control changes come with it. perf-phase-attribution now finds committed calibration artifacts structurally instead of by one hardcoded path — naming a single file meant the check went blind exactly when a record was withheld and the remaining evidence most needed checking. And pair integrity is enforced on every artifact carrying a reproducibility verdict, not only on reports of record, because exploratory evidence makes a two-run claim too. An observation recorded and deliberately not acted on: every cell that failed in either pair is a rust cell with a median between 2.5 and 9.2 ms, while the Python cells at 76-111 ms all held. One relative tolerance is applied uniformly across cells spanning roughly fiftyfold in duration and bites first where they are shortest. Whether the policy should be magnitude-aware is a measurement-design decision for the owner, taken with these numbers visible. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01CYLNQy6tLXqsV1CbuNqsSb
Run A and run B diverged on tree_sha AND python_reference_commit in both committed pairs. A was recorded at 6c33e2a, then COMMITTED, and B was measured afterwards — and python_reference_commit() is git rev-parse HEAD, so committing the evidence moved the recorded reference identity while ownlang had not changed at all. That is the single-reference rule broken by the procedure rather than by the code. The pair-integrity check did not catch it. It compared harness digest, repetitions and warmup, and said nothing about the reference or the candidate binary. Those pairs happen to share one candidate, but nothing required it, so a future pair could have compared two different binaries and still reported reproduced true. PAIR_IDENTITY_FIELDS now names every axis two halves must share: tree_sha, tree_dirty, python_reference_commit, workload_manifest_sha256, harness_digest, calibration_repetitions, warmup_discards, candidate_sha256, candidate_bytes. reproduce() REFUSES to produce a verdict when any of them differ, so a mismatched pair cannot be made rather than merely being detectable afterwards; the control is the second lock on evidence that already exists. The procedure is corrected with it: both halves measured on one clean source commit, written outside the repository, verified, then copied in and committed together. A does not need to be committed before B is measured — only by the time the evidence ships, which is a different moment. The existing failed pairs are re-recorded under this procedure in the next commit. Their qualitative finding is expected to survive; their provenance currently does not satisfy the contract they are meant to support. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01CYLNQy6tLXqsV1CbuNqsSb
…cedure Both halves of both pairs measured on ONE clean source commit (b64ecee), written outside the repository under their final committed names, verified on all nine PAIR_IDENTITY_FIELDS, then copied in and committed together. tree_sha, python_reference_commit, candidate sha256 and every other axis now agree between A and B, which the previous pairs did not manage. Two corrected attempts were needed, because earlier_run.path records the basename it was handed: scratch files under working names left run B pointing at a file that would never ship. The scratch names are now the final names. The results disagree with each other, and that is the finding. first corrected attempt n=5 passed, n=15 passed second corrected attempt n=5 passed, n=15 FAILED (one cell, 0.361) Running tally on this machine: n=5 has failed twice and passed four times; n=15 has failed twice and passed twice. Neither count reproduces reliably and which one "works" depends on when it ran. So no calibration of record is restored. Re-running was authorised to repair provenance, not to obtain a pass, and promoting whichever pair happened to succeed would be selection regardless of why the re-run happened. Both pairs ship as sizing evidence, pass and fail alike, and p022-263a-calibration.linux.json stays absent until the variance characterisation determines the measurement model. The four pre-contract artifacts are preserved byte-identical under docs/evidence/historical/, outside the glob the live controls scan. They fail the tightened pair check by construction — their earlier_run blocks predate the fields it requires, and their halves genuinely were measured under two different reference commits. Not deleted, because they record a real measurement and the defect that produced it; not left in place, because a control that must exempt named files is a control with a list of excuses. The move is byte-preserving and reversible if the owner prefers an explicit exemption instead. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01CYLNQy6tLXqsV1CbuNqsSb
The frozen §6 defines C1 as an immutable payload COMMIT and C2 as a detached attestation in a DESCENDANT commit. The implementation had used the same two letters for two sections inside one file, with "c2" holding threshold values. The word had changed profession, and an equivalent-looking scheme is not the frozen scheme. Five parts of one P0, all open on 388598e: 1. C1's observation omitted the mandatory #263-A tree sha and harness version. It now carries both, alongside the Python reference sha and tree, and the workload manifest digest, and every one is compared against observation. 2. The detached object had no payload_blob_sha at all. It now carries one and the verifier checks it against the payload's git blob at payload_commit_sha. 3. payload_sha256 was computed over a canonical JSON re-serialisation. §6 says the sha256 of the EXACT payload bytes, which is the only version that identifies a file rather than a file up to reformatting. A control now feeds the old canonical hash and requires refusal. 4. Nothing proved C2 lives in a descendant commit. The verifier now requires strict descent and a control builds C2 on a branch forked before C1. 5. python_reference_commit() was git rev-parse HEAD. That is not the reference; it is wherever the repository is standing. Every commit moved it, evidence commits included, which is how one pair recorded two reference identities while ownlang had not changed. At D7 it is worse: the freeze creates C1 and then C2, so a correct freeze would invalidate the binding it had just written. An alarm that counts its own installation as a break-in. The reference commit is now the last commit that touched the reference source, with python_reference_tree recording its content-addressed tree object beside it. Both are in PAIR_IDENTITY_FIELDS. Controls: 18 damaged D7 freezes, each matched to the check that owns it, built as real two-commit repositories. Notable cases — a canonical-JSON hash where an exact-byte hash is required; a payload reformatted after commit with the attestation updated to match the new exact hash, so only blob identity can catch it; and C2 on a sibling branch that forked before C1. One guard is defensive and untested, and is marked so in source and note: a same-commit state cannot be constructed with ordinary git, because the attestation would have to contain the sha of a commit whose sha depends on the attestation. Recorded as unconstructible rather than faked with a mocked git. Also applied: the owner's ruling on the "predicted result" wording. The earlier text claimed more than the evidence carries — nothing predicted that n=15 would fail on three cells or fail worse than n=5. The note now states the pre-run alternative it is consistent with, the narrower operational hypothesis it falsifies, and that it establishes no cause. IMPLEMENTED FROM THE OWNER'S QUOTATION OF §6. prompt-D7-preregistration-4.md is not in this session's possession, so fidelity to the full frozen text is unverified and the quotation may not be all of §6. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01CYLNQy6tLXqsV1CbuNqsSb
…entity repair The digest moved with 637a619, so the previous artifacts certify an instrument that no longer exists. Both halves of both pairs re-measured on that one clean commit, written outside the repository under their final names, verified on all ten PAIR_IDENTITY_FIELDS, then copied in and committed together. The tenth axis is new and is the point: python_reference_tree. The reference now reads 20879db while the tree reads 637a619 — a reference identity and a repository position are finally two different things, so an evidence commit no longer moves the reference and a D7 freeze no longer invalidates the binding it just wrote. Both pairs passed this time. Running tally: n=5 has failed twice and passed five times; n=15 has failed twice and passed three times. The note records that every re-record was forced by a provenance repair and never sought for a better answer, because the raw counts drift toward passing simply because repairs keep requiring fresh runs, and that drift is not evidence about the workload. No calibration of record is restored. Round 6 has not run, and a passing pair does not change what the record requires. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01CYLNQy6tLXqsV1CbuNqsSb
…three strings Two P0s from the owner's review of 2128ec4, both the same defect class this PR keeps re-learning: a check that observes something other than the thing it claims to verify. P0.1 — instrument_tree_sha = git rev-parse HEAD made a real freeze self-destroying. The lifecycle is S -> C1 -> C2 -> #263-B, so by the time the decisive run happens HEAD is at least C2 and can never equal the accepted tree S. A freeze satisfying that check would have to name a future content-addressed commit inside the commit that determines it. The twin defect was fixed one round earlier for python_reference_commit, with the argument spelled out in its docstring, and the identical sentence about instrument_tree_sha went unread three lines away. The anchor is now DECLARED by C1 as instrument_anchor_commit and PROVED by content: S exists, the instrument at S hashes to the harness_digest C1 froze, and S is an ancestor of C1. C1 is already a strict ancestor of C2 and C2 is found in HEAD's history, so S <= C1 < C2 <= HEAD falls out without asserting a position anywhere. Because the workload manifest is one of INSTRUMENT_SOURCES, proving the harness digest at S also proves the manifest at S. harness_digest() and harness_digest_at() share one formula, since the anchor proof compares them and two implementations would make their agreement meaningless. The fixtures could not see any of this: they committed a README and two JSON files to a throwaway repository while observe() read the real one, so the payload was built from the same HEAD it was compared against. _build_freeze now commits the real instrument sources at S, does ordinary work, freezes C1 and C2, and keeps committing; perf-gate-arms-by-data asserts anchor, C1, C2 and HEAD are four distinct commits. Reinstating "anchor must equal HEAD" makes a valid freeze be refused — verified by mutation. P0.2 — C1's protocol was three key names checked for presence, and the fixture proving that check worked armed the gate with three copies of the string "<frozen by D7, never read by the gate>". Two immaculate commits, three celebratory strings, decisive firewall open. The gate still never READS a threshold, but that was never the same question as whether a preregistration is there. C1 must now carry per cell: all four dimensions; bound, pass_fail_rule, inconclusive_band, comparison_statistic, rss_policy, allocation_policy; a repetition_ladder with initial_n, escalation_stages, transition_predicates, max_n and terminal_outcome; or an explicit not_applicable with a stated reason. Roll-ups must carry workload_class, phase and overall_g3. _present() inspects presence and container shape only — nothing compares, orders or records a value. Fixture cells use deliberately non-production placeholders, because a fixture that only passed with plausible numbers would prove the gate reads thresholds. Damaged D7 freezes: 18 -> 33, each matched to the phrase the check that owns it must produce. Verified by mutation: blinding the protocol verifier fails ten cases, not one. Also, self-found while unifying the digest readers: load_manifest() hashed RAW bytes while the harness digest normalized them — the Round 3 defect still alive in a sibling path, on a value C1 also freezes. Normalization now lives in one function. The value is unchanged on Linux, so this closes a Windows exposure without moving any number. Also: the note's pair-identity paragraph listed nine axes and omitted python_reference_tree, which the code has carried since the reference identity was separated from HEAD. It is generated from the tuple now and reads ten. The owner has read prompt-D7-preregistration-4.md and closed the previous "§6 fidelity unverified" caveat. The cell and roll-up key names here come from the owner's review message, not from the frozen file, which is still not in this session's possession. The gate does NOT check cell-set completeness: that needs the decisive cell population, and a completeness check written from a guess would be a threshold decision wearing a schema's coat. ruff 0, mypy Success (43 files), selftest 0, controls 13/13, run_tests 0. The harness digest moves with this commit, so both sizing pairs are stale and are re-recorded separately. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01CYLNQy6tLXqsV1CbuNqsSb
CI failure on aaf0d42, ubuntu leg: OSError: [Errno 39] Directory not empty: '/tmp/perf-d7-anchor-not-ancestor-opknvimr/repo/.git' Mine, and new. `git commit` spawns `gc --auto`, which keeps writing inside .git after the commit returns; `TemporaryDirectory` teardown then races it and rmtree dies. The anchor fixtures added in the previous commit commit more often than the old ones — an orphan branch, a drift commit, post-freeze work — which made the race likely enough to bite on a runner. It did not reproduce locally, which is the whole character of the thing. Fixed at the cause: fixture repositories now run with gc.auto=0 and maintenance.auto=false. They live for one check and are deleted, so they have nothing to gain from housekeeping and everything to lose from a background writer. Also ignore_cleanup_errors on the two fixture temp directories. Disposing of a throwaway repository is not a property under test, and it must never be able to read as a gate failure — the check that failed here had already passed. Verified: 4 consecutive runs of the control file, no leftover fixture directories, no stray git processes. ruff 0, controls 13/13. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01CYLNQy6tLXqsV1CbuNqsSb
…pair The anchor and protocol repairs move the harness digest, so every pair recorded before them describes an instrument that no longer exists. Both are re-recorded at f8e7112 on a clean tree, both halves written outside the repository under their final committed names, run B naming run A by path and sha256. Both pairs are admissible: 40 of 40 cells timed, no cell refused, outcomes_reproduced, timings_reproduced and environment_valid all true. Tally after this round: n=5 has failed twice and passed six times; n=15 has failed twice and passed four times. Four consecutive passes are recorded and explicitly NOT banked. The last two re-records were forced by provenance repairs, not sought for a better answer, but a tally where passes outnumber failures three to one is exactly where selection starts to look like evidence. The failures stay in the ledger and in docs/evidence/historical/ rather than being summarised away, and p022-263a-calibration.linux.json stays absent. Note also corrected: the pair-identity paragraph is generated from the tuple and reads ten axes, not the nine it listed while the code carried python_reference_tree. One recording defect of mine, caught by the instrument rather than by me: the first re-record refused 8 cells and then 4. I passed --candidate as a relative path, so OWEN_RUST_CORE did not resolve for the launcher surface and those invocations exited 127 or 2. The rung outcome contract from round 2 refused to time them instead of reporting a command-not-found path as a measurement. With an absolute path it is 40 of 40 again. The environment was never at fault. ruff 0, mypy Success (43 files), selftest 0, controls 13/13, run_tests 0. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01CYLNQy6tLXqsV1CbuNqsSb
…l universe Both halves of the owner's P0.2, found on 055b8e4. P0.2a — not_applicable: true was ACCEPTED, and the PR description said it was refused. _present() ended in `value is not None`, and in Python a bool is an int, so _present(True) and _present(False) were both True. Both {"not_applicable": true} and {"not_applicable": false} passed a check whose whole purpose was to demand a stated reason. I wrote the claim that a bare true was refused and never tested it — an untested assertion about a guard, published, in the same document that argues a green gate proves only what it was asked to prove. Fixed as a class, not an instance: _present() rejects bool outright (checked before int, since a bool is one), so a yes/no can never stand in for a preregistered value in ANY field. not_applicable must additionally be a non-empty string — a reason in words. A cell carrying not_applicable alongside any rule field is refused too: applicable and inapplicable at once is not a decision. P0.2b — the verifier checked every cell it was handed and never asked whether it had been handed them all. D7 says "for every applicable (phase x workload-id x platform x cold/warm regime) cell", so a payload with one immaculate cell and three roll-up strings passed. The old fixture demonstrated the hole instead of catching it: _valid_payload carried three synthetic cells and armed the gate. expected_d7_cells() enumerates 11 phases x 12 canonical decisive workloads x 2 platforms x 2 regimes = 528 cells, from the frozen taxonomy and the frozen manifest. Missing, unknown, or duplicated dimension tuples are all refused, and each expected cell needs a complete rule or an explicit not_applicable with a reason. The alias earns no second vote: large-solution-control is alias_of oss-ShareX.sln, declared under its own id because #262's gate names it and recorded as an alias so it is never counted twice in a denominator. canonical_workload_id() resolves it, and a payload naming both for the same phase/platform/regime is refused as a duplicate. A completeness check that missed this would have turned a double-counting safeguard into a double-counting machine. The universe covers BOTH platforms wherever the gate runs. #263-B runs on both, so a Linux gate demanding only Linux cells would let half the experiment through unfrozen. This selects nothing. No budget, no N, no tolerance; no value is read, compared or recorded. Applicability is expressed by the payload — the gate only proves the owner decided something about every cell. The previous note argued completeness "needs the decisive cell population" and recorded it as not implementable. That was wrong, and wrong in the convenient direction: the brief permits enumerating, hashing and availability-checking decisive workloads before D7 and forbids only performance exposure, and it has to, because a decisive population unknown before D7 could not be written into D7. Damaged D7 freezes: 33 -> 39, with catchers for boolean-true N/A, boolean-false N/A, N/A beside a rule, a missing cell, an unknown cell and the ShareX duplicate. Mutation-verified in both directions: restoring the old _present arms the gate on both boolean N/A cases, and skipping the completeness block arms it on all three cell-set cases. ruff 0, mypy Success (43 files), selftest 0, controls 13/13, run_tests 0. The harness digest moves again, so both pairs are re-recorded separately. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01CYLNQy6tLXqsV1CbuNqsSb
The N/A and completeness repairs move the harness digest again, so both pairs were stale. Re-recorded at 43b6a4a on a clean tree, both halves outside the repository under their final committed names. n=15 reproduced. n=5 did NOT: one cell, core-full-sarif|rust|cal-facts-tiny|process-cold, at 0.508 against the 0.35 policy. Both runs are valid, 40 of 40 cells timed, outcomes reproduced in both pairs, environments valid; only the timings axis failed, on that one cell. The failing pair ships exactly as recorded. It is NOT re-run: obtaining a passing partner for a failed run is the single move this instrument exists to make impossible to hide, and the permitted repetition counts are still only 5 and 15 with nothing to escalate to. The previous commit's note warned that four consecutive passes were exactly where selection starts to look like evidence. The very next attempt failed. That is not a vindication of the caution so much as a demonstration of what it was about: the streak was never evidence, and neither is its ending. Tally: n=5 has failed three times and passed six; n=15 has failed twice and passed five. Five DISTINCT cells have now fallen outside tolerance at least once, and every one of them is a rust cell with a median under 10 ms — the newest at 2.93 ms. The magnitude observation from earlier rounds is unchanged and still deliberately not acted on: whether the relative tolerance should be magnitude-aware is a measurement-design decision for the owner. Still no calibration of record. p022-263a-calibration.linux.json stays absent. ruff 0, mypy Success (43 files), selftest 0, controls 13/13, run_tests 0. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01CYLNQy6tLXqsV1CbuNqsSb
P0.2c from the owner's review of 34f7ef3. Round 9's completeness check was mechanically sound and proved the completeness of the wrong set. expected_d7_cells() enumerated sorted(PHASES). PHASES is the ATTRIBUTION taxonomy: it answers what a timed interval contains, so a rung can be described honestly. It is not D7's vocabulary, and using one as the other was wrong in both directions at once. It demanded a preregistered rule for cli-argv-parse, cli-usage-refusal, ownir-door-refusal and frontend-extraction on every workload, platform and regime. It omitted the end-to-end C# run entirely — one of the metrics the G3 verdict weighs most — because that is a rung, not a part, and has no entry in PHASES at all. The code had already said this and was not listened to. The accept fixture keyed its not-applicable slice on phase == "frontend-extraction", auto-excusing all 48 of those cells: the instrument admitting they were never D7's business while the universe went on requiring them. A check that manufactures an obligation and then ceremonially forgives it is paperwork, not a check. D7_PHASES is now a separate constant — process-startup-core, process-startup-launcher, ownir-parse, bridge-lowering, analysis, render-human, render-sarif, end-to-end-csharp — and the relationship between the two vocabularies is data rather than coincidence. D7_NON_METRIC_PHASES records each attribution component that is not a gate together with the reason it is not. D7_METRICS_WITHOUT_PHASE records end-to-end-csharp as a metric whose interval no single attribution phase names. d7_vocabulary_problems() enforces that every attribution phase is either promoted to a metric or explicitly excluded WITH a reason; neither is a default, so a phase added to PHASES later cannot silently join or silently miss the D7 universe. It runs in --selftest as well as in the new perf-d7-phase-universe control, because the failure mode here is drift rather than a wrong line. Universe: 8 metrics x 12 canonical decisive workloads x 2 platforms x 2 regimes = 384 cells, down from a 528 that was neither complete nor correct. Controls 13 -> 14; damaged freezes 39 -> 40 with proto-attribution-phase, which dresses cli-argv-parse as a G3 metric and is refused as a cell D7 does not cover. Mutation-verified three ways: dropping end-to-end-csharp, promoting cli-argv-parse, and rebuilding the universe from PHASES are each caught, and the first two fail --selftest as well. NAMES ARE PROVISIONAL, and that is the actual lesson. D7_PHASES comes from the owner's review and the live #262 performance-gate list, not from the frozen file. The mapping between these ids and the frozen #262/#263-A vocabulary needs one explicit ratification, because this defect is precisely what happens when an implementation label is left to become a normative contract by default. RSS and allocation deliberately stay per-cell policies rather than becoming a phase axis; the frozen schema already models them that way. ruff 0, mypy Success (43 files), selftest 0, controls 14/14, run_tests 0. The harness digest moves again, so both pairs are re-recorded separately. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01CYLNQy6tLXqsV1CbuNqsSb
Separating D7_PHASES from PHASES moves the harness digest, so both pairs were stale again. Re-recorded at 843d6fe on a clean tree, both halves outside the repository under their final committed names. Both pairs reproduced: 40 of 40 cells timed, no cell refused, outcomes and timings reproduced, environments valid, both admissible. Tally: n=5 has failed three times and passed seven; n=15 has failed twice and passed six. Six attempts in, n=5 reads pass, pass, fail, pass, fail, pass. Nothing about that is a trend, and the ledger exists so no revision of the note can quietly turn it into one. Five distinct cells have fallen outside tolerance at least once across the whole history, every one a rust cell with a median under 10 ms. Still no calibration of record. p022-263a-calibration.linux.json stays absent, and the failing pairs stay in the ledger and in docs/evidence/historical/. ruff 0, mypy Success (43 files), selftest 0, controls 14/14, run_tests 0. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01CYLNQy6tLXqsV1CbuNqsSb
Two records, no source change. The harness digest is byte-identical (2d6e52fe4352…), so both committed pairs stay valid and nothing is re-recorded. RATIFICATION. The owner has ratified the eight D7 metrics and the four exclusions against the frozen prompt-263A-measurement-design-4.md, the final D7 brief and live #262, at exact head 0374793 whose CI completed success on both runs. Metrics: process-startup-core and process-startup-launcher as two, per the frozen text; ownir-parse, bridge-lowering, analysis; render-human and render-sarif as the two concrete surfaces of the frozen CLI/SARIF rendering item; and end-to-end-csharp as the whole user-visible rung. Not metrics: cli-argv-parse, cli-usage-refusal and ownir-door-refusal are attribution only, and frontend-extraction is diagnostic population B excluded by frozen #263-A itself. The denominator is ratified with them: 8 x 12 x 2 x 2 = 384, with 12 canonical workloads because large-solution-control aliases oss-ShareX.sln. Ratifying 384 does not mean 384 numeric budgets — an explicit not_applicable with a reason stays a legitimate decision. It means the owner must decide about each cell rather than lose it between two lists. The source comment still reads NAMES ARE PROVISIONAL, deliberately. At the commit that wrote it they were. Editing it now would move the harness digest, stale both pairs and force an eleventh re-record — relabelling the fire extinguisher and re-calibrating the laboratory because of it. The ratification binds in the ledger and to the exact head it was given at. PREREGISTRATION. docs/notes/p022-263a-round6-preregistration.md fixes the Round 6 analysis before any Round 6 data exists: a synthetic deterministic helper that is neither Rust nor Python product code, fresh process per iteration, eight target durations from 1 to 128 ms, n in {5,15}, warmup 2, five independent A/B sessions per (scale, n) — 160 measurement runs. Every raw observation, MAD, IQR, absolute and relative median shift, scheduler metadata, CPU affinity and a context-switch signal are recorded. No pass/fail tolerance is computed or applied and 0.35 is not consulted. Three outcomes are named in advance with what each licenses. O1, absolute drift roughly constant while relative shift grows at short durations, licenses DESIGNING a v2 envelope |delta| <= A + R*t with A and R from the metrology dataset only. O2, a stable ladder at 2-8 ms while real Rust invocations wander, forbids an absolute floor. O3, neither shape, is inconclusive and stops. O3 is named deliberately: without it an ambiguous dataset gets read as whichever of O1 or O2 the reader was hoping for. Forbidden and written down: deriving A or R from the Owen cells, using 0.508 or any observed failing value toward a threshold, an absolute floor under O2, choosing T_min/N_min/N_max, moving 0.35 or warmup, any repetition count other than 5 and 15, and every decisive surface — no decisive workload or timing, no resource exposure of decisive work, no D7 threshold or budget, no C1/C2 freeze, no #263-B, no Stage 3. Existing calibration history may be read as calibration evidence, which frozen #263-A permits as an input to sizing the instrument and the owner has ratified. The observation that all five historical failures sit on rust cells under 10 ms may be investigated; it may not become a number in a D7 acceptance rule. ruff 0, selftest 0, controls 14/14. No measurement has been taken. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01CYLNQy6tLXqsV1CbuNqsSb
The Round 6 apparatus, committed before it has produced a number.
scripts/round6/spin.c is the synthetic helper: deterministic work, no I/O, no
allocation, neither Rust nor Python product code. scripts/round6/metrology.py
drives the preregistered ladder — setup pass fixes one iteration count per rung
and records the helper's sha256; measurement pass runs 8 rungs x {5,15}
repetitions x 5 independent A/B sessions and records every raw observation, MAD,
IQR, absolute and relative median shift, plus CPU affinity, load average and
context-switch counters either side of each session.
AMENDMENT 1 to the preregistration, made before any measurement pass ran and
recorded there with its reason: the helper is compiled rather than interpreted.
The instrument's timed interval includes process spawn, and on this machine
python3 -c pass has a floor of 11.8 ms against /bin/true at 1.14 ms — so a
Python helper could not reach the 1, 2, 4 or 8 ms rungs at all. The compiled
helper's floor is ~1.25 ms. The 1 ms rung therefore sits at or below the spawn
floor; it is kept and flagged rather than dropped, because it bounds what a
relative tolerance can mean at that scale, which is close to the question this
round asks.
Timing goes through perf_baseline.Harness._run_once deliberately — that is the
instrument's own interval, perf_counter_ns around Popen and wait4. Round 6
characterises THAT interval, so a re-implementation here would characterise a
different one and let the resemblance do the arguing. Writing one formula twice
is the defect class this PR keeps finding.
No verdict, no tolerance, 0.35 never consulted, no decisive workload spawned or
named. The firewall is never invoked because no workload is passed to it.
ruff 0. No measurement has been taken.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CYLNQy6tLXqsV1CbuNqsSb
The setup pass fixes one iteration count per rung and records it, so the helper is not re-deriving its own duration on every run and measuring a moving target. Run at 8c61a7b on a clean tree; the counts were written into the preregistration BEFORE the measurement pass started, which is what that document promised. Helper f937b36bd4da, spawn floor 1.332 ms, ~690k iterations per ms. | target | iterations | achieved | |--------|-----------:|---------:| | 1 ms | 0 | 1.448 ms (at/below spawn floor) | 2 ms | 460,280 | 2.013 ms | 4 ms | 1,838,383 | 4.090 ms | 8 ms | 4,594,589 | 8.499 ms | 16 ms | 10,107,001 | 16.496 ms | 32 ms | 21,131,826 | 31.691 ms | 64 ms | 43,181,474 | 63.326 ms | 128 ms | 87,280,771 | 127.313 ms The ladder tracks its targets from 2 ms upward. The 1 ms rung IS the spawn floor: it is kept and flagged rather than dropped, and read as a bound rather than a duration, because what a relative tolerance can mean at that scale is close to the question this round asks. The artifact lives in docs/evidence/round6/ rather than beside the calibration reports. That keeps it outside the flat report-of-record glob by construction instead of by key shape — the live controls would skip it either way, but a directory boundary does not depend on a schema staying the way it is today. It is metrology, not a calibration report, so it does not belong in that glob at all. No verdict, no tolerance, 0.35 not consulted, no decisive workload involved. ruff 0, selftest 0, controls 14/14. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01CYLNQy6tLXqsV1CbuNqsSb
80 A/B sessions: 8 rungs x {5,15} repetitions x 5 independent sessions, every
raw observation kept, run at 8c61a7b on a clean tree with helper f937b36bd4da.
OUTCOME O2, read off the preregistered table rather than chosen after the fact.
The synthetic ladder is stable exactly where the real Rust invocations wander,
so an absolute floor must NOT be introduced and the cause is elsewhere.
A measurement-scale effect is real but small: relative shift rises from ~0.006
at 8-16 ms to ~0.04-0.05 at 1-2 ms, a fixed cost over a shrinking denominator.
That much of O1's description holds.
O1 is nevertheless NOT licensed, because its precondition fails. It required
absolute drift roughly constant across the ladder; absolute drift instead grows
from 0.056 ms to 1.126 ms at n=5, about twentyfold.
The decisive comparison, which is what the round was built to make. In the
2-10 ms band — every rust cell that has ever failed lives there — 27 A/B
sessions of the ladder produced a median relative shift of 0.0131, a worst
relative shift of 0.1564, and a worst absolute drift of 0.3203 ms. The smallest
drift any failing cell required is 1.23 ms (0.396 on a 3.12 ms median); the
largest is 4.41 ms. Worst-case ladder against minimum-case failure is short by
3.9x. The ladder never crossed 0.35 at any rung at either repetition count.
Short duration alone therefore does not produce the instability. Something the
real invocations do, and this helper does not, is producing it.
Recorded because someone will eventually be tempted the other way: reading the
table charitably as O1 and fitting A + R*t anyway yields an envelope TIGHTER
than 0.35 at these durations, rejecting the Rust cells more often rather than
fewer. This dataset cannot be used to loosen anything.
BOUNDARY OF THE CONCLUSION, stated rather than buried. The helper is a small
compiled binary doing pure arithmetic — no file reads, no parsing, no large
dynamic image. own-cli is ~1.8 MB and opens files. The ladder isolates DURATION
and does not isolate WHAT A REAL INVOCATION DOES; dynamic loading, page-cache
state and file I/O are unexamined and are the obvious next suspects. O2's
conclusion is "the cause is elsewhere", and a reader is entitled to know how
much elsewhere was actually searched.
Unchanged, per the preregistration and standing rulings: no absolute floor, no
v2 envelope, 0.35 unmoved, no T_min/N_min/N_max, no threshold derived from any
number here, no calibration of record restored.
ruff 0, mypy Success (43 files), selftest 0, controls 14/14, run_tests 0.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CYLNQy6tLXqsV1CbuNqsSb
Docs only. The harness digest is unchanged (2d6e52fe4352), the dataset is untouched, and Round 6 is NOT re-run. Its provenance is sound and re-running a measurement to obtain a tidier verdict is the move this instrument exists to prevent. The owner rejected the formal O2 verdict and was right on all three counts. 1. The reading consulted 0.35, which the preregistration forbids in as many words: "No pass/fail tolerance is computed, applied, or reported ... 0.35 is not consulted." The withdrawn reading observed that the ladder "never once crossed 0.35" and argued from it. The measurement driver never reads 0.35, so the dataset is uncontaminated — the interpretation failed, not the collection. 2. O1 and O2 were never operationalised as numbers. "Roughly constant" and "stable at 2-8 ms" were English, not decision rules. The quantities that ended up deciding — twentyfold drift growth, 0.3203 ms against 1.23 ms, a factor of 3.9 — were all chosen after seeing the data. They are legitimate descriptive findings and they are not a preregistered boundary, so "O2, read off the preregistered table" claimed more than the plan could support. 3. The A + R*t fit was performed in a branch that did not license it. The plan permits deriving A and R only under O1; the reading selected O2 and reported the fit anyway. No coefficients were published, but the analysis ran where the plan disallowed it. Formal outcome is therefore O3, inconclusive, in the owner's words: the dataset is descriptively O2-like, but the preregistration did not operationalise the O1/O2 boundary tightly enough to support a formal O2 verdict, so the preregistered outcome is conservatively O3; the observed separation is retained as exploratory calibration evidence and may motivate a separately preregistered follow-up. The numbers survive as observations licensing nothing: drift grows ~20x across the ladder so O1's own written precondition does not hold; the scale effect is visible (relative ~0.006 at 8-16 ms rising to ~0.04-0.05 at 1-2 ms); in the 2-10 ms band the ladder's worst absolute drift was 0.3203 ms while the historical failures implied 1.23 to 4.41 ms. That separation is a reason to keep looking, not a verdict. The envelope fit is quarantined under an explicit POST-HOC / NOT PART OF THE PREREGISTERED OUTCOME / NOT POLICY INPUT heading rather than deleted, because deleting an analysis that was actually performed is its own dishonesty. NEXT-SUSPECT LIST CORRECTED, from the owner. It led with file I/O and should not have: core-usage|rust|cal-facts-tiny|process-cold is one of the historical failures and that rung never opens an OwnIR document at all — it starts the Rust core, parses argv, writes a usage refusal. Input-document I/O is NOT a necessary condition for the instability, and this round's own evidence said so while the reading looked past it. What survives: executable/library mapping and page faults, dynamic loader/runtime startup, scheduler context switches and CPU migration, CPU-time against wall-time, and the output/refusal path. core-usage is where a follow-up should start — the minimal real Rust process path, already a witness, with document parsing removed from the picture. No new measurement is authorised, and none was taken. ruff 0, selftest 0, controls 14/14. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01CYLNQy6tLXqsV1CbuNqsSb
…ssertion Step 2 of the owner's order. The harness digest is unchanged (2d6e52fe4352) — the control lives in tests/, which is not an instrument source, so no pair is staled and nothing is re-recorded. python_reference_commit and python_reference_tree content-address exactly ownlang/. That identity is only honest if the reference's behaviour is confined to what it names; otherwise the reference could change behaviour without changing its own sha, which is the defect class of an identity that names the checkout. Settled by observing. An audit hook records every file the reference OPENS and every process it spawns, on both render surfaces. Result: ZERO repository files opened outside ownlang/, no process spawned, 18 ownlang modules loaded. Three further routes checked and closed. Static imports: ownlang imports the standard library and itself, nothing else in the tree. Config: config.py states it has no auto-discovery, no environment variables and no per-path overrides — a config enters only via an explicit --config, which no rung passes, so there is no consulted-if-present file to slip past an audit that only sees opens. Subprocesses: the fix_* repair pipeline does spawn, but none of it loads on the measured path. perf-reference-boundary keeps this proven rather than asserted. Controls 14 -> 15. Mutation-verified both ways: a probe that opens pyproject.toml fails it, and one that spawns /bin/true fails it. ONE RESIDUAL, RECORDED AND DELIBERATELY NOT FIXED. OWNLANG_DEBUG is read by ownlang/__main__.py and the instrument passes ambient environment through unchanged. It affects only the internal-error path — on a successful run the branch is never taken, so it cannot alter analysis output — and when it fires it returns 70, which no rung declares, so the outcome contract refuses the cell rather than timing it. Fail-closed, but still an uncontrolled input the reference identity does not cover. Pinning it edits perf_baseline.py, moves the digest and stales both pairs; re-recording is not currently authorised, so this is worth batching with any other digest-moving change rather than spending a re-record on it alone. ruff 0, mypy Success (43 files), selftest 0, controls 15/15, run_tests 0. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01CYLNQy6tLXqsV1CbuNqsSb
…v input
Step 3 of the owner's order, batched with the OWNLANG_DEBUG residual from step
2. Both move the harness digest, so doing them together costs one staleness
event instead of two.
MYPY. scripts/perf_baseline.py was outside mypy's file list, so "Success, 43
files" never covered the file that produces the numbers. It is in the list now —
44 files — and passes --strict.
The 42 errors it exposed were almost entirely one root cause: JSON-parsed values
are object, and the code called .get(), int() or float() on them behind `x or
{}` idioms with type: ignore comments that had drifted to the wrong error codes.
Four narrowing helpers — _as_obj, _as_list, _as_int, _as_float — say it once, to
the reader and the checker at the same time.
They narrow fail-closed. A non-object where an object belongs becomes empty, so
the caller's own completeness check refuses it rather than raising an
AttributeError three frames from the cause. _as_int and _as_float reject bool
for the same reason _present does: a yes/no is not a measurement.
Six errors were not narrowing at all. ctypes.windll and Popen._handle exist on
Windows and are absent from the stubs mypy checks against on Linux; those carry
targeted type: ignore[attr-defined] comments, one API each with the reason
recorded, rather than a blanket suppression that would also hide the next real
defect.
ENV PIN. REFERENCE_ENV_PINNED_UNSET names the environment inputs the reference
reads that its content-addressed identity does not cover, and every measured
invocation clears them. OWNLANG_DEBUG is the only current member. It affects
only the internal-error path so it cannot change analysis output on a successful
run, but an identity that content-addresses ownlang/ does not cover an
environment variable, and the same reference source should not behave two ways
depending on the shell it was launched from.
NO CONTROL CHANGED AND NONE WAS WEAKENED: 15/15 throughout, full suite green at
every step. That is the only reason to believe a mechanical rewrite of this size
did not quietly alter behaviour.
BOTH PAIRS ARE NOW STALE AND ARE NOT RE-RECORDED. The digest moved from
2d6e52fe4352 to 6713e7300c7c, so all four committed halves describe an
instrument that no longer exists. They are left in place and the note says so:
re-recording is a measurement, and new measurements are not currently
authorised. Stale evidence that admits it beats fresh evidence nobody asked for.
ruff 0, mypy Success (44 files), selftest 0, controls 15/15, run_tests 0.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CYLNQy6tLXqsV1CbuNqsSb
…, not run
Step 5 of the owner's order. Docs only; the harness digest is unchanged
(6713e7300c7c) and no measurement exists.
THE QUESTION, narrowed by the owner's correction. Round 6 was formally O3 and
descriptively showed short duration alone looks insufficient. The follow-up must
NOT lead with file I/O: core-usage|rust|cal-facts-tiny|process-cold is a witness
and that rung never opens an OwnIR document — it starts the Rust core, parses
argv, writes a usage refusal. So core-usage is the minimal real Rust process
path, already a witness, with document parsing removed from the picture.
DESIGN. Round 6 varied duration and held shape fixed; Round 7 does the opposite.
Three arms at one duration: A, the Round 6 C helper at its 2 ms rung; B, the
same helper with byte-identical work padded to own-cli's ~1.84 MB, isolating
image size and mapping; C, own-cli ownir with no arguments, the real witness.
D(C)-D(A) is the excess to explain, D(B)-D(A) the part from image size, and
D(C)-D(B) the part from everything else the real binary does — loader symbol
resolution, runtime startup, argv handling, the refusal write. Ten independent
A/B sessions per arm and repetition count, n in {5,15}, warmup 2: 120 runs.
QUANTITATIVE BOUNDARIES, which is what Round 6 lacked. P1 through P5 are
arithmetic on D(arm,n), the median across sessions of |median_B - median_A|:
P1 image dominates if D(B) >= 3*D(A) and D(C) <= 1.5*D(B); P2 not image if D(B)
<= 1.5*D(A) and D(C) >= 3*D(A); P3 both if D(B) >= 3*D(A) and D(C) >= 2*D(B);
P4 witness did not reproduce if D(C) <= 1.5*D(A); P5 otherwise. A rule must hold
at BOTH repetition counts or the outcome is P5. Mechanism attribution is
likewise fixed in advance as ratios over utime+stime, wall, context switches and
page faults against arm A.
Ratios rather than absolute milliseconds on purpose: an absolute boundary would
have to come from the failing Owen cells, which are forbidden as input, or from
a number invented to fit, which is worse. A ratio against a baseline arm
measured in the same sessions cannot import a threshold by accident.
DECLARED UPFRONT, NOT DISCOVERED LATER: Harness._run_once discards ru_utime,
ru_stime, ru_minflt, ru_majflt, ru_nvcsw and ru_nivcsw — exactly the fields that
separate the candidate causes. Round 7 needs it to return the full rusage. That
is a source change, it moves the harness digest, and it stales any pair recorded
before it. APPROVING THIS PLAN MEANS APPROVING THAT CHANGE AND THE RE-RECORD IT
IMPLIES; if the change is not wanted, the round cannot run as designed.
0.35 is not consulted anywhere, including in the reading. That was Round 6's
defect and it is not repeated. Reporting an outcome the table did not produce is
forbidden: suggestive numbers with no rule firing are P5 plus an exploratory
note, not a verdict.
Two design choices are left open for the owner rather than assumed: whether a
statically linked fourth arm should separate the dynamic loader from image size,
and whether both regimes are measured or only process-cold where every witness
lives.
Step 6 is a separate review and authorisation. Step 7 is execution. Neither has
happened.
ruff 0, digest unchanged, no measurement taken.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CYLNQy6tLXqsV1CbuNqsSb
Four decisions and two corrections, none of them mine to make: * full POSIX rusage in _run_once — YES * a fourth, statically linked arm — NO, and not carried as a footnote * arms A/B/C — kept * process-cold only, or both regimes — BOTH * run Round 7 now — NO Arm B claimed too much. It said it isolated "image size, mapping and page-fault cost"; Linux demand-pages an image, so padding that is mapped and never read need not fault in at all. B now claims the padded image footprint and its mapping metadata, D(B) - D(A) is stated as a LOWER bound, and a B that reproduces nothing reads as "a padding-only image effect is insufficient" rather than "image mapping and page faults are excluded". A four-part structural preflight (padding survived linking, sits in a PT_LOAD, matches arm C's mapped size to within a page, work bytes identical to arm A) is a stop condition rather than a description written afterwards. The outcome table was not mutually exclusive: A=1, B=3, C=1.4 fired P1 and P4 at once. Ten rounds of insisting a check must read the thing, and I preregistered the numbers and forgot to preregister the logic. The owner's ratified set replaces it, with P1 and P2 renamed to claim only what arm B can see. Every remaining pairwise overlap reduces to D(A)=0 or D(B)=0 exactly; a degenerate-zero guard is PROPOSED and flagged as NOT YET RATIFIED rather than folded in quietly. Also corrected: the draft said every witness lives in process-cold. One of the five, core-full-sarif|rust|cal-facts-small|warm, does not. Both regimes are measured and classified independently; a cold/warm split is reported as a cache-sensitivity signal, not collapsed into P5. The regimes are stated as the instrument implements them — cold 0 discarded iterations, warm 2, both a fresh process per iteration — instead of the draft's self-contradiction. Session halves are no longer also called A and B. Still a draft. Execution is not authorised. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01CYLNQy6tLXqsV1CbuNqsSb
_run_once reaped every child through os.wait4, which returns the kernel's
complete per-child rusage, and read one field of it. User CPU, system CPU, minor
and major faults, and voluntary and involuntary context switches were fetched
and dropped on the floor — and those six are exactly what separates "the work
took longer" from "the scheduler moved it" from "the image cost more to map".
Without them a variance round can only report that something got slower.
Four constraints, each a way this could have been made wrong:
* ru_utime/ru_stime are ROUNDED to nanoseconds. Truncation biases every
sample the same direction, and a systematic half-tick error survives
averaging in a way a symmetric one does not.
* the interval is unchanged in meaning: t0, Popen, wait4, elapsed. The
rusage is parsed after the clock stops.
* the two pre-existing lines between wait4 and elapsed stay where they are.
Moving them would TIGHTEN the interval, and a tightened interval silently
un-compares every future number against every recorded one.
* off POSIX all six are null WITH A REASON. A reported 0 minor-fault count
for a platform nobody asked would be the most confident possible lie, and a
median over such zeros would look exactly like a flat measurement.
The fields reach the cell too, not only the private driver: accounting
(summarized in ns for the durations and count for the counters),
accounting_unavailable_reason, and raw_accounting per iteration. Collecting the
accounting and then dropping it one layer up would be the same defect wearing a
different hat.
perf-child-accounting is the sixteenth control. It checks the claim that would
otherwise be invisible — that the parse sits outside the clock — by wrapping
os.wait4 so every rusage field costs 10 ms to read, then asking whether
elapsed_ns grew. Seven mutations, each caught by the check that owns it. The
last two needed a second attempt: both make measure_cell raise KeyError partway
through a cell, so the first version recorded its finding and then died before
reporting it, which in CI is a traceback naming the crash site and no FAIL line
naming the cause. Those checks are terminal now.
Harness digest 6713e7300c7c -> 562a7f7232da. The four committed sizing halves
were already stale and are still not re-recorded.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CYLNQy6tLXqsV1CbuNqsSb
The owner did the algebra properly and mine was wrong in both directions. I claimed every remaining overlap reduced to "A=0 or B=0". It does not. Overlap occurs only when A=0 AND B=0 together, and P1 is disjoint from every other rule unconditionally — its C > 1.5A clause excludes P4 outright, and every other intersection forces B=0 and then C=0, which contradicts that same clause. So "I checked every pair" was again a stronger claim than the checking behind it, which is the specific habit this PR exists to make expensive. The guard I proposed was correspondingly too wide. Keying it on B would have refused A=1, B=0, C=4 — a clean P2 where the padded arm shows no excess and the real binary does, one of the cleanest results this round can produce. The ratified rule keys on A alone, and not because of exclusivity: A is the multiplicative reference for 1.5A, 3A and 6A, and against a baseline of zero any positive drift is infinitely large. Definite, and metrologically degenerate. No epsilon. A tolerance like A < 0.01 ms would become an absolute threshold with no preregistered basis for its value, which is the move this round forbids. Exact zero is a structural degenerate case, not another knob. The exclusivity claim is no longer asserted anywhere: a control evaluates it. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01CYLNQy6tLXqsV1CbuNqsSb
One classifier, five controls, four structural checks. No clock anywhere in it, and none may be added: the calibration pass itself is not authorised. scripts/round7/classify.py is the ONLY implementation of P1-P5 and the zero-A guard. A second copy would be a second opinion and the reading could then choose between them. It is pure — numbers in, an outcome out — so it is exercised exhaustively before any measurement exists: 50,653 exact rational triples, in sixths so the ratified edges 3/2, 3, 2/3 and 2 land exactly on grid points, with Fraction arithmetic throughout because a binary-float 2/3 would decide the (2/3)B <= C edge by rounding direction rather than by the rule. The census runs twice. Against the real rules, which must overlap only at A=B=0; and against the previous draft's P1, which must be reported broken. A control that has only ever seen correct input is a control nobody has tested. Six mutations of the real classifier, each caught by the check that owns it: the old overlapping P1; the zero guard widened to A-or-B; the guard removed; ambiguity resolved by precedence instead of raising; P4's edge made strict; and an epsilon smuggled into the zero comparison. The second is the guard the owner rejected, and the control now refuses to let it back in. The crash-before-reporting defect recurred, in a file written after I recorded it. Removing the guard makes A=B=C=0 fire three rules, the classifier raises, and the control died on a traceback: non-zero exit, mutation scored as caught, no FAIL line naming anything. Reading the output rather than the exit code is what caught it both times. The preflight reads the ELF rather than a rendering of it — program and section headers parsed directly, and the padding located through st_shndx, which names its section, instead of by matching addresses, an inference that finds nothing at all for a non-allocated section whose address is zero. The four damage cases are built, not simulated; the first attempt at the unmapped one failed to produce any damage, which is why it was checked before being believed. B1-B4 all pass. Padding 1,626,024 bytes in .rodata inside the PT_LOAD at 0x2000; arm B maps 1,628,917 bytes against arm C's 1,628,892, +25 against a 4,096-byte allowance; main is 107 bytes and byte-identical in both arms, so B4 decided on raw bytes and the disassembly fallback was not needed. The apparatus is not the instrument: the harness digest is unchanged at 562a7f7232da across the whole addition. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01CYLNQy6tLXqsV1CbuNqsSb
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
p022-263a-calibration.linux.jsonhas been deleted, not replaced, and this head does not restore it.Both pairs reproduced at the head they were recorded against. Across the six recorded pair attempts below,
n=5reads fail, pass, pass, pass, fail, pass — and the running tally counts more observations than the table does. Nothing in that sequence is a trend, and the ledger exists so that no revision of this description can quietly turn it into one.n=5n=15core-full-sarif|rust|cal-facts-medium|process-coldat 0.478core-usage|rust|cal-facts-tiny|process-coldat 0.516388598e)core-full-sarif|rust|cal-facts-medium|process-coldat 0.361core-full-sarif|rust|cal-facts-tiny|process-coldat 0.508That table is the six pair attempts recorded under successive identity procedures; it is not the whole ledger.
Running tally on this machine:
n=5has failed three times and passed seven times;n=15has failed twice and passed six times. Neither count reproduces reliably, and which one "works" depends on when it was run.No calibration of record is restored, and that is the whole point. Every re-record was forced by a provenance repair — a changed digest, a changed reference identity, a changed pair-identity contract — never sought for a better answer. Promoting whichever pair happened to pass would be selection on the outcome regardless of why the re-run happened. By now the coin has been watched land often enough to know it is a coin. No failing pair has ever been re-run to obtain a passing partner. That is the single move this instrument exists to make impossible to hide, and the permitted counts are still only 5 and 15 with nothing to escalate to. Both pairs ship as sizing evidence, pass and fail alike, and the report-of-record glob stays empty.
What the
n=15failure did and did not establishThis outcome is consistent with the pre-run alternative that increasing N need not restore reproducibility when the dominant variability is not sampling uncertainty of the median. It falsifies the narrower operational hypothesis tested by the preregistered escalation: that moving from
n=5ton=15would reliably produce an admissible calibration on this setup. It does not establish the cause of the remaining variability, nor distinguish intrinsic workload spread from between-run environmental variation.An earlier revision of this description called that outcome predicted. It was not, and the claim is withdrawn. Nothing stated in advance predicted that
n=15would fail on three cells, or fail worse thann=5. What was stated in advance was only the alternative above and the stop rule. A provenance claim about when a belief was formed is exactly the kind of claim this instrument exists to make checkable, so it does not get to be sloppy here.No knob was turned in response. The permitted escalation was exactly one step,
5 → 15, with a mandatory stop on failure. There is non=25, non=45,--warmupstays at 2, and0.35has not moved. A tolerance adjusted after seeing which runs it rejects is a threshold fitted to a result, which is the single thing this instrument exists to prevent.An observation offered as a finding. Every cell that has ever fallen outside tolerance is a
rustcell with a median between 2.5 ms and 9.2 ms. No Python cell has ever fallen outside it. Of the 40 cells, the 16 with a median under 10 ms are allrust, and all five distinct failures to date sit in that group. The population spans roughly 1.8 ms to 3.2 s, and one relative tolerance is applied uniformly across all of it.A variance characterisation was authorised to test whether that is a duration artefact. It came back formally inconclusive — see below. The magnitude-aware question remains a measurement-design decision for the owner, and nothing in this PR advances it.
Variance characterisation: formal outcome O3 (inconclusive)
(The owner calls this "Round 6". It is unrelated to review Round 6 below, which was the A/B pair defect. Two different numbering schemes met in one document; the collision is noted rather than silently renamed.)
Preregistered in
docs/notes/p022-263a-round6-preregistration.md, committed before the first measurement existed, naming three outcomes and what each licenses. Amended once, also before any measurement and recorded with its reason: the helper had to be compiled rather than interpreted, because the timed interval includes process spawn and a Python floor of 11.8 ms sits above half the ladder.A synthetic C helper — neither Rust nor Python product code — ran through the instrument's own timed interval (
Harness._run_once, not a re-implementation) across 8 rungs from 1 to 128 ms, atnin {5, 15}, warmup 2, five independent A/B sessions each: 80 sessions, every raw observation kept.An earlier revision of this description claimed outcome O2. That claim is withdrawn.
The owner rejected it on three counts and was right on all three.
0.35, which the plan forbids in as many words — "No pass/fail tolerance is computed, applied, or reported ...0.35is not consulted." The withdrawn reading argued from the ladder never crossing it. The measurement driver never reads0.35, so the dataset is uncontaminated; the interpretation failed, not the collection.A + R*tfit ran in a branch that did not license it. The plan permits derivingAandRonly under O1; the reading selected O2 and reported the fit anyway.The formal outcome is O3, inconclusive: the dataset is descriptively suggestive, but the preregistration did not operationalise the O1/O2 boundary tightly enough to support a formal verdict. The separation is retained as exploratory calibration evidence and may motivate a separately preregistered follow-up.
What the numbers descriptively show, licensing nothing
n=5, roughly twentyfold — so O1's own written precondition does not hold. That much is a direct reading of the plan's text.That separation is a reason to keep looking, not a verdict. The envelope fit is quarantined in the note under an explicit post-hoc, not part of the preregistered outcome, not policy input heading rather than deleted, because deleting an analysis that was actually performed is its own dishonesty.
The boundary, and a correction to it. The helper does pure arithmetic;
own-cliis ~1.8 MB. The ladder isolates duration, not what a real invocation does. The original next-suspect list led with file I/O and should not have:core-usage|rust|cal-facts-tiny|process-coldis a witness and that rung never opens an OwnIR document. Input-document I/O is not a necessary condition. What survives: executable/library mapping and page faults, dynamic loader and runtime startup, scheduler context switches and CPU migration, CPU-time against wall-time, and the output/refusal path.Nothing changed as a result: no floor, no envelope,
0.35unmoved, noT_min/N_min/N_max, no threshold derived from any number in the dataset.Follow-up work: the plan amended, the instrument extended, the round not run
docs/notes/p022-263a-round7-preregistration.mdis now a RATIFIED DESIGN whose measurements are still not authorised. Two reviews. The first ruled CHANGES REQUIRED and decided four questions — full POSIXrusageyes, a statically linked fourth arm no, arms A/B/C kept, both regimes measured. The second ruled PASS WITH ONE MECHANICAL CORRECTION, ratified a narrower zero rule than the one proposed here, and authorised building the apparatus and running the structural preflight — nothing with a clock in it. Four corrections came out of the two, and they matter more than the decisions.A=1, B=3, C=1.4fires P1 and P4 at once. Ten review rounds spent insisting that a check must read the thing, and I preregistered the numbers and forgot to preregister the logic. The owner's ratified set replaces it, with P1 and P2 renamed to claim only what arm B can see.A=0orB=0. It does not. Overlap occurs only whenA=0andB=0together, and P1 is disjoint from every other rule unconditionally — itsC > 1.5Aclause excludes P4 outright, and every other intersection forcesB=0and thenC=0, contradicting that same clause. My proposed guard was correspondingly too wide: keying it onBwould have refusedA=1, B=0, C=4, a clean P2 where the padded arm shows no excess and the real binary does, which is one of the cleanest results this round can produce. The owner's ratified guard keys onAalone, and not for exclusivity —Ais the multiplicative reference for1.5A,3Aand6A, and against a baseline of zero any positive drift is infinitely large. No epsilon: a tolerance likeA < 0.01 mswould become an absolute threshold with no preregistered basis, which is the move this round forbids.D(B) − D(A)is stated as a lower bound; and a B that reproduces nothing reads as "a padding-only image effect is insufficient", never as "image mapping and page faults are excluded". A four-part structural preflight — padding survived linking, sits inside aPT_LOAD, matches arm C's mapped size to within a page, work bytes identical to arm A — is a stop condition, not a caveat written afterwards.process-cold. One of the five,core-full-sarif|rust|cal-facts-small|warm, does not. Both regimes are now measured and classified independently, and a cold/warm split is reported as a cache-sensitivity signal rather than collapsed into P5. The regimes are also stated as the instrument implements them — cold 0 discarded iterations, warm 2, both a fresh process per iteration — instead of the draft's own contradiction._run_oncenow keeps the accountingwait4was already handing back. It reaped every child throughos.wait4, which returns the kernel's complete per-childrusage, and read one field of it. User CPU, system CPU, minor and major faults, and voluntary and involuntary context switches were fetched and dropped — and those six are exactly what separates "the work took longer" from "the scheduler moved it" from "the image cost more to map". The interval is unchanged in meaning:t0,Popen,wait4,elapsed, with therusageparsed after the clock stops. The two pre-existing lines betweenwait4andelapsedstay where they are, because moving them would tighten the interval and silently un-compare every future number against every recorded one.ru_utime/ru_stimeare rounded to nanoseconds, since truncation biases every sample the same direction. Off POSIX all six arenullwith a reason, never0.perf-child-accountingis the sixteenth control, and it checks the claim that would otherwise be invisible — that the parse sits outside the clock — by wrappingos.wait4so everyrusagefield costs 10 ms to read, then asking whetherelapsed_nsgrew. Seven mutations, each caught by the check that owns it.ownlang/reference boundary is closed by observation, andperf-reference-boundarykeeps it proven.perf_baseline.pyis in mypy's file list and passes--strict, now alongside the Round 7 apparatus at 47 files.The Round 7 apparatus, built and preflighted — still no measurement
scripts/round7/is a classifier and a structural preflight. No clock appears anywhere in it, and the calibration pass itself remains unauthorised.classify.pyis the only implementation of P1–P5 and the zero-A guard; a second copy would be a second opinion and the reading could then choose between them. Being pure, it is exercised exhaustively before any measurement exists: 50,653 exact rational triples, in sixths so the ratified edges3/2,3,2/3and2land exactly on grid points, withFractionarithmetic throughout because a binary-float2/3would decide the(2/3)B ≤ Cedge by rounding direction rather than by the rule.A=B=0; and against the previous draft's P1, which must be reported broken. A control that has only ever seen correct input is a control nobody has tested.A-or-B; the guard removed; ambiguity resolved by precedence instead of raising; P4's edge made strict; an epsilon smuggled into the zero comparison. The second is the guard the owner rejected, and the control now refuses to let it back in.st_shndx, which names its section, rather than by matching addresses — an inference that finds nothing at all for a non-allocated section, whose address is zero. The four damage cases are built, not simulated: a real linker does drop an unreferenced non-volatileconstant under--gc-sections, and a section emitted with empty flags really is absent from everyPT_LOAD..rodatainside thePT_LOADat0x2000; arm B maps 1,628,917 bytes against arm C's 1,628,892,+25against a 4,096-byte allowance;mainis 107 bytes and byte-identical in both arms, so B4 decided on raw bytes and the disassembly fallback was not needed. Arm A and arm B are linked from the samespin.o, which removes the compiler-determinism question rather than answering it.562a7f7232daacross the whole addition.All four committed sizing halves are stale. The digest moved to
6713e7300c7cwhen the instrument was typed and the reference environment pinned, and again to562a7f7232daunder the accounting change above. They are deliberately not re-recorded: re-recording is a measurement, and measurements are not authorised. Stale evidence that admits it beats fresh evidence nobody asked for.What ships: all four halves of both pairs, each recorded on a clean tree, each run B naming its run A by path and sha256 so the verdicts can be recomputed rather than trusted. The four pre-contract artifacts — including both failing recordings — are preserved byte-identical under
docs/evidence/historical/, outside the report-of-record glob. Nothing was deleted to make the record look better.Что и зачем
#263-A: the frozen instrument for #263's P-022 performance baselines — how a number is produced, never what the number should be. It permits calibration and structurally refuses the decisive measurement, which is #263-B behind the D7 freeze.
No decisive numbers, no thresholds, no engine verdict on any decisive workload. No decisive
N, no escalation ladder, no Rust-vs-Python comparison statistic, no D7 freeze payload.Тип изменения
No production behaviour changes.
rust/,frontend/andownlang/are untouched by this branch.Review rounds
Eleven review rounds. Nearly every finding was invisible to a passing CI. (The variance characterisation reported above is separately numbered by the owner and is not one of these.)
Round 1 — the C1/C2 verifier read a proxy. It armed on any JSON echoing three values anyone with the repository can compute, proving the instrument was the instrument and calling that a freeze. Also:
core-usageclaimed to be startup; it became a labelled lower bound.Round 2 — four blockers. (1) The production freeze could never have armed: the verifier treated the payload's directory as the repository, so it asked git for a bare basename, while
<rev>:<path>resolves from the tree root. Fixtures passed because they put both objects at the repository root, where the wrong model is accidentally right. (2) Signed ratification readgit verify-commit --rawfrom stdout when git writes it to stderr; repairing it would still leave the accept direction untestable, so it is withdrawn. (3).NETwas installed but never on PATH, so 12 of 44 cells were rc127 command-not-found, timed and reported as reproduced; each rung now declares the exit codes that mean it did its job plus a post-condition, verified once, untimed. (4)launcher-extractclaimed "no core runs" while--emit-factslets Stage 2 run anyway; withdrawn.Round 3 — the harness digest named the checkout. The same tree hashed differently under
core.autocrlf. D7's C1 freezes that value and #263-B runs on both platforms, so a freeze would have armed on one and been refused on the other. Now content-addressed. Found by a control written in this PR, not by a reviewer.Round 4 — a phase list that meant nothing.
cli-argv-refusalbundled argv parsing with the usage-refusal message, so it appeared on three rungs that never write one; the report copied the taxonomy and the control confirmed the copy was faithful. Split intocli-argv-parse,cli-usage-refusal,ownir-door-refusalwith equivalence rules in both directions. Reproducibility split into three axes —outcomes_reproduced,timings_reproduced,environment_valid— withreproducedstill the conjunction rather than a renamed subset.Round 5 — provenance of the calibration itself. (1)
--warmupsilently moved 2→3 alongside the authorised repetitions escalation, so two knobs were turned when one was permitted; warmup is now a policy default a control enforces. (2) Only run B of each pair was committed, so "reproduced" could not be recomputed from evidence; run B now records run A's path, sha256, tree, digest and both knobs, and a control requires A committed beside it. (3)reproduce()'s docstring and the note still taught the old false model of the tolerance while the implementation was already correct. (4) The rationale forn=15was restated: a preregistered bounded escalation, not "a better-estimated median" — that argument picks 25, then 100.Round 6 — an A/B pair that was two experiments. Both halves must be one experiment or their difference measures the wrong thing.
PAIR_IDENTITY_FIELDSnow pins ten axes — reference commit and tree, harness digest and version, workload manifest digest, interpreter, platform, both knobs, and the tag — andreproduce()raises rather than reporting a verdict when any of them differ. Fail-closed, because a pair that quietly compares across a source change is worse than no pair at all.Round 7 — the C1/C2 implementation was not the frozen C1/C2. D7 §6 defines two git commits, not two sections of one file. The earlier scheme used the same two letters for something else, with
c2holding threshold values inside the payload; the word had quietly changed profession, and an equivalent-looking scheme is not the frozen scheme. It had no equivalent at all ofpayload_blob_sha, of exact-byte hashing, or of the ancestry requirement. Also in this round:python_reference_commit()wasgit rev-parse HEAD— it named the repository's position, not the reference source — and is now the last commit touchingownlang, withpython_reference_treerecording its content-addressed tree object. Both joined pair identity, and they are distinct values from the repository's own position.Round 8 — two checks that read something other than the thing. (1)
instrument_tree_shawasgit rev-parse HEAD, so C1 had to equal the current position. The lifecycle is S → C1 → C2 → #263-B, meaning HEAD is at least C2 by the decisive run and can never equal the accepted tree S; a freeze satisfying that check would have to name a future content-addressed commit inside the commit that determines it. The twin defect had been fixed one round earlier forpython_reference_commit, with the argument spelled out in its own docstring, three lines away. The anchor is now declared by C1 and proved by content. The fixtures could not see any of it: they built a freeze in a throwaway repository while the gate observed the real one, so HEAD could never move past S. (2) C1's protocol was three key names checked for presence, and the fixture proving that check worked armed the gate with three copies of the string<frozen by D7, never read by the gate>— two immaculate commits, three celebratory strings, decisive firewall open. The gate still never reads a threshold; it now verifies that a preregistration is structurally there. Damaged freezes went 18 → 33. Also self-found in this round:load_manifest()hashed raw bytes while the harness digest normalized them, which is the Round 3 defect still alive in a sibling path, on a value C1 also freezes.Round 9 — the protocol gate had a boolean escape hatch and no sense of scale. (1)
not_applicable: truewas accepted, and so wasfalse._present()ended invalue is not None, and in Python aboolis anint. Both passed a check whose entire purpose was to demand a stated reason, and the previous revision of this description asserted that a baretruewas refused — an untested claim about a guard, published, in the document that argues a green gate proves only what it was asked to prove._present()now rejectsbooloutright,not_applicablemust be a non-empty string, and a cell that is bothnot_applicableand carries a rule is refused. (2) The verifier checked every cell it was handed and never asked whether it had been handed them all, so one immaculate cell and three roll-up strings passed; the fixture demonstrated the hole rather than catching it.expected_d7_cells()now enumerates the universe and the gate refuses any cell missing, unknown, or named twice. Damaged freezes 33 → 39.Round 10 — the completeness check proved the wrong set. Round 9's mechanism was sound and its universe was not:
expected_d7_cells()enumeratedsorted(PHASES), the attribution taxonomy that answers what a timed interval contains. Using it as D7's vocabulary was wrong in both directions at once. It demanded a preregistered rule forcli-argv-parse,cli-usage-refusal,ownir-door-refusalandfrontend-extractionon every workload, platform and regime; and it omitted the end-to-end C# run entirely, one of the metrics the G3 verdict weighs most, because that is a rung rather than a part and appears nowhere inPHASES. The code had already said so and was ignored: the accept fixture keyed its not-applicable slice onfrontend-extraction, auto-excusing all 48 of those cells — the instrument admitting they were never D7's while the universe went on requiring them. A check that manufactures an obligation and then ceremonially forgives it is paperwork.D7_PHASESis now separate, and the relationship between the two vocabularies is stored as data:D7_NON_METRIC_PHASESrecords each attribution component that is not a gate with its reason, andD7_METRICS_WITHOUT_PHASErecordsend-to-end-csharpas a metric no single phase names.d7_vocabulary_problems()requires every attribution phase to be promoted or excluded explicitly — neither is a default — and runs in--selftestas well as a control, because the failure mode is drift. Universe 528 → 384.Round 11 — a preregistration whose numbers were arithmetic and whose logic was not. The Round 7 draft's five outcomes overlapped:
A=1, B=3, C=1.4satisfies P1 and P4 simultaneously, so the same dataset could have licensed "the padded image explains it" and "the phenomenon never reproduced" at once. Every earlier round in this list was about a check reading a proxy instead of the thing; this one is the same defect one level up, in a document written specifically to stop me choosing after seeing the data. Two more in the same round: arm B was advertised as isolating page-fault cost, which a demand-paged image whose padding is never touched cannot show, and the draft asserted that every witness lives inprocess-coldwhencore-full-sarif|rust|cal-facts-small|warmdoes not. The first was a claim the experiment could not support; the second was a claim the repository's own evidence contradicted, one table away. And the replacement had the same defect as the thing it replaced: my statement that the corrected rules overlapped only whereA=0orB=0was asserted, not checked, and the owner's algebra showed both halves wrong — overlap needsA=0andB=0, and P1 overlaps nothing at all. The fix this time is not a better assertion. A control evaluates the rules over fifty thousand exact triples and is itself tested against the broken version.Three policy constants, one historical value
NOISE_PROBE_MAX_RELATIVE_IQRNOISE_PROBE_MAX_DRIFTREPRODUCIBILITY_MAX_MEDIAN_CHANGEThree different statistical quantities that were one constant. All three are
0.35, deliberately unchanged. It is chosen; what makes it admissible is that it was chosen before the runs it judges.The honest limit
core-usagecore-parse-refusedcore-full-human/core-full-sariflauncher-e2ebridge-lowering,analysisand launcher-scoped extraction isolation are marked not separately observable. Derived views are labelled derived bounds.What the gate actually verifies
cellsandrollups, the instrument anchor, the Python reference commit and tree, harness digest and version, the workload manifest digest, and the owner's ratificationinstrument_anchor_commitS; the gate proves the instrument at S hashes to theharness_digestC1 froze, and that S is an ancestor of C1. HEAD is never required to equal anythinginitial_n,escalation_stages,transition_predicates,max_n,terminal_outcome— or an explicitnot_applicablewhose reason is a non-empty string, never a boolean, and never alongside a rule; roll-ups carryworkload_class,phase,overall_g3D7_PHASES, not the instrument's attribution taxonomy, socli-argv-parseandfrontend-extractionare excluded with recorded reasons whileend-to-end-csharpis included.large-solution-controlresolves tooss-ShareX.sln, so the alias earns no second voteThreshold values are checked for shape and never read. Because the workload manifest is one of the instrument sources, proving the harness digest at S also proves the manifest at S — there is no second anchor check to write and none to get wrong.
The four carried obligations
perf-gate-dormant: unarmed, fails closed on an unreadable file and a bare identity echoperf-gate-arms-by-databuilds a real S → C1 → C2 → further-work lifecycle in a throwaway repository in the nested production layout, and asserts those four are four distinct commits.perf-gate-payload-identitydamages one property at a time across 40 cases, each asserted against the phrase the check that owns it must produce — because without that, one broad refusal masquerades as forty working checksperf-notary-outsidecounts digest computations and git invocations inside a measured interval, requiring zeroperf-session-drift: a changed candidate refuses the remainder and names how many cells preceded itКак проверено
ruff check .exit 0;mypySuccess, 47 files (the instrument and the Round 7 apparatus included)python scripts/perf_baseline.py --selftestexit 0python tests/test_perf_instrument.py— 16/16, exit 0python tests/test_round7_apparatus.py— 5/5, exit 0; 6 classifier mutations caughtpython scripts/round7/preflight.py— B1–B4 all pass, recorded as committed evidencepython tests/run_tests.pyexit 0ruff check . >/dev/null && echo okonce turned a lint failure into silence that read as successСвязанные issue
Refs #263 (this is #263-A only), #262. Closes nothing.
Findings recorded rather than fixed
prompt-D7-preregistration-4.mdand confirmed the two-commit structure against it, so the earlier "fidelity unverified" caveat is closed. The cell and roll-up key names come from the owner's review messages rather than from the frozen file, which is still not in the implementing session's possession.prompt-263A-measurement-design-4.md, the final D7 brief and live P-022 step 8: Rust-default cutover, rollback gate, and Python distribution removal #262, together with the 384-cell denominator. Ratifying 384 does not mean 384 numeric budgets — an explicitnot_applicablewith a reason remains a legitimate decision. The source comment still readsNAMES ARE PROVISIONALdeliberately, since editing it would move the harness digest and force another re-record for a change of wording; the ratification binds in the ledger and to the head it was given at.not_applicable: truewas refused. It was not —_present()accepted every boolean — and the owner found it by reading the code rather than the prose. The guard is fixed, with catchers fortrue,false, an empty reason, and N/A beside a rule. Recorded because it is the exact failure mode this PR exists to argue against: an assertion that sounded like a check.ownlang/reference boundary is CLOSED, by observation. An audit hook recorded every file opened and process spawned on both render surfaces: zero repository files outsideownlang/, nothing spawned. Static imports are the standard library and itself;config.pyhas no auto-discovery and no environment variables, so there is no consulted-if-present file to slip past an audit that only sees opens; thefix_*subprocess pipeline never loads on the measured path.perf-reference-boundarykeeps it proven and is mutation-verified both ways. The one residual found —OWNLANG_DEBUG, an ambient variable the reference reads — is now cleared for every measured invocation viaREFERENCE_ENV_PINNED_UNSET.scripts/perf_baseline.pywas outsidemypy's file list, so "Success, 43 files" never covered the file that produces the numbers. It is in the list and passes--strict. The 42 errors it exposed were almost entirely one root cause — JSON values typedobjectbehindx or {}idioms withtype: ignorecodes that had drifted — replaced by four narrowing helpers that narrow fail-closed. Six were real Windows APIs absent from the Linux stubs and carry targeted ignores with reasons, not blanket suppression./usr/bin/timeis absent on the development container, henceos.wait4.n=5artifact is lost. It was reverted before the preservation rule existed. What survives of that first observation is narrative, and narrative is not evidence. Every failure since is committed.5d6dab4,35befb2) show as Unverified and are deliberately not rewritten. Review verdicts bind to exact head SHAs.perf-child-accounting: two structural checks recorded their findings in a list and then ran on into the code those very faults break, dying onKeyErrorinsidemeasure_cell. Then again inround7-zero-guard, in a file written after that finding was recorded: removing the zero guard makesA=B=C=0fire three rules, the classifier raises, and the control died before reporting. Both times the mutation scored as caught because the process exited non-zero; both times reading the output rather than the exit code is what exposed it. Twice in two rounds says the habit is to check that a mutation fails rather than that the check speaks.A=0orB=0" was hand-checked and wrong in both directions. Preregistration constrains the analyst only as far as its logic is checkable, and prose about logic is not checkable. The rules are now evaluated by a control, and the control is tested against the broken version.STOP
Stopped, awaiting an owner ruling. No merge, no thresholds, no decisive corpus, no paired timed measurement, no D7 freeze payload or attestation committed to this repository, no #263-B, no Stage 3, no Stage 4. The variance characterisation was authorised, executed under its preregistration, and reported a formal O3 — inconclusive without changing a single policy value. The Round 7 design is ratified, its apparatus is built, and the B1–B4 structural preflight passes. No Round 7 measurement exists and none is authorised: the calibration pass is a separate decision, due after a short exact-head implementation review.
The D7 metric mapping and the 384-cell universe are ratified by the owner against the frozen briefs. The source comment still reads
NAMES ARE PROVISIONAL, deliberately: it was true when written, and editing it would move the harness digest and force another re-record for a change of wording.🤖 Generated with Claude Code
https://claude.ai/code/session_01CYLNQy6tLXqsV1CbuNqsSb