Skip to content

feat(perf): #263-A — the measurement instrument (CALIBRATION_ONLY; no admissible baseline yet) - #351

Open
PhysShell wants to merge 42 commits into
mainfrom
claude/p-022-analysis-wiring-mwhqlw
Open

feat(perf): #263-A — the measurement instrument (CALIBRATION_ONLY; no admissible baseline yet)#351
PhysShell wants to merge 42 commits into
mainfrom
claude/p-022-analysis-wiring-mwhqlw

Conversation

@PhysShell

@PhysShell PhysShell commented Sep 10, 2026

Copy link
Copy Markdown
Owner

⚠️ Status: BLOCKED — there is no admissible calibration of record at this head

p022-263a-calibration.linux.json has been deleted, not replaced, and this head does not restore it.

Both pairs reproduced at the head they were recorded against. Across the six recorded pair attempts below, n=5 reads fail, pass, pass, pass, fail, pass — and the running tally counts more observations than the table does. Nothing in that sequence is a trend, and the ledger exists so that no revision of this description can quietly turn it into one.

attempt n=5 n=15
pre-contract (quarantined) failed — 1 of 40 cells, worst core-full-sarif|rust|cal-facts-medium|process-cold at 0.478 failed — 3 of 40, worst core-usage|rust|cal-facts-tiny|process-cold at 0.516
corrected identity procedure (388598e) passed failed — 1 cell, core-full-sarif|rust|cal-facts-medium|process-cold at 0.361
under the D7 §6 + reference-identity repair passed passed
under the anchor + protocol repair passed passed
under the N/A + completeness repair failed — 1 cell, core-full-sarif|rust|cal-facts-tiny|process-cold at 0.508 passed
under the D7-vocabulary repair passed passed

That table is the six pair attempts recorded under successive identity procedures; it is not the whole ledger.

Running tally on this machine: n=5 has failed three times and passed seven times; n=15 has failed twice and passed six times. Neither count reproduces reliably, and which one "works" depends on when it was run.

No calibration of record is restored, and that is the whole point. Every re-record was forced by a provenance repair — a changed digest, a changed reference identity, a changed pair-identity contract — never sought for a better answer. Promoting whichever pair happened to pass would be selection on the outcome regardless of why the re-run happened. By now the coin has been watched land often enough to know it is a coin. No failing pair has ever been re-run to obtain a passing partner. That is the single move this instrument exists to make impossible to hide, and the permitted counts are still only 5 and 15 with nothing to escalate to. Both pairs ship as sizing evidence, pass and fail alike, and the report-of-record glob stays empty.

What the n=15 failure did and did not establish

This outcome is consistent with the pre-run alternative that increasing N need not restore reproducibility when the dominant variability is not sampling uncertainty of the median. It falsifies the narrower operational hypothesis tested by the preregistered escalation: that moving from n=5 to n=15 would reliably produce an admissible calibration on this setup. It does not establish the cause of the remaining variability, nor distinguish intrinsic workload spread from between-run environmental variation.

An earlier revision of this description called that outcome predicted. It was not, and the claim is withdrawn. Nothing stated in advance predicted that n=15 would fail on three cells, or fail worse than n=5. What was stated in advance was only the alternative above and the stop rule. A provenance claim about when a belief was formed is exactly the kind of claim this instrument exists to make checkable, so it does not get to be sloppy here.

No knob was turned in response. The permitted escalation was exactly one step, 5 → 15, with a mandatory stop on failure. There is no n=25, no n=45, --warmup stays at 2, and 0.35 has not moved. A tolerance adjusted after seeing which runs it rejects is a threshold fitted to a result, which is the single thing this instrument exists to prevent.

An observation offered as a finding. Every cell that has ever fallen outside tolerance is a rust cell with a median between 2.5 ms and 9.2 ms. No Python cell has ever fallen outside it. Of the 40 cells, the 16 with a median under 10 ms are all rust, and all five distinct failures to date sit in that group. The population spans roughly 1.8 ms to 3.2 s, and one relative tolerance is applied uniformly across all of it.

A variance characterisation was authorised to test whether that is a duration artefact. It came back formally inconclusive — see below. The magnitude-aware question remains a measurement-design decision for the owner, and nothing in this PR advances it.

Variance characterisation: formal outcome O3 (inconclusive)

(The owner calls this "Round 6". It is unrelated to review Round 6 below, which was the A/B pair defect. Two different numbering schemes met in one document; the collision is noted rather than silently renamed.)

Preregistered in docs/notes/p022-263a-round6-preregistration.md, committed before the first measurement existed, naming three outcomes and what each licenses. Amended once, also before any measurement and recorded with its reason: the helper had to be compiled rather than interpreted, because the timed interval includes process spawn and a Python floor of 11.8 ms sits above half the ladder.

A synthetic C helper — neither Rust nor Python product code — ran through the instrument's own timed interval (Harness._run_once, not a re-implementation) across 8 rungs from 1 to 128 ms, at n in {5, 15}, warmup 2, five independent A/B sessions each: 80 sessions, every raw observation kept.

An earlier revision of this description claimed outcome O2. That claim is withdrawn.

The owner rejected it on three counts and was right on all three.

  1. The reading consulted 0.35, which the plan forbids in as many words — "No pass/fail tolerance is computed, applied, or reported ... 0.35 is not consulted." The withdrawn reading argued from the ladder never crossing it. The measurement driver never reads 0.35, so the dataset is uncontaminated; the interpretation failed, not the collection.
  2. O1 and O2 were never operationalised as numbers. "Roughly constant" and "stable" were English. The quantities that ended up deciding were all chosen after seeing the data, so "read off the preregistered table" claimed more than the plan could support.
  3. The A + R*t fit ran in a branch that did not license it. The plan permits deriving A and R only under O1; the reading selected O2 and reported the fit anyway.

The formal outcome is O3, inconclusive: the dataset is descriptively suggestive, but the preregistration did not operationalise the O1/O2 boundary tightly enough to support a formal verdict. The separation is retained as exploratory calibration evidence and may motivate a separately preregistered follow-up.

What the numbers descriptively show, licensing nothing

  • Absolute drift is not constant across the ladder — 0.056 ms to 1.126 ms at n=5, roughly twentyfold — so O1's own written precondition does not hold. That much is a direct reading of the plan's text.
  • A measurement-scale effect is visible: relative shift rises from ~0.006 at 8–16 ms to ~0.04–0.05 at 1–2 ms.
  • In the 2–10 ms band, across 27 A/B sessions, the ladder's worst absolute drift was 0.3203 ms; the historical failures in that band implied 1.23 to 4.41 ms.

That separation is a reason to keep looking, not a verdict. The envelope fit is quarantined in the note under an explicit post-hoc, not part of the preregistered outcome, not policy input heading rather than deleted, because deleting an analysis that was actually performed is its own dishonesty.

The boundary, and a correction to it. The helper does pure arithmetic; own-cli is ~1.8 MB. The ladder isolates duration, not what a real invocation does. The original next-suspect list led with file I/O and should not have: core-usage|rust|cal-facts-tiny|process-cold is a witness and that rung never opens an OwnIR document. Input-document I/O is not a necessary condition. What survives: executable/library mapping and page faults, dynamic loader and runtime startup, scheduler context switches and CPU migration, CPU-time against wall-time, and the output/refusal path.

Nothing changed as a result: no floor, no envelope, 0.35 unmoved, no T_min/N_min/N_max, no threshold derived from any number in the dataset.

Follow-up work: the plan amended, the instrument extended, the round not run

docs/notes/p022-263a-round7-preregistration.md is now a RATIFIED DESIGN whose measurements are still not authorised. Two reviews. The first ruled CHANGES REQUIRED and decided four questions — full POSIX rusage yes, a statically linked fourth arm no, arms A/B/C kept, both regimes measured. The second ruled PASS WITH ONE MECHANICAL CORRECTION, ratified a narrower zero rule than the one proposed here, and authorised building the apparatus and running the structural preflight — nothing with a clock in it. Four corrections came out of the two, and they matter more than the decisions.

  • The outcome table was not mutually exclusive. A=1, B=3, C=1.4 fires P1 and P4 at once. Ten review rounds spent insisting that a check must read the thing, and I preregistered the numbers and forgot to preregister the logic. The owner's ratified set replaces it, with P1 and P2 renamed to claim only what arm B can see.
  • Then I got the replacement's algebra wrong too, in both directions. I claimed every remaining overlap reduced to A=0 or B=0. It does not. Overlap occurs only when A=0 and B=0 together, and P1 is disjoint from every other rule unconditionally — its C > 1.5A clause excludes P4 outright, and every other intersection forces B=0 and then C=0, contradicting that same clause. My proposed guard was correspondingly too wide: keying it on B would have refused A=1, B=0, C=4, a clean P2 where the padded arm shows no excess and the real binary does, which is one of the cleanest results this round can produce. The owner's ratified guard keys on A alone, and not for exclusivity — A is the multiplicative reference for 1.5A, 3A and 6A, and against a baseline of zero any positive drift is infinitely large. No epsilon: a tolerance like A < 0.01 ms would become an absolute threshold with no preregistered basis, which is the move this round forbids.
  • Arm B claimed more than it can see. It said it isolated "image size, mapping and page-fault cost". Linux demand-pages an image, so padding that is mapped and never read need not fault in at all. B now claims the padded image footprint and its mapping metadata; D(B) − D(A) is stated as a lower bound; and a B that reproduces nothing reads as "a padding-only image effect is insufficient", never as "image mapping and page faults are excluded". A four-part structural preflight — padding survived linking, sits inside a PT_LOAD, matches arm C's mapped size to within a page, work bytes identical to arm A — is a stop condition, not a caveat written afterwards.
  • A factual claim in the draft was false. It said every witness lives in process-cold. One of the five, core-full-sarif|rust|cal-facts-small|warm, does not. Both regimes are now measured and classified independently, and a cold/warm split is reported as a cache-sensitivity signal rather than collapsed into P5. The regimes are also stated as the instrument implements them — cold 0 discarded iterations, warm 2, both a fresh process per iteration — instead of the draft's own contradiction.
  • _run_once now keeps the accounting wait4 was already handing back. It reaped every child through os.wait4, which returns the kernel's complete per-child rusage, and read one field of it. User CPU, system CPU, minor and major faults, and voluntary and involuntary context switches were fetched and dropped — and those six are exactly what separates "the work took longer" from "the scheduler moved it" from "the image cost more to map". The interval is unchanged in meaning: t0, Popen, wait4, elapsed, with the rusage parsed after the clock stops. The two pre-existing lines between wait4 and elapsed stay where they are, because moving them would tighten the interval and silently un-compare every future number against every recorded one. ru_utime/ru_stime are rounded to nanoseconds, since truncation biases every sample the same direction. Off POSIX all six are null with a reason, never 0.
  • perf-child-accounting is the sixteenth control, and it checks the claim that would otherwise be invisible — that the parse sits outside the clock — by wrapping os.wait4 so every rusage field costs 10 ms to read, then asking whether elapsed_ns grew. Seven mutations, each caught by the check that owns it.
  • The ownlang/ reference boundary is closed by observation, and perf-reference-boundary keeps it proven.
  • The instrument is type-checkedperf_baseline.py is in mypy's file list and passes --strict, now alongside the Round 7 apparatus at 47 files.

The Round 7 apparatus, built and preflighted — still no measurement

scripts/round7/ is a classifier and a structural preflight. No clock appears anywhere in it, and the calibration pass itself remains unauthorised.

  • One classifier. classify.py is the only implementation of P1–P5 and the zero-A guard; a second copy would be a second opinion and the reading could then choose between them. Being pure, it is exercised exhaustively before any measurement exists: 50,653 exact rational triples, in sixths so the ratified edges 3/2, 3, 2/3 and 2 land exactly on grid points, with Fraction arithmetic throughout because a binary-float 2/3 would decide the (2/3)B ≤ C edge by rounding direction rather than by the rule.
  • The exclusivity claim is no longer asserted — it is measured. The census runs twice: against the real rules, which must overlap only at A=B=0; and against the previous draft's P1, which must be reported broken. A control that has only ever seen correct input is a control nobody has tested.
  • Six mutations of the real classifier, each caught by the check that owns it: the old overlapping P1; the zero guard widened to A-or-B; the guard removed; ambiguity resolved by precedence instead of raising; P4's edge made strict; an epsilon smuggled into the zero comparison. The second is the guard the owner rejected, and the control now refuses to let it back in.
  • The preflight reads the ELF, not a rendering of it. Program and section headers are parsed directly, and the padding is located through st_shndx, which names its section, rather than by matching addresses — an inference that finds nothing at all for a non-allocated section, whose address is zero. The four damage cases are built, not simulated: a real linker does drop an unreferenced non-volatile constant under --gc-sections, and a section emitted with empty flags really is absent from every PT_LOAD.
  • B1–B4 all pass. Padding 1,626,024 bytes in .rodata inside the PT_LOAD at 0x2000; arm B maps 1,628,917 bytes against arm C's 1,628,892, +25 against a 4,096-byte allowance; main is 107 bytes and byte-identical in both arms, so B4 decided on raw bytes and the disassembly fallback was not needed. Arm A and arm B are linked from the same spin.o, which removes the compiler-determinism question rather than answering it.
  • The apparatus is not the instrument, and the harness digest proves it: unchanged at 562a7f7232da across the whole addition.

All four committed sizing halves are stale. The digest moved to 6713e7300c7c when the instrument was typed and the reference environment pinned, and again to 562a7f7232da under the accounting change above. They are deliberately not re-recorded: re-recording is a measurement, and measurements are not authorised. Stale evidence that admits it beats fresh evidence nobody asked for.

What ships: all four halves of both pairs, each recorded on a clean tree, each run B naming its run A by path and sha256 so the verdicts can be recomputed rather than trusted. The four pre-contract artifacts — including both failing recordings — are preserved byte-identical under docs/evidence/historical/, outside the report-of-record glob. Nothing was deleted to make the record look better.


Что и зачем

#263-A: the frozen instrument for #263's P-022 performance baselines — how a number is produced, never what the number should be. It permits calibration and structurally refuses the decisive measurement, which is #263-B behind the D7 freeze.

No decisive numbers, no thresholds, no engine verdict on any decisive workload. No decisive N, no escalation ladder, no Rust-vs-Python comparison statistic, no D7 freeze payload.

Тип изменения

  • feat — новая возможность
  • fix — исправление бага
  • docs — документация
  • refactor / chore / test / ci — без изменения поведения

No production behaviour changes. rust/, frontend/ and ownlang/ are untouched by this branch.

Review rounds

Eleven review rounds. Nearly every finding was invisible to a passing CI. (The variance characterisation reported above is separately numbered by the owner and is not one of these.)

Round 1 — the C1/C2 verifier read a proxy. It armed on any JSON echoing three values anyone with the repository can compute, proving the instrument was the instrument and calling that a freeze. Also: core-usage claimed to be startup; it became a labelled lower bound.

Round 2 — four blockers. (1) The production freeze could never have armed: the verifier treated the payload's directory as the repository, so it asked git for a bare basename, while <rev>:<path> resolves from the tree root. Fixtures passed because they put both objects at the repository root, where the wrong model is accidentally right. (2) Signed ratification read git verify-commit --raw from stdout when git writes it to stderr; repairing it would still leave the accept direction untestable, so it is withdrawn. (3) .NET was installed but never on PATH, so 12 of 44 cells were rc127 command-not-found, timed and reported as reproduced; each rung now declares the exit codes that mean it did its job plus a post-condition, verified once, untimed. (4) launcher-extract claimed "no core runs" while --emit-facts lets Stage 2 run anyway; withdrawn.

Round 3 — the harness digest named the checkout. The same tree hashed differently under core.autocrlf. D7's C1 freezes that value and #263-B runs on both platforms, so a freeze would have armed on one and been refused on the other. Now content-addressed. Found by a control written in this PR, not by a reviewer.

Round 4 — a phase list that meant nothing. cli-argv-refusal bundled argv parsing with the usage-refusal message, so it appeared on three rungs that never write one; the report copied the taxonomy and the control confirmed the copy was faithful. Split into cli-argv-parse, cli-usage-refusal, ownir-door-refusal with equivalence rules in both directions. Reproducibility split into three axes — outcomes_reproduced, timings_reproduced, environment_valid — with reproduced still the conjunction rather than a renamed subset.

Round 5 — provenance of the calibration itself. (1) --warmup silently moved 2→3 alongside the authorised repetitions escalation, so two knobs were turned when one was permitted; warmup is now a policy default a control enforces. (2) Only run B of each pair was committed, so "reproduced" could not be recomputed from evidence; run B now records run A's path, sha256, tree, digest and both knobs, and a control requires A committed beside it. (3) reproduce()'s docstring and the note still taught the old false model of the tolerance while the implementation was already correct. (4) The rationale for n=15 was restated: a preregistered bounded escalation, not "a better-estimated median" — that argument picks 25, then 100.

Round 6 — an A/B pair that was two experiments. Both halves must be one experiment or their difference measures the wrong thing. PAIR_IDENTITY_FIELDS now pins ten axes — reference commit and tree, harness digest and version, workload manifest digest, interpreter, platform, both knobs, and the tag — and reproduce() raises rather than reporting a verdict when any of them differ. Fail-closed, because a pair that quietly compares across a source change is worse than no pair at all.

Round 7 — the C1/C2 implementation was not the frozen C1/C2. D7 §6 defines two git commits, not two sections of one file. The earlier scheme used the same two letters for something else, with c2 holding threshold values inside the payload; the word had quietly changed profession, and an equivalent-looking scheme is not the frozen scheme. It had no equivalent at all of payload_blob_sha, of exact-byte hashing, or of the ancestry requirement. Also in this round: python_reference_commit() was git rev-parse HEAD — it named the repository's position, not the reference source — and is now the last commit touching ownlang, with python_reference_tree recording its content-addressed tree object. Both joined pair identity, and they are distinct values from the repository's own position.

Round 8 — two checks that read something other than the thing. (1) instrument_tree_sha was git rev-parse HEAD, so C1 had to equal the current position. The lifecycle is S → C1 → C2 → #263-B, meaning HEAD is at least C2 by the decisive run and can never equal the accepted tree S; a freeze satisfying that check would have to name a future content-addressed commit inside the commit that determines it. The twin defect had been fixed one round earlier for python_reference_commit, with the argument spelled out in its own docstring, three lines away. The anchor is now declared by C1 and proved by content. The fixtures could not see any of it: they built a freeze in a throwaway repository while the gate observed the real one, so HEAD could never move past S. (2) C1's protocol was three key names checked for presence, and the fixture proving that check worked armed the gate with three copies of the string <frozen by D7, never read by the gate> — two immaculate commits, three celebratory strings, decisive firewall open. The gate still never reads a threshold; it now verifies that a preregistration is structurally there. Damaged freezes went 18 → 33. Also self-found in this round: load_manifest() hashed raw bytes while the harness digest normalized them, which is the Round 3 defect still alive in a sibling path, on a value C1 also freezes.

Round 9 — the protocol gate had a boolean escape hatch and no sense of scale. (1) not_applicable: true was accepted, and so was false. _present() ended in value is not None, and in Python a bool is an int. Both passed a check whose entire purpose was to demand a stated reason, and the previous revision of this description asserted that a bare true was refused — an untested claim about a guard, published, in the document that argues a green gate proves only what it was asked to prove. _present() now rejects bool outright, not_applicable must be a non-empty string, and a cell that is both not_applicable and carries a rule is refused. (2) The verifier checked every cell it was handed and never asked whether it had been handed them all, so one immaculate cell and three roll-up strings passed; the fixture demonstrated the hole rather than catching it. expected_d7_cells() now enumerates the universe and the gate refuses any cell missing, unknown, or named twice. Damaged freezes 33 → 39.

Round 10 — the completeness check proved the wrong set. Round 9's mechanism was sound and its universe was not: expected_d7_cells() enumerated sorted(PHASES), the attribution taxonomy that answers what a timed interval contains. Using it as D7's vocabulary was wrong in both directions at once. It demanded a preregistered rule for cli-argv-parse, cli-usage-refusal, ownir-door-refusal and frontend-extraction on every workload, platform and regime; and it omitted the end-to-end C# run entirely, one of the metrics the G3 verdict weighs most, because that is a rung rather than a part and appears nowhere in PHASES. The code had already said so and was ignored: the accept fixture keyed its not-applicable slice on frontend-extraction, auto-excusing all 48 of those cells — the instrument admitting they were never D7's while the universe went on requiring them. A check that manufactures an obligation and then ceremonially forgives it is paperwork. D7_PHASES is now separate, and the relationship between the two vocabularies is stored as data: D7_NON_METRIC_PHASES records each attribution component that is not a gate with its reason, and D7_METRICS_WITHOUT_PHASE records end-to-end-csharp as a metric no single phase names. d7_vocabulary_problems() requires every attribution phase to be promoted or excluded explicitly — neither is a default — and runs in --selftest as well as a control, because the failure mode is drift. Universe 528 → 384.

Round 11 — a preregistration whose numbers were arithmetic and whose logic was not. The Round 7 draft's five outcomes overlapped: A=1, B=3, C=1.4 satisfies P1 and P4 simultaneously, so the same dataset could have licensed "the padded image explains it" and "the phenomenon never reproduced" at once. Every earlier round in this list was about a check reading a proxy instead of the thing; this one is the same defect one level up, in a document written specifically to stop me choosing after seeing the data. Two more in the same round: arm B was advertised as isolating page-fault cost, which a demand-paged image whose padding is never touched cannot show, and the draft asserted that every witness lives in process-cold when core-full-sarif|rust|cal-facts-small|warm does not. The first was a claim the experiment could not support; the second was a claim the repository's own evidence contradicted, one table away. And the replacement had the same defect as the thing it replaced: my statement that the corrected rules overlapped only where A=0 or B=0 was asserted, not checked, and the owner's algebra showed both halves wrong — overlap needs A=0 and B=0, and P1 overlaps nothing at all. The fix this time is not a better assertion. A control evaluates the rules over fifty thousand exact triples and is itself tested against the broken version.

Three policy constants, one historical value

policy bounds
NOISE_PROBE_MAX_RELATIVE_IQR dispersion within one probe
NOISE_PROBE_MAX_DRIFT opening probe against closing probe
REPRODUCIBILITY_MAX_MEDIAN_CHANGE one cell's median, run A against run B

Three different statistical quantities that were one constant. All three are 0.35, deliberately unchanged. It is chosen; what makes it admissible is that it was chosen before the runs it judges.

The honest limit

rung expected exit interval contains
core-usage 2 core startup + argv parse + usage refusal — a lower bound on startup
core-parse-refused 2 startup + argv parse + read + parse + door refusal
core-full-human / core-full-sarif 0 or 1 startup + argv parse + parse + bridge + analysis + render
launcher-e2e 0 launcher startup + extraction + core startup + argv parse + core + render

bridge-lowering, analysis and launcher-scoped extraction isolation are marked not separately observable. Derived views are labelled derived bounds.

What the gate actually verifies

C1 the immutable payload commit — the preregistered cells and rollups, the instrument anchor, the Python reference commit and tree, harness digest and version, the workload manifest digest, and the owner's ratification
C2 a detached attestation in a descendant commit — the payload's commit sha, its blob sha, the sha256 of its exact bytes, the instrument/reference/workload bindings, and the ratification binding
anchor C1 declares instrument_anchor_commit S; the gate proves the instrument at S hashes to the harness_digest C1 froze, and that S is an ancestor of C1. HEAD is never required to equal anything
protocol each cell carries all four dimensions plus bound, pass/fail rule, inconclusive band, comparison statistic, RSS and allocation policy, and a ladder with initial_n, escalation_stages, transition_predicates, max_n, terminal_outcome — or an explicit not_applicable whose reason is a non-empty string, never a boolean, and never alongside a rule; roll-ups carry workload_class, phase, overall_g3
completeness C1 must decide every cell D7 covers — 8 G3 metrics × 12 canonical decisive workloads × 2 platforms × 2 regimes = 384 — with none missing, none unknown, and no dimension tuple named twice. The metrics are D7_PHASES, not the instrument's attribution taxonomy, so cli-argv-parse and frontend-extraction are excluded with recorded reasons while end-to-end-csharp is included. large-solution-control resolves to oss-ShareX.sln, so the alias earns no second vote

Threshold values are checked for shape and never read. Because the workload manifest is one of the instrument sources, proving the harness digest at S also proves the manifest at S — there is no second anchor check to write and none to get wrong.

The four carried obligations

how it is proved
(a) gate exists now, dormant perf-gate-dormant: unarmed, fails closed on an unreadable file and a bare identity echo
(b) arming is data only perf-gate-arms-by-data builds a real S → C1 → C2 → further-work lifecycle in a throwaway repository in the nested production layout, and asserts those four are four distinct commits. perf-gate-payload-identity damages one property at a time across 40 cases, each asserted against the phrase the check that owns it must produce — because without that, one broad refusal masquerades as forty working checks
(c) fail-closed before any interval perf-notary-outside counts digest computations and git invocations inside a measured interval, requiring zero
(d) two identity domains perf-session-drift: a changed candidate refuses the remainder and names how many cells preceded it

Как проверено

  • ruff check . exit 0; mypy Success, 47 files (the instrument and the Round 7 apparatus included)
  • python scripts/perf_baseline.py --selftest exit 0
  • python tests/test_perf_instrument.py16/16, exit 0
  • python tests/test_round7_apparatus.py5/5, exit 0; 6 classifier mutations caught
  • python scripts/round7/preflight.pyB1–B4 all pass, recorded as committed evidence
  • python tests/run_tests.py exit 0
  • every gate's exit code read directly, after ruff check . >/dev/null && echo ok once turned a lint failure into silence that read as success

Связанные issue

Refs #263 (this is #263-A only), #262. Closes nothing.

Findings recorded rather than fixed

  1. A green gate proves what it was asked to prove. Most blockers here were invisible to a passing CI: fixtures exercising a layout production never has, a signature branch nothing could test, an exit code nobody read, a phase list the report copied faithfully, and a digest that named the checkout.
  2. §6 topology is confirmed; the protocol key names are not. The owner has read the frozen prompt-D7-preregistration-4.md and confirmed the two-commit structure against it, so the earlier "fidelity unverified" caveat is closed. The cell and roll-up key names come from the owner's review messages rather than from the frozen file, which is still not in the implementing session's possession.
  3. The D7 metric names are RATIFIED. They were recorded as provisional when written, because leaving an implementation label to become a normative contract by default is exactly the defect Round 10 fixed. The owner has since ratified the eight metrics and the four exclusions against the frozen prompt-263A-measurement-design-4.md, the final D7 brief and live P-022 step 8: Rust-default cutover, rollback gate, and Python distribution removal #262, together with the 384-cell denominator. Ratifying 384 does not mean 384 numeric budgets — an explicit not_applicable with a reason remains a legitimate decision. The source comment still reads NAMES ARE PROVISIONAL deliberately, since editing it would move the harness digest and force another re-record for a change of wording; the ratification binds in the ledger and to the head it was given at.
  4. A claim about a guard was published without being tested. Round 8's description said a bare not_applicable: true was refused. It was not — _present() accepted every boolean — and the owner found it by reading the code rather than the prose. The guard is fixed, with catchers for true, false, an empty reason, and N/A beside a rule. Recorded because it is the exact failure mode this PR exists to argue against: an assertion that sounded like a check.
  5. Completeness was recorded as not implementable, and that was wrong. Round 8 argued the cell universe "needs the decisive cell population" and left the check out. The brief permits enumerating, hashing and availability-checking decisive workloads before D7 and forbids only performance exposure — and it must, because a population unknown before D7 could not be written into D7. The error was in the convenient direction.
  6. The ownlang/ reference boundary is CLOSED, by observation. An audit hook recorded every file opened and process spawned on both render surfaces: zero repository files outside ownlang/, nothing spawned. Static imports are the standard library and itself; config.py has no auto-discovery and no environment variables, so there is no consulted-if-present file to slip past an audit that only sees opens; the fix_* subprocess pipeline never loads on the measured path. perf-reference-boundary keeps it proven and is mutation-verified both ways. The one residual found — OWNLANG_DEBUG, an ambient variable the reference reads — is now cleared for every measured invocation via REFERENCE_ENV_PINNED_UNSET.
  7. One guard is defensive and untested, and is marked so in both source and note: a same-commit attestation cannot be constructed with ordinary git, because it would have to contain the sha of a commit whose sha depends on it.
  8. Git object identity is not a signature. Arming costs two reviewed commits rather than one text editor, but anyone with write access could author both.
  9. The instrument is type-checked now. scripts/perf_baseline.py was outside mypy's file list, so "Success, 43 files" never covered the file that produces the numbers. It is in the list and passes --strict. The 42 errors it exposed were almost entirely one root cause — JSON values typed object behind x or {} idioms with type: ignore codes that had drifted — replaced by four narrowing helpers that narrow fail-closed. Six were real Windows APIs absent from the Linux stubs and carry targeted ignores with reasons, not blanket suppression.
  10. A GitHub-hosted Windows runner is not reliably measurement-grade. Two runs minutes apart on one commit disagreed about their own environment. The decisive Windows measurement needs a single-tenant machine.
  11. /usr/bin/time is absent on the development container, hence os.wait4.
  12. The original failing n=5 artifact is lost. It was reverted before the preservation rule existed. What survives of that first observation is narrative, and narrative is not evidence. Every failure since is committed.
  13. Two commits (5d6dab4, 35befb2) show as Unverified and are deliberately not rewritten. Review verdicts bind to exact head SHAs.
  14. A control that crashes before reporting is a traceback, not a finding — and it happened twice. First with perf-child-accounting: two structural checks recorded their findings in a list and then ran on into the code those very faults break, dying on KeyError inside measure_cell. Then again in round7-zero-guard, in a file written after that finding was recorded: removing the zero guard makes A=B=C=0 fire three rules, the classifier raises, and the control died before reporting. Both times the mutation scored as caught because the process exited non-zero; both times reading the output rather than the exit code is what exposed it. Twice in two rounds says the habit is to check that a mutation fails rather than that the check speaks.
  15. The preregistered outcome rules were not checked for mutual exclusivity, and neither was the correction. The draft satisfied every other discipline this PR argues for — written before the data, quantitative boundaries, a named "neither" outcome, ratios rather than imported thresholds — and still admitted two contradictory verdicts on one dataset. The repair then repeated the defect one level up: "every remaining overlap reduces to A=0 or B=0" was hand-checked and wrong in both directions. Preregistration constrains the analyst only as far as its logic is checkable, and prose about logic is not checkable. The rules are now evaluated by a control, and the control is tested against the broken version.

STOP

Stopped, awaiting an owner ruling. No merge, no thresholds, no decisive corpus, no paired timed measurement, no D7 freeze payload or attestation committed to this repository, no #263-B, no Stage 3, no Stage 4. The variance characterisation was authorised, executed under its preregistration, and reported a formal O3 — inconclusive without changing a single policy value. The Round 7 design is ratified, its apparatus is built, and the B1–B4 structural preflight passes. No Round 7 measurement exists and none is authorised: the calibration pass is a separate decision, due after a short exact-head implementation review.

The D7 metric mapping and the 384-cell universe are ratified by the owner against the frozen briefs. The source comment still reads NAMES ARE PROVISIONAL, deliberately: it was true when written, and editing it would move the harness digest and force another re-record for a change of wording.

🤖 Generated with Claude Code

https://claude.ai/code/session_01CYLNQy6tLXqsV1CbuNqsSb

…alled

The instrument for #263's P-022 performance baselines: how a number is
produced, never what the number should be. It permits calibration and refuses
the decisive measurement, which is #263-B behind the D7 freeze. No thresholds,
no repetition count N, no engine-comparison statistic — those are D7's, and a
harness that offered any of them would be choosing them.

Preflight, against primary source. Live #263 and #262's Performance-gates
reconcile as: the brief's phase list is #262's seven gate phases plus frontend
extraction, recorded and explicitly NOT D7-gated. Both populations are frozen —
A, the G3 cutover subset, feeds D7; B, the IDE-foundation remainder, does not
unless separately ratified. What this instrument does NOT measure is written
down rather than left implied, so completing #263-A cannot narrow #263's own
acceptance by omission: the .own workloads, OwnIR serialization, no-op and
single-file-edit recheck, the latency proxy, allocation profiles and IDE
budgets are owed on B's track.

The honest limit, recorded rather than routed around. Neither engine exposes
per-phase timing through its production surface and #263-A may not add
production instrumentation, so phases are measured by a LADDER of real
production invocations, each recorded as the composed interval it actually is
with the phases inside it named. bridge-lowering and analysis are marked NOT
SEPARATELY OBSERVABLE. Derived views — parse by subtraction, core work as
end-to-end minus extraction — are labelled derived and never presented as
measurements. #263 asks for unavailable stages to be marked, not imputed.

The firewall is structural, not clerical. Exactly one function starts a clock
and it refuses a decisive workload unless the identity gate is armed, which
needs a D7 attestation that does not exist. Deliberately stricter than "no
paired runs": a single-engine timed run of a decisive workload is refused too,
because a stopwatch appears nowhere in the permitted list of enumerate / hash /
fetch / availability-check / non-timed smoke. "We did not save the delta" is
accounting-clean and epistemically worthless — whoever watched the two numbers
has already peeked.

The four carried obligations live in the harness. The identity gate exists NOW
and dormant, because D7's C1 freezes the harness digest and a gate added after
that freeze makes this a different instrument. Arming is a JSON file: a control
proves the digest does not move when the gate arms and that the firewall then
opens. Verification completes before any clock starts, and a control counts
digest calls INSIDE a measured interval and requires zero — a gate that pays
for itself out of the startup benchmark is a defect wearing a safeguard's coat.
The candidate binary is a separate domain: frozen at session start, its drift
refuses the remainder rather than measuring half the cells on another binary.

Peak RSS comes from os.wait4's per-child ru_maxrss rather than /usr/bin/time,
which is a package that can simply be absent — it is absent here, and "the tool
was missing" is not a memory measurement. Windows uses a Job Object. Where
nothing is available the value is null WITH A REASON, never a confident zero.

Two defects of my own, found while building this and worth recording. The
comparison guard first grepped the source for "ratio" and refused its own
report — the sentence explaining the prohibition matched it, and then so did
"calibration" and "iterations". It now reads the report's structure and matches
whole tokens. And an aggregate of bytes was being emitted under a median_ns
key: a unit lie in a report whose entire purpose is measurement.

Calibrated: 44 cells, both engines, all six rungs, run twice on one machine with
every cell's median reproducing within the noise floor that run itself recorded.
CI calibrates on Linux and Windows and asserts the same, plus that no decisive
workload was timed and the gate stayed dormant.

scripts/perf_baseline.py is a new wrapper that reaches the launcher, so it is
added to the Stage-2 census entry points and the job classified — the census
enumerates wrappers, which makes that list a maintenance obligation, now stated
in the ledger.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CYLNQy6tLXqsV1CbuNqsSb
Recorded on a clean tree at f09fa00, dirty:false, with the D7 gate
dormant and no decisive workload timed.

  44 cells        both engines, all six rungs, process-cold and warm
  reproduced      44/44 medians within the noise floor the run itself recorded
  RSS             os.wait4 ru_maxrss, per child, on every cell
  noise           not invalidated; opening and closing probes agree
  decisive timed  none

The report validates the INSTRUMENT and nothing else. It is not evidence about
either engine, it cannot choose a threshold, a repetition count or a comparison
statistic, and it carries no engine comparison — the harness refuses to emit
one, checked on the way out as well as on the way in.

The Python reference commit is recorded beside the interpreter version so a
later artifact cannot silently combine comparisons bound to two different
reference states. The environment is a development container, recorded as such
in the fingerprint (runner_class: local); CI calibrates on ubuntu and windows
runners and asserts the same properties there.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CYLNQy6tLXqsV1CbuNqsSb
@coderabbitai

coderabbitai Bot commented Sep 10, 2026

Copy link
Copy Markdown

Important

  • 🔍 Trigger review

This repository does not receive automatic reviews because it has fewer than 10 stars.

⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: CHILL

Plan: Advanced

Run ID: 77e28ee7-8b95-4e49-9008-2956369304d5


Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

claude and others added 18 commits September 10, 2026 01:46
The Windows leg refused its own calibration on the first CI run: the opening
noise probe measured a relative IQR of 1.807 against the 0.35 floor, with 0.870
drift between the opening and closing probes. That is the instrument doing
exactly what §4 and §7 ask — refusing to report numbers taken on a machine that
was not holding still — and the job read it as a failure.

So CI now asks two questions instead of one. Did the HARNESS stand up and
report honestly: tagged CALIBRATION_ONLY, gate dormant, zero decisive workloads
timed, cells produced, raw per-iteration data retained, a named RSS mechanism,
complete provenance. That is the gate, on both platforms. And separately: is
this ENVIRONMENT measurement-grade — recorded per run, and allowed to be "no",
with the reasons required when it is.

The noise floor was deliberately NOT raised. An instrument-validity constant
chosen after seeing which runs it rejects is a threshold fitted to a result,
and fitting thresholds to results is the entire failure mode #263-A exists to
prevent. The number that would have made this green is the one number I am
least entitled to pick.

The finding itself matters more than the fix: a GitHub-hosted Windows runner is
not measurement-grade, so the decisive Windows measurement will need a
single-tenant machine. #263-B needed to know that before it started, and now it
does — from the instrument's first run rather than from a surprise halfway
through a decisive session.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CYLNQy6tLXqsV1CbuNqsSb
On ONE commit, two Windows runs minutes apart disagreed about their own
environment. The first refused it — relative IQR 1.807 against the 0.35 floor,
0.870 drift between probes. The second accepted it and completed a full 44-cell
calibration that reproduced within 0.35. Same commit, same runner class,
opposite validity verdicts.

The risk is therefore not that a hosted-Windows measurement would be noisy. It
is that whether it is ADMISSIBLE turns on scheduling luck, so a decisive
session started on the lucky run and continued into the unlucky one would be
half a measurement. D7 should not preregister a Windows protocol that assumes a
hosted runner.

Both legs are otherwise green and complete: 44 cells each, 44 reproduced,
peak RSS through os.wait4 on Linux and the Job Object on Windows, gate dormant,
zero decisive workloads timed.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CYLNQy6tLXqsV1CbuNqsSb
…t back

Two findings from review, both confirmed directly in the code.

P0 — the C1/C2 verifier read a proxy, not the thing. IdentityGate.load()
observed harness_digest, workload_manifest_sha256 and python_reference_commit,
and armed on any JSON carrying those three values. They are values anyone
holding this repository can compute in one line, so the gate proved the
instrument was the instrument and called that a freeze: it answered "is this
the harness?" when the question is "did D7 happen?". There was no check for a
C1 section, no payload commit, no git blob identity, no payload hash and no
ratification. The control was no better — it built a temp JSON out of exactly
those three fields and required the firewall to open.

The freeze is now two objects, because a payload cannot name the commit that
contains it: that sha would have to be inside the bytes hashed into it, which
is the reason detached signatures exist. A payload carries kind/schema, a C1
section and a C2 section, and hashes over its own canonical body. A separately
committed ratification names that body hash and the commit the payload is
frozen at. Arming requires all of: the discriminator, both sections complete,
the recomputed body hash, C1 matching observation, a ratification of that same
hash, the payload's working-tree bytes being the blob at the named commit, and
the ratification itself committed and unmodified. C2 is checked for presence
and never read — its values are thresholds, and an artifact from this stage
that quoted one would leak the number the module exists to keep out.

Arming still costs no source patch, so the harness digest D7 freezes does not
move when the gate arms. It now costs two reviewed commits instead of one text
editor. What it still is not is a signature: anyone with write access could
author both objects, so a declared key is verified against the ratification
commit and a declared "none" is recorded as unsigned. Both stated in the note's
limitations rather than implied away.

P1 — the ladder floor claimed to be a phase it merely contains. core-usage was
recorded as observability "direct" with the note that the interval "IS core
startup", while the invocation it times starts the process, parses argv, writes
a usage refusal and exits. It is now composed over process-startup-core plus a
named cli-argv-refusal phase, and described as a lower bound. The consequence
is corrected too: subtracting it from core-parse-refused does not leave parse,
it leaves read+parse under an assumption this instrument never measures, so the
difference is a derived bound.

Controls: perf-gate-payload-identity damages one property of a real committed
freeze at a time, thirteen cases, each required to be refused by the check that
owns it. One case reformats the committed payload without changing its
canonical hash, so only git object identity can catch it; if that case passes,
the blob check is decorative. perf-gate-arms-by-data now builds a real freeze
in a throwaway repository and additionally asserts no C2 value is recorded.
perf-phase-attribution checks the rung table against itself and against the
shipped report, which turns "re-record after changing a rung" into a gate.
perf-notary-outside now counts git invocations as well as digests: the gate
forks, and obligation (c) is violated by forking inside an interval just as
much as by hashing inside one.

The harness digest changes, so the calibration report no longer certifies the
shipped instrument and is re-recorded in the next commit.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CYLNQy6tLXqsV1CbuNqsSb
…ment

The gate repair and the startup-attribution fix both change perf_baseline.py,
so the harness digest moves from 34fa2717f24d to 5c1cdf0c69a9. The previous
report certified an instrument that no longer exists; a report whose provenance
names a tree that has been superseded is not evidence, it is a receipt.

Re-recorded on a clean tree at 5d6dab4: 44 cells, gate dormant, no decisive
workload timed, environment valid, 44/44 cells reproduced within the 0.35 noise
floor the earlier run recorded. The report now also carries gate_requirements
(the eleven properties a D7 freeze must satisfy) and gate_pinned, so the
evidence describes its own verifier instead of asking a reader to trust the
harness source.

perf-phase-attribution checks the shipped report against the shipped rung
table, so this re-recording is enforced rather than remembered: the control
failed on the stale report and passes on this one.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CYLNQy6tLXqsV1CbuNqsSb
…ong things

CI was green on the previous head and proved exactly what it had been asked to
prove. Four blockers, found by reading rather than by running.

1. The production D7 freeze could never have armed. The verifier took the
payload's own directory as the repository, so it asked git for
`<commit>:p022-263a-d7-ratification.json` for a file that lives at
`docs/evidence/...`. `<rev>:<path>` resolves from the TREE ROOT; only a path
starting `./` or `../` is read relative to the current directory. Every control
passed because the throwaway fixtures put both objects at the repository root,
where the wrong model is accidentally right — a control proving the verifier on
a layout production never has. The root now comes from
`git rev-parse --show-toplevel`, every object path is resolved against it, and
the fixtures use the nested production layout. The ratification must also name
the exact repository-relative path this instrument reads its payload from,
rather than being free to point at any blob with the same bytes.

2. The advertised signed-ratification path was written wrong and could not have
worked. It read `git verify-commit --raw` from stdout; git writes that raw
status to stderr, and the helper discards stderr. A correctly signed commit
would have verified cryptographically and then been refused for not naming its
own key. No control caught it because no environment here holds a signing key.
Rather than fix a path whose accept direction still could not be exercised,
signed mode is withdrawn from the accepted contract: `"signature"` must be
exactly `"none"` and the freeze is recorded as unsigned. An advertised path that
is observably wrong is worse than an absent one.

3. The calibration timed failure and called it reproducible. With no .NET on
PATH the launcher rungs exited 127 before doing anything, and nothing looked at
exit codes: 12 of 44 committed cells reproducibly measured a command-not-found
path, and the CI gate did not check either, since it asked for cells, raw
timings, RSS, provenance and reproducibility but never whether a cell had done
its work. Each rung now declares the exit codes that mean it did its job and a
post-condition proving it, verified once, untimed, with output captured. A cell
that fails is not timed at all, the calibration is refused, the CI gate reads
the verdict, and reproducibility compares outcome identity as well as medians —
two runs agreeing about the wrong path is the worst possible reassurance.

The exit codes and evidence were probed against both engines before being
written down. The first version of the floor check asserted the banner contains
the word "usage"; it does not, and asserting instead of measuring is the exact
habit this contract exists to break.

4. P1 was not closed. `launcher-extract` invoked the launcher with
`--emit-facts` and was documented as "no core runs, so this isolates the
frontend stage". The launcher copies the intermediate facts and then runs Stage
2 anyway, so the interval contained the whole pipeline while recording itself as
launcher startup plus extraction — the same defect as the old `core-usage`
label, one level up, hidden locally because rc 127 never reached Stage 2.
Changing the production launcher is out of scope and a direct extractor process
is not the launcher's extraction stage, so the rung is withdrawn and
launcher-scoped extraction isolation is recorded unavailable.
`frontend-extraction` stays measured as a member of `launcher-e2e`.

Also: `perf-provenance-complete` now requires the committed report's harness and
manifest digests to EQUAL the shipped ones, not merely to be present. Presence
made a stale report tick every box.

Controls: 16 damaged freezes, each matched to the check that owns it, including
a basename `payload_path`, a ratification naming another file, and a signed
ratification. New `perf-rung-outcome` proves 127 is never success, a silent
zero-exit is refused, and the real production floor invocation is accepted.

The harness digest moves again, so the calibration is re-recorded next commit —
this time with a toolchain on PATH and every cell doing its rung's work.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CYLNQy6tLXqsV1CbuNqsSb
…ally there

The previous report was recorded with .NET absent from PATH — installed all
along at ~/.dotnet, never on PATH, which is also why the environment fingerprint
recorded an empty dotnet_sdk and nobody looked. Twelve of its forty-four cells
were command-not-found exits, timed, summarised, and counted as reproduced.

Re-recorded at 55aad67 on a clean tree with the toolchain present:

  40 cells, every one timed and every outcome valid
  launcher-e2e now exits 0 having actually run the pipeline, not 127
  dotnet_sdk 8.0.425 recorded rather than empty
  40/40 reproduced within the 0.35 floor the run itself recorded
  no cell whose outcome changed between runs
  gate dormant, no decisive workload timed

Forty rather than forty-four because launcher-extract is withdrawn: its four
cells measured an interval that contained the whole pipeline while claiming to
isolate extraction.

Two controls now make this re-recording enforced rather than remembered.
perf-provenance-complete requires the report's harness and manifest digests to
equal the shipped ones, and perf-phase-attribution requires the report's rung
table to equal the instrument's. Both failed on the stale report and pass on
this one.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CYLNQy6tLXqsV1CbuNqsSb
Caught by the control added in the previous commit, on the Windows CI leg: the
same tree hashed af32f04ddbad on Linux and 51ef2ca2422a on windows-latest, where
core.autocrlf rewrites source on checkout. One instrument, two identities.

That matters more than a mismatched string. D7's C1 freezes the harness digest
and #263-B has to run on both platforms, so a freeze would have armed on one and
been refused on the other — and the refusal would have looked like tampering.

This repository already knew the defect class. .gitattributes pins
docs/evidence/*.json to LF because the mutation campaigns' definition hashes hit
it first, which is exactly why the workload manifest digest matched across
platforms while the harness digest did not. An attribute only governs files git
checks out under it, though: a tree that predates the rule, a zip download or a
different local config all still differ. An identity D7 will freeze should
depend on content, so it is normalized in the digest rather than delegated to a
checkout setting.

sha256_file stays raw. The candidate binary's identity is its actual bytes, and
a hasher that normalized them would be a different kind of wrong; the control
asserts both directions.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CYLNQy6tLXqsV1CbuNqsSb
…ed digest

The digest fix changes the harness identity a third time, from af32f04ddbad to
69a3a3f8b074 — this time to a value that is the same on both platforms, which is
the point.

Re-recorded at 2feb496 on a clean tree with the toolchain present: 40 cells,
every outcome valid and every cell timed, launcher-e2e exiting 0 having run the
pipeline, dotnet_sdk 8.0.425, 40/40 reproduced with no cell whose outcome
changed, environment not invalidated, gate dormant, no decisive workload timed.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CYLNQy6tLXqsV1CbuNqsSb
One phase name covered two different actions. `cli-argv-refusal` bundled argv
parsing with the usage-refusal message, so it appeared on core-parse-refused,
core-full-human and core-full-sarif — none of which write a usage refusal. The
instrument's own outcome contract already distinguishes them: usage-help versus
door-refusal. The committed calibration carried the same wrong taxonomy, and
perf-phase-attribution compared the report against the rung table and agreed
they matched, which is how a schema stays perfectly self-consistent while saying
something false. Two identical tables are one claim written twice, not evidence.

The methodology note was already more accurate than the code: its table never
put a usage refusal in the successful rungs.

The vocabulary now separates three things. cli-argv-parse is in every real core
invocation. cli-usage-refusal is only the floor rung's driver banner.
ownir-door-refusal is only the strict door's version refusal. A refusal phase
may appear exactly on the rung whose outcome evidence proves that refusal
happened: cli-usage-refusal iff usage-help, ownir-door-refusal iff door-refusal,
and neither on any rung that reaches a verdict. Stated as an equivalence in both
directions, so a rung can neither claim a refusal it did not perform nor omit
one it did.

Applying the same rule the other way, launcher-e2e now also names
process-startup-core and cli-argv-parse. The launcher spawns the core, so those
actions are inside that interval and a list that means "this interval did these
things" has to include them. Over-claiming was the reported defect;
under-claiming is the same rule read in the other direction.

The rules were verified against the defect itself: re-adding cli-usage-refusal
to core-full-human makes perf-phase-attribution fail on both the equivalence and
the reaches-a-verdict rule.

Harness digest moves again, so the calibration is re-recorded next commit.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CYLNQy6tLXqsV1CbuNqsSb
… taxonomy

Splitting cli-argv-refusal into cli-argv-parse, cli-usage-refusal and
ownir-door-refusal changes perf_baseline.py, so the harness digest moves to
84b3e98fc4d9 and the previous report certifies an instrument that no longer
exists.

Re-recorded on a clean tree at 889c285: 40 cells, all outcomes valid and all
timed, 40/40 reproduced within the noise floor the run itself recorded, no cell
whose outcome changed, environment valid, gate dormant, no decisive workload
timed. Every cell's phase list now names only what its interval actually did:
core-full-human no longer claims a usage refusal it never performs, and
launcher-e2e names the core startup and argv parse that happen inside it.

perf-phase-attribution fails on a stale report by design, so this re-recording
is enforced rather than remembered.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CYLNQy6tLXqsV1CbuNqsSb
The Windows leg failed on dfdce55 with two cells outside tolerance:
core-full-human|rust|cal-facts-medium|warm moved 0.355 and
core-usage|python|cal-facts-tiny|process-cold moved 0.362, against 0.35. Every
outcome was identical between the two runs, every exit code matched, and the
noise probe passed at both ends.

That is not a defect in the harness. It is the harness correctly detecting that
a hosted Windows runner does not hold still — the exact risk this PR already
records as its most important finding for #263-B — and the gate converting that
detection into "the instrument is broken" threw the signal away.

Reproducibility now answers the two questions separately. The INSTRUMENT
reproduces when both runs did the same work: same cells, identical outcomes. A
disagreement there is a defect and still fails. The ENVIRONMENT reproduces when
the timings agree within the policy; a disagreement there is a property of the
machine, recorded with its numbers and allowed to be "no".

The tolerance did not move. 0.35 is still 0.35, still derived from the noise
floor the run itself recorded rather than chosen, and the failing cells are
still named with their measured drift. Raising it to make a red run green would
be a threshold fitted to a result, and it is the one number I am least entitled
to pick.

The shipped evidence gets stricter, not looser. A CI leg may now record "this
environment is not measurement-grade" and pass; the COMMITTED calibration may
not. perf-provenance-complete now requires the report of record to have
reproduced on both counts and to carry a run that was never invalidated. Only
the diagnosis of a hosted runner became more honest; the artifact that feeds D7
became harder to produce.

Stated plainly because the timing invites suspicion: this changes a check that
had just gone red. What it changes is which of two questions a disagreement
answers, not the threshold that decides it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CYLNQy6tLXqsV1CbuNqsSb
…nstant

Owner review of the proposed two-axis split. All three points were right.

1. Two axes were not enough, and the second was named for a conclusion rather
than an observation. Calling a timing disagreement "the environment" asserts one
of three possible causes without evidence: a contended runner, a genuinely
variable workload, and a median estimate that is uncertain at this repetition
count are indistinguishable from where the harness stands. The contract is now
outcomes_reproduced, timings_reproduced and environment_valid, with an
admissibility block requiring all three for a report of record. A CI leg exists
to prove the instrument stands up and may pass while recording
timings_reproduced false as a diagnostic; the committed calibration may not, and
perf-provenance-complete checks all three on the shipped artifact.

reproducibility.reproduced keeps the meaning it has always had — everything
agreed — and remains the conjunction rather than a renamed subset. A red result
must not disappear because a word changed profession.

2. The report's claim that the tolerance was "the noise floor recorded by the
earlier run, not a chosen number" was false twice over. The recorded limit IS
the hard-coded constant, so reading it back was reading the constant through a
JSON detour, and nothing about it tightened on a quiet machine. Worse, one
constant was bounding three different statistical quantities: dispersion within
a probe, drift between the opening and closing probes, and a cell's median
change between runs. They are now three named policies. The VALUE is 0.35 for
all three and is deliberately unchanged — it is chosen, and what makes it
admissible is that it was chosen before the runs it judges and is not adjusted
after seeing which of them it rejects.

3. The sample-size argument was a hypothesis wearing a proof's clothes. More
samples reduce the uncertainty of the median ESTIMATE; they do not narrow the
intrinsic spread of the workload's distribution. "Observed IQR 0.404 exceeds the
tolerance, therefore too few samples" does not follow. The escalation is
therefore bounded in advance and recorded in the note: n=15 is the only
permitted sizing step from the observed n=5, the tolerance does not move, the
n=5 evidence is preserved as exploratory and non-admissible, and if n=15 also
fails timings_reproduced the answer is to stop rather than to keep climbing
until CI turns green.

Also recorded, because it is a hole in the provenance rather than a tidy story:
the specific n=5 report that failed was reverted from the working tree before
the preservation rule existed, since a control forbids committing a
non-reproducing calibration and reverting was the reflex. Its numbers survive in
the note; the file does not. The exploratory evidence is a fresh n=5 pair taken
under the shipped instrument, not the run that prompted the change.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CYLNQy6tLXqsV1CbuNqsSb
… is intermittent

The owner's rule was that failed n=5 evidence must survive the sizing decision
it influenced. It nearly did not: the specific report that failed was reverted
from the working tree before the rule existed, because a control forbids
committing a non-reproducing calibration and reverting was the reflex. That hole
is recorded in the note rather than papered over.

What is preserved here is a fresh n=5 pair taken under the shipped instrument —
and it changes the conclusion. It reproduced: admissible on all three axes,
nothing outside tolerance, noise not invalidated. So the original n=5 failure
was intermittent rather than systematic, and "n=5 is too few samples" is no
longer merely unproven, it is weakly contradicted.

That makes the sizing choice precautionary rather than demonstrated, and says so.
The report of record will use n=15 because a report of record should carry the
better-estimated median, which was true before anything went red and does not
depend on n=15 having been the only green option. It was not: both counts
reproduced. What n=5 showed is that its admissibility is luck-dependent on this
machine — a reason to prefer more samples for the artifact, and not evidence
about the workload's intrinsic spread.

The file is named p022-263a-sizing-n5.linux.json so that it cannot match the
glob identifying a report of record. Exploratory evidence that could be mistaken
for canonical evidence is worse than none.

CI keeps n=5: its question is whether the instrument stands up, it may record
timings_reproduced false as a diagnostic, and tripling every leg's runtime would
buy an artifact property that job does not produce.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CYLNQy6tLXqsV1CbuNqsSb
Recorded at ad4c302 with the tree clean, harness digest 2b36de7a7bc9 matching
the shipped instrument: 40 cells, all timed, all outcomes valid, admissible on
all three axes — outcomes_reproduced, timings_reproduced and environment_valid.

n=15 rather than n=5, and not because n=5 was red. A fresh n=5 pair reproduced
cleanly and is preserved alongside as sizing evidence, so both counts are
admissible on this machine. The record uses more samples because a report of
record should carry the better-estimated median, which was true before anything
went red. What the intermittent n=5 failure showed is that its admissibility is
luck-dependent here, not that the workload's spread is wide.

The first attempt at this report was discarded rather than committed: it carried
tree_dirty true, because writing the sizing evidence had modified the tree
before the run. A report of record with dirty provenance is not a report of
record.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CYLNQy6tLXqsV1CbuNqsSb
Four findings, all provenance rather than architecture.

1. A second knob moved. The authorised escalation was exactly repetitions
5 -> 15 with every other rule unchanged. The canonical report was recorded with
--warmup 3 while the shipped default and the n=5 sizing run both use 2, so two
parameters were turned after seeing which runs went red, not one. Warmup is now
a named policy default, DEFAULT_WARMUP_DISCARDS, and a control refuses a
committed calibration that does not use it. Repetitions may differ, because that
is the sanctioned escalation; warmup may not, because changing it should be a
deliberate policy edit in one place rather than a flag someone passes.

2. Only half of each pair was committed. "This pair reproduced" could not be
recomputed from evidence: run B's raw samples were committed, run A's existed
only inside the program that had already reached the verdict. reproduce() now
records the earlier run's path, sha256, byte length, tree, harness digest and
both calibration knobs; a control requires that file to be committed beside its
mate, to hash to the recorded value, and to agree with it on digest, repetitions
and warmup. A pair measured under different settings is not a reproducibility
check.

3. The corrected truth about the tolerance was shipping next to the old lie. The
implementation was right and the emitted report was right, but reproduce()'s own
docstring still said the tolerance was "the noise floor the run itself recorded"
and "tightens on a quiet machine", and the methodology note said the same. Both
are false: the recorded limit is the constant, read back out of JSON and
returning disguised as a measurement, and it never tightened. Source and note
now say it is chosen, and that what makes it admissible is having been chosen
before the runs it judges.

4. The rationale for 15 was the one the owner had already rejected. "A report of
record should carry the better-estimated median" does not pick 15 — it picks 25,
then 100. n=15 is used because a review given before any result was seen
permitted exactly one bounded escalation with policies unchanged and a mandatory
stop on failure. That it came back green is a result, not the reason. The note
now says that; the historical commit messages are left alone, since they are the
record of how the reasoning changed.

The note also stops softening the destroyed evidence: the original failing n=5
artifact is lost, what is preserved is a fresh pair, and the original
observations survive only as narrative, which is not evidence.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CYLNQy6tLXqsV1CbuNqsSb
…tree

Run A of each pair, recorded at 6c33e2a with tree_dirty false. Written outside
the repository and copied in afterwards, because writing run A into the working
tree is precisely what gave the previous attempt dirty provenance: run B then
measured a tree whose sha no longer described what was on disk.

An earlier attempt produced both pairs admissible, with matching hashes, digests
and knobs, and was discarded anyway — both reports carried tree_dirty true, and
a report of record whose tree_sha does not identify what was measured is not a
report of record. That standard was set two rounds ago and it costs nothing to
keep except a re-run.

Neither filename matches the report-of-record glob p022-263a-calibration.*.json,
so a run-A half can never be mistaken for a canonical artifact.

Run B of each pair follows in the next commit, measured against these committed
files on a clean tree.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CYLNQy6tLXqsV1CbuNqsSb
…s withdrawn

The preregistered escalation is spent and it did not work.

  n=5   1 of 40 cells outside REPRODUCIBILITY_MAX_MEDIAN_CHANGE, worst 0.478
  n=15  3 of 40 cells outside it, worst 0.516

Both pairs reproduced their outcomes exactly and both environments were valid.
Only the timings disagreed, and tripling the repetition count produced more
disagreement rather than less.

That is the predicted result, predicted before the run: more samples tighten the
uncertainty of the median estimate, they do not narrow the intrinsic spread of
the distribution being sampled. Reading a single n=5 failure as "too few
samples" was a hypothesis, and this is the evidence against it. Across four
observed n=5 pairs it has now failed twice and passed twice, so n=5 is
intermittent rather than systematically insufficient.

So there is no admissible calibration of record at this head, and none is
manufactured. The permitted escalation was exactly one step with a mandatory
stop on failure. There is no n=25, no n=45, and 0.35 does not move: a tolerance
adjusted after seeing which runs it rejects is a threshold fitted to a result.

p022-263a-calibration.linux.json is DELETED rather than left certifying a
superseded instrument. What ships instead is all four halves of both failed
pairs as exploratory evidence, each recorded on a clean tree, each run B naming
its run A by path and sha256 so the verdicts can be recomputed rather than
trusted. None of the four matches the report-of-record glob.

Two control changes come with it. perf-phase-attribution now finds committed
calibration artifacts structurally instead of by one hardcoded path — naming a
single file meant the check went blind exactly when a record was withheld and
the remaining evidence most needed checking. And pair integrity is enforced on
every artifact carrying a reproducibility verdict, not only on reports of
record, because exploratory evidence makes a two-run claim too.

An observation recorded and deliberately not acted on: every cell that failed in
either pair is a rust cell with a median between 2.5 and 9.2 ms, while the
Python cells at 76-111 ms all held. One relative tolerance is applied uniformly
across cells spanning roughly fiftyfold in duration and bites first where they
are shortest. Whether the policy should be magnitude-aware is a
measurement-design decision for the owner, taken with these numbers visible.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CYLNQy6tLXqsV1CbuNqsSb
Run A and run B diverged on tree_sha AND python_reference_commit in both
committed pairs. A was recorded at 6c33e2a, then COMMITTED, and B was measured
afterwards — and python_reference_commit() is git rev-parse HEAD, so committing
the evidence moved the recorded reference identity while ownlang had not changed
at all. That is the single-reference rule broken by the procedure rather than by
the code.

The pair-integrity check did not catch it. It compared harness digest,
repetitions and warmup, and said nothing about the reference or the candidate
binary. Those pairs happen to share one candidate, but nothing required it, so a
future pair could have compared two different binaries and still reported
reproduced true.

PAIR_IDENTITY_FIELDS now names every axis two halves must share: tree_sha,
tree_dirty, python_reference_commit, workload_manifest_sha256, harness_digest,
calibration_repetitions, warmup_discards, candidate_sha256, candidate_bytes.
reproduce() REFUSES to produce a verdict when any of them differ, so a
mismatched pair cannot be made rather than merely being detectable afterwards;
the control is the second lock on evidence that already exists.

The procedure is corrected with it: both halves measured on one clean source
commit, written outside the repository, verified, then copied in and committed
together. A does not need to be committed before B is measured — only by the
time the evidence ships, which is a different moment.

The existing failed pairs are re-recorded under this procedure in the next
commit. Their qualitative finding is expected to survive; their provenance
currently does not satisfy the contract they are meant to support.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CYLNQy6tLXqsV1CbuNqsSb
@PhysShell PhysShell changed the title feat(perf): #263-A — the measurement instrument, calibrated (CALIBRATION_ONLY) feat(perf): #263-A — the measurement instrument (CALIBRATION_ONLY; no admissible baseline yet) Sep 10, 2026
…cedure

Both halves of both pairs measured on ONE clean source commit (b64ecee), written
outside the repository under their final committed names, verified on all nine
PAIR_IDENTITY_FIELDS, then copied in and committed together. tree_sha,
python_reference_commit, candidate sha256 and every other axis now agree between
A and B, which the previous pairs did not manage.

Two corrected attempts were needed, because earlier_run.path records the
basename it was handed: scratch files under working names left run B pointing at
a file that would never ship. The scratch names are now the final names.

The results disagree with each other, and that is the finding.

  first corrected attempt   n=5 passed, n=15 passed
  second corrected attempt  n=5 passed, n=15 FAILED (one cell, 0.361)

Running tally on this machine: n=5 has failed twice and passed four times; n=15
has failed twice and passed twice. Neither count reproduces reliably and which
one "works" depends on when it ran.

So no calibration of record is restored. Re-running was authorised to repair
provenance, not to obtain a pass, and promoting whichever pair happened to
succeed would be selection regardless of why the re-run happened. Both pairs
ship as sizing evidence, pass and fail alike, and
p022-263a-calibration.linux.json stays absent until the variance
characterisation determines the measurement model.

The four pre-contract artifacts are preserved byte-identical under
docs/evidence/historical/, outside the glob the live controls scan. They fail
the tightened pair check by construction — their earlier_run blocks predate the
fields it requires, and their halves genuinely were measured under two different
reference commits. Not deleted, because they record a real measurement and the
defect that produced it; not left in place, because a control that must exempt
named files is a control with a list of excuses. The move is byte-preserving and
reversible if the owner prefers an explicit exemption instead.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CYLNQy6tLXqsV1CbuNqsSb
The frozen §6 defines C1 as an immutable payload COMMIT and C2 as a detached
attestation in a DESCENDANT commit. The implementation had used the same two
letters for two sections inside one file, with "c2" holding threshold values.
The word had changed profession, and an equivalent-looking scheme is not the
frozen scheme.

Five parts of one P0, all open on 388598e:

1. C1's observation omitted the mandatory #263-A tree sha and harness version.
   It now carries both, alongside the Python reference sha and tree, and the
   workload manifest digest, and every one is compared against observation.
2. The detached object had no payload_blob_sha at all. It now carries one and
   the verifier checks it against the payload's git blob at payload_commit_sha.
3. payload_sha256 was computed over a canonical JSON re-serialisation. §6 says
   the sha256 of the EXACT payload bytes, which is the only version that
   identifies a file rather than a file up to reformatting. A control now feeds
   the old canonical hash and requires refusal.
4. Nothing proved C2 lives in a descendant commit. The verifier now requires
   strict descent and a control builds C2 on a branch forked before C1.
5. python_reference_commit() was git rev-parse HEAD. That is not the reference;
   it is wherever the repository is standing. Every commit moved it, evidence
   commits included, which is how one pair recorded two reference identities
   while ownlang had not changed. At D7 it is worse: the freeze creates C1 and
   then C2, so a correct freeze would invalidate the binding it had just
   written. An alarm that counts its own installation as a break-in. The
   reference commit is now the last commit that touched the reference source,
   with python_reference_tree recording its content-addressed tree object
   beside it. Both are in PAIR_IDENTITY_FIELDS.

Controls: 18 damaged D7 freezes, each matched to the check that owns it, built
as real two-commit repositories. Notable cases — a canonical-JSON hash where an
exact-byte hash is required; a payload reformatted after commit with the
attestation updated to match the new exact hash, so only blob identity can catch
it; and C2 on a sibling branch that forked before C1.

One guard is defensive and untested, and is marked so in source and note: a
same-commit state cannot be constructed with ordinary git, because the
attestation would have to contain the sha of a commit whose sha depends on the
attestation. Recorded as unconstructible rather than faked with a mocked git.

Also applied: the owner's ruling on the "predicted result" wording. The earlier
text claimed more than the evidence carries — nothing predicted that n=15 would
fail on three cells or fail worse than n=5. The note now states the pre-run
alternative it is consistent with, the narrower operational hypothesis it
falsifies, and that it establishes no cause.

IMPLEMENTED FROM THE OWNER'S QUOTATION OF §6. prompt-D7-preregistration-4.md is
not in this session's possession, so fidelity to the full frozen text is
unverified and the quotation may not be all of §6.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CYLNQy6tLXqsV1CbuNqsSb
…entity repair

The digest moved with 637a619, so the previous artifacts certify an instrument
that no longer exists. Both halves of both pairs re-measured on that one clean
commit, written outside the repository under their final names, verified on all
ten PAIR_IDENTITY_FIELDS, then copied in and committed together.

The tenth axis is new and is the point: python_reference_tree. The reference now
reads 20879db while the tree reads 637a619 — a reference identity and a
repository position are finally two different things, so an evidence commit no
longer moves the reference and a D7 freeze no longer invalidates the binding it
just wrote.

Both pairs passed this time. Running tally: n=5 has failed twice and passed five
times; n=15 has failed twice and passed three times. The note records that every
re-record was forced by a provenance repair and never sought for a better
answer, because the raw counts drift toward passing simply because repairs keep
requiring fresh runs, and that drift is not evidence about the workload.

No calibration of record is restored. Round 6 has not run, and a passing pair
does not change what the record requires.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CYLNQy6tLXqsV1CbuNqsSb
…three strings

Two P0s from the owner's review of 2128ec4, both the same defect class this PR
keeps re-learning: a check that observes something other than the thing it
claims to verify.

P0.1 — instrument_tree_sha = git rev-parse HEAD made a real freeze
self-destroying. The lifecycle is S -> C1 -> C2 -> #263-B, so by the time the
decisive run happens HEAD is at least C2 and can never equal the accepted tree
S. A freeze satisfying that check would have to name a future content-addressed
commit inside the commit that determines it. The twin defect was fixed one
round earlier for python_reference_commit, with the argument spelled out in its
docstring, and the identical sentence about instrument_tree_sha went unread
three lines away.

The anchor is now DECLARED by C1 as instrument_anchor_commit and PROVED by
content: S exists, the instrument at S hashes to the harness_digest C1 froze,
and S is an ancestor of C1. C1 is already a strict ancestor of C2 and C2 is
found in HEAD's history, so S <= C1 < C2 <= HEAD falls out without asserting a
position anywhere. Because the workload manifest is one of INSTRUMENT_SOURCES,
proving the harness digest at S also proves the manifest at S. harness_digest()
and harness_digest_at() share one formula, since the anchor proof compares them
and two implementations would make their agreement meaningless.

The fixtures could not see any of this: they committed a README and two JSON
files to a throwaway repository while observe() read the real one, so the
payload was built from the same HEAD it was compared against. _build_freeze now
commits the real instrument sources at S, does ordinary work, freezes C1 and
C2, and keeps committing; perf-gate-arms-by-data asserts anchor, C1, C2 and
HEAD are four distinct commits. Reinstating "anchor must equal HEAD" makes a
valid freeze be refused — verified by mutation.

P0.2 — C1's protocol was three key names checked for presence, and the fixture
proving that check worked armed the gate with three copies of the string
"<frozen by D7, never read by the gate>". Two immaculate commits, three
celebratory strings, decisive firewall open. The gate still never READS a
threshold, but that was never the same question as whether a preregistration is
there. C1 must now carry per cell: all four dimensions; bound, pass_fail_rule,
inconclusive_band, comparison_statistic, rss_policy, allocation_policy; a
repetition_ladder with initial_n, escalation_stages, transition_predicates,
max_n and terminal_outcome; or an explicit not_applicable with a stated reason.
Roll-ups must carry workload_class, phase and overall_g3. _present() inspects
presence and container shape only — nothing compares, orders or records a
value. Fixture cells use deliberately non-production placeholders, because a
fixture that only passed with plausible numbers would prove the gate reads
thresholds.

Damaged D7 freezes: 18 -> 33, each matched to the phrase the check that owns it
must produce. Verified by mutation: blinding the protocol verifier fails ten
cases, not one.

Also, self-found while unifying the digest readers: load_manifest() hashed RAW
bytes while the harness digest normalized them — the Round 3 defect still alive
in a sibling path, on a value C1 also freezes. Normalization now lives in one
function. The value is unchanged on Linux, so this closes a Windows exposure
without moving any number.

Also: the note's pair-identity paragraph listed nine axes and omitted
python_reference_tree, which the code has carried since the reference identity
was separated from HEAD. It is generated from the tuple now and reads ten.

The owner has read prompt-D7-preregistration-4.md and closed the previous
"§6 fidelity unverified" caveat. The cell and roll-up key names here come from
the owner's review message, not from the frozen file, which is still not in
this session's possession. The gate does NOT check cell-set completeness: that
needs the decisive cell population, and a completeness check written from a
guess would be a threshold decision wearing a schema's coat.

ruff 0, mypy Success (43 files), selftest 0, controls 13/13, run_tests 0.
The harness digest moves with this commit, so both sizing pairs are stale and
are re-recorded separately.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CYLNQy6tLXqsV1CbuNqsSb
CI failure on aaf0d42, ubuntu leg:

    OSError: [Errno 39] Directory not empty:
      '/tmp/perf-d7-anchor-not-ancestor-opknvimr/repo/.git'

Mine, and new. `git commit` spawns `gc --auto`, which keeps writing inside
.git after the commit returns; `TemporaryDirectory` teardown then races it and
rmtree dies. The anchor fixtures added in the previous commit commit more often
than the old ones — an orphan branch, a drift commit, post-freeze work — which
made the race likely enough to bite on a runner. It did not reproduce locally,
which is the whole character of the thing.

Fixed at the cause: fixture repositories now run with gc.auto=0 and
maintenance.auto=false. They live for one check and are deleted, so they have
nothing to gain from housekeeping and everything to lose from a background
writer.

Also ignore_cleanup_errors on the two fixture temp directories. Disposing of a
throwaway repository is not a property under test, and it must never be able to
read as a gate failure — the check that failed here had already passed.

Verified: 4 consecutive runs of the control file, no leftover fixture
directories, no stray git processes. ruff 0, controls 13/13.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CYLNQy6tLXqsV1CbuNqsSb
…pair

The anchor and protocol repairs move the harness digest, so every pair recorded
before them describes an instrument that no longer exists. Both are re-recorded
at f8e7112 on a clean tree, both halves written outside the repository under
their final committed names, run B naming run A by path and sha256.

Both pairs are admissible: 40 of 40 cells timed, no cell refused,
outcomes_reproduced, timings_reproduced and environment_valid all true.

Tally after this round: n=5 has failed twice and passed six times; n=15 has
failed twice and passed four times. Four consecutive passes are recorded and
explicitly NOT banked. The last two re-records were forced by provenance
repairs, not sought for a better answer, but a tally where passes outnumber
failures three to one is exactly where selection starts to look like evidence.
The failures stay in the ledger and in docs/evidence/historical/ rather than
being summarised away, and p022-263a-calibration.linux.json stays absent.

Note also corrected: the pair-identity paragraph is generated from the tuple and
reads ten axes, not the nine it listed while the code carried
python_reference_tree.

One recording defect of mine, caught by the instrument rather than by me: the
first re-record refused 8 cells and then 4. I passed --candidate as a relative
path, so OWEN_RUST_CORE did not resolve for the launcher surface and those
invocations exited 127 or 2. The rung outcome contract from round 2 refused to
time them instead of reporting a command-not-found path as a measurement. With
an absolute path it is 40 of 40 again. The environment was never at fault.

ruff 0, mypy Success (43 files), selftest 0, controls 13/13, run_tests 0.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CYLNQy6tLXqsV1CbuNqsSb
…l universe

Both halves of the owner's P0.2, found on 055b8e4.

P0.2a — not_applicable: true was ACCEPTED, and the PR description said it was
refused. _present() ended in `value is not None`, and in Python a bool is an
int, so _present(True) and _present(False) were both True. Both
{"not_applicable": true} and {"not_applicable": false} passed a check whose
whole purpose was to demand a stated reason. I wrote the claim that a bare true
was refused and never tested it — an untested assertion about a guard, published,
in the same document that argues a green gate proves only what it was asked to
prove.

Fixed as a class, not an instance: _present() rejects bool outright (checked
before int, since a bool is one), so a yes/no can never stand in for a
preregistered value in ANY field. not_applicable must additionally be a
non-empty string — a reason in words. A cell carrying not_applicable alongside
any rule field is refused too: applicable and inapplicable at once is not a
decision.

P0.2b — the verifier checked every cell it was handed and never asked whether it
had been handed them all. D7 says "for every applicable (phase x workload-id x
platform x cold/warm regime) cell", so a payload with one immaculate cell and
three roll-up strings passed. The old fixture demonstrated the hole instead of
catching it: _valid_payload carried three synthetic cells and armed the gate.

expected_d7_cells() enumerates 11 phases x 12 canonical decisive workloads x 2
platforms x 2 regimes = 528 cells, from the frozen taxonomy and the frozen
manifest. Missing, unknown, or duplicated dimension tuples are all refused, and
each expected cell needs a complete rule or an explicit not_applicable with a
reason.

The alias earns no second vote: large-solution-control is alias_of
oss-ShareX.sln, declared under its own id because #262's gate names it and
recorded as an alias so it is never counted twice in a denominator.
canonical_workload_id() resolves it, and a payload naming both for the same
phase/platform/regime is refused as a duplicate. A completeness check that
missed this would have turned a double-counting safeguard into a double-counting
machine.

The universe covers BOTH platforms wherever the gate runs. #263-B runs on both,
so a Linux gate demanding only Linux cells would let half the experiment through
unfrozen.

This selects nothing. No budget, no N, no tolerance; no value is read, compared
or recorded. Applicability is expressed by the payload — the gate only proves
the owner decided something about every cell. The previous note argued
completeness "needs the decisive cell population" and recorded it as not
implementable. That was wrong, and wrong in the convenient direction: the brief
permits enumerating, hashing and availability-checking decisive workloads before
D7 and forbids only performance exposure, and it has to, because a decisive
population unknown before D7 could not be written into D7.

Damaged D7 freezes: 33 -> 39, with catchers for boolean-true N/A, boolean-false
N/A, N/A beside a rule, a missing cell, an unknown cell and the ShareX
duplicate. Mutation-verified in both directions: restoring the old _present arms
the gate on both boolean N/A cases, and skipping the completeness block arms it
on all three cell-set cases.

ruff 0, mypy Success (43 files), selftest 0, controls 13/13, run_tests 0.
The harness digest moves again, so both pairs are re-recorded separately.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CYLNQy6tLXqsV1CbuNqsSb
The N/A and completeness repairs move the harness digest again, so both pairs
were stale. Re-recorded at 43b6a4a on a clean tree, both halves outside the
repository under their final committed names.

n=15 reproduced. n=5 did NOT: one cell,
core-full-sarif|rust|cal-facts-tiny|process-cold, at 0.508 against the 0.35
policy. Both runs are valid, 40 of 40 cells timed, outcomes reproduced in both
pairs, environments valid; only the timings axis failed, on that one cell.

The failing pair ships exactly as recorded. It is NOT re-run: obtaining a
passing partner for a failed run is the single move this instrument exists to
make impossible to hide, and the permitted repetition counts are still only 5
and 15 with nothing to escalate to.

The previous commit's note warned that four consecutive passes were exactly
where selection starts to look like evidence. The very next attempt failed.
That is not a vindication of the caution so much as a demonstration of what it
was about: the streak was never evidence, and neither is its ending.

Tally: n=5 has failed three times and passed six; n=15 has failed twice and
passed five. Five DISTINCT cells have now fallen outside tolerance at least
once, and every one of them is a rust cell with a median under 10 ms — the
newest at 2.93 ms. The magnitude observation from earlier rounds is unchanged
and still deliberately not acted on: whether the relative tolerance should be
magnitude-aware is a measurement-design decision for the owner.

Still no calibration of record. p022-263a-calibration.linux.json stays absent.

ruff 0, mypy Success (43 files), selftest 0, controls 13/13, run_tests 0.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CYLNQy6tLXqsV1CbuNqsSb
P0.2c from the owner's review of 34f7ef3. Round 9's completeness check was
mechanically sound and proved the completeness of the wrong set.

expected_d7_cells() enumerated sorted(PHASES). PHASES is the ATTRIBUTION
taxonomy: it answers what a timed interval contains, so a rung can be described
honestly. It is not D7's vocabulary, and using one as the other was wrong in
both directions at once. It demanded a preregistered rule for cli-argv-parse,
cli-usage-refusal, ownir-door-refusal and frontend-extraction on every workload,
platform and regime. It omitted the end-to-end C# run entirely — one of the
metrics the G3 verdict weighs most — because that is a rung, not a part, and has
no entry in PHASES at all.

The code had already said this and was not listened to. The accept fixture keyed
its not-applicable slice on phase == "frontend-extraction", auto-excusing all 48
of those cells: the instrument admitting they were never D7's business while the
universe went on requiring them. A check that manufactures an obligation and
then ceremonially forgives it is paperwork, not a check.

D7_PHASES is now a separate constant — process-startup-core,
process-startup-launcher, ownir-parse, bridge-lowering, analysis, render-human,
render-sarif, end-to-end-csharp — and the relationship between the two
vocabularies is data rather than coincidence. D7_NON_METRIC_PHASES records each
attribution component that is not a gate together with the reason it is not.
D7_METRICS_WITHOUT_PHASE records end-to-end-csharp as a metric whose interval no
single attribution phase names.

d7_vocabulary_problems() enforces that every attribution phase is either
promoted to a metric or explicitly excluded WITH a reason; neither is a default,
so a phase added to PHASES later cannot silently join or silently miss the D7
universe. It runs in --selftest as well as in the new perf-d7-phase-universe
control, because the failure mode here is drift rather than a wrong line.

Universe: 8 metrics x 12 canonical decisive workloads x 2 platforms x 2 regimes
= 384 cells, down from a 528 that was neither complete nor correct.

Controls 13 -> 14; damaged freezes 39 -> 40 with proto-attribution-phase, which
dresses cli-argv-parse as a G3 metric and is refused as a cell D7 does not
cover. Mutation-verified three ways: dropping end-to-end-csharp, promoting
cli-argv-parse, and rebuilding the universe from PHASES are each caught, and the
first two fail --selftest as well.

NAMES ARE PROVISIONAL, and that is the actual lesson. D7_PHASES comes from the
owner's review and the live #262 performance-gate list, not from the frozen
file. The mapping between these ids and the frozen #262/#263-A vocabulary needs
one explicit ratification, because this defect is precisely what happens when an
implementation label is left to become a normative contract by default. RSS and
allocation deliberately stay per-cell policies rather than becoming a phase
axis; the frozen schema already models them that way.

ruff 0, mypy Success (43 files), selftest 0, controls 14/14, run_tests 0.
The harness digest moves again, so both pairs are re-recorded separately.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CYLNQy6tLXqsV1CbuNqsSb
Separating D7_PHASES from PHASES moves the harness digest, so both pairs were
stale again. Re-recorded at 843d6fe on a clean tree, both halves outside the
repository under their final committed names.

Both pairs reproduced: 40 of 40 cells timed, no cell refused, outcomes and
timings reproduced, environments valid, both admissible.

Tally: n=5 has failed three times and passed seven; n=15 has failed twice and
passed six. Six attempts in, n=5 reads pass, pass, fail, pass, fail, pass.
Nothing about that is a trend, and the ledger exists so no revision of the note
can quietly turn it into one. Five distinct cells have fallen outside tolerance
at least once across the whole history, every one a rust cell with a median
under 10 ms.

Still no calibration of record. p022-263a-calibration.linux.json stays absent,
and the failing pairs stay in the ledger and in docs/evidence/historical/.

ruff 0, mypy Success (43 files), selftest 0, controls 14/14, run_tests 0.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CYLNQy6tLXqsV1CbuNqsSb
Two records, no source change. The harness digest is byte-identical
(2d6e52fe4352…), so both committed pairs stay valid and nothing is re-recorded.

RATIFICATION. The owner has ratified the eight D7 metrics and the four
exclusions against the frozen prompt-263A-measurement-design-4.md, the final D7
brief and live #262, at exact head 0374793 whose CI completed success on both
runs. Metrics: process-startup-core and process-startup-launcher as two, per the
frozen text; ownir-parse, bridge-lowering, analysis; render-human and
render-sarif as the two concrete surfaces of the frozen CLI/SARIF rendering
item; and end-to-end-csharp as the whole user-visible rung. Not metrics:
cli-argv-parse, cli-usage-refusal and ownir-door-refusal are attribution only,
and frontend-extraction is diagnostic population B excluded by frozen #263-A
itself. The denominator is ratified with them: 8 x 12 x 2 x 2 = 384, with 12
canonical workloads because large-solution-control aliases oss-ShareX.sln.
Ratifying 384 does not mean 384 numeric budgets — an explicit not_applicable
with a reason stays a legitimate decision. It means the owner must decide about
each cell rather than lose it between two lists.

The source comment still reads NAMES ARE PROVISIONAL, deliberately. At the
commit that wrote it they were. Editing it now would move the harness digest,
stale both pairs and force an eleventh re-record — relabelling the fire
extinguisher and re-calibrating the laboratory because of it. The ratification
binds in the ledger and to the exact head it was given at.

PREREGISTRATION. docs/notes/p022-263a-round6-preregistration.md fixes the Round
6 analysis before any Round 6 data exists: a synthetic deterministic helper that
is neither Rust nor Python product code, fresh process per iteration, eight
target durations from 1 to 128 ms, n in {5,15}, warmup 2, five independent A/B
sessions per (scale, n) — 160 measurement runs. Every raw observation, MAD, IQR,
absolute and relative median shift, scheduler metadata, CPU affinity and a
context-switch signal are recorded. No pass/fail tolerance is computed or
applied and 0.35 is not consulted.

Three outcomes are named in advance with what each licenses. O1, absolute drift
roughly constant while relative shift grows at short durations, licenses
DESIGNING a v2 envelope |delta| <= A + R*t with A and R from the metrology
dataset only. O2, a stable ladder at 2-8 ms while real Rust invocations wander,
forbids an absolute floor. O3, neither shape, is inconclusive and stops. O3 is
named deliberately: without it an ambiguous dataset gets read as whichever of O1
or O2 the reader was hoping for.

Forbidden and written down: deriving A or R from the Owen cells, using 0.508 or
any observed failing value toward a threshold, an absolute floor under O2,
choosing T_min/N_min/N_max, moving 0.35 or warmup, any repetition count other
than 5 and 15, and every decisive surface — no decisive workload or timing, no
resource exposure of decisive work, no D7 threshold or budget, no C1/C2 freeze,
no #263-B, no Stage 3.

Existing calibration history may be read as calibration evidence, which frozen
#263-A permits as an input to sizing the instrument and the owner has ratified.
The observation that all five historical failures sit on rust cells under 10 ms
may be investigated; it may not become a number in a D7 acceptance rule.

ruff 0, selftest 0, controls 14/14. No measurement has been taken.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CYLNQy6tLXqsV1CbuNqsSb
The Round 6 apparatus, committed before it has produced a number.

scripts/round6/spin.c is the synthetic helper: deterministic work, no I/O, no
allocation, neither Rust nor Python product code. scripts/round6/metrology.py
drives the preregistered ladder — setup pass fixes one iteration count per rung
and records the helper's sha256; measurement pass runs 8 rungs x {5,15}
repetitions x 5 independent A/B sessions and records every raw observation, MAD,
IQR, absolute and relative median shift, plus CPU affinity, load average and
context-switch counters either side of each session.

AMENDMENT 1 to the preregistration, made before any measurement pass ran and
recorded there with its reason: the helper is compiled rather than interpreted.
The instrument's timed interval includes process spawn, and on this machine
python3 -c pass has a floor of 11.8 ms against /bin/true at 1.14 ms — so a
Python helper could not reach the 1, 2, 4 or 8 ms rungs at all. The compiled
helper's floor is ~1.25 ms. The 1 ms rung therefore sits at or below the spawn
floor; it is kept and flagged rather than dropped, because it bounds what a
relative tolerance can mean at that scale, which is close to the question this
round asks.

Timing goes through perf_baseline.Harness._run_once deliberately — that is the
instrument's own interval, perf_counter_ns around Popen and wait4. Round 6
characterises THAT interval, so a re-implementation here would characterise a
different one and let the resemblance do the arguing. Writing one formula twice
is the defect class this PR keeps finding.

No verdict, no tolerance, 0.35 never consulted, no decisive workload spawned or
named. The firewall is never invoked because no workload is passed to it.

ruff 0. No measurement has been taken.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CYLNQy6tLXqsV1CbuNqsSb
The setup pass fixes one iteration count per rung and records it, so the helper
is not re-deriving its own duration on every run and measuring a moving target.
Run at 8c61a7b on a clean tree; the counts were written into the preregistration
BEFORE the measurement pass started, which is what that document promised.

Helper f937b36bd4da, spawn floor 1.332 ms, ~690k iterations per ms.

| target | iterations | achieved |
|--------|-----------:|---------:|
|   1 ms |          0 |  1.448 ms  (at/below spawn floor)
|   2 ms |    460,280 |  2.013 ms
|   4 ms |  1,838,383 |  4.090 ms
|   8 ms |  4,594,589 |  8.499 ms
|  16 ms | 10,107,001 | 16.496 ms
|  32 ms | 21,131,826 | 31.691 ms
|  64 ms | 43,181,474 | 63.326 ms
| 128 ms | 87,280,771 | 127.313 ms

The ladder tracks its targets from 2 ms upward. The 1 ms rung IS the spawn
floor: it is kept and flagged rather than dropped, and read as a bound rather
than a duration, because what a relative tolerance can mean at that scale is
close to the question this round asks.

The artifact lives in docs/evidence/round6/ rather than beside the calibration
reports. That keeps it outside the flat report-of-record glob by construction
instead of by key shape — the live controls would skip it either way, but a
directory boundary does not depend on a schema staying the way it is today. It
is metrology, not a calibration report, so it does not belong in that glob at
all.

No verdict, no tolerance, 0.35 not consulted, no decisive workload involved.

ruff 0, selftest 0, controls 14/14.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CYLNQy6tLXqsV1CbuNqsSb
80 A/B sessions: 8 rungs x {5,15} repetitions x 5 independent sessions, every
raw observation kept, run at 8c61a7b on a clean tree with helper f937b36bd4da.

OUTCOME O2, read off the preregistered table rather than chosen after the fact.
The synthetic ladder is stable exactly where the real Rust invocations wander,
so an absolute floor must NOT be introduced and the cause is elsewhere.

A measurement-scale effect is real but small: relative shift rises from ~0.006
at 8-16 ms to ~0.04-0.05 at 1-2 ms, a fixed cost over a shrinking denominator.
That much of O1's description holds.

O1 is nevertheless NOT licensed, because its precondition fails. It required
absolute drift roughly constant across the ladder; absolute drift instead grows
from 0.056 ms to 1.126 ms at n=5, about twentyfold.

The decisive comparison, which is what the round was built to make. In the
2-10 ms band — every rust cell that has ever failed lives there — 27 A/B
sessions of the ladder produced a median relative shift of 0.0131, a worst
relative shift of 0.1564, and a worst absolute drift of 0.3203 ms. The smallest
drift any failing cell required is 1.23 ms (0.396 on a 3.12 ms median); the
largest is 4.41 ms. Worst-case ladder against minimum-case failure is short by
3.9x. The ladder never crossed 0.35 at any rung at either repetition count.

Short duration alone therefore does not produce the instability. Something the
real invocations do, and this helper does not, is producing it.

Recorded because someone will eventually be tempted the other way: reading the
table charitably as O1 and fitting A + R*t anyway yields an envelope TIGHTER
than 0.35 at these durations, rejecting the Rust cells more often rather than
fewer. This dataset cannot be used to loosen anything.

BOUNDARY OF THE CONCLUSION, stated rather than buried. The helper is a small
compiled binary doing pure arithmetic — no file reads, no parsing, no large
dynamic image. own-cli is ~1.8 MB and opens files. The ladder isolates DURATION
and does not isolate WHAT A REAL INVOCATION DOES; dynamic loading, page-cache
state and file I/O are unexamined and are the obvious next suspects. O2's
conclusion is "the cause is elsewhere", and a reader is entitled to know how
much elsewhere was actually searched.

Unchanged, per the preregistration and standing rulings: no absolute floor, no
v2 envelope, 0.35 unmoved, no T_min/N_min/N_max, no threshold derived from any
number here, no calibration of record restored.

ruff 0, mypy Success (43 files), selftest 0, controls 14/14, run_tests 0.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CYLNQy6tLXqsV1CbuNqsSb
Docs only. The harness digest is unchanged (2d6e52fe4352), the dataset is
untouched, and Round 6 is NOT re-run. Its provenance is sound and re-running a
measurement to obtain a tidier verdict is the move this instrument exists to
prevent.

The owner rejected the formal O2 verdict and was right on all three counts.

1. The reading consulted 0.35, which the preregistration forbids in as many
words: "No pass/fail tolerance is computed, applied, or reported ... 0.35 is not
consulted." The withdrawn reading observed that the ladder "never once crossed
0.35" and argued from it. The measurement driver never reads 0.35, so the
dataset is uncontaminated — the interpretation failed, not the collection.

2. O1 and O2 were never operationalised as numbers. "Roughly constant" and
"stable at 2-8 ms" were English, not decision rules. The quantities that ended
up deciding — twentyfold drift growth, 0.3203 ms against 1.23 ms, a factor of
3.9 — were all chosen after seeing the data. They are legitimate descriptive
findings and they are not a preregistered boundary, so "O2, read off the
preregistered table" claimed more than the plan could support.

3. The A + R*t fit was performed in a branch that did not license it. The plan
permits deriving A and R only under O1; the reading selected O2 and reported the
fit anyway. No coefficients were published, but the analysis ran where the plan
disallowed it.

Formal outcome is therefore O3, inconclusive, in the owner's words: the dataset
is descriptively O2-like, but the preregistration did not operationalise the
O1/O2 boundary tightly enough to support a formal O2 verdict, so the
preregistered outcome is conservatively O3; the observed separation is retained
as exploratory calibration evidence and may motivate a separately preregistered
follow-up.

The numbers survive as observations licensing nothing: drift grows ~20x across
the ladder so O1's own written precondition does not hold; the scale effect is
visible (relative ~0.006 at 8-16 ms rising to ~0.04-0.05 at 1-2 ms); in the
2-10 ms band the ladder's worst absolute drift was 0.3203 ms while the
historical failures implied 1.23 to 4.41 ms. That separation is a reason to keep
looking, not a verdict.

The envelope fit is quarantined under an explicit POST-HOC / NOT PART OF THE
PREREGISTERED OUTCOME / NOT POLICY INPUT heading rather than deleted, because
deleting an analysis that was actually performed is its own dishonesty.

NEXT-SUSPECT LIST CORRECTED, from the owner. It led with file I/O and should not
have: core-usage|rust|cal-facts-tiny|process-cold is one of the historical
failures and that rung never opens an OwnIR document at all — it starts the Rust
core, parses argv, writes a usage refusal. Input-document I/O is NOT a necessary
condition for the instability, and this round's own evidence said so while the
reading looked past it. What survives: executable/library mapping and page
faults, dynamic loader/runtime startup, scheduler context switches and CPU
migration, CPU-time against wall-time, and the output/refusal path. core-usage
is where a follow-up should start — the minimal real Rust process path, already
a witness, with document parsing removed from the picture.

No new measurement is authorised, and none was taken.

ruff 0, selftest 0, controls 14/14.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CYLNQy6tLXqsV1CbuNqsSb
…ssertion

Step 2 of the owner's order. The harness digest is unchanged (2d6e52fe4352) —
the control lives in tests/, which is not an instrument source, so no pair is
staled and nothing is re-recorded.

python_reference_commit and python_reference_tree content-address exactly
ownlang/. That identity is only honest if the reference's behaviour is confined
to what it names; otherwise the reference could change behaviour without
changing its own sha, which is the defect class of an identity that names the
checkout.

Settled by observing. An audit hook records every file the reference OPENS and
every process it spawns, on both render surfaces. Result: ZERO repository files
opened outside ownlang/, no process spawned, 18 ownlang modules loaded.

Three further routes checked and closed. Static imports: ownlang imports the
standard library and itself, nothing else in the tree. Config: config.py states
it has no auto-discovery, no environment variables and no per-path overrides — a
config enters only via an explicit --config, which no rung passes, so there is
no consulted-if-present file to slip past an audit that only sees opens.
Subprocesses: the fix_* repair pipeline does spawn, but none of it loads on the
measured path.

perf-reference-boundary keeps this proven rather than asserted. Controls 14 ->
15. Mutation-verified both ways: a probe that opens pyproject.toml fails it, and
one that spawns /bin/true fails it.

ONE RESIDUAL, RECORDED AND DELIBERATELY NOT FIXED. OWNLANG_DEBUG is read by
ownlang/__main__.py and the instrument passes ambient environment through
unchanged. It affects only the internal-error path — on a successful run the
branch is never taken, so it cannot alter analysis output — and when it fires it
returns 70, which no rung declares, so the outcome contract refuses the cell
rather than timing it. Fail-closed, but still an uncontrolled input the
reference identity does not cover. Pinning it edits perf_baseline.py, moves the
digest and stales both pairs; re-recording is not currently authorised, so this
is worth batching with any other digest-moving change rather than spending a
re-record on it alone.

ruff 0, mypy Success (43 files), selftest 0, controls 15/15, run_tests 0.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CYLNQy6tLXqsV1CbuNqsSb
…v input

Step 3 of the owner's order, batched with the OWNLANG_DEBUG residual from step
2. Both move the harness digest, so doing them together costs one staleness
event instead of two.

MYPY. scripts/perf_baseline.py was outside mypy's file list, so "Success, 43
files" never covered the file that produces the numbers. It is in the list now —
44 files — and passes --strict.

The 42 errors it exposed were almost entirely one root cause: JSON-parsed values
are object, and the code called .get(), int() or float() on them behind `x or
{}` idioms with type: ignore comments that had drifted to the wrong error codes.
Four narrowing helpers — _as_obj, _as_list, _as_int, _as_float — say it once, to
the reader and the checker at the same time.

They narrow fail-closed. A non-object where an object belongs becomes empty, so
the caller's own completeness check refuses it rather than raising an
AttributeError three frames from the cause. _as_int and _as_float reject bool
for the same reason _present does: a yes/no is not a measurement.

Six errors were not narrowing at all. ctypes.windll and Popen._handle exist on
Windows and are absent from the stubs mypy checks against on Linux; those carry
targeted type: ignore[attr-defined] comments, one API each with the reason
recorded, rather than a blanket suppression that would also hide the next real
defect.

ENV PIN. REFERENCE_ENV_PINNED_UNSET names the environment inputs the reference
reads that its content-addressed identity does not cover, and every measured
invocation clears them. OWNLANG_DEBUG is the only current member. It affects
only the internal-error path so it cannot change analysis output on a successful
run, but an identity that content-addresses ownlang/ does not cover an
environment variable, and the same reference source should not behave two ways
depending on the shell it was launched from.

NO CONTROL CHANGED AND NONE WAS WEAKENED: 15/15 throughout, full suite green at
every step. That is the only reason to believe a mechanical rewrite of this size
did not quietly alter behaviour.

BOTH PAIRS ARE NOW STALE AND ARE NOT RE-RECORDED. The digest moved from
2d6e52fe4352 to 6713e7300c7c, so all four committed halves describe an
instrument that no longer exists. They are left in place and the note says so:
re-recording is a measurement, and new measurements are not currently
authorised. Stale evidence that admits it beats fresh evidence nobody asked for.

ruff 0, mypy Success (44 files), selftest 0, controls 15/15, run_tests 0.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CYLNQy6tLXqsV1CbuNqsSb
…, not run

Step 5 of the owner's order. Docs only; the harness digest is unchanged
(6713e7300c7c) and no measurement exists.

THE QUESTION, narrowed by the owner's correction. Round 6 was formally O3 and
descriptively showed short duration alone looks insufficient. The follow-up must
NOT lead with file I/O: core-usage|rust|cal-facts-tiny|process-cold is a witness
and that rung never opens an OwnIR document — it starts the Rust core, parses
argv, writes a usage refusal. So core-usage is the minimal real Rust process
path, already a witness, with document parsing removed from the picture.

DESIGN. Round 6 varied duration and held shape fixed; Round 7 does the opposite.
Three arms at one duration: A, the Round 6 C helper at its 2 ms rung; B, the
same helper with byte-identical work padded to own-cli's ~1.84 MB, isolating
image size and mapping; C, own-cli ownir with no arguments, the real witness.
D(C)-D(A) is the excess to explain, D(B)-D(A) the part from image size, and
D(C)-D(B) the part from everything else the real binary does — loader symbol
resolution, runtime startup, argv handling, the refusal write. Ten independent
A/B sessions per arm and repetition count, n in {5,15}, warmup 2: 120 runs.

QUANTITATIVE BOUNDARIES, which is what Round 6 lacked. P1 through P5 are
arithmetic on D(arm,n), the median across sessions of |median_B - median_A|:
P1 image dominates if D(B) >= 3*D(A) and D(C) <= 1.5*D(B); P2 not image if D(B)
<= 1.5*D(A) and D(C) >= 3*D(A); P3 both if D(B) >= 3*D(A) and D(C) >= 2*D(B);
P4 witness did not reproduce if D(C) <= 1.5*D(A); P5 otherwise. A rule must hold
at BOTH repetition counts or the outcome is P5. Mechanism attribution is
likewise fixed in advance as ratios over utime+stime, wall, context switches and
page faults against arm A.

Ratios rather than absolute milliseconds on purpose: an absolute boundary would
have to come from the failing Owen cells, which are forbidden as input, or from
a number invented to fit, which is worse. A ratio against a baseline arm
measured in the same sessions cannot import a threshold by accident.

DECLARED UPFRONT, NOT DISCOVERED LATER: Harness._run_once discards ru_utime,
ru_stime, ru_minflt, ru_majflt, ru_nvcsw and ru_nivcsw — exactly the fields that
separate the candidate causes. Round 7 needs it to return the full rusage. That
is a source change, it moves the harness digest, and it stales any pair recorded
before it. APPROVING THIS PLAN MEANS APPROVING THAT CHANGE AND THE RE-RECORD IT
IMPLIES; if the change is not wanted, the round cannot run as designed.

0.35 is not consulted anywhere, including in the reading. That was Round 6's
defect and it is not repeated. Reporting an outcome the table did not produce is
forbidden: suggestive numbers with no rule firing are P5 plus an exploratory
note, not a verdict.

Two design choices are left open for the owner rather than assumed: whether a
statically linked fourth arm should separate the dynamic loader from image size,
and whether both regimes are measured or only process-cold where every witness
lives.

Step 6 is a separate review and authorisation. Step 7 is execution. Neither has
happened.

ruff 0, digest unchanged, no measurement taken.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CYLNQy6tLXqsV1CbuNqsSb
Four decisions and two corrections, none of them mine to make:

  * full POSIX rusage in _run_once — YES
  * a fourth, statically linked arm — NO, and not carried as a footnote
  * arms A/B/C — kept
  * process-cold only, or both regimes — BOTH
  * run Round 7 now — NO

Arm B claimed too much. It said it isolated "image size, mapping and page-fault
cost"; Linux demand-pages an image, so padding that is mapped and never read
need not fault in at all. B now claims the padded image footprint and its
mapping metadata, D(B) - D(A) is stated as a LOWER bound, and a B that
reproduces nothing reads as "a padding-only image effect is insufficient" rather
than "image mapping and page faults are excluded". A four-part structural
preflight (padding survived linking, sits in a PT_LOAD, matches arm C's mapped
size to within a page, work bytes identical to arm A) is a stop condition rather
than a description written afterwards.

The outcome table was not mutually exclusive: A=1, B=3, C=1.4 fired P1 and P4 at
once. Ten rounds of insisting a check must read the thing, and I preregistered
the numbers and forgot to preregister the logic. The owner's ratified set
replaces it, with P1 and P2 renamed to claim only what arm B can see. Every
remaining pairwise overlap reduces to D(A)=0 or D(B)=0 exactly; a degenerate-zero
guard is PROPOSED and flagged as NOT YET RATIFIED rather than folded in quietly.

Also corrected: the draft said every witness lives in process-cold. One of the
five, core-full-sarif|rust|cal-facts-small|warm, does not. Both regimes are
measured and classified independently; a cold/warm split is reported as a
cache-sensitivity signal, not collapsed into P5. The regimes are stated as the
instrument implements them — cold 0 discarded iterations, warm 2, both a fresh
process per iteration — instead of the draft's self-contradiction. Session
halves are no longer also called A and B.

Still a draft. Execution is not authorised.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CYLNQy6tLXqsV1CbuNqsSb
_run_once reaped every child through os.wait4, which returns the kernel's
complete per-child rusage, and read one field of it. User CPU, system CPU, minor
and major faults, and voluntary and involuntary context switches were fetched
and dropped on the floor — and those six are exactly what separates "the work
took longer" from "the scheduler moved it" from "the image cost more to map".
Without them a variance round can only report that something got slower.

Four constraints, each a way this could have been made wrong:

  * ru_utime/ru_stime are ROUNDED to nanoseconds. Truncation biases every
    sample the same direction, and a systematic half-tick error survives
    averaging in a way a symmetric one does not.
  * the interval is unchanged in meaning: t0, Popen, wait4, elapsed. The
    rusage is parsed after the clock stops.
  * the two pre-existing lines between wait4 and elapsed stay where they are.
    Moving them would TIGHTEN the interval, and a tightened interval silently
    un-compares every future number against every recorded one.
  * off POSIX all six are null WITH A REASON. A reported 0 minor-fault count
    for a platform nobody asked would be the most confident possible lie, and a
    median over such zeros would look exactly like a flat measurement.

The fields reach the cell too, not only the private driver: accounting
(summarized in ns for the durations and count for the counters),
accounting_unavailable_reason, and raw_accounting per iteration. Collecting the
accounting and then dropping it one layer up would be the same defect wearing a
different hat.

perf-child-accounting is the sixteenth control. It checks the claim that would
otherwise be invisible — that the parse sits outside the clock — by wrapping
os.wait4 so every rusage field costs 10 ms to read, then asking whether
elapsed_ns grew. Seven mutations, each caught by the check that owns it. The
last two needed a second attempt: both make measure_cell raise KeyError partway
through a cell, so the first version recorded its finding and then died before
reporting it, which in CI is a traceback naming the crash site and no FAIL line
naming the cause. Those checks are terminal now.

Harness digest 6713e7300c7c -> 562a7f7232da. The four committed sizing halves
were already stale and are still not re-recorded.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CYLNQy6tLXqsV1CbuNqsSb
The owner did the algebra properly and mine was wrong in both directions.

I claimed every remaining overlap reduced to "A=0 or B=0". It does not. Overlap
occurs only when A=0 AND B=0 together, and P1 is disjoint from every other rule
unconditionally — its C > 1.5A clause excludes P4 outright, and every other
intersection forces B=0 and then C=0, which contradicts that same clause. So
"I checked every pair" was again a stronger claim than the checking behind it,
which is the specific habit this PR exists to make expensive.

The guard I proposed was correspondingly too wide. Keying it on B would have
refused A=1, B=0, C=4 — a clean P2 where the padded arm shows no excess and the
real binary does, one of the cleanest results this round can produce. The
ratified rule keys on A alone, and not because of exclusivity: A is the
multiplicative reference for 1.5A, 3A and 6A, and against a baseline of zero any
positive drift is infinitely large. Definite, and metrologically degenerate.

No epsilon. A tolerance like A < 0.01 ms would become an absolute threshold with
no preregistered basis for its value, which is the move this round forbids.
Exact zero is a structural degenerate case, not another knob.

The exclusivity claim is no longer asserted anywhere: a control evaluates it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CYLNQy6tLXqsV1CbuNqsSb
One classifier, five controls, four structural checks. No clock anywhere in it,
and none may be added: the calibration pass itself is not authorised.

scripts/round7/classify.py is the ONLY implementation of P1-P5 and the zero-A
guard. A second copy would be a second opinion and the reading could then choose
between them. It is pure — numbers in, an outcome out — so it is exercised
exhaustively before any measurement exists: 50,653 exact rational triples, in
sixths so the ratified edges 3/2, 3, 2/3 and 2 land exactly on grid points, with
Fraction arithmetic throughout because a binary-float 2/3 would decide the
(2/3)B <= C edge by rounding direction rather than by the rule.

The census runs twice. Against the real rules, which must overlap only at
A=B=0; and against the previous draft's P1, which must be reported broken. A
control that has only ever seen correct input is a control nobody has tested.

Six mutations of the real classifier, each caught by the check that owns it: the
old overlapping P1; the zero guard widened to A-or-B; the guard removed;
ambiguity resolved by precedence instead of raising; P4's edge made strict; and
an epsilon smuggled into the zero comparison. The second is the guard the owner
rejected, and the control now refuses to let it back in.

The crash-before-reporting defect recurred, in a file written after I recorded
it. Removing the guard makes A=B=C=0 fire three rules, the classifier raises, and
the control died on a traceback: non-zero exit, mutation scored as caught, no
FAIL line naming anything. Reading the output rather than the exit code is what
caught it both times.

The preflight reads the ELF rather than a rendering of it — program and section
headers parsed directly, and the padding located through st_shndx, which names
its section, instead of by matching addresses, an inference that finds nothing
at all for a non-allocated section whose address is zero. The four damage cases
are built, not simulated; the first attempt at the unmapped one failed to
produce any damage, which is why it was checked before being believed.

B1-B4 all pass. Padding 1,626,024 bytes in .rodata inside the PT_LOAD at 0x2000;
arm B maps 1,628,917 bytes against arm C's 1,628,892, +25 against a 4,096-byte
allowance; main is 107 bytes and byte-identical in both arms, so B4 decided on
raw bytes and the disassembly fallback was not needed.

The apparatus is not the instrument: the harness digest is unchanged at
562a7f7232da across the whole addition.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CYLNQy6tLXqsV1CbuNqsSb
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants