Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
63 commits
Select commit Hold shift + click to select a range
f09fa00
feat(perf): #263-A — the measurement instrument, calibrated and firew…
claude Sep 10, 2026
3118095
evidence(perf): the CALIBRATION_ONLY validation report
claude Sep 10, 2026
e9d3bce
fix(perf): a contended runner is a result, not a job failure
claude Sep 10, 2026
790c872
docs(perf): the Windows finding is sharper than "the runners are noisy"
claude Sep 10, 2026
5d6dab4
fix(perf): the D7 gate must verify a freeze, not a file that echoes i…
PhysShell Sep 10, 2026
35befb2
evidence(perf): re-record the calibration against the repaired instru…
PhysShell Sep 10, 2026
55aad67
fix(perf): four review findings — the instrument was measuring the wr…
claude Sep 10, 2026
0590ea7
evidence(perf): re-record the calibration on a toolchain that is actu…
claude Sep 10, 2026
2feb496
fix(perf): the harness digest named the checkout, not the instrument
claude Sep 10, 2026
438a567
evidence(perf): re-record the calibration against the content-address…
claude Sep 10, 2026
889c285
fix(perf): a phase list must mean the interval did those things
claude Sep 10, 2026
dfdce55
evidence(perf): re-record the calibration against the corrected phase…
claude Sep 10, 2026
b3080e4
fix(perf): a runner that will not hold still is not a broken instrument
claude Sep 10, 2026
6aacbc5
fix(perf): three axes for reproducibility, and three names for one co…
claude Sep 10, 2026
ad4c302
evidence(perf): preserve the n=5 sizing evidence, and record that n=5…
claude Sep 10, 2026
63e57db
evidence(perf): the calibration of record, at n=15 on a clean tree
claude Sep 10, 2026
6c33e2a
fix(perf): warmup is not a sizing knob, and a pair needs both its halves
claude Sep 10, 2026
742dad7
evidence(perf): run A of both calibration pairs, measured on a clean …
claude Sep 10, 2026
29deb17
evidence(perf): both sizing pairs failed; the calibration of record i…
claude Sep 10, 2026
b64ecee
fix(perf): a pair is one experiment, or it is two
claude Sep 10, 2026
388598e
evidence(perf): re-record both pairs under the corrected identity pro…
claude Sep 10, 2026
637a619
fix(perf): implement D7 §6 literally — two commits, not two sections
claude Sep 10, 2026
2128ec4
evidence(perf): re-record both pairs under the D7 §6 and reference-id…
claude Sep 10, 2026
aaf0d42
fix(perf): the D7 anchor is provenance, and a preregistration is not …
claude Sep 10, 2026
f8e7112
fix(perf): stop git housekeeping racing the fixture teardown
claude Sep 10, 2026
055b8e4
evidence(perf): re-record both pairs under the anchor and protocol re…
claude Sep 10, 2026
43b6a4a
fix(perf): close the boolean N/A escape hatch and check the whole cel…
claude Sep 10, 2026
34f7ef3
evidence(perf): re-record at 43b6a4a; n=5 failed and ships as it landed
claude Sep 10, 2026
843d6fe
fix(perf): D7's metrics are not the instrument's attribution taxonomy
claude Sep 10, 2026
0374793
evidence(perf): re-record at 843d6fe under the D7-vocabulary repair
claude Sep 10, 2026
71c5832
docs(perf): ratify the D7 metric mapping, and preregister Round 6
claude Sep 10, 2026
8c61a7b
feat(perf): Round 6 metrology helper and driver, no measurement yet
claude Sep 10, 2026
4f7171c
evidence(perf): Round 6 setup pass — the calibrated duration ladder
claude Sep 10, 2026
406bc85
evidence(perf): Round 6 metrology dataset and its reading — outcome O2
claude Sep 10, 2026
de43401
docs(perf): withdraw the formal O2 verdict; Round 6 is formally O3
claude Sep 10, 2026
1a6990d
test(perf): close the Python reference boundary by observation, not a…
claude Sep 10, 2026
1a49bc8
fix(perf): put the instrument under mypy --strict, and pin the one en…
claude Sep 10, 2026
e57322e
docs(perf): draft the Round 7 preregistration — DRAFT, not authorised…
claude Sep 10, 2026
48cfd38
docs(perf): amend the Round 7 preregistration under the owner's review
claude Sep 10, 2026
3c055a6
feat(perf): keep the per-child accounting wait4 was already handing back
claude Sep 10, 2026
99795cd
docs(perf): correct the overlap algebra and take the ratified zero rule
claude Sep 11, 2026
8db808b
feat(perf): build the Round 7 apparatus and pass the B1-B4 preflight
claude Sep 11, 2026
9d15d36
feat(perf): close the three Round 7 execution-contract gaps, still no…
claude Sep 11, 2026
9a0779d
feat(perf): bind Round 7's timed path to the work its arms exist to do
claude Sep 11, 2026
7ad4a0b
fix(perf): make a refused Round 7 run leave a black box, not a traceback
claude Sep 11, 2026
1128ddd
feat(perf): run the one authorised Round 7 pass; outcome (P4, P5)
claude Sep 11, 2026
6c58dc1
fix(perf): withdraw an unsupported provenance claim; state the spawn …
claude Sep 11, 2026
6cf4bf4
docs(perf): propose the #263-A calibration reproducibility policy, wi…
claude Sep 11, 2026
c7c5e88
docs(perf): revision 2 of the calibration policy - a computable funct…
claude Sep 11, 2026
eae1807
docs(perf): revision 3 of the calibration policy - two P0, two P1, st…
claude Sep 11, 2026
87a66f7
fix(perf): close the zero-versus-zero hole in Round 7 mechanism attri…
claude Sep 11, 2026
1754cf4
feat(perf): step 2 of the calibration policy - a pure implementation …
claude Sep 11, 2026
8df8514
fix(perf): enforce the calibration policy's exact domain at every ent…
claude Sep 12, 2026
b4f657a
fix(perf): close the calibration policy's public arithmetic surface
claude Sep 12, 2026
2a037b8
freeze(perf): step 4 — the calibration policy digest freeze
claude Sep 12, 2026
eedf6d3
record(perf): step 5 — the owner's ratified design constants
claude Sep 12, 2026
9dd2359
prereg(perf): step 6 — the training collection preregistration
claude Sep 12, 2026
ea75061
fix(perf): step 6 — stop CI taking observations the preregistration f…
claude Sep 12, 2026
4fef796
fix(perf): step 6 — the measurement guard was fail-OPEN on its author…
claude Sep 12, 2026
9de3ab4
fix(perf): step 6 — the measurement guard was blind to a line break
claude Sep 12, 2026
44a4df2
test(perf): step 6 — make the multiline bypass a standing regression …
claude Sep 12, 2026
78460f4
feat(perf): step 7 — the measurement-free environment identity manifest
claude Sep 12, 2026
fd3f09f
fix(perf): step 7 — three defects the first Windows CI run exposed
claude Sep 12, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
114 changes: 114 additions & 0 deletions .github/workflows/ci.yml
Original file line number Diff line number Diff line change
Expand Up @@ -2894,6 +2894,120 @@ jobs:
[ "$rc" -ne 0 ] || { echo "FAIL: a broken candidate exited 0"; exit 1; }
echo "OK: a broken candidate is a visible failure (exit $rc), never a Python rescue"

# #263-A — the performance measurement INSTRUMENT, calibrated. Not a
# baseline: this job proves the instrument stands up, on both platforms, and
# that it refuses to measure anything decisive before the D7 freeze.
#
# It publishes no engine comparison and can publish none: the harness holds
# one engine's samples per cell and refuses a report carrying a comparison
# key. What it asserts here is that the refusal is real, not that Rust is
# anything.
perf-instrument:
name: "#263-A measurement instrument (CALIBRATION_ONLY)"
strategy:
fail-fast: false
matrix:
os: [ubuntu-latest, windows-latest]
runs-on: ${{ matrix.os }}
defaults:
run:
shell: bash
steps:
- uses: actions/checkout@34e114876b0b11c390a56381ad16ebd13914f8d5 # v4
with:
# The step 4 freeze controls read git objects at the frozen source
# commit and prove it is an ancestor of this head. A depth-1 checkout
# cannot answer either question, and those controls refuse rather than
# pass vacuously — so the history is the thing that makes them real.
fetch-depth: 0
- uses: actions/setup-dotnet@67a3573c9a986a3f9c594539f4ab511d57bb3ce9 # v4
with:
dotnet-version: "8.0.x"
- uses: actions/setup-python@a26af69be951a213d495a4c3e4e4022e16d87065 # v5
with:
python-version: "3.13"
- uses: dtolnay/rust-toolchain@fa04a1451ff1842e2626ccb99004d0195b455a88 # master, 2026-07-10
with:
toolchain: stable
- name: Build the production own-cli candidate
working-directory: rust
run: cargo build -p own-cli --release
- name: The instrument's own selftest (no measurement taken)
run: python scripts/perf_baseline.py --selftest
- name: The instrument controls (firewall, both identity domains, provenance)
run: python tests/test_perf_instrument.py
# The Round 7 apparatus is a classifier and a structural preflight, both
# exercisable without a clock. The exclusivity control is the one that
# matters: the preregistration's first outcome table let one dataset
# satisfy two contradictory verdicts, and nothing in the process that
# produced it would have noticed.
- name: Round 7 apparatus controls (outcome exclusivity, zero guard, B1-B4)
run: python tests/test_round7_apparatus.py
# The calibration policy is pure: no clock, no files, no randomness, and no
# constant values at all. It is step 2 of the ratified sequence, written
# before any training corpus exists, so every claim it makes is checkable
# here on both platforms and none of it needs a measurement.
- name: Calibration policy controls (no defaults, symmetry, exact fit)
run: python tests/test_calibration_policy.py
# Step 4 froze WHICH BYTES the policy is, and nothing else. These controls
# recompute that digest from git blobs at the frozen commit, and refuse the
# artifact outright if it ever grows a field that is not identity.
- name: Calibration policy freeze controls (step 4 digest identity)
run: python tests/test_calibration_freeze.py
# Step 5 records the five constants the owner chose, bound to the step 4
# identity. The frozen implementation is the judge of its own ranges, so
# these controls hand it the committed pairs rather than re-stating them.
- name: Calibration design constants controls (step 5 ratified values)
run: python tests/test_calibration_constants.py
# Step 6 preregisters a training collection that has not happened. These
# controls prove the protocol is complete and bound, and that the scope
# combiner reads the frozen selector's own margin instead of copying it.
- name: Training preregistration controls (step 6 protocol + scope binding)
run: python tests/test_training_prereg.py
# Step 7 needs two dedicated measurement hosts that do not exist yet. This
# captures WHAT a machine is and never how fast it is: no clock, no resource
# observation, proved by walking the AST rather than the prose. It also
# proves the tool sits outside all three frozen source sets, so preparing
# step 7 cannot move a digest steps 4, 5 and 6 are already bound to.
- name: Step 7 environment capture controls (measurement-free identity manifest)
run: python tests/test_step7_envcapture.py
# The same tool run for real on this runner. It writes nothing, starts no
# clock, and is NOT a binding input: a hosted runner is not admissible for
# step 7, and the manifest records the runner and CI flag that say so.
- name: Step 7 environment capture selftest (writes nothing, starts nothing)
run: python scripts/step7/envcapture.py --selftest
# Plan mode builds the arms, preflights THOSE bytes, freezes their
# identities and fixes the execution order — then returns before the
# timing function is reachable. Linux only: the arms are ELF binaries.
- name: Round 7 runner — plan mode, no clock started
if: matrix.os == 'ubuntu-latest'
run: |
python scripts/round7/runner.py \
--candidate "$PWD/rust/target/release/own-cli" \
--plan --build-dir "$RUNNER_TEMP/round7-arms" \
--out "$RUNNER_TEMP/round7-plan.json"
# Decisive workloads: availability and pins only. No clock, no resource
# sampling — this step is what "enumerated, hashed, availability-checked"
# looks like when it is the only thing permitted.
- name: Decisive workloads — non-timed structural smoke
run: python scripts/perf_baseline.py --smoke-decisive
# The timed legacy calibration pair USED TO RUN HERE, and it is gone while
# step 7 is closed. It started a real stopwatch, produced forty timed cells
# per platform on every push and uploaded them, which made the step 6
# preregistration's own claim to take no new observation false on the very
# commit that made it. A preregistration that fixes "exactly one collection"
# while CI quietly gathers numbers every twelve minutes has built an
# epistemic side channel, whatever anyone intends to do with the output.
#
# What still proves the apparatus, all of it untimed: --selftest above, the
# instrument controls, the round 7 apparatus controls, the round 7 runner in
# plan mode, the decisive structural smoke, and the calibration policy,
# freeze, constants and training preregistration controls.
#
# training-no-incidental-measurement fails this suite if any workflow
# regains a path to a measurement entrypoint while the preregistration
# still records step 7 as unauthorised.

# P-014 Tier B: external-reference resolution. The SAME sample, run two ways, must give two
# verdicts — proving the extractor binds a THIRD-PARTY event only when its DLL is referenced:
# A (no refs) -> ObservableObject is an error type -> OWN050 (honest skip), no leak
Expand Down
26 changes: 26 additions & 0 deletions docs/evidence/calibration/p022-263a-design-constants.json
Original file line number Diff line number Diff line change
@@ -0,0 +1,26 @@
{
"artifact": "p022-263a-calibration-design-constants",
"bound_measurement_harness_digest": "562a7f7232dad2f4c79c6adfe0e1e7e25680b4b6f3444d54824bf0405e3c14b3",
"bound_policy_implementation_digest": "c3068ed7fa880a7083866ead25fe8bf65c87889d242d8af1f7eee01582cd5cbf",
"constants": {
"G": [
1,
10
],
"M": [
2,
1
],
"N_ladder": [
5,
15,
45
],
"R_runs": 5,
"q": [
19,
20
]
},
"ratified_by": "owner"
}
17 changes: 17 additions & 0 deletions docs/evidence/calibration/p022-263a-policy-freeze.json
Original file line number Diff line number Diff line change
@@ -0,0 +1,17 @@
{
"artifact": "p022-263a-calibration-policy-freeze",
"measurement_harness_digest": "562a7f7232dad2f4c79c6adfe0e1e7e25680b4b6f3444d54824bf0405e3c14b3",
"policy_implementation_digest": "c3068ed7fa880a7083866ead25fe8bf65c87889d242d8af1f7eee01582cd5cbf",
"policy_implementation_digest_framing": "sha256 over the source set ordered by the UTF-8 bytes of each repo-relative POSIX path. Each file contributes, with no header and no separator: its path byte length as an 8-byte big-endian unsigned integer, its path's exact UTF-8 bytes, its blob byte length as an 8-byte big-endian unsigned integer, and its exact git blob bytes.",
"policy_source_commit": "b4f657a0abdfdfaae199cbc7eee0c47acd8b0057",
"policy_source_files": [
{
"blob_sha1": "ec3fab8bbd8d5e0d0a0e1c6f60e7249d97f3ef4f",
"path": "scripts/calibration/policy.py",
"size_bytes": 25486
}
],
"policy_source_root": "scripts/calibration/",
"policy_source_selector": "all committed *.py recursively under policy_source_root at policy_source_commit",
"policy_source_tree": "2f49574433fea23f6158f4d8544a3b1543c718d4"
}
170 changes: 170 additions & 0 deletions docs/evidence/calibration/p022-263a-training-preregistration.json
Original file line number Diff line number Diff line change
@@ -0,0 +1,170 @@
{
"abort_semantics": {
"durable_record": "every abort writes a structured record naming stratum, rung, run index, cell where applicable, and the condition",
"failed_invocation_or_exit_contract": "invalidates the run",
"identity_drift": "invalidates the stratum",
"missing_cell_or_incomplete_rung": "invalidates the run; a partial rung is not a smaller rung",
"no_automatic_retry": "a refusal is recorded and the collection stops; there is no retry because CI blinked",
"noise_probe_refusal": "invalidates the run that carried it",
"repeat_after_abort": "permitted only by a rule already written here, or by a new owner ruling",
"rss_mechanism_unavailable": "recorded and the affected field is null with a reason, never zero",
"timeout_or_runner_loss": "invalidates the run"
},
"admissibility_of_observations": {
"quarantined": [
{
"labels": [
"POST_PREREG_LEGACY_DIAGNOSTIC",
"INADMISSIBLE_FOR_TRAINING",
"NEVER_FIT",
"NEVER_SELECT_N"
],
"produced": "two CALIBRATION_ONLY reports of 40 timed cells per platform, uploaded as workflow artifacts",
"source": "CI run 2038, the legacy calibration pair, ubuntu-latest and windows-latest"
}
],
"rule": "Every timing/resource observation produced outside the single owner-authorised Step-7 training collection is permanently inadmissible as fitting, N-selection, holdout, or D7 evidence, regardless of whether it predates or postdates this preregistration.",
"scope": "The labels are a record, not the mechanism. The mechanism is the rule above, which is unconditional, plus training-no-incidental-measurement, which stops a workflow regaining a path to a measurement entrypoint while step 7 is closed.",
"why_not_only_pre_step_6": "An earlier wording excluded only data recorded before this preregistration, which is narrower than the rule needs to be. It would have readmitted anything a later rerun produced, on the entirely sincere ground that it is just a diagnostic."
},
"anchor_commit": "eedf6d3ed44ecf7bda69fa509dcd23b45960cf76",
"artifact": "p022-263a-calibration-training-preregistration",
"bindings": {
"design_constants_blob_sha1": "d03c8455cbcaf71446f88aeea0957b8e0c56c880",
"measurement_harness_digest": "562a7f7232dad2f4c79c6adfe0e1e7e25680b4b6f3444d54824bf0405e3c14b3",
"policy_implementation_digest": "c3068ed7fa880a7083866ead25fe8bf65c87889d242d8af1f7eee01582cd5cbf",
"training_scope_implementation_digest": "614bf9efe6ba2bb10e26a251c10d6c4a8fdaa99985d43ea03d76c6774447d25b",
"training_scope_root": "scripts/training/"
},
"collection_protocol": {
"environment_rule": "the whole ladder and all runs of one stratum belong to ONE preregistered environment identity for that stratum; measuring one rung on one runner and another rung on a different runner makes N not the only variable that changed",
"observation_rule": "consecutive run pairs (r1,r2), (r2,r3), (r3,r4), (r4,r5), never a chosen five pairs after the outcome is visible",
"observations_per_cell_per_rung": 4,
"runs_per_rung_per_stratum": 5,
"runs_per_stratum": 15,
"runs_total": 30,
"warmup_discards": 2
},
"design_constants": {
"G": [
1,
10
],
"M": [
2,
1
],
"N_ladder": [
5,
15,
45
],
"R_runs": 5,
"q": [
19,
20
]
},
"exactly_one_collection": {
"identity": "the collection carries a durable identity from its first run",
"outcomes": [
"ADMISSIBLE",
"INVALID"
],
"rule": "exactly one training collection is run; it is not the first satisfactory one of several",
"stratum_failure": "if either stratum invalidates, the cross-platform collection yields no admissible step 8 input; the surviving stratum's numbers remain durable evidence of what happened and must never become a single-stratum fallback model"
},
"forbidden_in_step_6": [
"any clock",
"any new timing or resource observation",
"any fitted A_abs or R_rel",
"any selected N",
"any fitted width",
"any change to the frozen policy or the measurement harness"
],
"future_procedure": {
"no_common_n": "an orchestration stop, not a fifth reproducibility verdict; the frozen four describe cells and run pairs and are untouched. No fallback rung, no stricter-stratum tie-break, no second margin, no second fit",
"note": "described here so it is fixed before data; step 6 produces no numeric result of it",
"rejected_alternatives": {
"averaged_widths": "lets an unstable stratum buy a stable one a wider bound",
"max_of_per_stratum_picks": "can name a rung admissible on neither stratum",
"one_stratum_as_reference": "puts the other into D7 with no measurement model of its own"
},
"steps": [
"per stratum, pool its cells within a rung and fit the envelope by the frozen fitter",
"per stratum, compute each rung's width and the admissible rung set by the frozen selector",
"intersect the strata's admissible sets and take the smallest ladder rung in the intersection",
"if the intersection is empty, stop with NO_COMMON_N"
]
},
"holdout": "not authorised by this preregistration and not described by it; the validation pair is a separate owner decision",
"identity_drift_rule": "any change to a required identity field after collection starts invalidates that stratum's training collection; two epochs are never glued into one dataset",
"order": {
"seed_derivation": "the first 16 hex digits of policy_implementation_digest, frozen at step 4 and owner-accepted before this preregistration existed, so it is not a number that could have been re-rolled until the order looked tidy",
"seed_hex": "c3068ed7fa880a70",
"stratum_order": "linux then windows, fixed",
"within_stratum_order": "ladder rungs ascending; within a rung the five runs are sequential; cell order within a run is the instrument's own fixed order and is not re-shuffled"
},
"ratified_by": "owner",
"run_identity_required_fields": [
"tree_sha",
"tree_dirty",
"python_reference_commit",
"python_reference_tree",
"workload_manifest_sha256",
"harness_digest",
"calibration_repetitions",
"warmup_discards",
"candidate_sha256",
"candidate_bytes",
"provenance.environment",
"provenance.timestamp_utc",
"raw samples per cell"
],
"step_7_collection": "NOT AUTHORISED",
"stratification": {
"across_strata": "observations, medians, envelopes and widths are never pooled, averaged or summed across strata",
"cell_identity": [
"rung",
"engine",
"workload",
"regime"
],
"platform_is_a_cell_axis": false,
"regime_rule": "cold and warm are distinct cells of the frozen four-part identity and pool within a stratum",
"strata": [
"linux",
"windows"
],
"within_stratum": "every cell of U_p pools into one fit per ladder rung, by the frozen fitter unchanged"
},
"universe": {
"applicability_rule": "the universe is exactly the cell set the frozen instrument emits for the six calibration workloads; a rung that does not apply to a workload is absent rather than N/A, and the count is asserted by a control rather than trusted",
"calibration_workloads": [
"cal-facts-tiny",
"cal-facts-small",
"cal-facts-medium",
"cal-facts-refused",
"cal-source-small",
"cal-source-medium"
],
"decisive_workloads": "excluded, and the CALIBRATION_ONLY firewall refuses them structurally",
"engines": [
"python",
"rust"
],
"expected_cells_per_stratum": 40,
"pre_step6_data": "excluded from U, and so is every observation taken outside the single authorised step 7 collection whenever it was taken; see admissibility_of_observations",
"regimes": [
"process-cold",
"warm"
],
"rungs": [
"core-usage",
"core-parse-refused",
"core-full-human",
"core-full-sarif",
"launcher-e2e"
]
}
}
Loading
Loading