You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Master roadmap after the completion of P-022 step 4 in #214 / PR #249.
Reconciled at 206e9c71ec62ceef2668e74d1c62228a402a03a2 (PR #347 merge — #261's 261.B, the production Rust OwnIR executable own-cli ownir, landed: the own-cli crate with its single ownir subcommand, a Python-authored CLI fixture replayed with zero Python on Linux and Windows CI, both failure-mode rulings measured under an off-by-default fault-injection feature, and the ownir_version Version family fixed to byte parity — V1/V2/V4 declared and V3 reproduced per #262's parser-domain rulings — carried onto P-022 row 7b, the proposals index and rust/README.md; #261 subsequently closed completed). Reconciled at 48b8799ee54b6152161d4b90a20d03de41170f5f (PR #346 merge — the #261 decision-packet ratification of 2026-09-08, C-1..C-5 with the exit-code rulings and the owner's rulings on the four readings, carried onto P-022 rows 7b and 8, the proposals index and rust/README.md; this body, #261, #262 and the new #345 moved with it), at f48780630780b7f38f239469a1d7d0a87f30d356 (PR #344 merge, the CI re-record of #260's sweep) and at 4520a543e0c47a886217d065b3d80920199d4f93 (PR #343 merge), carrying #260's final acceptance — the sweep onto this surface together with docs/proposals/P-022-rust-core-migration.md and the proposals index. Previous reconciliations: b05b38a (PR #342, #260's acceptance surfaces over the committed corpus), 21fb0c3 (PR #341, #259 final acceptance), 834f295 (PR #340, which also carried #337, #338 and #339), 3fc1246 (PR #336), 0738d29 (PR #325), 984de7d (PR #324) and fdcb222 (PR #322). Statuses are checkpoint-level: each step names its completed checkpoints, its remaining acceptance, its normative blocker (what its acceptance actually requires) and this roadmap's preferred sequencing (what order is cheapest). Conflating the last two is what made the P-022 table drift twice.
How to keep parity work from drifting again is now written down: P-022 § Parity-work discipline. It is the single home for those rules — deliberately not copied elsewhere, because two copies of one law drift.
Since 3fc1246 no count is typed on a status surface: the census, the surface inventory and every mutation campaign are rendered from the tree into docs/generated/ by scripts/render_checkpoint_status.py, and the Python suite fails while a fragment is stale, a campaign result no longer matches its definition, was taken on a dirty tree, missed a required catcher, names a catcher that no longer exists, or names a commit the tree does not descend from.
Two independent outcomes must now be delivered:
Owen Alpha is actually published and verified from clean external consumers.
Rust participates in the real C# → OwnIR → verdict path, first in shadow mode and then through an explicit cutover gate.
The release track must not wait for the full Rust migration. The Rust track must not use the release as permission to weaken parity.
OwnIR defensive limits (feat(ownir): bound source coordinates and nesting depth (Python-first) #326, Python-first, on main): signed-64 source coordinates and a 32-level nesting limit, normative in spec/OwnIR.md §4.2 and carried by spec/ownir.schema.json. The reference used to accept coordinates and depths no other consumer could represent — not generosity, but an accident of CPython's unbounded integers and deep stack leaking into the contract. Closed in the contract, not by widening Rust.
Generated checkpoint evidence (PR docs(P-022): generate the cp4 census and mutation evidence from the tree #337): the verdict-ledger census and every mutation campaign are facts in the tree (manifest + goldens; a machine-readable campaign definition + its raw recorded result), interpreted once (tests/verdict_census.py, scripts/mutate_campaign.py) and projected into docs/generated/*.md by scripts/render_checkpoint_status.py; tests/test_checkpoint_status.py runs the --check in the suite. The campaign runner keeps the M00 honesty control, refuses fail-fast, labels every catcher by package and target, refuses to run on a dirty tree, and a caught mutation whose named catchers did not all fail voids the evidence at the gate. Every later checkpoint added its fragments to the same pipeline; since PR feat(p022): #260 acceptance over the committed corpus — byte-attested same input, the verdict layer in scope, the derived SARIF gate, the compare driver #342--validate also resolves every named catcher and refuses a Python mutation that does not parse.
Bridge full fact-to-verdict parity (P-022 step 6b: implement Rust own-bridge with layered OwnIR parity #259 cp5, PR feat(bridge): #259 cp5 — the full Finding, the refusal text and the rendered surfaces, proven #339):complete at the cp5 surface. A surface inventory (tests/verdict_surface_inventory.py) names every BR-V4 wording branch with who owns the string, every BR-V5 evidence family and degradation rule, and every BR-V9 rule, with coverage computed from the goldens; own_bridge::Finding carries message, related and flow, and the replay compares every member; own_cfg::Diag carries the resolver's message so refusals compare in full (removing that cut exposed a py_repr quoting defect the cut had hidden — fixed and pinned against CPython); ownlang/renders.py + tests/fixtures/verdict_renders/ freeze the rendered surfaces and a Rust replay compares the bytes. Branches no facts document can reach get their expected text from the reference's own recorded output (tests/test_unreachable_branch_probe.py), re-runnable. No pre-existing golden was regenerated or edited.
Bridge protocol analysis (P-022 step 6b: implement Rust own-bridge with layered OwnIR parity #259 4b, PR feat(p022): port the obligation-protocol analysis and wire it through the bridge (#259 cp4b) #340):complete — Layer 3 parity over the measured set, protocol family included.ownlang/obligations.py is ported into own-analysis (BR-B1); its typed values come from the one grammar in own-ir the strict door already delegated to, so door and analysis cannot drift into two readings; a third analysis-level fact-parity family freezes every violation member with zero Python; the bridge maps BR-P3 in its BR-V1 place; refuse_protocols is gone and both protocol documents are promoted out of the exclusion ledger without regenerating either golden. Two things measured rather than claimed: BR-V5's two-step rule is not applied on the protocol path (a Python-first decision still owed), and the family's append position is unobservable end to end.
Shadow-mode acceptance surfaces over the committed corpus (P-022 step 7a: add dual-engine shadow mode and zero-diff reproduction artifacts #260, PR feat(p022): #260 acceptance over the committed corpus — byte-attested same input, the verdict layer in scope, the derived SARIF gate, the compare driver #342): the two owner decisions P-022 step 7a: add dual-engine shadow mode and zero-diff reproduction artifacts #260's acceptance waited on were ratified as D-4..D-7, B-2, B-3, R-1, R-2 (verbatim in the owner-decision ledger) and landed: artifact v3 carries the byte-exact input and each engine's independently recorded consumed identity, both equal to input.raw, every claim from an actual run (B-2/B-3, REPRO_VERSION 3); the verdict layer enters reduction scope with REDUCTION_SCOPE aliased to LAYER_ORDER (REDUCTION_VERSION 2, D-4); observation kind and acceptance are orthogonal fields, missing-layer replaces unexplained as a kind, every content difference is acceptance-unexplained, and a status/projection observation is declared-boundary only by an exact (layer, kind, class) entry of a frozen policy — three OD-1 typed-door entries, declared structurally by the refusing engine, never matched by text (D-5); canonical SARIF is compared as a derived zero-diff surface under one named configuration, not as a trace layer (D-6); the BR-V8 address is the pairing address with the duplicate-address ordinal part of it, pinned by an adversarial control on both sides (D-7); the dev-only own-shadow-engine adapter (stdin bytes in, one capture out, no paths, no fallback) and scripts/shadow_compare.py drive both engines over one byte sequence, and a crash, timeout or non-zero exit is a run-level hard failure with a failure report, never a refusal and never a synthetic capture (R-1/R-2). Both readers were measured on ten byte-level variants, one column per engine: CRLF, whitespace, key order and trailing newline accepted with one canonical identity and distinct raw identities; BOM, invalid UTF-8 and truncated JSON refused by both. Two CI jobs gate it: shadow compare (committed corpus) and shadow compare (C# samples), the latter fed by the wpf-extractor job's own OwnIR through one upload-artifact handshake — no second extraction — and its one sample document compared clean. What compare mode reports over the committed corpus: zero acceptance-unexplained at all three layers and on the derived SARIF, on byte-attested same input, with the two OD-1 typed-door boundaries declared by policy. Not shadow mode, not P-022 step 7a: add dual-engine shadow mode and zero-diff reproduction artifacts #260's acceptance: the five-repository sweep and the large-solution controls were still owed. Counts only in docs/generated/p022-shadow-census.md and p022-shadow-mutations.md; the record is docs/notes/p022-shadow-acceptance.md.
Shadow-mode final acceptance — the sweep (P-022 step 7a: add dual-engine shadow mode and zero-diff reproduction artifacts #260, PR feat(shadow): #260 final acceptance — the five-repository sweep, the large-solution controls, the examples and the path forms #343): the part of P-022 step 7a: add dual-engine shadow mode and zero-diff reproduction artifacts #260's test matrix the committed corpus could not supply, taken and recorded: the five pinned OSS repositories of docs: full 5-repo oracle remeasure after the complete #218-#240 batch #243 at their verified pins (drift is a failed target, never a newer measurement), the largest classic solution of every target that has one (a second extractor path over the same checkout, and measurably a differently ordered document rather than a subset), and the examples/ tree — each document extracted once through own-check.sh --emit-facts and compared from those bytes; the Windows path forms measured on a Windows host with an adapter built there. The compare driver gained its version-2 surfaces for it: every result and failure report names the adapter by sha256 and byte length, taken from the file that ran; a --manifest run verifies every document against its facts_sha256 before any engine starts and reports the denominators per target; a run that compared zero documents fails, and so does a declared target nothing reached. The sweep is a committed definition plus one recorded run, interpreted once by tests/shadow_sweep.py, rendered to docs/generated/p022-shadow-sweep.md and gated with the other fragments; .github/workflows/shadow-sweep.yml is the scheduled/manual gate; it executed after the issue closed, and the record on file is now a CI run of it (PR docs(p022): re-record the #260 sweep from a CI run of its workflow #344): workflow_run_url set, every leg on GitHub-hosted runners, each leg's document named by id rather than by a runner path — the workflow's very first execution had recorded the runner's temp path as every document's source and was superseded without being committed. Taking the measurement found defects in the harnesses and none in either engine (the sweep note's §5). What compare mode reports over P-022 step 7a: add dual-engine shadow mode and zero-diff reproduction artifacts #260's full test matrix: zero acceptance-unexplained at all three layers and on the derived SARIF, on byte-attested same input, with the two OD-1 typed-door boundaries declared by policy. Not "shadow mode achieved" in any production sense — nothing runs in production shadow, Python remains the public engine, no production behaviour changed. Counts only in docs/generated/p022-shadow-sweep.md, p022-shadow-census.md and p022-shadow-mutations.md; the record is docs/notes/p022-shadow-sweep.md.
Still missing:
repository license and real release-owner configuration;
Explicitly not missing, because it was struck rather than deferred: a Rust .ownreport.json. See the #256 entry below.
Non-negotiable migration rules
Python remains the oracle until a separate cutover decision.
A Rust/Python divergence is a Rust bug unless behavior changes in a separate Python-first PR.
Migration PRs do not add diagnostic rules, change severity, broaden Roslyn heuristics, or weaken fixtures.
Every layer owns a frozen Python-authored parity surface; steady-state Rust tests run with zero Python.
Production crate edges remain CI-enforced.
own-codegen remains independent of own-analysis and own-diagnostics.
own-bridge feeds facts and maps verdicts; it must not duplicate analysis algorithms.
Owen Alpha publication does not depend on Rust-default cutover.
Parity work follows P-022 § Parity-work discipline — oracle over reviewer prose, mutation over plausible tests, no fail-fast during mutation campaigns, insertion-stable generated goldens.
.ownreport.json — described as carrying "schema/version, tool and run metadata, diagnostics and ordered Evidence"; measured, it is {module, buffers[]}, a buffer storage report where diagnostics contribute only four boolean checks. A faithful port needs ast_nodes + buffers.resolve, which the same issue's guardrail forbids own-diagnostics from reaching — and the project already refused that shape («НЕ перегружать .ownreport.json», AGENTS.execution-surfaces.md; "build_report untouched" in docs/tasks/evidence-coverage.md). Struck, not deferred: no cutover step needs a Rust buffer report.
DI004/DI005 control — produced by ownir.build_sarif, the OwnIR path excluded by the same issue's "no OwnIR bridge". Partly landed with P-022 step 6b: implement Rust own-bridge with layered OwnIR parity #259: DI004/DI005 are frozen at Layer 3 from real facts with their messages and registration related (cp4/cp5), and cp5 replays build_sarif on the bridge path byte for byte (the BR-V9 family) — but that family is listed, not swept, and carries no DI004/DI005 case yet. Adding one is a small open item, named here rather than assumed covered.
"GitHub upload of Rust-generated SARIF" — the Rust workspace has no binary target and P-022 step 5b: port the SARIF projection with canonical parity #256 forbids CLI migration, so nothing can write a log. Replaced by asserting the structural rules an ingest enforces over the Rust-produced value, on top of byte-identity with the Python log CI already uploads.
1 — typed OwnIR validation: complete via PR feat(ownir): make the strict door accept the same language as the reference #325 — no known strict-door divergence. Three censuses. The first froze 77 controls and read 0/0/0 — then review found seven divergences the ledger could not express, because the same author wrote the ledger and the port and one gap in reading BR-D1 produced a matching gap in each. The second, derived from load() and obligations.py line by line, is 193 controls and opened a further 58 permissive documents and 9 category mismatches; 47 of the 58 were the obligation acceptance grammar, which load() calls and therefore owns. Closing them was architectural: a sequential raw-document validator reproducing BR-D1's per-section interleaving, with serde demoted to typed constructor and a guard asserting nothing escapes into it. The third admitted the two divergence families the second had measured and deliberately excluded, once feat(ownir): bound source coordinates and nesting depth (Python-first) #326 closed them Python-first — opening 7 more permissive documents and 8 more category mismatches. The defect under the mismatches is the one worth keeping: the ledger had been reading its category off the reference's diagnostic rather than off the mechanism, because _check_column raises one message for a bool, a string, a float, an out-of-range integer and a zero alike. Taxonomy is seven categories on two axes — Shape is "no representable primitive or container form", Location is "a representable coordinate violating its domain rule", WellFormedness covers records that are typed and vocabulary-legal and still cannot mean anything. Final at cp1: 216 controls, 35/181, 0/0/0, nothing excluded; 48 mutations across the three rounds, all caught. A fourth census landed with final acceptance (PR feat(ownir): bound source coordinates to int32 and mirror it in the bridge (#259 final acceptance) #341), generated rather than typed: docs/generated/p022-cp1-census.md.
2 — fact lowering: complete, 27/27 byte-exact.
3 — interprocedural MOS: complete for the stage-1 scalar-metadata domain, 35 goldens byte-exact. Container-valued metadata explicitly outside the domain.
4 — analysis wiring: complete at the checkpoint-4 surface via PR feat(bridge): wire the lowered module through the core analyses (#259 cp4) #336 — not "verdict parity complete". check_facts feeds real OwnIR facts to ownership, lifetime, buffer policy, and to the DI and effect finders, then maps ERROR-tier verdicts to fact handles through the verdict subject with the reference's map-or-raise refusal and the analysis-selected anchors preserved (DI004 call site, DI005 store site, OWN025 view site). The Layer 3 fixture family is built and replayed by own-bridge/tests/verdicts.rs with zero Python; the cp4 comparison was identity, anchor, kind and tiering, asserted over the replayed set — every divergence collected without fail-fast, any one of them a red build. The complement is named, not hidden — an exclusion ledger the replay executes: at cp4 the protocol-bearing documents (promoted at 4b), the u32 line-domain controls (promoted at final acceptance), and the Bridge contract OD-1/2/3: pin the tolerant door — direct check_facts() diverges from load() (unknown kind fallback, line coercion) #294 OD-1 controls unreachable through the typed Rust constructor (the only entries left). The one declared comparison boundary of the time — refusal text compared up to its message= member — was removed at cp5.2. No count for this checkpoint is typed on a status surface — the census and the recorded mutation campaign are generated from the ledger and the campaign evidence by scripts/render_checkpoint_status.py, live in docs/generated/p022-cp4-census.md and docs/generated/p022-cp4-mutations.md, and a Python gate fails while either is stale.
4b — protocol analysis (OBL001–005): complete via PR feat(p022): port the obligation-protocol analysis and wire it through the bridge (#259 cp4b) #340 — Layer 3 parity over the measured set, protocol family included. The gap in P-022 step 6b: implement Rust own-bridge with layered OwnIR parity #259's own checkpoint list is closed: ownlang/obligations.py has a port in own-analysis/src/obligation.rs (BR-B1 — the analysis owns its verdict: the {OPEN, CLOSED} set lattice with min-line provenance, the opens → closes → barriers leaf order with allow beating barrier, the never-invent asymmetry of an opaque write, exits anchored at the acquire, the loop's silent fixpoint and single emitting pass, the close-line evidence and the four-part sort key). One grammar, two consumers: own-ir/src/protocol.rs grew from validate-only to validate-and-construct, so the strict door and the analysis read the same typed Protocol/MethodEvents values and no strict-door error text or category moved; since PR feat(ownir): bound source coordinates to int32 and mirror it in the bridge (#259 final acceptance) #341 it takes a Door so that only the coordinate domain follows the door while type and representability stay grammar. A third analysis-level fact-parity family (tests/test_obligation_fact_parity.py → tests/fixtures/obligation_fact_parity.json → own-analysis/tests/obligation_parity.rs) freezes every violation whole from the reference's own check_protocols/unmatched_scopes and replays the raw documents with zero Python. The bridge maps BR-P3 in its BR-V1 place; refuse_protocols is gone; both protocol documents are promoted out of rust_replay_excluded without regenerating either golden, and synthetic Layer 3 and rendered controls close every row the corpus had never reached (it had reached exactly one shape). Measured rather than claimed: BR-V5's two-step rule is not applied on the protocol path — a leak off the end of a method carries a one-step slice, the port reproduces it, and whether the spec sentence or the code moves is a Python-first decision still owed (cp4b note §6.7; it does not block anything); the family's append position is unobservable end to end because the BR-V8 sort key's code component decides first. One golden family regenerated for a stated reason: the shadow artifact and trace of the protocol document, whose Rust verdicts layer moves from refused to produced. Counts live only in docs/generated/p022-cp4-census.md, docs/generated/p022-cp5-inventory.md and docs/generated/p022-cp4b-mutations.md.
5 — full fact-to-verdict parity: complete at the cp5 surface via PR feat(bridge): #259 cp5 — the full Finding, the refusal text and the rendered surfaces, proven #339 — Layer 3 parity over the measured set at the full Finding and the rendered surfaces. Its comparison set already existed in full — message, severity, subject, resource kind and ordered Evidence from P-022 step 5a: port diagnostic messages and ordered Evidence with Python parity #255, plus the canonical SARIF surface from P-022 step 5b: port the SARIF projection with canonical parity #256 — and the goldens already carried message, related and flow, so none was regenerated. 5.0 the surface inventory (tests/verdict_surface_inventory.py → docs/generated/p022-cp5-inventory.md): every BR-V4 wording branch with who owns the string, every BR-V5 family and degradation rule, every BR-V9 rule, coverage computed from the goldens, a zero row a branch that must get a control. 5.1own_bridge::Finding carries message/related/flow, the BR-V4 matrix and BR-V5 slice builders are ported, and the replay compares every member; the branches no facts document can reach are pinned by controls reading the reference's recorded answer from tests/fixtures/unreachable_branches.json (tests/test_unreachable_branch_probe.py, re-runnable). 5.2own_cfg::Diag carries the resolver's message BR-V3 interpolates, the message= cut is gone and refusals compare in full — which exposed the py_repr quote-switch defect the cut had hidden; fixed and pinned against CPython. 5.3ownlang/renders.py + tests/fixtures/verdict_renders/ + a Rust replay comparing the bytes (BR-V9; codeFlows reuses own_diagnostics::code_flow, relatedLocations deliberately does not, a golden pins why). 5.4 the surfaces and the campaigns (docs/generated/p022-cp5-mutations.md); the cp4 campaign re-run against the cp5 tree showed that a comparison surface gaining a member can lose controls for the members it subsumes, and those controls now drive dedup directly.
#345, the residual .own/dev CLI split out of #261, hangs off #257 for its emit slice and is not on the #262 path; the figure above is unchanged because it draws the cutover's normative dependencies only.
Preferred queue:#262 is next.#261's 261.B production OwnIR executable landed (PR #347, 206e9c7); #261 is closed completed. #260 is closed at final acceptance and off this queue, with cp5, 4b and the coordinate-domain decision. In parallel and off the critical chain: #257 (without it #345's emit slice cannot close), #263 (without it #262's performance gates have no baseline), #345 now that #261's own-cli skeleton exists, and #269's reconciliation as independent cleanup.
The defensive limits that used to head this queue landed in #326, and their position was load-bearing rather than tidy: they changed what the reference accepts, so they had to land Python-first and cp1 had to be re-measured against them rather than merged beside them. The coordinate-domain decision was the same kind of item and landed the same way (PR #341); #260's two decisions were decided the same way and landed in PR #342.
Production implementation remains blocked until the runtime marker/helper/escape-hatch contract is finalized. Do not duplicate call-site use-after-dispose rules.
Recommended agent allocation
Owner decision (not delegable)
#269 reconciliation: what is delivered (trace, first-divergence reduction) and what is deferred (the bounded minimizer, a named explain-divergence command)
BR-V5 on the protocol path (spec sentence vs `_protocol_findings`; Python-first, blocks nothing)
#260's decisions are resolved (D-4..D-7, B-2, B-3, R-1, R-2), #261's packet is ratified (C-1..C-5, 2026-09-08) and its 261.B implementation landed (PR #347, 206e9c7); #269's reconciliation heads this list now.
Local/corpus-capable agent
#263 measurements
Strong agent
#345: the residual subcommands on the same `own-cli` binary that #261's 261.B built (PR #347); `emit` only after #257
#255, #256, #258, #259 and now #261's 261.B (PR #347) are complete and drop out of this chain. #259 is closed at final acceptance — do not re-open it or any of its checkpoints; the declared OD-1 boundary is on the exclusion ledger, not open work.
Separate strong or medium-strong agent
#257
Keep codegen isolated from analysis.
Medium agent
#252 verification packet
#253/#254 release evidence and external checks
Status-drift rule
A step is never described by a single Implemented/Missing bit. Any status edit to this issue or to docs/proposals/P-022-rust-core-migration.md must state, per step: completed checkpoints, remaining acceptance, normative blocker, and preferred sequencing. Both surfaces are updated in the same change — a reconciliation that touches only one of them replaces a stale pair with a contradictory pair. The proposals index row counts as a third surface for the same fact.
A child issue's own body is a fourth surface when its acceptance turns out to be wrong. #256 is the worked example: its requirements described a .ownreport.json the project does not have and had already refused to build, so the correction belongs in the issue, in this roadmap, in P-022 and in the index — together, or not at all. #259 is the second: its checkpoint list had no row for the obligation-protocol analysis while its final acceptance required it, so 4b was added to the child issue, to P-022 and here in the same move.
A child issue's open/closed bit is a fifth surface, and #259 is its worked example too: GitHub closed it when PR #339 merged, because its parser read the body's "does not close #259" as close #259 — the negation is not parsed — while the sentence meant the opposite and this queue still read "→ #259 final acceptance". A closed issue with an open acceptance is a contradictory pair with the roadmap. It was reopened, and closed by hand at the #341 merge once the acceptance was reached. Checkpoint PRs reference their issue as Refs #N and never put a closing keyword before an issue number, negated or not; nothing but the final-acceptance change closes it, and the owner does that by hand together with the body update.
Parity checkpoints carry one more surface, and it is not a status surface: the frozen ledger. A checkpoint whose acceptance is "two implementations agree" is proved by an artifact that can share the implementation's blind spots, and a green matrix over an incomplete ledger is indistinguishable from a green matrix over a complete one. #259 cp1 is the worked example, three times over — 0/0/0 over 77 controls, then 58 permissive documents and 9 category mismatches once the ledger was rebuilt from the reference instead of from the author's reading of it, then 7 more permissive documents and 8 more category mismatches once the two deliberately excluded families were admitted.
The third round added a second failure mode worth naming separately: a ledger can carry the right controls and still take its category from the wrong place. _check_column raises one message for five distinct mechanisms, and the ledger inherited one message as one category — so the classification was correct about accept/reject and wrong about why, for a year, in a file whose entire purpose is to be right about why.
cp5 added a third: a comparison surface that gains a member can lose controls for the members it subsumes. Putting message into the BR-V7 dedup key made several key members unobservable at the output, and the cp4 campaign — re-run, not trusted — turned those mutations from caught to survived. The fix was controls that drive the production dedup directly, and the rule is that every earlier campaign is re-run against the new tree, which is what the generated evidence pipeline exists to make cheap.
Final acceptance added a fourth, about the evidence pipeline itself: a campaign's expected catchers can rot between runs without the gate noticing.--validate re-anchored every mutation against the current tree, but it did not ask whether the tests a mutation names still exist or still fail; shadow-cp4's M61 carried two that had stopped failing on main after 4b's promotion. Closed in PR #342: --validate now resolves every named catcher, and the re-run rule is written down.
#260's acceptance work added a fifth, about the comparison itself: a green gate over an empty set is worse than a red one, because a red gate at least says it is awake. The compare driver therefore says out loud how many documents it compared and how many agreed, and the C# samples gate was read from its log rather than from its tick.
#260's sweep added a sixth, about provenance rather than comparison: the environment that records evidence can shape it. A checkout's line-ending setting changes the byte digest of a definition without changing a character; a machine's git identity rides into every commit made from it; a console codepage rewrites a commit message on the way in. None of it is visible to a reader of text on a screen, and all of it is caught only by comparing bytes and commit metadata. The rule: recorded artifacts are produced and compared as bytes; work committed from a local machine enters the tree under a repository identity; a pre-push report scans metadata as well as content; and the fix for the line-ending half is a researched .gitattributes, not an operator's global configuration. PR #344 added the CI half of the same lesson: a workflow's manifest generator wrote the runner's temp path into every document's source — a constant dressed as provenance — and it was caught by the same check on the record before anything was committed; the generator now names a document by its id.
The distinction matters: a ledger cannot report a normative blocker, a preferred sequencing or a remaining acceptance. It carries no project state at all. It is evidence for a claim the status surfaces make, so it is reviewed as evidence — is it derived from the reference or from the port, can it express absence, does every category have a control, and is any family excluded — and when it turns out to be incomplete, the correction is a further census with every result on the record, not a fix to the port.
For how to keep parity work honest — not just its status — see P-022 § Parity-work discipline.
Global PR acceptance packet
Every substantive child PR must state:
Scope:
Explicit non-goals:
Python source of truth:
Frozen fixture:
Fixture regeneration command:
Steady-state test command:
Production dependency changes:
Behavior changes:
Acceptance changes:
Local commands:
GitHub Actions links:
Known deferred cases:
Release smoke tests use project references/rebuilds instead of the packed .nupkg.
Python and Rust receive different OwnIR bytes in compare mode.
SARIF normalization deletes semantic fields merely to remove a diff.
Performance work begins without a baseline/profile.
A new abstraction layer has no vertical consumer.
Schema, semantics, output and packaging are mixed into one supposedly small PR.
A parity result is reported as complete while known divergence families sit outside the measured set. Excluding them is legitimate; calling the remainder "parity" is not.
A parity category is taken from where the reference raises its error rather than from the mechanism the document violated. One diagnostic covering several mechanisms is normal in a reference written for humans; inheriting it as one category makes the taxonomy decorative.
Note on (1) versus a corrected acceptance: striking a requirement because the tree proves it describes something that does not exist is not weakening it. The distinction is evidence — #256's strikes each carry a measurement, and the surfaces that stated the old acceptance were all corrected in the same move. Narrowing what the reference accepts as a Python-first contract decision with the ledger re-measured against it (#326, then #341) is not weakening either: the reference moved first, and the census followed.
Milestone completion
This roadmap reaches its next major milestone when all are true:
Owen Alpha
installable from nuget.org on a clean machine;
owen check works on Windows and Linux;
immutable GitHub Action tag works from an external repository;
Status
Master roadmap after the completion of P-022 step 4 in #214 / PR #249.
Two independent outcomes must now be delivered:
The release track must not wait for the full Rust migration. The Rust track must not use the release as permission to weaken parity.
Current baseline
Completed:
Owen.Cli,owen, Owen Action/SARIF identity (feat(owen): public facade rebrand for the CLI, Action, and SARIF identity #246).own-ir,own-syntax,own-cfg,own-diagnostics,own-analysis,own-lowered,own-bridge, andown-shadow(dev/test, step 7a; since PR feat(p022): #260 acceptance over the committed corpus — byte-attested same input, the verdict layer in scope, the derived SARIF gate, the compare driver #342 also the dev-onlyown-shadow-engineadapter binary).render/render_prettytext, the emission ordering contract (a stable sort on(line, code)), and a self-policing ledger over all 47TITLEScodes with missing/orphan/stale guards.own_diagnostics::sarifportsdiag_sarif.py+ the twoevidence.pybuilders as typed structs; 16 cases / 21 results replayed with zero Python, canonical comparison stripping zero volatile fields (empty list proven by census + determinism check), six mutations each caught under--no-fail-fast.spec/Bridge.md+spec/BridgeBehaviorMatrix.mdnormative and onmainvia PR spec(bridge): #258 — executable own-bridge contract (Bridge.md + behavior matrix) #297.rust_replaycases byte-exact.*.summaries.jsongoldens byte-identical, over the declared scalar-metadata parity domain.main): signed-64 source coordinates and a 32-level nesting limit, normative inspec/OwnIR.md§4.2 and carried byspec/ownir.schema.json. The reference used to accept coordinates and depths no other consumer could represent — not generosity, but an accident of CPython's unbounded integers and deep stack leaking into the contract. Closed in the contract, not by widening Rust.own_bridge::check_factsdrives the ported analyses from real OwnIR facts and maps ERROR-tier verdicts back to fact handles through the verdictsubject, preserving the analysis-selected anchors. Complete at the checkpoint-4 surface — identity, anchor, kind and tiering over the frozen Layer 3 ledger, with the replay asserting equality on those members rather than reporting a number — which is deliberately not "verdict parity complete": three divergence families sat outside the measured set at the time and are carried by an executable exclusion ledger (protocol documents — since promoted at 4b; theu32line-domain boundary — since promoted at final acceptance; the Bridge contract OD-1/2/3: pin the tolerant door — direct check_facts() diverges from load() (unknown kind fallback, line coercion) #294 OD-1 door controls). One core change was required:own-analysisnow stampssubjectwhereanalysis.py/lifetimes.pydo.tests/verdict_census.py,scripts/mutate_campaign.py) and projected intodocs/generated/*.mdbyscripts/render_checkpoint_status.py;tests/test_checkpoint_status.pyruns the--checkin the suite. The campaign runner keeps the M00 honesty control, refuses fail-fast, labels every catcher by package and target, refuses to run on a dirty tree, and a caught mutation whose named catchers did not all fail voids the evidence at the gate. Every later checkpoint added its fragments to the same pipeline; since PR feat(p022): #260 acceptance over the committed corpus — byte-attested same input, the verdict layer in scope, the derived SARIF gate, the compare driver #342--validatealso resolves every named catcher and refuses a Python mutation that does not parse.tests/verdict_surface_inventory.py) names every BR-V4 wording branch with who owns the string, every BR-V5 evidence family and degradation rule, and every BR-V9 rule, with coverage computed from the goldens;own_bridge::Findingcarriesmessage,relatedandflow, and the replay compares every member;own_cfg::Diagcarries the resolver's message so refusals compare in full (removing that cut exposed apy_reprquoting defect the cut had hidden — fixed and pinned against CPython);ownlang/renders.py+tests/fixtures/verdict_renders/freeze the rendered surfaces and a Rust replay compares the bytes. Branches no facts document can reach get their expected text from the reference's own recorded output (tests/test_unreachable_branch_probe.py), re-runnable. No pre-existing golden was regenerated or edited.ownlang/obligations.pyis ported intoown-analysis(BR-B1); its typed values come from the one grammar inown-irthe strict door already delegated to, so door and analysis cannot drift into two readings; a third analysis-level fact-parity family freezes every violation member with zero Python; the bridge maps BR-P3 in its BR-V1 place;refuse_protocolsis gone and both protocol documents are promoted out of the exclusion ledger without regenerating either golden. Two things measured rather than claimed: BR-V5's two-step rule is not applied on the protocol path (a Python-first decision still owed), and the family's append position is unobservable end to end.ownlang/repro.pyis the reference emitter andown-shadowthe port's half — canonical same-inputOwnIRidentity and the reproduction-artifact format (cp1), the engine protocol where each engine declares what it could produce (cp2), theAnalysisTracewith stable-ID normalization and per-layer ordering semantics declared rather than normalized away (cp3, P-022 diagnostics: normalized AnalysisTrace and first-divergence minimizer for Python↔Rust shadow mode #269), and first-divergence reduction over the lowered/MOS layers naming layer, step and minimal difference (cp4). Four findings closed as contract decisions and three departures from the slice's brief ratified in the owner-decision ledger.spec/OwnIR.md§4.2 bounds everylineto[0, 2147483647]and everycolumnto[1, 2147483647]— int32 because int32 is the line type of every consumer this project feeds,0kept legal as the reference's own absent sentinel, a negative line rejected because no producer emits one; the strict door refuses an out-of-domain coordinate as Location on every line-bearing field, including the two §4.2 had recorded as validated nowhere; the tolerant door degrades to0/ absent, never clamps;own-irandown-bridgemirror all of it; the fourverdict_boundary_*controls are promoted out of the exclusion ledger with goldens regenerated by Python, and the ledger names only the two Bridge contract OD-1/2/3: pin the tolerant door — direct check_facts() diverges from load() (unknown kind fallback, line coercion) #294 OD-1 door controls. The cp1 ledger took its fourth census, and every earlier campaign was re-run against the new tree. Counts only indocs/generated/p022-cp1-census.md,p022-coord-census.mdandp022-coord-mutations.md; the record isdocs/notes/p022-bridge-verdict-final-acceptance.md.consumedidentity, both equal toinput.raw, every claim from an actual run (B-2/B-3,REPRO_VERSION3); the verdict layer enters reduction scope withREDUCTION_SCOPEaliased toLAYER_ORDER(REDUCTION_VERSION2, D-4); observation kind and acceptance are orthogonal fields,missing-layerreplacesunexplainedas a kind, every content difference is acceptance-unexplained, and astatus/projectionobservation isdeclared-boundaryonly by an exact(layer, kind, class)entry of a frozen policy — three OD-1 typed-door entries, declared structurally by the refusing engine, never matched by text (D-5); canonical SARIF is compared as a derived zero-diff surface under one named configuration, not as a trace layer (D-6); the BR-V8 address is the pairing address with the duplicate-address ordinal part of it, pinned by an adversarial control on both sides (D-7); the dev-onlyown-shadow-engineadapter (stdin bytes in, one capture out, no paths, no fallback) andscripts/shadow_compare.pydrive both engines over one byte sequence, and a crash, timeout or non-zero exit is a run-level hard failure with a failure report, never a refusal and never a synthetic capture (R-1/R-2). Both readers were measured on ten byte-level variants, one column per engine: CRLF, whitespace, key order and trailing newline accepted with one canonical identity and distinct raw identities; BOM, invalid UTF-8 and truncated JSON refused by both. Two CI jobs gate it:shadow compare (committed corpus)andshadow compare (C# samples), the latter fed by thewpf-extractorjob's own OwnIR through one upload-artifact handshake — no second extraction — and its one sample document compared clean. What compare mode reports over the committed corpus: zero acceptance-unexplained at all three layers and on the derived SARIF, on byte-attested same input, with the two OD-1 typed-door boundaries declared by policy. Not shadow mode, not P-022 step 7a: add dual-engine shadow mode and zero-diff reproduction artifacts #260's acceptance: the five-repository sweep and the large-solution controls were still owed. Counts only indocs/generated/p022-shadow-census.mdandp022-shadow-mutations.md; the record isdocs/notes/p022-shadow-acceptance.md.examples/tree — each document extracted once throughown-check.sh --emit-factsand compared from those bytes; the Windows path forms measured on a Windows host with an adapter built there. The compare driver gained its version-2 surfaces for it: every result and failure report names the adapter bysha256and byte length, taken from the file that ran; a--manifestrun verifies every document against itsfacts_sha256before any engine starts and reports the denominators per target; a run that compared zero documents fails, and so does a declared target nothing reached. The sweep is a committed definition plus one recorded run, interpreted once bytests/shadow_sweep.py, rendered todocs/generated/p022-shadow-sweep.mdand gated with the other fragments;.github/workflows/shadow-sweep.ymlis the scheduled/manual gate; it executed after the issue closed, and the record on file is now a CI run of it (PR docs(p022): re-record the #260 sweep from a CI run of its workflow #344):workflow_run_urlset, every leg on GitHub-hosted runners, each leg's document named by id rather than by a runner path — the workflow's very first execution had recorded the runner's temp path as every document'ssourceand was superseded without being committed. Taking the measurement found defects in the harnesses and none in either engine (the sweep note's §5). What compare mode reports over P-022 step 7a: add dual-engine shadow mode and zero-diff reproduction artifacts #260's full test matrix: zero acceptance-unexplained at all three layers and on the derived SARIF, on byte-attested same input, with the two OD-1 typed-door boundaries declared by policy. Not "shadow mode achieved" in any production sense — nothing runs in production shadow, Python remains the public engine, no production behaviour changed. Counts only indocs/generated/p022-shadow-sweep.md,p022-shadow-census.mdandp022-shadow-mutations.md; the record isdocs/notes/p022-shadow-sweep.md.Still missing:
Owen.Cliand Action releases;own-codegen(P-022 step 5c: port own-codegen as an analysis-independent Rust sibling #257);dirtybit and adefinition_sha256check in the sweep evidence, a researched.gitattributesfor byte-sensitive evidence, the definition-level counts in the P-022 row's prose — small, listed under P-022 step 7a: add dual-engine shadow mode and zero-diff reproduction artifacts #260's entry below;own-cli ownir) — command, output and exit-code parity behind the existing launcher #261's 261.B measurement):UnicodeDecodeErroron non-UTF-8 facts input →OwnIRError→ rc 2 — today the reference'sload()catchesOSError/JSONDecodeErroronly, so invalid UTF-8 escapes to the exit-70 catch-all; recorded in P-022 step 8: Rust-default cutover, rollback gate, and Python distribution removal #262's Known differences, not a feat(own-cli): the production Rust OwnIR executable, behind the unchanged launcher (#261 261.B) #347 blocker;own-cli ownir) — command, output and exit-code parity behind the existing launcher #261's 261.B measurement): Rust emits canonical UTF-8 portably, and native-Windows Python byte parity is not claimed — recorded in P-022 step 8: Rust-default cutover, rollback gate, and Python distribution removal #262 as a declared behaviour change rather than a parity gap, with P-022 step 8: Rust-default cutover, rollback gate, and Python distribution removal #262 the tracker of record;.own/dev CLI (P-022 step 7b residual: Rust CLI for the.ownand dev surfaces —cfg,summaries,explain,.own check;emitbehind #257 #345) and the Rust-default cutover (P-022 step 8: Rust-default cutover, rollback gate, and Python distribution removal #262) — the production OwnIR executable itself (261.B) landed in PR feat(own-cli): the production Rust OwnIR executable, behind the unchanged launcher (#261 261.B) #347;Explicitly not missing, because it was struck rather than deferred: a Rust
.ownreport.json. See the #256 entry below.Non-negotiable migration rules
own-codegenremains independent ofown-analysisandown-diagnostics.own-bridgefeeds facts and maps verdicts; it must not duplicate analysis algorithms.Child issues and execution order
A. Owen Alpha release
Owen.Clialpha. Blocked by Owen Alpha gate: choose license, configure protected release environments, and approve version #252.Release DAG:
The release track is independent of the Rust-default cutover.
B. P-022 Rust production vertical
Each entry states its normative blocker first; preferred ordering is marked as such and is advisory.
.ownflow-diagnostic SARIF projection, byte-identical under a canonical comparison that normalizes only object key order and leaves every ordered array in place. Three of the issue's original requirements were struck as unbuildable as written, each verified by running the reference:.ownreport.json— described as carrying "schema/version, tool and run metadata, diagnostics and ordered Evidence"; measured, it is{module, buffers[]}, a buffer storage report where diagnostics contribute only four booleanchecks. A faithful port needsast_nodes+buffers.resolve, which the same issue's guardrail forbidsown-diagnosticsfrom reaching — and the project already refused that shape («НЕ перегружать.ownreport.json»,AGENTS.execution-surfaces.md; "build_reportuntouched" indocs/tasks/evidence-coverage.md). Struck, not deferred: no cutover step needs a Rust buffer report.DI004/DI005control — produced byownir.build_sarif, the OwnIR path excluded by the same issue's "no OwnIR bridge". Partly landed with P-022 step 6b: implement Rust own-bridge with layered OwnIR parity #259: DI004/DI005 are frozen at Layer 3 from real facts with their messages and registrationrelated(cp4/cp5), and cp5 replaysbuild_sarifon the bridge path byte for byte (the BR-V9 family) — but that family is listed, not swept, and carries no DI004/DI005 case yet. Adding one is a small open item, named here rather than assumed covered.own-codegen. Ready, independent of the analysis path; parallelizable.spec/Bridge.md+spec/BridgeBehaviorMatrix.mdmerged via PR spec(bridge): #258 — executable own-bridge contract (Bridge.md + behavior matrix) #297, onmain.own-bridge. Final acceptance reached via PR feat(ownir): bound source coordinates to int32 and mirror it in the bridge (#259 final acceptance) #341; closed by hand at that merge. Normative blocker P-022 step 6a: formalize OwnIR bridge semantics before the Rust port #258 satisfied; landed ahead of the preferred "after P-022 step 5a: port diagnostic messages and ordered Evidence with Python parity #255/P-022 step 5b: port the SARIF projection with canonical parity #256" ordering, which was advisory. (For the record: the issue was closed by automation at the feat(bridge): #259 cp5 — the full Finding, the refusal text and the rendered surfaces, proven #339 merge — GitHub's keyword parser read feat(bridge): #259 cp5 — the full Finding, the refusal text and the rendered surfaces, proven #339's own "does not close P-022 step 6b: implement Rust own-bridge with layered OwnIR parity #259" asclose #259, negation unparsed — reopened while its acceptance was open, and closed by hand once it was reached.) Checkpoints:load()andobligations.pyline by line, is 193 controls and opened a further 58 permissive documents and 9 category mismatches; 47 of the 58 were the obligation acceptance grammar, whichload()calls and therefore owns. Closing them was architectural: a sequential raw-document validator reproducing BR-D1's per-section interleaving, withserdedemoted to typed constructor and a guard asserting nothing escapes into it. The third admitted the two divergence families the second had measured and deliberately excluded, once feat(ownir): bound source coordinates and nesting depth (Python-first) #326 closed them Python-first — opening 7 more permissive documents and 8 more category mismatches. The defect under the mismatches is the one worth keeping: the ledger had been reading its category off the reference's diagnostic rather than off the mechanism, because_check_columnraises one message for a bool, a string, a float, an out-of-range integer and a zero alike. Taxonomy is seven categories on two axes —Shapeis "no representable primitive or container form",Locationis "a representable coordinate violating its domain rule",WellFormednesscovers records that are typed and vocabulary-legal and still cannot mean anything. Final at cp1: 216 controls, 35/181, 0/0/0, nothing excluded; 48 mutations across the three rounds, all caught. A fourth census landed with final acceptance (PR feat(ownir): bound source coordinates to int32 and mirror it in the bridge (#259 final acceptance) #341), generated rather than typed:docs/generated/p022-cp1-census.md.check_factsfeeds real OwnIR facts to ownership, lifetime, buffer policy, and to the DI and effect finders, then maps ERROR-tier verdicts to fact handles through the verdictsubjectwith the reference's map-or-raise refusal and the analysis-selected anchors preserved (DI004 call site, DI005 store site, OWN025 view site). The Layer 3 fixture family is built and replayed byown-bridge/tests/verdicts.rswith zero Python; the cp4 comparison was identity, anchor, kind and tiering, asserted over the replayed set — every divergence collected without fail-fast, any one of them a red build. The complement is named, not hidden — an exclusion ledger the replay executes: at cp4 the protocol-bearing documents (promoted at 4b), theu32line-domain controls (promoted at final acceptance), and the Bridge contract OD-1/2/3: pin the tolerant door — direct check_facts() diverges from load() (unknown kind fallback, line coercion) #294 OD-1 controls unreachable through the typed Rust constructor (the only entries left). The one declared comparison boundary of the time — refusal text compared up to itsmessage=member — was removed at cp5.2. No count for this checkpoint is typed on a status surface — the census and the recorded mutation campaign are generated from the ledger and the campaign evidence byscripts/render_checkpoint_status.py, live indocs/generated/p022-cp4-census.mdanddocs/generated/p022-cp4-mutations.md, and a Python gate fails while either is stale.ownlang/obligations.pyhas a port inown-analysis/src/obligation.rs(BR-B1 — the analysis owns its verdict: the{OPEN, CLOSED}set lattice with min-line provenance, the opens → closes → barriers leaf order with allow beating barrier, the never-invent asymmetry of an opaque write, exits anchored at the acquire, the loop's silent fixpoint and single emitting pass, the close-line evidence and the four-part sort key). One grammar, two consumers:own-ir/src/protocol.rsgrew from validate-only to validate-and-construct, so the strict door and the analysis read the same typedProtocol/MethodEventsvalues and no strict-door error text or category moved; since PR feat(ownir): bound source coordinates to int32 and mirror it in the bridge (#259 final acceptance) #341 it takes aDoorso that only the coordinate domain follows the door while type and representability stay grammar. A third analysis-level fact-parity family (tests/test_obligation_fact_parity.py→tests/fixtures/obligation_fact_parity.json→own-analysis/tests/obligation_parity.rs) freezes every violation whole from the reference's owncheck_protocols/unmatched_scopesand replays the raw documents with zero Python. The bridge maps BR-P3 in its BR-V1 place;refuse_protocolsis gone; both protocol documents are promoted out ofrust_replay_excludedwithout regenerating either golden, and synthetic Layer 3 and rendered controls close every row the corpus had never reached (it had reached exactly one shape). Measured rather than claimed: BR-V5's two-step rule is not applied on the protocol path — a leak off the end of a method carries a one-step slice, the port reproduces it, and whether the spec sentence or the code moves is a Python-first decision still owed (cp4b note §6.7; it does not block anything); the family's append position is unobservable end to end because the BR-V8 sort key's code component decides first. One golden family regenerated for a stated reason: the shadow artifact and trace of the protocol document, whose Rustverdictslayer moves fromrefusedtoproduced. Counts live only indocs/generated/p022-cp4-census.md,docs/generated/p022-cp5-inventory.mdanddocs/generated/p022-cp4b-mutations.md.Findingand the rendered surfaces. Its comparison set already existed in full — message, severity, subject, resource kind and ordered Evidence from P-022 step 5a: port diagnostic messages and ordered Evidence with Python parity #255, plus the canonical SARIF surface from P-022 step 5b: port the SARIF projection with canonical parity #256 — and the goldens already carriedmessage,relatedandflow, so none was regenerated. 5.0 the surface inventory (tests/verdict_surface_inventory.py→docs/generated/p022-cp5-inventory.md): every BR-V4 wording branch with who owns the string, every BR-V5 family and degradation rule, every BR-V9 rule, coverage computed from the goldens, a zero row a branch that must get a control. 5.1own_bridge::Findingcarriesmessage/related/flow, the BR-V4 matrix and BR-V5 slice builders are ported, and the replay compares every member; the branches no facts document can reach are pinned by controls reading the reference's recorded answer fromtests/fixtures/unreachable_branches.json(tests/test_unreachable_branch_probe.py, re-runnable). 5.2own_cfg::Diagcarries the resolver's message BR-V3 interpolates, themessage=cut is gone and refusals compare in full — which exposed thepy_reprquote-switch defect the cut had hidden; fixed and pinned against CPython. 5.3ownlang/renders.py+tests/fixtures/verdict_renders/+ a Rust replay comparing the bytes (BR-V9;codeFlowsreusesown_diagnostics::code_flow,relatedLocationsdeliberately does not, a golden pins why). 5.4 the surfaces and the campaigns (docs/generated/p022-cp5-mutations.md); the cp4 campaign re-run against the cp5 tree showed that a comparison surface gaining a member can lose controls for the members it subsumes, and those controls now drivededupdirectly.linein[0, 2147483647]with0legal as the reference's own absent sentinel and a negative line rejected, everycolumnin[1, 2147483647]or absent — int32 because int32 is the line type of every consumer this project feeds, never "Rust holdsu32, so the reference is wrong"; the strict door refuses out-of-domain coordinates asLocationon every line-bearing field, the two §4.2 had recorded as validated nowhere included; the tolerant door degrades to0/ absent and never clamps;own-irandown-bridgemirror it; the four boundary controls are promoted and the exclusion ledger names only the two OD-1 door controls (Bridge contract OD-1/2/3: pin the tolerant door — direct check_facts() diverges from load() (unknown kind fallback, line coercion) #294), a declared and measured boundary of the typed constructor, not open work. The BR-V5 protocol-path sentence (cp4b note §6.7) stays a Python-first item on its own, blocking nothing. The three review tails were carried to P-022 step 7a: add dual-engine shadow mode and zero-diff reproduction artifacts #260 and closed there by PR feat(p022): #260 acceptance over the committed corpus — byte-attested same input, the verdict layer in scope, the derived SARIF gate, the compare driver #342; the fourth (the final-acceptance note'sStatus:header) closed in PR feat(shadow): #260 final acceptance — the five-repository sweep, the large-solution controls, the examples and the path forms #343.examples/tree, the five pinned OSS repositories of docs: full 5-repo oracle remeasure after the complete #218-#240 batch #243 at their verified pins and the large-solution controls: zero acceptance-unexplained at all three layers and on the derived SARIF, on byte-attested same input, with the two Bridge contract OD-1/2/3: pin the tolerant door — direct check_facts() diverges from load() (unknown kind fallback, line coercion) #294 OD-1 typed-door boundaries declared by policy. Every document of the sweep definition was extracted exactly once throughown-check.sh --emit-factsand compared from those bytes; the.slnfan-out is a second extractor path over the same checkout and measurably a differently ordered document rather than a subset. Coverage is defined so it cannot be faked: a run that compared zero documents fails, a declared target nothing reached fails, and the denominators are recorded per target. The driver names the adapter bysha256and byte length in every result and failure report (shadow_compare_version2; artifact v3 untouched) and verifies every manifest document against itsfacts_sha256before any engine starts. Taking the measurement found harness defects and no engine divergence. Still not shadow mode, not "P-022 done" and not "Rust is the default" — that is P-022 step 8: Rust-default cutover, rollback gate, and Python distribution removal #262's cutover behind P-022 step 7b: the production Rust OwnIR executable (own-cli ownir) — command, output and exit-code parity behind the existing launcher #261; Python stays the public engine and no production behaviour changed. Records:docs/notes/p022-shadow-sweep.mdanddocs/notes/p022-shadow-acceptance.md; counts only indocs/generated/p022-shadow-sweep.md,p022-shadow-census.mdandp022-shadow-mutations.md. Tails carried out at the close, none blocking: the sweep evidence carries nodirtybit and its interpreter does not verifydefinition_sha256(both should mirror the campaign harness); byte-sensitive evidence depends on the operator'score.autocrlf(a researched.gitattributesfordocs/evidence,docs/generatedandtests/fixtures, never a global config change); the sweep was re-recorded from a CI run of.github/workflows/shadow-sweep.ymlin PR docs(p022): re-record the #260 sweep from a CI run of its workflow #344, closing that tail (a re-run replaces the result whole); the P-022 row 7a and the proposals index carry definition-level counts in prose;.slnxis an extractor gap outside P-022; and P-022 diagnostics: normalized AnalysisTrace and first-divergence minimizer for Python↔Rust shadow mode #269's body is to be reconciled against the tree — the trace and the first-divergence reduction are delivered, the bounded minimizer and a namedexplain-divergencecommand are not, and are deferred until a first real divergence exists.own-cli ownir) — command, output and exit-code parity behind the existing launcher #261 — the production Rust OwnIR executable (own-cli ownir): command/output/exit-code parity behind the existing launcher. Normative blockers satisfied (P-022 step 6b: implement Rust own-bridge with layered OwnIR parity #259 and P-022 step 7a: add dual-engine shadow mode and zero-diff reproduction artifacts #260 reached). 261.A decision packet ratified 2026-09-08 (C-1..C-5 with the exit-code rulings, recorded verbatim in the issue); 261.B implementation landed 2026-09-08 via PR feat(own-cli): the production Rust OwnIR executable, behind the unchanged launcher (#261 261.B) #347, merged as206e9c7— theown-clicrate with its singleownirsubcommand, a Python-authored CLI fixture replayed with zero Python on Linux and Windows CI, both failure-mode rulings measured under an off-by-defaultfault-injectionfeature, and theownir_versionVersion family fixed to byte parity (V1/V2/V4 declared and V3 reproduced per P-022 step 8: Rust-default cutover, rollback gate, and Python distribution removal #262's parser-domain rulings). P-022 step 7b: the production Rust OwnIR executable (own-cli ownir) — command, output and exit-code parity behind the existing launcher #261 closed completed 2026-09-08 (PR feat(own-cli): the production Rust OwnIR executable, behind the unchanged launcher (#261 261.B) #347,206e9c7). What was decided: the oracle is split by surface and behaviour class —python -m ownlangfor the core semantics ofownirthrough Python-authored fixtures replayed with zero Python, the publicowenconvention for the top-level shell only — help, version, the empty invocation, an unknown command — whose text becomes a cross-implementation parity surface rather than an invented Python contract, and afterownireverything as the reference measures, its usage errors included (C-1, boundary fixed 2026-09-08);reportstruck, not deferred, P-022 step 5b: port the SARIF projection with canonical parity #256 unreopened (C-2); the production seam — the one core invocationown-check.*, the Action andowen checkmake — separated from the residual.own/dev CLI, which is P-022 step 7b residual: Rust CLI for the.ownand dev surfaces —cfg,summaries,explain,.own check;emitbehind #257 #345 (C-3); engine selection outside the executable — one engine, no silent fallback, orchestration in the dev driver now and in the launcher at P-022 step 8: Rust-default cutover, rollback gate, and Python distribution removal #262's stages (C-4); the production slice alone as P-022 step 8: Rust-default cutover, rollback gate, and Python distribution removal #262's CLI prerequisite (C-5). Rulings: stdin is not part of the contract; a catchable panic is exactly one actionable stderr diagnostic and rc 70 underpanic = "unwind"with a top-levelcatch_unwind; an uncatchable death is a visible hard failure with no OS exit number contracted; SIGINT is measured on both reference platforms before it is contracted. The executable lives behind the unchangedowen; nothing is wired, published or defaulted — that is P-022 step 8: Rust-default cutover, rollback gate, and Python distribution removal #262. 261.B is completed, merged code (PR feat(own-cli): the production Rust OwnIR executable, behind the unchanged launcher (#261 261.B) #347,206e9c7); the residual.own/dev CLI is P-022 step 7b residual: Rust CLI for the.ownand dev surfaces —cfg,summaries,explain,.own check;emitbehind #257 #345 and the cutover is P-022 step 8: Rust-default cutover, rollback gate, and Python distribution removal #262, both still open..ownand dev surfaces —cfg,summaries,explain,.own check;emitbehind #257 #345 — the residual Rust CLI:cfg,summaries,explain,.own check, andemitbehind P-022 step 5c: port own-codegen as an analysis-independent Rust sibling #257. Split out of P-022 step 7b: the production Rust OwnIR executable (own-cli ownir) — command, output and exit-code parity behind the existing launcher #261 by C-3; a migration obligation under C-5 for as long as P-022 is the full surface migration, and not on the P-022 step 8: Rust-default cutover, rollback gate, and Python distribution removal #262 path. Sameown-clibinary as P-022 step 7b: the production Rust OwnIR executable (own-cli ownir) — command, output and exit-code parity behind the existing launcher #261's, subcommands added after its skeleton exists; C-1's oracle split and the exit-code rulings apply unchanged (summaries: expected OwnIR refusal 2, unexpected exception 70;emit: 1 on ordinary diagnostics or a refusal before codegen, 70 when aCodegenErrorescapesgenerate()or on any other unexpected exception).own-cli ownir) — command, output and exit-code parity behind the existing launcher #261's production OwnIR executable alone (C-5; P-022 step 7a: add dual-engine shadow mode and zero-diff reproduction artifacts #260 reached; P-022 step 7b residual: Rust CLI for the.ownand dev surfaces —cfg,summaries,explain,.own check;emitbehind #257 #345 is not on this path) — that prerequisite is now satisfied: P-022 step 7b: the production Rust OwnIR executable (own-cli ownir) — command, output and exit-code parity behind the existing launcher #261 / 261.B landed via PR feat(own-cli): the production Rust OwnIR executable, behind the unchanged launcher (#261 261.B) #347 (206e9c7) and P-022 step 7b: the production Rust OwnIR executable (own-cli ownir) — command, output and exit-code parity behind the existing launcher #261 is closed completed. Its performance gates need baselines, which is IDE foundation: establish cold/warm/incremental latency and memory baselines #263's output — an evidence prerequisite for the cutover decision, not a normative blocker. The launcher rulings ratified with P-022 step 7b: the production Rust OwnIR executable (own-cli ownir) — command, output and exit-code parity behind the existing launcher #261's packet are recorded in the issue: engine selection is the launcher's, explicit, never a silent fallback; an unexpected Rust child exit code outside the legal set takes the public internal-error path with the raw child status retained in the evidence; SIGINT is measured before it is contracted.Rust DAG (unchanged — normative dependencies only):
#345, the residual
.own/dev CLI split out of #261, hangs off #257 for itsemitslice and is not on the #262 path; the figure above is unchanged because it draws the cutover's normative dependencies only.Preferred queue: #262 is next. #261's 261.B production OwnIR executable landed (PR #347,
206e9c7); #261 is closed completed. #260 is closed at final acceptance and off this queue, with cp5, 4b and the coordinate-domain decision. In parallel and off the critical chain: #257 (without it #345'semitslice cannot close), #263 (without it #262's performance gates have no baseline), #345 now that #261'sown-cliskeleton exists, and #269's reconciliation as independent cleanup.The defensive limits that used to head this queue landed in #326, and their position was load-bearing rather than tidy: they changed what the reference accepts, so they had to land Python-first and cp1 had to be re-measured against them rather than merged beside them. The coordinate-domain decision was the same kind of item and landed the same way (PR #341); #260's two decisions were decided the same way and landed in PR #342.
C. IDE/incremental path
.ownRust LSP MVP with broken-code tolerance and cancellation. P-022 step 5a: port diagnostic messages and ordered Evidence with Python parity #255 and P-022 step 5b: port the SARIF projection with canonical parity #256 satisfied; informed by IDE foundation: establish cold/warm/incremental latency and memory baselines #263.IDE DAG:
D. P-034
Production implementation remains blocked until the runtime marker/helper/escape-hatch contract is finalized. Do not duplicate call-site use-after-dispose rules.
Recommended agent allocation
Owner decision (not delegable)
#260's decisions are resolved (D-4..D-7, B-2, B-3, R-1, R-2),#261's packet is ratified (C-1..C-5, 2026-09-08) and its 261.B implementation landed (PR #347,206e9c7); #269's reconciliation heads this list now.Local/corpus-capable agent
Strong agent
#255,#256,#258,#259and now#261's 261.B (PR #347) are complete and drop out of this chain.#259is closed at final acceptance — do not re-open it or any of its checkpoints; the declared OD-1 boundary is on the exclusion ledger, not open work.Separate strong or medium-strong agent
Keep codegen isolated from analysis.
Medium agent
Status-drift rule
A step is never described by a single Implemented/Missing bit. Any status edit to this issue or to
docs/proposals/P-022-rust-core-migration.mdmust state, per step: completed checkpoints, remaining acceptance, normative blocker, and preferred sequencing. Both surfaces are updated in the same change — a reconciliation that touches only one of them replaces a stale pair with a contradictory pair. The proposals index row counts as a third surface for the same fact.A child issue's own body is a fourth surface when its acceptance turns out to be wrong. #256 is the worked example: its requirements described a
.ownreport.jsonthe project does not have and had already refused to build, so the correction belongs in the issue, in this roadmap, in P-022 and in the index — together, or not at all. #259 is the second: its checkpoint list had no row for the obligation-protocol analysis while its final acceptance required it, so 4b was added to the child issue, to P-022 and here in the same move.A child issue's open/closed bit is a fifth surface, and #259 is its worked example too: GitHub closed it when PR #339 merged, because its parser read the body's "does not close #259" as
close #259— the negation is not parsed — while the sentence meant the opposite and this queue still read "→ #259 final acceptance". A closed issue with an open acceptance is a contradictory pair with the roadmap. It was reopened, and closed by hand at the #341 merge once the acceptance was reached. Checkpoint PRs reference their issue asRefs #Nand never put a closing keyword before an issue number, negated or not; nothing but the final-acceptance change closes it, and the owner does that by hand together with the body update.Acceptance-evidence surface (parity checkpoints only)
Parity checkpoints carry one more surface, and it is not a status surface: the frozen ledger. A checkpoint whose acceptance is "two implementations agree" is proved by an artifact that can share the implementation's blind spots, and a green matrix over an incomplete ledger is indistinguishable from a green matrix over a complete one. #259 cp1 is the worked example, three times over — 0/0/0 over 77 controls, then 58 permissive documents and 9 category mismatches once the ledger was rebuilt from the reference instead of from the author's reading of it, then 7 more permissive documents and 8 more category mismatches once the two deliberately excluded families were admitted.
The third round added a second failure mode worth naming separately: a ledger can carry the right controls and still take its category from the wrong place.
_check_columnraises one message for five distinct mechanisms, and the ledger inherited one message as one category — so the classification was correct about accept/reject and wrong about why, for a year, in a file whose entire purpose is to be right about why.cp5 added a third: a comparison surface that gains a member can lose controls for the members it subsumes. Putting
messageinto the BR-V7 dedup key made several key members unobservable at the output, and the cp4 campaign — re-run, not trusted — turned those mutations from caught to survived. The fix was controls that drive the productiondedupdirectly, and the rule is that every earlier campaign is re-run against the new tree, which is what the generated evidence pipeline exists to make cheap.Final acceptance added a fourth, about the evidence pipeline itself: a campaign's expected catchers can rot between runs without the gate noticing.
--validatere-anchored every mutation against the current tree, but it did not ask whether the tests a mutation names still exist or still fail; shadow-cp4's M61 carried two that had stopped failing onmainafter 4b's promotion. Closed in PR #342:--validatenow resolves every named catcher, and the re-run rule is written down.#260's acceptance work added a fifth, about the comparison itself: a green gate over an empty set is worse than a red one, because a red gate at least says it is awake. The compare driver therefore says out loud how many documents it compared and how many agreed, and the C# samples gate was read from its log rather than from its tick.
#260's sweep added a sixth, about provenance rather than comparison: the environment that records evidence can shape it. A checkout's line-ending setting changes the byte digest of a definition without changing a character; a machine's git identity rides into every commit made from it; a console codepage rewrites a commit message on the way in. None of it is visible to a reader of text on a screen, and all of it is caught only by comparing bytes and commit metadata. The rule: recorded artifacts are produced and compared as bytes; work committed from a local machine enters the tree under a repository identity; a pre-push report scans metadata as well as content; and the fix for the line-ending half is a researched
.gitattributes, not an operator's global configuration. PR #344 added the CI half of the same lesson: a workflow's manifest generator wrote the runner's temp path into every document'ssource— a constant dressed as provenance — and it was caught by the same check on the record before anything was committed; the generator now names a document by its id.The distinction matters: a ledger cannot report a normative blocker, a preferred sequencing or a remaining acceptance. It carries no project state at all. It is evidence for a claim the status surfaces make, so it is reviewed as evidence — is it derived from the reference or from the port, can it express absence, does every category have a control, and is any family excluded — and when it turns out to be incomplete, the correction is a further census with every result on the record, not a fix to the port.
For how to keep parity work honest — not just its status — see P-022 § Parity-work discipline.
Global PR acceptance packet
Every substantive child PR must state:
Migration PRs additionally report:
The final value must be
0.Stop conditions
Stop and redesign the child scope if any of these occur:
own-bridgereimplements ownership/lifetime/effect/DI algorithms.own-codegenstarts consuming analysis verdicts..nupkg.Note on (1) versus a corrected acceptance: striking a requirement because the tree proves it describes something that does not exist is not weakening it. The distinction is evidence — #256's strikes each carry a measurement, and the surfaces that stated the old acceptance were all corrected in the same move. Narrowing what the reference accepts as a Python-first contract decision with the ledger re-measured against it (#326, then #341) is not weakening either: the reference moved first, and the census followed.
Milestone completion
This roadmap reaches its next major milestone when all are true:
Owen Alpha
owen checkworks on Windows and Linux;Rust vertical
Findingand the rendered surfaces; declared boundary: the two OD-1 door controls (Bridge contract OD-1/2/3: pin the tolerant door — direct check_facts() diverges from load() (unknown kind fallback, line coercion) #294), measured, not open work;own-cli ownir, P-022 step 7b: the production Rust OwnIR executable (own-cli ownir) — command, output and exit-code parity behind the existing launcher #261's 261.B) landed (PR feat(own-cli): the production Rust OwnIR executable, behind the unchanged launcher (#261 261.B) #347) — command/output/exit-code parity behind the unchanged launcher, replaying with zero Python on Linux and Windows; the Rust CLI is ready for the explicit default-engine decision (P-022 step 8: Rust-default cutover, rollback gate, and Python distribution removal #262).IDE foundation
.ownLSP proves cancellation and stale-result correctness;Generated by Claude Code