Conversation
`--policy`, `--min-confidence`, and `--input` for the policy commands, and one consumer added to the existing `--max-steps` declaration rather than a second flag under the same name. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The neutral half of the policy feature: the candidate projection a policy head chooses between, a content digest that detects an action which changed nothing, and the caller-supplied text map a field is filled from. The digest excludes refs on purpose. Refs are reissued per snapshot generation, so including them would make every screen look changed. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
One provider behind a named abstraction, over global fetch with no new dependency. The credential comes from TYPESAFE_API_KEY in the environment only, so it never reaches argv, a recorded script, or an MCP tool schema. Every field of a response is validated against the candidate set rather than trusted, and an unreachable head fails with the status as a typed reason plus the documented fallback: the agent's own snapshot-and-choose loop. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Snapshot, decide, act, verify, repeat. Three device operations behind a narrow port so the loop is testable without a device. What the loop had to learn from a real iOS sign-in flow: - A settled observation only proves the local UI went quiet, so a transition waiting on the network still reads as the old screen. One grace re-read separates a slow transition from an action that did nothing. - A masked field reformats what it receives, so the runner reports TEXT_ENTRY_MISMATCH on text that arrived. The screen decides, not the write. - A digit-box code field never reflects its value at all. A confirmed value is sufficient evidence of a write but not necessary; a screen that moved on is the other half, and the on-screen keypad is the fallback when neither holds. - A press on a text field only focuses it, so the supplied entry rather than the field's emptiness decides whether to write. - A head asked whether progress is blocked says yes on any sign-in screen. A concrete target at 0.9 confidence outranks that flag, or the loop stalls on step one of every sign-in. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
`suggest <goal>` prints one typed decision without acting. `act <goal>` runs the loop. Both are local-cli commands: they compose snapshot, press, and fill from the client process and own no daemon route, which is what keeps them out of the public catalog that requires one. Every nested call is an ordinary daemon command, so a policy-driven run carries the same claims, ref frames, and recording behaviour as a run an agent drives by hand, and each mutation is pinned to the generation of the snapshot that issued its ref. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Both are pure delegators, so they join the `none` device-claim list. `act` joins the daemon-preserving and unbounded-envelope lists: a run is a sequence of ordinary commands that each carry their own envelope, and an outer one would abort a run still making progress while a reset would destroy the session the remaining steps need. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
A command-reference section, a reference page covering the decision shape, the text-supply rule, the post-action checks, and the limitations, plus one paragraph on the skill routing card. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The changed-code quality gate flagged the step loop as critically complex and six exports as unconsumed. The loop now reads as snapshot, decide, resolve step, record, with the write and press branches and the keypad walk as their own functions, and the candidate projection resolves a node's role separately from building the candidate. The internal constants are module-local. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The hint read "Without it, Fall back to agent-driven policy" wherever it was interpolated. The clause now lives in its own module, lower-cased, so both the provider and the composer can name it without one importing the other. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
A policy asked whether progress is blocked says yes on anything that looks like a wall. In live runs on an iOS sign-in flow the flag was set twice while the policy also named the element that led onward: the phone field on the sign-in screen, and "I'll start fresh" on an onboarding page offering a file import. The previous rule acted through the flag only above 0.9 confidence, so the second one ended the run three screens short of the goal. The flag is now terminal only together with nothing worth acting on: no target, or one below --min-confidence. A named target above the floor is acted on. Being wrong there costs one step that the same-screen check catches and the step budget bounds; believing the flag costs the run. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The CLI startup closure must not grow, and a static edge from the client facade pulled the whole policy subtree into every invocation. The policy head is optional, so its client loads inside the two methods that use it, the way debug symbols already does. Also replaces a `as string` on the chosen target with a discriminated judgment, and reaches the client fixture's device group through the spread idiom that file already uses for observability, so the oversized-test-file ratchet holds. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…ady reaches The eager-closure gate holds src/cli.ts to the module count at the merge-base, and a command family is reached from it through the family registry, so a new family directory grows it by every module the family index evaluates. The gate names the remedy: give the new code a home in a module the closure already evaluates. suggest and act join the interaction family, whose vocabulary they already share: find, get, and is are reads that name an element, and act is a loop over the verbs beside it. Their CLI defaults move next to the flags that publish them, so the help prose and the loop read one declaration. The policy runtime stays under commands/policy and is still reached only through the client's on-demand import. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The move left both descriptors naming a deleted module, which the explain-command table caught, and left three exports with no consumers. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…nges Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The changed-line coverage gate passed at 83%, but the two largest gaps were the seam that pins a mutation to its snapshot generation and the text both commands print. A run that loses the pin fails silently, on whichever element inherited the ref, so it is worth its own test rather than coverage by a live run. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
"Gated on TYPESAFE_API_KEY" read as though the commands disappear without one. They do not: the catalog and the MCP tool list are the same either way, because neither reads the environment. What the credential gates is whether a run proceeds. The surface-identity test now asserts exactly that, by building both projections with the key set and unset and comparing them, and by pinning that neither command owns a daemon route. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The exposed-name list proves the commands are registered, not that a usable tool reaches a model. The assertion now reads the rendered tool: it takes a goal, its description names the variable to set, and no input in its schema can ever carry the credential's value. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
Reviewed at b4150c0. The second matcher in The keypad fallback in
Not blocking: the provider's error path can leak the API key into an error message when the key has a control character (reproduced with plain Node fetch, undici's header-value error), Nothing in the loop needs core internals — it only uses This PR drives a device, so it needs live evidence: an I did not run tests, gates, or any device; these findings come from reading the code, except the header-value leak, which I reproduced with plain Node fetch. The Android finding assumes the client snapshot keeps Next: settle where this loop belongs. If it stays in core, fix the ref-generation, secret-leak, recording, Android-candidate and double-entry issues above, then attach the live evidence described. |
|
Thank you for this PR, and for the care in it. The live-run notes on masked fields, digit-box code fields, network-backed transitions and the Maintainer decision at b4150c0: we will not merge this shape into core. The reason is scope, not quality. A vendor HTTP client, a vendor API key, a pricing constant, and two always-listed CLI commands and MCP tools that refuse without the key put model handling inside core. agent-device is the device side of an agent; we keep model handling outside core, in We want to keep the idea, in a different home. We opened two issues that carry the direction:
We also checked the speed premise without a vendor. A schema-forced choice among candidate refs on a small general model picked correctly 16/16 on your three sign-in screens plus a real iOS artifact, at 620-730 ms median and about 200 input tokens per decision. So the loop keeps its gain with whatever model the host brings. Two We are closing this PR. Smaller PRs against #2656, and later #2657, are very welcome, and we would be glad to review them. |
Summary
An agent driving a device spends most of its wall clock deciding, not acting. A snapshot returns in
~250 ms and a settled press costs 1-10 s, but the model reading that snapshot and choosing a ref
takes seconds and a full inference, every step.
This adds an optional policy head: one HTTP call answering only "which of these elements advances
this goal", as a typed decision with calibrated probabilities. Two commands, always listed and
always offered over MCP, that refuse to run without
TYPESAFE_API_KEY— a typedINVALID_ARGScarrying
reason: policy-provider-unconfiguredand the fallback hint, raised before any snapshot,press or fill reaches the device. Neither owns a daemon route, and no existing command's schema
moves.
suggestsnapshots once and prints.actloops snapshot → decide → press/fill → re-snapshot,logging snapshot, decide and action milliseconds plus token cost per step. A
PolicyProviderinterface with one implementation,
jev, over globalfetch; no new dependency. Both composesnapshot,pressandfillthrough the client rather than owning a daemon route, so a runcarries the same claims, ref frames and recording behaviour as an agent driving by hand, each
mutation pinned to its snapshot generation. Text is never generated — a chosen field is filled from
--input key=valueorAGENT_DEVICE_INPUT_<KEY>, and escalates otherwise. Every action is followedby a content-digest comparison; one that changed nothing is a dead action, and three unproductive
steps end the run.
They join the
interactionfamily, whose vocabulary they share (find,getandisare readsthat name an element;
actloops the verbs beside them), and which keeps the startup closure flat:the policy runtime lives under
commands/policyand the client imports it on demand.Validation
Tested at
b4150c0. Green locally:pnpm typecheck,pnpm gate lint,pnpm gate format,pnpm test:unit(10,461 passed, 1 skipped, 0 failed),pnpm check:affected --run, and the PR's CIgates run individually —
di-seams,layering,depgraph,gate-manifest,gate-manifest-model,affected-selector,tmpdir-leaks-model,mcp-metadata,xctest-selection,maestro-conformance,command-docs,agent-guidance,fallow,production-exports,replay-compat,daemon-wire-compat,freerange,wire-compat-model,coverage-model,build,bundle-owner-files, and the changed-linecoveragegate.74 new unit tests over a mocked HTTP layer and a fake device port. They pin the credential's exact
blast radius: the CLI catalog and the rendered MCP tools are identical with the key set and unset,
each tool takes a goal and no input in its schema can carry the credential, neither descriptor
carries a daemon route, no other descriptor changed, and with no key both commands reject against a
device surface that fails the test if it is touched. Changed-line coverage
is 314/346 (90.75%) against a 70% threshold.
Live on an iOS simulator with a sign-in app, erased between runs. Goal: sign in by phone and reach
the main list through onboarding. Seven device actions either way.
Deciding is 12× faster per action. Wall clock only 1.6×, because the device dominates and that cost
is identical either way.
Four loop behaviours exist because a live run found them, each commented at its decision site: a
masked phone field reports
TEXT_ENTRY_MISMATCHon text that arrived, so the screen decides ratherthan the write; a digit-box code field never reports its value, so a screen that moved on is the
other half of the evidence, with the on-screen keypad as fallback; a network-backed transition lands
after the UI settles, so one grace re-read separates it from a dead action; and
blockeddescribesthe screen, not the goal — it was set on the sign-in screen and again on an onboarding import page
while naming the element that led onward, so it is terminal only when nothing is worth acting on.
Limitations. Needs an API key; no local head. Chooses between elements only: no text, no
multi-screen plan. Concurrency above ~4 invites rate limiting. 401/402/429 fail with the status as a
typed
reasonand the documented recovery,Jev unavailable: <status>; fall back to agent-driven policy. The keypad fallback has unit coverage only — the dev build used here pre-fills its code.Scope. 34 files, +3,087/-20: 1,648 production lines, 1,269 test, 167 docs. Above the 1,000-line
budget. If the shape is right but the surface is too large,
suggestalone is a coherent smallercut and the provider abstraction could collapse to its one implementation.
🤖 Generated with Claude Code