Skip to content

feat: optional Jev policy head (suggest/act) for fast typed next-action decisions - #2654

Closed
aneym wants to merge 17 commits into
callstack:mainfrom
aneym:feat/jev-policy
Closed

aneym wants to merge 17 commits into
callstack:mainfrom
aneym:feat/jev-policy

Conversation

@aneym

@aneym aneym commented Sep 16, 2026 •

Copy link
Copy Markdown

Summary

An agent driving a device spends most of its wall clock deciding, not acting. A snapshot returns in
~250 ms and a settled press costs 1-10 s, but the model reading that snapshot and choosing a ref
takes seconds and a full inference, every step.

This adds an optional policy head: one HTTP call answering only "which of these elements advances
this goal", as a typed decision with calibrated probabilities. Two commands, always listed and
always offered over MCP, that refuse to run without TYPESAFE_API_KEY — a typed INVALID_ARGS
carrying reason: policy-provider-unconfigured and the fallback hint, raised before any snapshot,
press or fill reaches the device. Neither owns a daemon route, and no existing command's schema
moves.

agent-device suggest "sign in and reach the main list"
# Next: fill @e7 at confidence 0.99 · Probabilities: @e7 99.0%, @e9 0.6% · decide 198ms

agent-device act "sign in and reach the main list" --input phone=5005550700 --max-steps 10

suggest snapshots once and prints. act loops snapshot → decide → press/fill → re-snapshot,
logging snapshot, decide and action milliseconds plus token cost per step. A PolicyProvider
interface with one implementation, jev, over global fetch; no new dependency. Both compose
snapshot, press and fill through the client rather than owning a daemon route, so a run
carries the same claims, ref frames and recording behaviour as an agent driving by hand, each
mutation pinned to its snapshot generation. Text is never generated — a chosen field is filled from
--input key=value or AGENT_DEVICE_INPUT_<KEY>, and escalates otherwise. Every action is followed
by a content-digest comparison; one that changed nothing is a dead action, and three unproductive
steps end the run.

They join the interaction family, whose vocabulary they share (find, get and is are reads
that name an element; act loops the verbs beside them), and which keeps the startup closure flat:
the policy runtime lives under commands/policy and the client imports it on demand.

Validation

Tested at b4150c0. Green locally: pnpm typecheck, pnpm gate lint, pnpm gate format,
pnpm test:unit (10,461 passed, 1 skipped, 0 failed), pnpm check:affected --run, and the PR's CI
gates run individually — di-seams, layering, depgraph, gate-manifest, gate-manifest-model,
affected-selector, tmpdir-leaks-model, mcp-metadata, xctest-selection,
maestro-conformance, command-docs, agent-guidance, fallow, production-exports,
replay-compat, daemon-wire-compat, freerange, wire-compat-model, coverage-model, build,
bundle-owner-files, and the changed-line coverage gate.

74 new unit tests over a mocked HTTP layer and a fake device port. They pin the credential's exact
blast radius: the CLI catalog and the rendered MCP tools are identical with the key set and unset,
each tool takes a goal and no input in its schema can carry the credential, neither descriptor
carries a daemon route, no other descriptor changed, and with no key both commands reject against a
device surface that fails the test if it is touched. Changed-line coverage
is 314/346 (90.75%) against a 70% threshold.

Live on an iOS simulator with a sign-in app, erased between runs. Goal: sign in by phone and reach
the main list through onboarding. Seven device actions either way.

act ×4 agent by hand
status done ×4 done
steps / actions 8 / 7, every run 8 / 7
wrong picks · dead actions · escalations 0 · 0 · 0 0 · 0 · 0
decide per action median 405 ms (n=32) median 5,145 ms (n=7)
decide per run 3.1 / 3.4 / 3.2 / 3.1 s 45.5 s
device action time 41.3 / 42.8 / 44.7 / 36.6 s 34.1 s
wall clock 49.0 / 49.9 / 51.0 / 43.9 s 81.6 s
cost $0.00046 per run one inference per step

Deciding is 12× faster per action. Wall clock only 1.6×, because the device dominates and that cost
is identical either way.

Four loop behaviours exist because a live run found them, each commented at its decision site: a
masked phone field reports TEXT_ENTRY_MISMATCH on text that arrived, so the screen decides rather
than the write; a digit-box code field never reports its value, so a screen that moved on is the
other half of the evidence, with the on-screen keypad as fallback; a network-backed transition lands
after the UI settles, so one grace re-read separates it from a dead action; and blocked describes
the screen, not the goal — it was set on the sign-in screen and again on an onboarding import page
while naming the element that led onward, so it is terminal only when nothing is worth acting on.

Limitations. Needs an API key; no local head. Chooses between elements only: no text, no
multi-screen plan. Concurrency above ~4 invites rate limiting. 401/402/429 fail with the status as a
typed reason and the documented recovery, Jev unavailable: <status>; fall back to agent-driven policy. The keypad fallback has unit coverage only — the dev build used here pre-fills its code.

Scope. 34 files, +3,087/-20: 1,648 production lines, 1,269 test, 167 docs. Above the 1,000-line
budget. If the shape is right but the surface is too large, suggest alone is a coherent smaller
cut and the provider abstraction could collapse to its one implementation.

🤖 Generated with Claude Code

alexneyman and others added 15 commits September 16, 2026 18:06
`--policy`, `--min-confidence`, and `--input` for the policy commands, and one
consumer added to the existing `--max-steps` declaration rather than a second
flag under the same name.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The neutral half of the policy feature: the candidate projection a policy head
chooses between, a content digest that detects an action which changed nothing,
and the caller-supplied text map a field is filled from.

The digest excludes refs on purpose. Refs are reissued per snapshot generation,
so including them would make every screen look changed.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
One provider behind a named abstraction, over global fetch with no new
dependency. The credential comes from TYPESAFE_API_KEY in the environment only,
so it never reaches argv, a recorded script, or an MCP tool schema.

Every field of a response is validated against the candidate set rather than
trusted, and an unreachable head fails with the status as a typed reason plus
the documented fallback: the agent's own snapshot-and-choose loop.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Snapshot, decide, act, verify, repeat. Three device operations behind a narrow
port so the loop is testable without a device.

What the loop had to learn from a real iOS sign-in flow:

- A settled observation only proves the local UI went quiet, so a transition
  waiting on the network still reads as the old screen. One grace re-read
  separates a slow transition from an action that did nothing.
- A masked field reformats what it receives, so the runner reports
  TEXT_ENTRY_MISMATCH on text that arrived. The screen decides, not the write.
- A digit-box code field never reflects its value at all. A confirmed value is
  sufficient evidence of a write but not necessary; a screen that moved on is
  the other half, and the on-screen keypad is the fallback when neither holds.
- A press on a text field only focuses it, so the supplied entry rather than the
  field's emptiness decides whether to write.
- A head asked whether progress is blocked says yes on any sign-in screen. A
  concrete target at 0.9 confidence outranks that flag, or the loop stalls on
  step one of every sign-in.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
`suggest <goal>` prints one typed decision without acting. `act <goal>` runs the
loop. Both are local-cli commands: they compose snapshot, press, and fill from
the client process and own no daemon route, which is what keeps them out of the
public catalog that requires one.

Every nested call is an ordinary daemon command, so a policy-driven run carries
the same claims, ref frames, and recording behaviour as a run an agent drives by
hand, and each mutation is pinned to the generation of the snapshot that issued
its ref.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Both are pure delegators, so they join the `none` device-claim list. `act` joins
the daemon-preserving and unbounded-envelope lists: a run is a sequence of
ordinary commands that each carry their own envelope, and an outer one would
abort a run still making progress while a reset would destroy the session the
remaining steps need.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
A command-reference section, a reference page covering the decision shape, the
text-supply rule, the post-action checks, and the limitations, plus one
paragraph on the skill routing card.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The changed-code quality gate flagged the step loop as critically complex and
six exports as unconsumed. The loop now reads as snapshot, decide, resolve step,
record, with the write and press branches and the keypad walk as their own
functions, and the candidate projection resolves a node's role separately from
building the candidate. The internal constants are module-local.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The hint read "Without it, Fall back to agent-driven policy" wherever it was
interpolated. The clause now lives in its own module, lower-cased, so both the
provider and the composer can name it without one importing the other.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
A policy asked whether progress is blocked says yes on anything that looks like
a wall. In live runs on an iOS sign-in flow the flag was set twice while the
policy also named the element that led onward: the phone field on the sign-in
screen, and "I'll start fresh" on an onboarding page offering a file import. The
previous rule acted through the flag only above 0.9 confidence, so the second
one ended the run three screens short of the goal.

The flag is now terminal only together with nothing worth acting on: no target,
or one below --min-confidence. A named target above the floor is acted on. Being
wrong there costs one step that the same-screen check catches and the step budget
bounds; believing the flag costs the run.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The CLI startup closure must not grow, and a static edge from the client facade
pulled the whole policy subtree into every invocation. The policy head is
optional, so its client loads inside the two methods that use it, the way debug
symbols already does.

Also replaces a `as string` on the chosen target with a discriminated judgment,
and reaches the client fixture's device group through the spread idiom that file
already uses for observability, so the oversized-test-file ratchet holds.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…ady reaches

The eager-closure gate holds src/cli.ts to the module count at the merge-base,
and a command family is reached from it through the family registry, so a new
family directory grows it by every module the family index evaluates. The gate
names the remedy: give the new code a home in a module the closure already
evaluates.

suggest and act join the interaction family, whose vocabulary they already
share: find, get, and is are reads that name an element, and act is a loop over
the verbs beside it. Their CLI defaults move next to the flags that publish
them, so the help prose and the loop read one declaration. The policy runtime
stays under commands/policy and is still reached only through the client's
on-demand import.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The move left both descriptors naming a deleted module, which the explain-command
table caught, and left three exports with no consumers.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…nges

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The changed-line coverage gate passed at 83%, but the two largest gaps were the
seam that pins a mutation to its snapshot generation and the text both commands
print. A run that loses the pin fails silently, on whichever element inherited
the ref, so it is worth its own test rather than coverage by a live run.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@aneym
aneym marked this pull request as ready for review September 16, 2026 23:50
alexneyman and others added 2 commits September 16, 2026 19:56
"Gated on TYPESAFE_API_KEY" read as though the commands disappear without one.
They do not: the catalog and the MCP tool list are the same either way, because
neither reads the environment. What the credential gates is whether a run
proceeds. The surface-identity test now asserts exactly that, by building both
projections with the key set and unset and comparing them, and by pinning that
neither command owns a daemon route.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The exposed-name list proves the commands are registered, not that a usable tool
reaches a model. The assertion now reads the rendered tool: it takes a goal, its
description names the variable to set, and no input in its schema can ever carry
the credential's value.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@thymikee

Copy link
Copy Markdown
Member

Reviewed at b4150c0.

The second matcher in src/commands/policy/text-inputs.ts#L51 accepts any field whose identifier or label contains the key as a substring. --input pin=1234 matches a "Shipping address" field, and code matches "Promo code" or "Zip code". An env value from AGENT_DEVICE_INPUT_* matches the same way in any app the goal reaches. textReadyRefs (act-loop.ts#L187) then marks those fields ready to fill, so supplied text, including secrets from the env channel, can land in the wrong field. Because the loop snapshots again after, that text also gets sent to the third-party endpoint. A supplied value should reach only a field it names exactly: keep exact identifier/label equality, or match whole words only, in both textReadyRefs and resolveTextEntry. Can we get a test for the "pin" vs "Shipping" case?

The keypad fallback in src/commands/policy/act-loop.ts#L318 runs whenever valueLanded is false and the text is all digits. valueLanded is always false for an iOS secure text field, whose AX value is bullets that sameText normalizes to empty, and for a field with no identifier or label, since candidate.name falls back to the node type and never matches entry.label. In both cases a successful fill is followed by digit-by-digit keypad presses, so the field ends up holding the PIN twice, and the endsWith check in sameText (line 246) hides this on later steps. A PIN or passcode can be entered twice into secure or unlabeled numeric fields while the step reports as acted. The keypad fallback should run only when the loop has positive evidence the write did not arrive: a readable value that differs from the text, not an unreadable one. Match the verification node by ref or position from the same snapshot, skip the fallback for secure fields, and drop or anchor the endsWith equivalence to the mask.

LEAF_ROLES, STRUCTURAL_TYPES and the StaticText/Cell/Key checks in src/commands/policy/candidate-elements.ts#L19 are XCUI type names. On Android, node.type is the raw class name (packages/platform-android/src/ui-hierarchy.ts#L303), so toPolicyCandidates returns nothing and buildJevRequest throws policy-no-candidates on every screen, even though the commands are listed for every platform. Candidates should be projected from the platform-neutral role vocabulary the snapshot presenter already uses, not from raw types; otherwise gate suggest/act to Apple targets with a typed unsupported-platform error and say so in the docs.

suggest in src/client/client-policy.ts#L75 takes a complete snapshot, which activates a new ref frame per ADR 0014, and returns refs from that frame without refsGeneration. The MCP pin store's REF_ISSUING_TOOLS (src/mcp/tool-ref-pins.ts) doesn't include suggest, so snapshot -> suggest -> press <ref> over MCP sends a stale frame and gets rejected with ref_generation_mismatch. That's the documented flow, and it fails whenever a snapshot happened earlier in the session, which is the normal case. Every response that issues refs should carry refsGeneration and be merged by the MCP pin store; add suggest to the ref-issuing set, ideally derived from a descriptor trait rather than a second hand-kept list.

state.elements in src/commands/policy/jev-composer.ts#L61 sends every candidate to the third-party host, including every label, identifier, accessibility value, and text field value. After the loop fills a non-secure field from --input or AGENT_DEVICE_INPUT_*, the next snapshot's candidate value carries that text into the next decide call to the vendor. Neither policy-head.md nor the tool descriptions disclose this, and the textReadyRefs comment ("values are never included") says the opposite. Values users were told to keep out of argv, plus any on-screen PII, end up sent to an external vendor undisclosed. Redact the value of any field the loop filled, or any value equal to a supplied input, before building the request, and document exactly which fields are sent and name the host in policy-head.md and the tool descriptions.

src/client/client-policy.ts#L131 calls fill without recordAs. The daemon only parameterizes a recorded fill literal when --record-as is given (session-action-recorder.ts, ADR 0017), so a session armed with --save-script records an env-sourced code or password value literally in the .ad script. That contradicts the docs, which call the env channel the way to keep secrets out of argv. Pass recordAs derived from the input key for every policy-sourced fill, through the existing --record-as seam.

Not blocking: the provider's error path can leak the API key into an error message when the key has a control character (reproduced with plain Node fetch, undici's header-value error), writeFailure swallows every fill rejection instead of only the named typed reasons production emits, a rejected step discards the whole act loop and the prior steps' mutations instead of returning a partial result with a per-step failure, AGENT_DEVICE_POLICY_ENDPOINT overrides the credentialed host without requiring https or being documented, the MCP tools are always offered without stating the egress and host in their descriptions, act's unbounded timeout policy and onStep have no real caller or cancellation, the act-loop tests use hand-written error codes no production path emits and never go through the real client/daemon, policyClientCalls re-implements client methods with a cast instead of using the typed client, the CLI's --min-confidence and step-count flags don't match the MCP bounds, a placeholder on an iOS field is read as its value and skips the "needs text" path, and the paragraph-long workaround comments plus the unbounded suggest candidates array and echoed response body should be trimmed; all of these can be taken or left.

Nothing in the loop needs core internals — it only uses capture.snapshot (with refsGeneration), interactions.press and interactions.fill, all already exposed by createAgentDeviceClient (see examples/sdk/client-session.ts). Would a separate package, like @agent-device/policy-jev, or an examples/sdk/policy-loop.ts, be a better home for the vendor, the credential, the pricing, the candidate projection and the loop, so core gains no commands, flags, MCP tools, registry entries or env contracts? If the loop stays in core, the smaller cut the author already offers, suggest alone with the provider collapsed to its one implementation, still needs the ref-generation, egress-disclosure and Android-candidate issues fixed. Either way, this needs a maintainer decision, or an ADR, on whether core should ship vendor-specific paid integrations that send screen content to a third party. Without that, a new provider seam with one vendor and two always-listed commands is growth without an owner.

This PR drives a device, so it needs live evidence: an agent-device act run on this commit with the --debug request log showing snapshot/press/fill requests carrying ~sN pinned refs and the step table. That should cover a fill into a numeric field whose value is unreadable, such as a secure PIN field, showing no double entry; an Android emulator run of suggest that returns a decision or a typed unsupported-platform refusal; an MCP session doing snapshot -> suggest -> press <suggested ref> that succeeds; and the keypad path, which the author says currently has unit coverage only. The PR body's iOS table has no logs or artifacts to back it.

I did not run tests, gates, or any device; these findings come from reading the code, except the header-value leak, which I reproduced with plain Node fetch. The Android finding assumes the client snapshot keeps node.type raw, based on ui-hierarchy.ts#L303 and existing fixtures, without tracing every presenter step. The ref-generation finding assumes an interactive-only snapshot activates a new complete frame per how ADR 0014 reads, without running it. The recording finding assumes recording is armed; I did not confirm the literal output in a saved .ad script. The PR body's live iOS numbers, test counts and gate results are unverified, and the packet shows zero checks. This is a cross-repo PR, so workflows likely need approval before they can run, and the registry, MCP-metadata and depgraph checks should run given the ledger, MCP surface and client changes here.

Next: settle where this loop belongs. If it stays in core, fix the ref-generation, secret-leak, recording, Android-candidate and double-entry issues above, then attach the live evidence described.

@thymikee

thymikee commented Sep 18, 2026 •

Copy link
Copy Markdown
Member

Thank you for this PR, and for the care in it. The live-run notes on masked fields, digit-box code fields, network-backed transitions and the blocked flag are exactly the kind of evidence we like to see, and they gave us the material for the two issues below.

Maintainer decision at b4150c0: we will not merge this shape into core. The reason is scope, not quality. A vendor HTTP client, a vendor API key, a pricing constant, and two always-listed CLI commands and MCP tools that refuse without the key put model handling inside core. agent-device is the device side of an agent; we keep model handling outside core, in agent-device/ai-sdk or in the host's own loop. That decision stands independently of the review findings above, and it applies to suggest alone as well.

We want to keep the idea, in a different home. We opened two issues that carry the direction:

We also checked the speed premise without a vendor. A schema-forced choice among candidate refs on a small general model picked correctly 16/16 on your three sign-in screens plus a real iOS artifact, at 620-730 ms median and about 200 input tokens per decision. So the loop keeps its gain with whatever model the host brings.

Two fill behaviors your loop worked around belong in fill itself. #2634 tracks the masked-field TEXT_ENTRY_MISMATCH. The digit-box code field that never takes a fill has no issue yet. If you can file one with the app, the command and the output, we will take it from there; the keypad path here has unit coverage only, so a repro would help a lot.

We are closing this PR. Smaller PRs against #2656, and later #2657, are very welcome, and we would be glad to review them.

@thymikee thymikee closed this Sep 18, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants