Skip to content

feat(eval): ondemand simulate — replay a dataset, evaluate synchronously - #2071

Merged
jariy17 merged 6 commits into
refactorfrom
feat/eval-ondemand-simulate
Aug 31, 2026
Merged

feat(eval): ondemand simulate — replay a dataset, evaluate synchronously#2071
jariy17 merged 6 commits into
refactorfrom
feat/eval-ondemand-simulate

Conversation

@jariy17

@jariy17 jariy17 commented Aug 22, 2026

Copy link
Copy Markdown
Contributor

What

eval ondemand simulate — replay a dataset against a runtime, then evaluate the sessions synchronously, client-side (scores print inline). The on-demand twin of batch-evaluation simulate.

Pipeline: invokeDataset (replay) → getTracesForAgent (CloudWatch) → evaluate (Evaluate API).

Follows the batch-evaluation simulate pattern

  • Same invoke flags + --ingestion-wait-ms (default 180000; 0 to skip)
  • Same SIGINT abort + invokeDataset composition
  • Refuses when 0 invoked, naming the first failure (failures[0])
  • Output adds sessions[] (exampleId ↔ sessionId) + failures[]

Differs by design: no --name / --kms-key-arn (no async job); output is inline scores, not a job id.

Output

{
  "sessionsEvaluated": 1,
  "results": [ { "evaluatorId": "Builtin.Helpfulness", "value": 0.83, "label": "Very Helpful", "explanation": "", "tokenUsage": {} } ],
  "examplesInvoked": 1,
  "examplesFailed": 0,
  "sessions": [ { "exampleId": "greet", "sessionId": "" } ],
  "failures": []
}

Tests (batch pattern)

  • Edge testsondemand.test.tsx (TestCoreClient): required-flag validation, refuse-when-nothing-invoked, --ingestion-wait-ms passthrough.
  • Fixture goldenondemand.fixture.test.tsx: real invoke → CloudWatch traces → Evaluate, recorded against a live agent, replayed offline via matchGolden. Deterministic via the injected newSessionId seam plus a now() clock seam — on-demand's trace-query window is otherwise Date.now()-based, so its StartQuery fixture key would drift between record and replay (batch has no client query; evaluate pins an explicit --start-time/--end-time).
  • Live-validated across 10 dataset scenarios against a real agent in the EXPLORE account (all real Builtin.Helpfulness scores; edge/negative paths handled).

Notes

  • Builtin.Helpfulness ignores ground-truth refs (ignoredReferenceInputFields); the handler still forwards them (adapter exercised).
  • Rebased onto current refactor (was stacked on the pre-merge batch-simulate work).

@github-actions github-actions Bot added size/m PR size: M agentcore-harness-reviewing AgentCore Harness review in progress and removed agentcore-harness-reviewing AgentCore Harness review in progress labels Aug 22, 2026
@codecov-commenter

codecov-commenter commented Aug 22, 2026

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.
✅ Project coverage is 97.16%. Comparing base (f2b73d6) to head (0cfacfe).

Additional details and impacted files
@@             Coverage Diff              @@
##           refactor    #2071      +/-   ##
============================================
+ Coverage     97.14%   97.16%   +0.02%     
============================================
  Files           494      495       +1     
  Lines         32550    32676     +126     
============================================
+ Hits          31622    31751     +129     
+ Misses          928      925       -3     

☔ View full report in Codecov by Harness.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.
  • 📦 JS Bundle Analysis: Save yourself from yourself by tracking and limiting bundle sizes in JS merges.

@jariy17
jariy17 force-pushed the feat/eval-ondemand-simulate branch from 634f6f9 to 94a16ac Compare August 22, 2026 16:23
Base automatically changed from feat/eval-invoke-dataset-pr to refactor August 24, 2026 22:48
@jariy17
jariy17 force-pushed the feat/eval-ondemand-simulate branch from 94a16ac to 623bbfc Compare August 27, 2026 20:53
@github-actions github-actions Bot added size/m PR size: M and removed size/m PR size: M labels Aug 27, 2026
@agentcore-devx-automation agentcore-devx-automation Bot added the claude-security-reviewing Claude Code /security-review in progress label Aug 27, 2026
@agentcore-devx-automation

Copy link
Copy Markdown
Contributor

Claude Security Review: no high-confidence findings. (run)

@agentcore-devx-automation agentcore-devx-automation Bot removed the claude-security-reviewing Claude Code /security-review in progress label Aug 27, 2026
@github-actions github-actions Bot added size/xl PR size: XL and removed size/m PR size: M labels Aug 27, 2026
@agentcore-devx-automation agentcore-devx-automation Bot added the claude-security-reviewing Claude Code /security-review in progress label Aug 27, 2026
@github-actions github-actions Bot added size/xl PR size: XL and removed size/xl PR size: XL labels Aug 27, 2026
@jariy17
jariy17 marked this pull request as ready for review August 27, 2026 21:16
@agentcore-devx-automation

Copy link
Copy Markdown
Contributor

Claude Security Review: no high-confidence findings. (run)

@agentcore-devx-automation agentcore-devx-automation Bot removed the claude-security-reviewing Claude Code /security-review in progress label Aug 27, 2026
@jariy17
jariy17 force-pushed the feat/eval-ondemand-simulate branch from ec6636c to f6776cc Compare August 27, 2026 21:17
@github-actions github-actions Bot added size/xl PR size: XL and removed size/xl PR size: XL labels Aug 27, 2026
@agentcore-devx-automation agentcore-devx-automation Bot added the claude-security-reviewing Claude Code /security-review in progress label Aug 27, 2026
@agentcore-devx-automation

Copy link
Copy Markdown
Contributor

Claude Security Review: no high-confidence findings. (run)

@agentcore-devx-automation agentcore-devx-automation Bot removed the claude-security-reviewing Claude Code /security-review in progress label Aug 27, 2026
@jariy17
jariy17 force-pushed the feat/eval-ondemand-simulate branch from f6776cc to c29ad31 Compare August 27, 2026 21:20
@github-actions github-actions Bot removed the size/xl PR size: XL label Aug 27, 2026
@github-actions github-actions Bot added size/l PR size: L and removed size/xl PR size: XL labels Aug 28, 2026
@agentcore-devx-automation agentcore-devx-automation Bot added the claude-security-reviewing Claude Code /security-review in progress label Aug 28, 2026
@agentcore-devx-automation

Copy link
Copy Markdown
Contributor

Claude Security Review: no high-confidence findings. (run)

@agentcore-devx-automation agentcore-devx-automation Bot removed the claude-security-reviewing Claude Code /security-review in progress label Aug 28, 2026
jariy17 pushed a commit that referenced this pull request Aug 31, 2026
… expectedResponse as a reference input

Addresses review on #2071: swap the hand-rolled SIGINT/AbortController for the shared
withUserCancellation helper, and include per-turn expectedResponse in the Evaluate
reference inputs so an expected-response-only dataset still contributes ground truth.
@jariy17
jariy17 force-pushed the feat/eval-ondemand-simulate branch from 1b0ca37 to 2c20882 Compare August 31, 2026 15:47
@github-actions github-actions Bot added size/l PR size: L and removed size/l PR size: L labels Aug 31, 2026
@agentcore-devx-automation agentcore-devx-automation Bot added the claude-security-reviewing Claude Code /security-review in progress label Aug 31, 2026
@agentcore-devx-automation

Copy link
Copy Markdown
Contributor

Claude Security Review: no high-confidence findings. (run)

@agentcore-devx-automation agentcore-devx-automation Bot removed the claude-security-reviewing Claude Code /security-review in progress label Aug 31, 2026
jariy17 added 6 commits August 31, 2026 17:17
… expectedResponse as a reference input

Addresses review on #2071: swap the hand-rolled SIGINT/AbortController for the shared
withUserCancellation helper, and include per-turn expectedResponse in the Evaluate
reference inputs so an expected-response-only dataset still contributes ground truth.
…late turn→traceId)

Bug bash against a live agent showed the Evaluate API rejects expectedResponse under a
session-only context. Correlate each turn's expectedResponse to its trace id instead, and
extend the fixture golden dataset with an expected_response turn graded by Builtin.Correctness
so the real Evaluate call validates the mapping at record time.
@jariy17
jariy17 force-pushed the feat/eval-ondemand-simulate branch from 2c20882 to 0cfacfe Compare August 31, 2026 17:18
@github-actions github-actions Bot added size/l PR size: L and removed size/l PR size: L labels Aug 31, 2026
@agentcore-devx-automation agentcore-devx-automation Bot added the claude-security-reviewing Claude Code /security-review in progress label Aug 31, 2026
@agentcore-devx-automation

Copy link
Copy Markdown
Contributor

Claude Security Review: no high-confidence findings. (run)

@agentcore-devx-automation agentcore-devx-automation Bot removed the claude-security-reviewing Claude Code /security-review in progress label Aug 31, 2026
@jariy17
jariy17 merged commit f431dce into refactor Aug 31, 2026
21 checks passed
@jariy17
jariy17 deleted the feat/eval-ondemand-simulate branch August 31, 2026 17:40
),
flag("header", "an ordered application header (repeatable)", z.array(z.string()).optional()),
flag(
"bearer-token",

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

this and --header should be marked sensitive

jariy17 pushed a commit that referenced this pull request Aug 31, 2026
Follow-up to #2071. These flags carry secrets (CUSTOM_JWT bearer token,
auth headers) but were logged in cleartext by the withLogging debug
middleware. The canonical 'runtime invoke' handler already marks both
sensitive; apply the same to the on-demand and batch-evaluation simulate
handlers.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size/l PR size: L

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants