Skip to content

direct: persist state before long-running waits in DoCreate/DoUpdate - #5391

Draft
denik wants to merge 32 commits into
mainfrom
denik/wait-method-removal
Draft

direct: persist state before long-running waits in DoCreate/DoUpdate#5391
denik wants to merge 32 commits into
mainfrom
denik/wait-method-removal

Conversation

@denik

@denik denik commented Jun 1, 2026

Copy link
Copy Markdown
Contributor

Changes

Replace the WaitAfterCreate/WaitAfterUpdate resource methods with inline
waits inside DoCreate/DoUpdate, and give both methods a *StateSaver
argument so a resource can persist intermediate state before a long-running
wait. This prevents orphaning a created resource if the deploy is interrupted
mid-wait: the state is already recorded, so the next deploy reconciles it
instead of leaking it.

StateSaver deduplicates writes, logs I/O failures without aborting the
deploy, and routes the final Create/Update save so an id mismatch is caught.
SaveStateWith temporarily overrides a field (e.g. published=false,
started=true) so the planner sees a real diff if the wait is interrupted.

Tests

Unit and acceptance tests, including new dashboard publish-failure/retry
scenarios that exercise the save-before-wait behavior.

@denik
denik temporarily deployed to test-trigger-is June 1, 2026 08:13 — with GitHub Actions Inactive
@denik
denik temporarily deployed to test-trigger-is June 1, 2026 08:13 — with GitHub Actions Inactive
@eng-dev-ecosystem-bot

eng-dev-ecosystem-bot commented Jun 1, 2026

Copy link
Copy Markdown
Collaborator

Integration test report

Commit: 060ad91

Run: 31008430713

Env ❌​FAIL 🟨​KNOWN 🔄​flaky 💚​RECOVERED 🙈​SKIP ✅​pass 🙈​skip Time
🔄​ aws linux 4 4 4 293 1095 43:47
❌​ aws windows 2 1 2 3 4 295 1093 37:52
🔄​ azure linux 1 4 4 289 1097 16:15
💚​ azure windows 4 4 292 1095 14:30
💚​ gcp linux 1 5 291 1097 12:20
🔄​ gcp windows 1 1 5 292 1095 12:34
14 interesting tests: 4 SKIP, 4 flaky, 3 RECOVERED, 2 FAIL, 1 KNOWN
Test Name aws linux aws windows azure linux azure windows gcp linux gcp windows
🟨​ TestAccept 💚​R 🟨​K 💚​R 💚​R 💚​R 💚​R
🙈​ TestAccept/bundle/invariant/no_drift 🙈​S 🙈​S 🙈​S 🙈​S 🙈​S 🙈​S
🔄​ TestAccept/bundle/resources/postgres_endpoints/update_autoscaling/DATABRICKS_BUNDLE_ENGINE=direct 🔄​f 🔄​f
🔄​ TestAccept/bundle/resources/postgres_endpoints/update_autoscaling/DATABRICKS_BUNDLE_ENGINE=terraform 🔄​f 🔄​f
❌​ TestAccept/bundle/resources/postgres_projects/update_display_name 🔄​f ❌​F 🙈​s 🙈​s 🙈​s 🙈​s
❌​ TestAccept/bundle/resources/postgres_projects/update_display_name/DATABRICKS_BUNDLE_ENGINE=direct 🔄​f ❌​F
🙈​ TestAccept/bundle/resources/vector_search_endpoints/drift/recreated_same_name 🙈​S 🙈​S 🙈​S 🙈​S 🙈​S 🙈​S
🙈​ TestAccept/bundle/resources/vector_search_indexes/recreate/embedding_dimension 🙈​S 🙈​S 🙈​S 🙈​S 🙈​S 🙈​S
🙈​ TestAccept/ssh/connection 🙈​S 🙈​S 🙈​S 🙈​S 🙈​S 🙈​S
🔄​ TestSyncIncrementalFileOverwritesFolder ✅​p ✅​p ✅​p ✅​p ✅​p 🔄​f
🔄​ TestSyncIncrementalSyncPythonNotebookToFile ✅​p ✅​p 🔄​f ✅​p ✅​p ✅​p
💚​ TestFetchRepositoryInfoAPI_FromRepo 💚​R 💚​R 💚​R 💚​R 🙈​S 🙈​S
💚​ TestFetchRepositoryInfoAPI_FromRepo/root 💚​R 💚​R 💚​R 💚​R
💚​ TestFetchRepositoryInfoAPI_FromRepo/subdir 💚​R 💚​R 💚​R 💚​R
Top 37 slowest tests (at least 2 minutes):
duration env testname
9:21 aws linux TestAccept/bundle/resources/postgres_endpoints/update_autoscaling/DATABRICKS_BUNDLE_ENGINE=terraform
8:22 aws windows TestAccept/bundle/resources/postgres_projects/update_display_name/DATABRICKS_BUNDLE_ENGINE=terraform
8:10 azure windows TestAccept
7:30 aws windows TestAccept/bundle/resources/postgres_endpoints/update_autoscaling/DATABRICKS_BUNDLE_ENGINE=terraform
7:05 aws linux TestAccept/bundle/resources/postgres_endpoints/update_autoscaling/DATABRICKS_BUNDLE_ENGINE=direct
6:57 aws windows TestAccept/bundle/resources/postgres_endpoints/update_autoscaling/DATABRICKS_BUNDLE_ENGINE=direct
6:51 azure windows TestFilerRecursiveDelete/workspace_files
6:33 aws linux TestAccept/bundle/resources/postgres_projects/update_display_name/DATABRICKS_BUNDLE_ENGINE=direct
6:13 gcp windows TestAccept
5:14 azure linux TestFilerReadDir/workspace_files_extensions
4:54 aws windows TestFilerReadWrite/workspace_files
4:12 azure windows TestImportDirWithOverwriteFlag
4:05 aws linux TestAccept/bundle/resources/postgres_projects/update_display_name/DATABRICKS_BUNDLE_ENGINE=terraform
3:26 azure windows TestFilerWorkspaceFilesExtensionsDelete
3:25 gcp windows TestExportDir
3:12 azure linux TestFilerWorkspaceFilesExtensionsStat
3:09 aws windows TestFilerWorkspaceFilesExtensionsReadDir
3:07 aws windows TestFilerWorkspaceNotebook/scalaNb.scala
3:01 azure windows TestFilerReadWrite/workspace_files_extensions
2:55 azure linux TestAccept
2:54 gcp linux TestAccept
2:42 gcp windows TestImportDirDoesNotOverwrite
2:42 gcp windows TestFilerReadWrite/workspace_files_extensions
2:38 aws linux TestFilerRecursiveDelete/workspace_files
2:37 azure linux TestFilerWorkspaceFilesExtensionsReadDir
2:33 aws windows TestFilerWorkspaceFilesExtensionsDelete
2:29 gcp linux TestImportDirWithOverwriteFlag
2:27 azure windows TestWorkspaceFilesExtensions_ExportFormatIsPreserved/source_sql
2:21 aws windows TestImportDirWithOverwriteFlag
2:19 azure linux TestFilerWorkspaceFilesExtensionsRead
2:16 aws windows TestFilerRecursiveDelete/workspace_files
2:13 azure linux TestImportDirDoesNotOverwrite
2:11 gcp windows TestFilerWorkspaceFilesExtensionsReadDir
2:08 azure linux TestExportDir
2:07 aws linux TestImportDirDoesNotOverwrite
2:06 aws windows TestFilerWorkspaceFilesExtensionsRead
2:01 gcp linux TestFilerRecursiveDelete/workspace_files_extensions

@denik
denik force-pushed the denik/wait-method-removal branch from b00a6f3 to 786f578 Compare June 1, 2026 13:05
@denik
denik temporarily deployed to test-trigger-is June 1, 2026 13:06 — with GitHub Actions Inactive
@denik
denik temporarily deployed to test-trigger-is June 1, 2026 13:06 — with GitHub Actions Inactive
@denik
denik temporarily deployed to test-trigger-is June 1, 2026 15:14 — with GitHub Actions Inactive
@denik
denik temporarily deployed to test-trigger-is June 1, 2026 15:14 — with GitHub Actions Inactive
@denik
denik temporarily deployed to test-trigger-is June 1, 2026 15:14 — with GitHub Actions Inactive
@denik
denik temporarily deployed to test-trigger-is June 1, 2026 15:14 — with GitHub Actions Inactive
@denik
denik temporarily deployed to test-trigger-is June 2, 2026 13:32 — with GitHub Actions Inactive
@denik
denik temporarily deployed to test-trigger-is June 2, 2026 13:32 — with GitHub Actions Inactive
@denik
denik force-pushed the denik/wait-method-removal branch from f0e65f7 to 9f7fc23 Compare June 2, 2026 19:10
@denik
denik temporarily deployed to test-trigger-is June 2, 2026 19:11 — with GitHub Actions Inactive
@denik
denik temporarily deployed to test-trigger-is June 2, 2026 19:11 — with GitHub Actions Inactive
@denik
denik temporarily deployed to test-trigger-is June 2, 2026 20:05 — with GitHub Actions Inactive
@denik
denik temporarily deployed to test-trigger-is June 2, 2026 20:05 — with GitHub Actions Inactive
@denik
denik temporarily deployed to test-trigger-is June 2, 2026 20:21 — with GitHub Actions Inactive
@denik
denik temporarily deployed to test-trigger-is June 2, 2026 20:21 — with GitHub Actions Inactive
@denik
denik temporarily deployed to test-trigger-is June 2, 2026 20:41 — with GitHub Actions Inactive
@denik
denik temporarily deployed to test-trigger-is June 2, 2026 20:41 — with GitHub Actions Inactive
@denik
denik temporarily deployed to test-trigger-is June 3, 2026 08:52 — with GitHub Actions Inactive
@denik
denik temporarily deployed to test-trigger-is June 3, 2026 08:52 — with GitHub Actions Inactive
@denik
denik force-pushed the denik/wait-method-removal branch from bb3c2e3 to cdd95c8 Compare June 3, 2026 15:15
@denik
denik temporarily deployed to test-trigger-is June 3, 2026 15:16 — with GitHub Actions Inactive
@denik
denik temporarily deployed to test-trigger-is June 3, 2026 15:16 — with GitHub Actions Inactive
@denik
denik temporarily deployed to test-trigger-is June 3, 2026 15:35 — with GitHub Actions Inactive
@denik
denik temporarily deployed to test-trigger-is June 3, 2026 15:35 — with GitHub Actions Inactive
@denik
denik force-pushed the denik/wait-method-removal branch from df19a64 to bf602f1 Compare June 4, 2026 09:41
@denik
denik temporarily deployed to test-trigger-is June 4, 2026 09:42 — with GitHub Actions Inactive
denik added 28 commits August 5, 2026 14:26
…nternally

Resource implementations cannot recover from a state-persistence failure —
the resource already exists on the server and aborting would not undo its
creation. Log the error via logdiag instead of propagating it, and drop the
error return so call sites are a single statement.

Co-authored-by: Denis Bilenko
All six resources with a gap between resource creation and a subsequent
long-running wait now call engine.SaveState immediately after the create
API returns, using waiter.Name() (available before Wait()) for the postgres
resources and createResp.DashboardId for the dashboard.

This ensures an interrupted deployment leaves a tracked resource rather
than an orphan the next plan must rediscover from remote state.

Co-authored-by: Denis Bilenko
dashboard.DoCreate now calls engine.SaveState immediately after the
dashboard is created (with etag persisted), before publishDashboard.
A failed publish leaves the draft tracked in state rather than
orphaned; the next deploy finds it via DoRead and re-publishes via
DoUpdate without recreating the dashboard.

The old trash-on-publish-failure cleanup is removed — it was a
fragile workaround for the lack of state persistence and is now
unnecessary.

Acceptance tests:
- publish-failure-cleans-up-dashboard: updated to reflect the new
  behavior per engine (direct: draft persists with URL; terraform:
  existing behavior, cleaned up). Output files split to per-engine
  variants (out.summary.*.txt, out.dashboardrequests.*.txt).
- publish-failure-retry (new, direct only): verifies end-to-end that
  a transient publish failure leaves the draft in state (summary shows
  URL, plan detects diff), and the subsequent deploy re-publishes
  without issuing a CREATE call.

Co-authored-by: Denis Bilenko
…cements

The local testserver uses ?o= and cloud uses ?w= for the workspace/org ID
in dashboard published URLs. Add a parent-level [[Repls]] rule that maps
both to ?[WSPARAM]= so the output files are environment-independent.

Co-authored-by: Denis Bilenko
Use per-test [[Repls]] rules (matching raw digits) in the two
publish-failure tests so the URL parameter is normalized to
?[WSPARAM]=[NUMID] regardless of whether the testserver or cloud
environment is used. Revert the over-broad parent-level rule that
broke detect-change's existing ?[ow]=... cleanup.

Co-authored-by: Denis Bilenko
macOS ships bash 3.2 which does not support &>> (bash 4+).
Replace with >> file 2>&1 which works on all bash versions.

Co-authored-by: Denis Bilenko
MSYS_NO_PATHCONV=1 (set in the parent test.toml) prevents MSYS2 from
converting POSIX paths to Windows paths, causing Python to receive a
broken path (/c/a/... instead of C:\a\...) when invoking fault.py.
Unset it before the fault.py call, matching the pattern used by other
dashboard scripts before their bin helper invocations.

Co-authored-by: Denis Bilenko
…d=false

Engine.SaveState:
- Accepts resourceKey for logging: "SaveState: resources.X id=Y N bytes: {...}"
- Skips WAL write if state is unchanged (structdiff.IsEqual), logging a skip message.
- Records the last saved value in e.lastSaved for subsequent comparisons.

Dashboard DoCreate:
- Saves intermediate state with Published=false (the actual draft state) instead of
  the user's Published=true. This ensures the planner sees a real diff (false→true)
  on the next deploy if publish is interrupted, rather than treating the resource as
  up-to-date and silently skipping the publish.

publish-failure-retry acceptance test:
- Adds `bundle plan -o json` and extracts the published change to verify
  old=false, new=true in the plan after a failed publish.

README.md: update stale SetID+SaveState reference to current SaveState(ctx, id, state) API.

Co-authored-by: Denis Bilenko
…reate/DoUpdate

Mergiraf reintroduced the standalone WaitAfterCreate and WaitAfterUpdate
methods during the rebase against main (sql_warehouse.go had a merge conflict).
Inline their wait logic directly into DoCreate and DoUpdate and remove
the standalone methods, consistent with the Engine callback pattern used
by other resources.

Co-authored-by: Denis Bilenko
…ublish-failure-retry

- Replace fixed hex regex [[Repls]] with replace_ids.py called after the first
  deploy; the real dashboard ID is captured from state as [DASHBOARD1_ID].
- Inline the plan -o json | jq output directly into output.txt instead of
  saving to out.plan_published_change.json and trace cat-ing it.

Co-authored-by: Denis Bilenko
…dashboard DoUpdate

Two more cases where state could be persisted earlier to avoid orphaning or
stale-etag conflicts:

sql_warehouse DoCreate: the warehouse is created, then DoCreate polls for RUNNING
(and may Stop it). If interrupted during the wait, the warehouse was orphaned.
Now SaveState is called right after Create returns the id, before the wait.

dashboard DoUpdate: Update() bumps the server-side etag, then publishDashboard()
can fail. Mirror the DoCreate fix — save the new etag with Published=false before
publishing. This keeps the etag in sync (a stale etag makes the next Update fail
with a conflict) and records published=false so the planner re-publishes next
deploy. Resolves the pre-existing TODO.

Add publish-failure-retry-on-update acceptance test: deploy, then trigger an
update with an injected publish failure, and verify the update issued a PATCH
(no CREATE) and the plan records published old=false.

Co-authored-by: Denis Bilenko
DoUpdate applies up to four sequential mutating calls (tags, AI gateway,
config, notifications) and then waits up to 35 minutes for the endpoint to
finish updating. If that wait was interrupted, none of the applied changes
were persisted to state, forcing them to be re-applied on the next deploy.

Save the new config right after the mutating calls and before the wait,
mirroring DoCreate which already saves before waitForEndpointReady. The id
(endpoint name) matches the one DoCreate persists, so there is no id mismatch.

Co-authored-by: Denis Bilenko
…, not resource id)

The previous attempt to save state before the async wait in postgres DoCreate
used engine.SaveState(ctx, waiter.Name(), config). waiter.Name() returns the
LRO operation name (e.g. .../branches/foo/operations/UUID), not the resource
name (…/branches/foo) that DoCreate ultimately returns. Using the operation
name as the id creates a dead WAL entry and does not help with orphan recovery.

Replace with a TODO comment explaining the two options for a proper fix:
  1. Derive the resource name from request inputs (Parent + resource-type + Id).
  2. Use waiter.Metadata() if it surfaces the resource name before Wait completes.

No panic in production: the framework saves state via db.SaveState (not through
the Engine) after DoCreate returns, so the mismatch between operation-name and
real-id never hits the Engine's id-mismatch check. The test panic was a test
artifact where the same Engine was reused across Create→Update.

Co-authored-by: Denis Bilenko
Previously apply.go called db.SaveState directly after DoCreate/DoUpdate
returned, bypassing the engine. This meant a DoCreate that called
engine.SaveState with a wrong id (e.g. an LRO operation name instead of
the real resource name) would go undetected — the final save would
silently use the correct id.

Route both final saves through the engine instead:

- If DoCreate/DoUpdate called engine.SaveState with a wrong id, the
  engine's id-mismatch check panics, catching the bug immediately in tests.
- If DoCreate/DoUpdate already saved identical state (e.g. the resource
  saved before a long wait), the engine's dedup skips the redundant write.
- State-save I/O errors are logged internally by the engine and no longer
  abort the deployment (the resource was already created/updated
  successfully; aborting would confuse the user).

UpdateWithID is left unchanged: DoUpdateWithID takes no Engine parameter.

Co-authored-by: Denis Bilenko
The type does exactly one thing: save state. StateSaver names that
directly. NewNopEngine -> NewNopStateSaver, engine.go -> state_saver.go.

Co-authored-by: Denis Bilenko
…aver

- SaveStateWith[F any]: type-safe helper that temporarily sets a field to
  an intermediate value before saving, then restores it. Used to save
  Published=false before publishDashboard and Lifecycle=nil before
  warehouse/cluster lifecycle management, so the planner sees a real diff
  if deployment is interrupted mid-way.

- Change saveFunc signature to func(id string, b json.RawMessage) error
  and add DeploymentState.SaveStateJSON to accept pre-marshaled bytes.
  StateSaver already marshals x to JSON for dedup; passing those bytes
  directly avoids a redundant marshal on every save.

- Fix lastSaved aliasing: store []byte (JSON snapshot) instead of any
  pointer. SaveStateWith modifies the config, saves, then restores; a
  pointer-based lastSaved would point at the restored value, making the
  subsequent final save look like a no-op and leaving Published=false in
  the WAL.

Co-authored-by: Isaac
If DoCreate is interrupted after the app is created but before
manageLifecycle completes (during waitForApp or the deploy step), the app
can reach ACTIVE on its own with no active deployment.

Previously, engine.SaveState saved lifecycle.started=true (the desired
value). On the next plan the planner would see no localDiff (state ==
desired) and no remoteDiff for lifecycle (remote is already ACTIVE), while
source_code_path drift is silently skipped by OverrideChangeDesc when the
remote has no active deployment. Result: the planner marks the resource as
Skip, and the app stays permanently un-deployed.

SaveStateWith(lifecycle=nil) records that the app exists but lifecycle has
not been applied yet. This creates a localDiff (nil→desired) that is not
remote-skippable, forcing DoUpdate → manageLifecycle → Deploy on the next
run.

Co-authored-by: Isaac
…per-test

The global ETAG replacement in acceptance/bundle/resources/dashboards/test.toml
replaced all 8+ digit numbers in all dashboard test outputs. This prevents
local tests from using add_repl.py to record specific etag values (e.g.
the bumped etag after a PATCH), which is necessary for tests that verify
etag-tracking behavior across SaveState calls.

Move the replacement to the three tests that actually need it for cloud
compatibility (non-deterministic real etags):
- detect-change: shows etag in bundle summary output
- change-name: etag appears in Terraform PATCH request body
- change-embed-credentials: etag appears in Terraform PATCH request body

Local-only tests now see the etag replaced by the global [NUMID] rule
(acceptance/test.toml, \d{8,}) instead of [ETAG]. The out.plan.direct.json
files update accordingly: [ETAG] (unquoted, invalid JSON) → "[NUMID]"
(quoted, valid JSON string).

Co-authored-by: Isaac
genie_space was added to main after this branch was cut, using the old
DoCreate/DoUpdate signatures that lack the *StateSaver parameter added
by this PR to the IResource interface.

Co-authored-by: Isaac
…nt test

SaveStateWith(Published=false) in DoUpdate makes the stale-published-content
bug permanently unrecoverable: after a PATCH+failed-POST, state has the new
etag + Published=false. On the next plan, the planner sees remote.Published=true
== desired=true and skips (remote_already_set), so neither a plain re-deploy
nor --force can fix the stale content.

Without SaveStateWith, state retains the pre-PATCH etag. The next plan detects
the etag mismatch as "modified remotely" and blocks — but --force recovers it
by forcing a full PATCH+POST cycle.

Also cherry-picks the publish-failure-stale-content acceptance test from
denik/dashboard-published-bug, which documents this pre-existing bug and
confirms --force recovers it. Includes the testserver change that bumps the
dashboard etag on every PATCH (matching cloud behavior), and test improvements
from that branch (explicit ETAG_1/ETAG_2 labels, add_repl.py calls, etc.).

Co-authored-by: Isaac
…le DoCreate/DoUpdate

Co-authored-by: Denis Bilenko <denis.bilenko@databricks.com>
…ignature

instance_pool and job_run were added on main after this branch was cut, so
their DoCreate/DoUpdate still had the pre-StateSaver signature and failed the
IResource interface check. Add the *StateSaver parameter (unused).

Also fix the app retry tests and dashboard publish-failure tests for the
post-rebase testserver: the apps testserver now keeps DELETING apps visible
(cloud-realistic), and the parent dashboards test.toml injects eventual-
consistency staleness on direct; opt the publish-failure/retry tests out of
that injection since they drive an explicit read-back.

Co-authored-by: Isaac
The Engine->StateSaver rename missed the async-APIs section of the dresources
README, which still called the second DoCreate/DoUpdate argument a *Engine.

Co-authored-by: Isaac
Revert the redaction added during the rebase conflict resolution: it depended
on structwalk.RedactSensitiveFields, which is being removed. State is marshaled
with plain json.Marshal in both SaveState and StateSaver.SaveState, as before.

Co-authored-by: Isaac
The scripts already capture each etag by value into ACC_REPLS (ETAG_1, ETAG_2,
ETAG), so the catch-all "\"[-0-9]{8,}\"" / "\"[0-9]{8,}\"" [[Repls]] were
redundant and risked masking unrelated quoted long integers. Remove them;
change-name and change-embed-credentials test.toml held nothing else, so drop
those files entirely.

Co-authored-by: Isaac
Now that DoRead derives published from the publish/update timestamps (#6119), the
reason DoUpdate avoided saving state before the publish is gone: a stale publish
reports remote published=false, so the remote_already_set skip that would have
stranded it can no longer happen.

Save the post-update etag and published=false before publishing. A failed publish
is now recoverable with a plain deploy: the next plan sees desired published=true
against a remote reported as false and republishes. Saving the post-update etag
also keeps state in sync with remote, so CheckDashboardsModifiedRemotely no longer
misreports this as an out-of-band edit and --force is no longer required.

Drops the Badness marker from publish-failure-stale-content and updates it to
assert recovery on a plain re-deploy.

Co-authored-by: Isaac
Dropping RedactSensitiveFields from the state save path also fixes the values
persisted for protobuf Duration fields. RedactSensitiveFields deep-clones the
struct field-by-field via reflection, which loses durationpb.Duration's internal
state, so its custom marshaler emitted the zero value: state recorded
history_retention_duration and suspend_timeout_duration as "0s" instead of the
real "604800s"/"300s". Plain json.Marshal preserves them.

Only the saved "old" values in these plan goldens change; there is no behavior
change beyond state now matching what the server returned.

Co-authored-by: Isaac
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants