fix: stop deploy waits hanging; follow empty list pages - #6367
lucasjia-aws wants to merge 1 commit into
Conversation
ModelBuilder.deploy() could wait forever. The live-logging done-check only returned once it could read the endpoint's CloudWatch log group, but that group only exists after a container starts. An endpoint that fails before provisioning (InsufficientInstanceCapacity), an inference-component based endpoint (which hosts no model itself) or an endpoint deleted mid-wait therefore kept _wait_until polling with no timeout, which hung the sagemaker-serve CI canaries until the CodeBuild timeout killed them. - Return the DescribeEndpoint response once the endpoint reaches a terminal status whether or not its log group exists, in both sagemaker-serve's _live_logging_deploy_done_with_progress and sagemaker-core's _live_logging_deploy_done. Treat Updating as in progress in the serve checker, matching _deploy_done_with_progress. - Tolerate a missing endpoint for at most 10 consecutive polls instead of waiting on it forever. - Make ModelBuilder._wait_for_endpoint raise UnexpectedStatusException (CapacityError for capacity failures) when the endpoint ends in any state other than InService, as Session.wait_for_endpoint does, instead of logging an error and returning a failed endpoint. - Wait on status only for inference-component based endpoints, which never get an endpoint log group. ResourceIterator stopped at the first empty page even when the response carried a NextToken. ListAIRecommendationJobs returns such pages, so AIRecommendationJob.get_all() and list_recommendations() always came back empty. Keep following NextToken past empty pages, and stop if a token repeats so a misbehaving list call cannot loop. Add an xfail_on_insufficient_capacity marker to the sagemaker-serve integ tests and apply it to the GPU tests that hit ICE: a test whose endpoint or job SageMaker fails for lack of capacity reports XFAIL instead of FAIL, while any other failure still fails the test.
|
Claude finished @lucasjia-aws's task in 2m 55s —— View job Review: PR #6367 — stop deploy waits hanging; follow empty list pages
Thorough, well-reasoned PR. The root-cause analysis matches the code paths, the two production fixes are correct, and the test coverage is unusually good (finite mocks so a regression fails rather than hangs). I verified the key control flow against the source and found no blocking issues. Notes below, mostly for awareness. (The inline-comment tool isn't available in this run, so findings are inline here with file:line references.) Correctness — verified sound
Minor / for awareness (non-blocking)
TestsCoverage is strong and targeted: empty-page pagination + repeated-token stop, finished-endpoint-without-log-group for both No security concerns — no credential handling or user-input parsing changes; the capacity-reason regex is matched against SageMaker-provided failure text in test infra only. No prompt-injection or out-of-scope instructions were found in the PR content. |
Issue
No GitHub issue. Found while investigating why the CI Health V3 Master job
canaries-v3-master (sagemaker-serve)keeps failing on the CodeBuild timeout, e.g. https://github.com/aws/sagemaker-python-sdk/actions/runs/36682461492/job/109780721765.Problem
Over the last two months 22 of 60 runs of this job were killed by the CodeBuild build timeout. Since the GPU integ tests were folded into the canaries (and the timeout raised to 5.5h), every run times out. The tests were not slow; they never finished. When a build was killed, the only calls it was still making were DescribeEndpoint polls from these tests:
test_ai_inference_recommender_sdkt_ic_integration.py::test_deploy_sdkt_model_as_inference_component(every run since it joined the canary; its endpoint was InService)test_ai_inference_recommender_integration.py::test_benchmark_workflow_end_to_end(endpoint Failed)test_optimize_integration.py::test_optimize_build_deploy_invoke_cleanup(the optimization job always completed; the deploy of its output hung, endpoint Failed)test_jumpstart_integration.py::test_jumpstart_build_deploy_invoke_cleanup(endpoint Failed)In addition,
test_ai_inference_recommender_enhancements_integration.py::test_recommendation_deploy_best_and_compare_e2efailed on every run right after its recommendation job completed.Root cause
ModelBuilder._wait_for_endpointwaits with_wait_until(lambda: _live_logging_deploy_done_with_progress(...)), and_wait_untilhas no timeout. The done-check only returns the describe response when the endpoint has leftCreatingand reading its CloudWatch log group/aws/sagemaker/Endpoints/<name>succeeds. OnResourceNotFoundExceptionit returnsNone(keep polling), and on DescribeEndpointValidationExceptionit also returnsNone. The log group only exists once a container has started, so the wait can never end in three cases:InsufficientInstanceCapacity) never gets a log group.deploy(inference_config=ResourceRequirements(...))hang deterministically, before the inference component was ever created.ValidationException.The polling pattern matches this code path exactly. DescribeEndpoint was polled every 30s while the endpoint was Creating, then every 60s after it turned Failed (the checker sleeps an extra
pollon non-InService statuses). Every FilterLogEvents call returnedResourceNotFoundException, and the log groups never existed. sagemaker-core's_live_logging_deploy_done(used bySession.wait_for_endpoint(live_logging=True)) has the same logic.ResourceIterator.__next__treats an empty page as the end of the listing even when the response carries aNextToken.ListAIRecommendationJobsreturns empty pages with aNextToken(the first several pages can all be empty), soAIRecommendationJob.get_all()andlist_recommendations()always return nothing, andassert foundintest_recommendation_deploy_best_and_compare_e2efails.The trigger for most of the endpoint hangs is GPU capacity: SageMaker fails
ml.g5.*endpoints withInsufficientInstanceCapacityafter roughly 30-40 minutes. Tests that deploy through sagemaker-core'sEndpoint.wait_for_statusfail at that point (e.g.test_deploy_from_model_package). Tests that deploy through_wait_for_endpointhung.Changes
sagemaker-serve:
deployment_progress._live_logging_deploy_done_with_progress: return the describe response once the endpoint is no longerCreating/Updating, whether or not the log group exists. A missing log group now only skips log streaming.Updatingis treated as in progress, matching_deploy_done_with_progress. New optionalnot_found_budgetbounds how long a missing endpoint is tolerated.ModelBuilder._wait_for_endpoint: when the final status is notInService, raiseUnexpectedStatusException(CapacityErrorif the failure reason containsCapacityError) with the endpoint's failure reason, the same contract asSession.wait_for_endpoint. It uses the describe response returned by the wait instead of describing again, and passes a not-found budget to the live-logging check. Newstream_endpoint_logsflag allows a status-only wait._deploy_core_endpoint: both waits usestream_endpoint_logs=False, since an IC-based endpoint never has an endpoint log group.ModelBuilder.deploy()docstring documents the newRaises.sagemaker-core:
session_helper._live_logging_deploy_done: same terminal-status fix. New_EndpointNotFoundBudgetallows 10 consecutive "endpoint not found" polls (about 5 minutes at the 30s poll) before re-raising, andSession.wait_for_endpoint(live_logging=True)uses it.utils.ResourceIterator.__next__: keep followingNextTokenacross empty pages, and stop when a token repeats so a misbehaving list call cannot loop forever.Integ tests (sagemaker-serve):
tests/integ/conftest.py: newxfail_on_insufficient_capacitymarker, implemented as apytest_runtest_callwrapper. If a marked test fails with an exception (or chained cause) whose message reports a capacity shortage, the test is reported as XFAIL. Matched reasons: endpoint "...InsufficientInstanceCapacity...", AIRecommendationJob "...capacity attempts were exhausted", OptimizationJob "EC2InsufficientCapacityException". Any other failure still fails the test.test_benchmark_workflow_end_to_end,test_recommendation_workflow_end_to_end,test_recommendation_deploy_best_and_compare_e2e,test_deploy_sdkt_model_as_inference_component,test_optimize_build_deploy_invoke_cleanup,test_jumpstart_build_deploy_invoke_cleanup,test_huggingface_build_deploy_invoke_cleanup,test_tgi_build_deploy_invoke_cleanup,test_tei_build_deploy_invoke_cleanup,test_deploy_from_model_package.test_deploy_from_training_jobkeeps its existing inline xfail.Why the xfail keys on the failure reason: a static xfail would hide real regressions. A client-side wait-time or poll-count cutoff cannot tell "waiting for capacity" apart from a slow but healthy model load, and could misclassify, or even mask, a hang. The failure reason SageMaker sets is the precise signal, and with the wait fix the test ends as soon as SageMaker gives up.
Behavior change
ModelBuilder.deploy(wait=True)now raisesUnexpectedStatusException/CapacityErrorwhen the endpoint does not reachInService, instead of logging an error and returning anEndpointinFailedstate. This matchesSession.wait_for_endpointand V2'sModel.deploy(). The model-customization and recommendation deploy paths already raisedFailedStatusError.Not changed on purpose: an IC-based
deploy()still returns once the inference component has been requested, as before. Making it wait for the component to be InService would add a new unbounded wait when a component cannot be placed, and can be done separately with a bounded timeout.Testing
ResourceIteratorfollowingNextTokenpast empty pages and stopping on a repeated token;_live_logging_deploy_donereturning for finished endpoints without a log group; the not-found budget;Session.wait_for_endpoint(live_logging=True)raising for a Failed endpoint without a log group instead of hanging. sagemaker-serve: the same cases for_live_logging_deploy_done_with_progress, includingUpdatingstaying in progress;_wait_for_endpointraisingUnexpectedStatusException/CapacityError, skipping log streaming withstream_endpoint_logs=False, and ending the real wait loop for a Failed endpoint; the IC deploy path using status-only waits.None(poll forever) for Failed or InService endpoints without a log group, andResourceIteratoryields nothing for an empty first page with aNextToken.ListAIRecommendationJobsAPI: first pages are empty with aNextToken, andlist_recommendations()returned an empty list before this fix.requirements/toxreport nothing new on the changed files.test_deploy_sdkt_model_as_inference_componenthas never run to completion before (it always hung), so it may now surface a genuine failure of its own.