Skip to content

fix: default Azure dual-pool autoscaling - #659

Open
adamrtalbot wants to merge 1 commit into
masterfrom
agent/fix-azure-dual-pool-autoscale
Open

fix: default Azure dual-pool autoscaling#659
adamrtalbot wants to merge 1 commit into
masterfrom
agent/fix-azure-dual-pool-autoscale

Conversation

@adamrtalbot

@adamrtalbot adamrtalbot commented Aug 7, 2026

Copy link
Copy Markdown
Contributor

Why

Azure Batch Forge dual-pool compute environments omitted headPool.autoScale and workerPool.autoScale when the corresponding disable flags were not supplied. This allowed the resulting pools to behave as fixed-size pools even though the CLI presents autoscaling as the default.

Closes #658.

What

  • Explicitly set autoscaling to true for head and worker pools by default.
  • Preserve --head-no-auto-scale and --worker-no-auto-scale behavior by setting the selected pool to false.
  • Add request-payload coverage for default and explicitly disabled autoscaling.
  • Update the existing dual-pool boot-disk payload expectation.

Impact

Dual-pool VM counts are now serialized as autoscaling maximums by default, matching the CLI help and single-pool behavior. Existing users who pass the disable flags retain fixed-size pools.

Explicitly enable autoscaling for Azure Batch Forge head and worker
pools when their disable flags are omitted.

Closes #658
@adamrtalbot
adamrtalbot marked this pull request as ready for review August 7, 2026 11:06
zhliUU added a commit to Aletechdev/ALE_Yeast that referenced this pull request Aug 31, 2026
`tw compute-envs add azure-batch forge --dual-pool` omits headPool.autoScale /
workerPool.autoScale from the request, so Azure builds enableAutoScale: False —
fixed-size pools billing 24/7 (~$66/day, measured 2026-08-07). Dual-pool CEs were
therefore being created by hand in the web UI.

They no longer are. `tw compute-envs export`/`import` round-trip autoScale, so
13_create_compute_env.sh imports a readback of a known-good CE and then runs the
verifier. No REST client was needed: the earlier "the CLI cannot express it"
conclusion was drawn from `add ... forge`'s flag set and over-generalised to tw.

Verified end to end with a throwaway CE: 12_verify_compute_env.sh 6/6, Azure
enableAutoScale True on both pools, drained to 0 + 0 after 17 minutes, deleted
with pools and disks disposed. Cost ~5 node-minutes.

Two corrections recorded alongside the original incident entry, which is kept
intact so the wrong conclusion is not re-derived:

  - `add ... forge` CAN set autoscale via explicit --head-no-auto-scale=false
    --worker-no-auto-scale=false (undocumented; seqeralabs/tower-cli#658).
    Kept as an escape hatch only — that route has no flags for jobMaxWallClockTime,
    deleteJobsOnCompletion, deleteTasksOnCompletion or terminateJobsOnCompletion,
    so it silently takes Platform defaults for the four settings our CEs pin.
  - autoScale is omitted from the payload, not sent as null; the null is
    Platform's readback of the missing field.

Upstream fix seqeralabs/tower-cli#659 is unmerged and 0.38.0 is the latest
release, so there is no version to upgrade to.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Azure dual-pool autoscale defaults are omitted

1 participant