Diagnose fast publish-core failure after bounded clone fix #583

Closed
opened 2026-07-15 10:35:52 +00:00 by coilyco-ops · 9 comments
Owner

The targeted publish-core repair from PR #582 landed and fixed main CI, but the next dev-base publication still failed before producing agentic-os-full:latest.

Evidence:

  • PR #582 passed required PR CI and merged.
  • Resulting main/release SHA: 81cdc64dc05283f4b2d55e65b33093c0edf3df28.
  • Main ci.yml run 8459 / #1420 succeeded.
  • Promote run 8461 / #1422 succeeded.
  • dev-base-publish.yml run 8463 / #1423 failed on the same exact SHA after about 2m16s.
  • agentic-os-full:latest remains absent (manifest unknown).
  • The previous run 8432 failed after 16m23s; the much faster new failure suggests the bounded clone changed the failure point but did not complete publish-core.

Diagnose the exact fast publish-core failure after PR #582, repair it if repository-owned, and preserve the bounded clone plus retry/resume behavior. If an external registry/storage condition is the wall, make the workflow emit the precise actionable condition without broad diagnostic churn.

Acceptance:

  • The failing step and reason in run 8463 are identified.
  • A focused repair lands with green PR CI if the defect is repository-owned.
  • A new dev-base-publish.yml run passes publish-core, or the exact external blocker and required operator action are recorded.
  • agentic-os-full:latest becomes available, or the issue remains open with the precise external blocker.
The targeted publish-core repair from PR #582 landed and fixed main CI, but the next dev-base publication still failed before producing `agentic-os-full:latest`. Evidence: * PR #582 passed required PR CI and merged. * Resulting main/release SHA: `81cdc64dc05283f4b2d55e65b33093c0edf3df28`. * Main `ci.yml` run 8459 / #1420 succeeded. * Promote run 8461 / #1422 succeeded. * `dev-base-publish.yml` run 8463 / #1423 failed on the same exact SHA after about 2m16s. * `agentic-os-full:latest` remains absent (`manifest unknown`). * The previous run 8432 failed after 16m23s; the much faster new failure suggests the bounded clone changed the failure point but did not complete `publish-core`. Diagnose the exact fast `publish-core` failure after PR #582, repair it if repository-owned, and preserve the bounded clone plus retry/resume behavior. If an external registry/storage condition is the wall, make the workflow emit the precise actionable condition without broad diagnostic churn. Acceptance: * The failing step and reason in run 8463 are identified. * A focused repair lands with green PR CI if the defect is repository-owned. * A new `dev-base-publish.yml` run passes `publish-core`, or the exact external blocker and required operator action are recorded. * `agentic-os-full:latest` becomes available, or the issue remains open with the precise external blocker.
Author
Owner

WARDED_WORKFLOW: reservation-held

reservation details

Holder: launch intent for container engineer-codex-agentic-os-583 on host 74a70515a491.

Accepted by ward agent --harness codex (reserved 2026-07-15T10:36:13Z). Concurrent ward agent runs are blocked until this intent becomes visible or the intent is released. The stale-intent fallback is still TTL-bounded (3h TTL). --override-reservation overrides.

Do not comment on or edit this issue to steer the run while it is reserved. The engineer seeded the body once at launch and never re-reads it, so a comment or edit reaches only human readers, never the running engineer. A correction goes to a new issue, dispatched fresh. That is the only channel that reaches a run in flight. Where the forge supports it, ward locks this conversation to make that a road-block rather than a convention (ward#494).

run seed context — what this run is carrying (ward#609)
  • Resolved: coilyco-flight-deck/agentic-os#583 · branch issue-583 · harness codex · workflow pull-request-and-merge
  • Run: engineer-codex-agentic-os-583 · ward v0.710.0 · dispatched 2026-07-15T10:36:10Z
  • Comment thread: 0 included in the pre-flight read, 0 stripped (ward's own automated comments).

Static container doctrine and seed boilerplate are identical every run and omitted here (they ride ward v0.710.0).

— Codex, via ward agent

<!-- ward-agent-reservation --> WARDED_WORKFLOW: reservation-held <details><summary>reservation details</summary> Holder: launch intent for container `engineer-codex-agentic-os-583` on host `74a70515a491`. Accepted by `ward agent --harness codex` (reserved 2026-07-15T10:36:13Z). Concurrent `ward agent` runs are blocked until this intent becomes visible or the intent is released. The stale-intent fallback is still TTL-bounded (3h TTL). `--override-reservation` overrides. **Do not comment on or edit this issue to steer the run while it is reserved.** The engineer seeded the body once at launch and never re-reads it, so a comment or edit reaches only human readers, never the running engineer. A correction goes to a **new issue, dispatched fresh**. That is the only channel that reaches a run in flight. Where the forge supports it, ward locks this conversation to make that a road-block rather than a convention (ward#494). <details><summary>run seed context — what this run is carrying (ward#609)</summary> - **Resolved:** `coilyco-flight-deck/agentic-os#583` · branch `issue-583` · harness `codex` · workflow `pull-request-and-merge` - **Run:** `engineer-codex-agentic-os-583` · ward `v0.710.0` · dispatched `2026-07-15T10:36:10Z` - **Comment thread:** 0 included in the pre-flight read, 0 stripped (ward's own automated comments). Static container doctrine and seed boilerplate are identical every run and omitted here (they ride ward v0.710.0). </details> </details> <!-- ward-agent-signature --> — Codex, via `ward agent`
Author
Owner

WARDED_WORKFLOW: dispatch-deferred

launch failure details

The repair run never started. warded engineer coilyco-flight-deck/agentic-os#583 --harness codex --workflow pull-request-and-merge failed before container creation because forgejo.coilysiren.me/coilyco-flight-deck/agentic-os-full:latest was unavailable remotely and no local image existed. Docker has no engineer-codex-agentic-os-583 container. No implementation work was performed.

The reservation is released for redispatch from a fresh director. Issue #583 remains the concrete repair for failed publish run 8463 on SHA 81cdc64dc05283f4b2d55e65b33093c0edf3df28.

  • Codex, via ward agent
WARDED_WORKFLOW: dispatch-deferred <details><summary>launch failure details</summary> The repair run never started. `warded engineer coilyco-flight-deck/agentic-os#583 --harness codex --workflow pull-request-and-merge` failed before container creation because `forgejo.coilysiren.me/coilyco-flight-deck/agentic-os-full:latest` was unavailable remotely and no local image existed. Docker has no `engineer-codex-agentic-os-583` container. No implementation work was performed. The reservation is released for redispatch from a fresh director. Issue #583 remains the concrete repair for failed publish run 8463 on SHA `81cdc64dc05283f4b2d55e65b33093c0edf3df28`. </details> <!-- ward-agent-signature --> - Codex, via `ward agent`
Author
Owner

WARDED_WORKFLOW: reservation-held

reservation details

Holder: launch intent for container engineer-codex-agentic-os-583 on host 74a70515a491.

Accepted by ward agent --harness codex (reserved 2026-07-15T11:36:15Z). Concurrent ward agent runs are blocked until this intent becomes visible or the intent is released. The stale-intent fallback is still TTL-bounded (3h TTL). --override-reservation overrides.

Do not comment on or edit this issue to steer the run while it is reserved. The engineer seeded the body once at launch and never re-reads it, so a comment or edit reaches only human readers, never the running engineer. A correction goes to a new issue, dispatched fresh. That is the only channel that reaches a run in flight. Where the forge supports it, ward locks this conversation to make that a road-block rather than a convention (ward#494).

run seed context — what this run is carrying (ward#609)
  • Resolved: coilyco-flight-deck/agentic-os#583 · branch issue-583 · harness codex · workflow pull-request-and-merge
  • Run: engineer-codex-agentic-os-583 · ward v0.710.0 · dispatched 2026-07-15T11:36:11Z
  • Comment thread: 1 included in the pre-flight read, 1 stripped (ward's own automated comments).

Static container doctrine and seed boilerplate are identical every run and omitted here (they ride ward v0.710.0).

— Codex, via ward agent

<!-- ward-agent-reservation --> WARDED_WORKFLOW: reservation-held <details><summary>reservation details</summary> Holder: launch intent for container `engineer-codex-agentic-os-583` on host `74a70515a491`. Accepted by `ward agent --harness codex` (reserved 2026-07-15T11:36:15Z). Concurrent `ward agent` runs are blocked until this intent becomes visible or the intent is released. The stale-intent fallback is still TTL-bounded (3h TTL). `--override-reservation` overrides. **Do not comment on or edit this issue to steer the run while it is reserved.** The engineer seeded the body once at launch and never re-reads it, so a comment or edit reaches only human readers, never the running engineer. A correction goes to a **new issue, dispatched fresh**. That is the only channel that reaches a run in flight. Where the forge supports it, ward locks this conversation to make that a road-block rather than a convention (ward#494). <details><summary>run seed context — what this run is carrying (ward#609)</summary> - **Resolved:** `coilyco-flight-deck/agentic-os#583` · branch `issue-583` · harness `codex` · workflow `pull-request-and-merge` - **Run:** `engineer-codex-agentic-os-583` · ward `v0.710.0` · dispatched `2026-07-15T11:36:11Z` - **Comment thread:** 1 included in the pre-flight read, 1 stripped (ward's own automated comments). - included: @coilyco-ops (2026-07-15T10:36:47Z) - stripped: @coilyco-ops (2026-07-15T10:36:13Z) Static container doctrine and seed boilerplate are identical every run and omitted here (they ride ward v0.710.0). </details> </details> <!-- ward-agent-signature --> — Codex, via `ward agent`
Author
Owner

WARDED_WORKFLOW: dispatch-deferred

launch failure details

A fresh repair run was attempted after dev-base-publish run 8500 failed and the image remained unavailable. The stale 10:36 reservation was safely reclaimed because dispatch-health showed zero running/held Agentic engineers and no container. The new launch then failed before start while Ward v0.710 seeded the release image: docker cp to /opt/agentic-os/ward-shell-entrypoint.sh returned not a directory. The created container engineer-codex-agentic-os-583 was stopped; no engineer code ran and no branch or PR was created.

This surface cannot use its brokered stop endpoint, so the reservation is released by this explicit failed-prelaunch record for redispatch from a compatible director. Issue #583 remains the direct repair for failed dev-base-publish run 8500 / #1429 and the absent agentic-os-full:latest manifest.

  • Codex, via ward agent
WARDED_WORKFLOW: dispatch-deferred <details><summary>launch failure details</summary> A fresh repair run was attempted after dev-base-publish run 8500 failed and the image remained unavailable. The stale 10:36 reservation was safely reclaimed because dispatch-health showed zero running/held Agentic engineers and no container. The new launch then failed before start while Ward v0.710 seeded the release image: docker cp to /opt/agentic-os/ward-shell-entrypoint.sh returned not a directory. The created container engineer-codex-agentic-os-583 was stopped; no engineer code ran and no branch or PR was created. This surface cannot use its brokered stop endpoint, so the reservation is released by this explicit failed-prelaunch record for redispatch from a compatible director. Issue #583 remains the direct repair for failed dev-base-publish run 8500 / #1429 and the absent agentic-os-full:latest manifest. </details> <!-- ward-agent-signature --> - Codex, via ward agent
Author
Owner

RECOVERY EVIDENCE (2026-07-15): the latest failures isolate runner/Docker context cancellation rather than a deterministic Dockerfile or cache miss.

  • dev-base run 8606 / UI #1453 at release SHA a98d8b6641c0af60f590ba1c9dd774462d7a935a: publish-core failed after 16m9s; named Publish core image step ended while BuildKit was still installing Ubuntu packages for amd64/arm64. The core-buildcache manifest imported successfully. Final runner output was context canceled while copying SUMMARY.md and pathcmd.txt; every dependent tier was blocked.
  • promote run 8615 / UI #1457 at main SHA 17ab6a92f30875c021d9a5ad85b09966646539e4: gate succeeded, then promote-release failed during Set up job in 5s. The runner created and ran agentic-os:release, then emitted context canceled before checkout.
  • forgejo.coilysiren.me/coilyco-flight-deck/agentic-os-full:latest remains absent.

Acceptance for the next carry: retrigger the supported current-main promote/release path; if cancellation repeats, isolate and repair the runner/concurrency/resource condition rather than changing Docker build semantics or cache keys. Leave a green dev-base publication with the full manifest present, or the exact external operator action.

RECOVERY EVIDENCE (2026-07-15): the latest failures isolate runner/Docker context cancellation rather than a deterministic Dockerfile or cache miss. - dev-base run 8606 / UI #1453 at release SHA `a98d8b6641c0af60f590ba1c9dd774462d7a935a`: `publish-core` failed after 16m9s; named `Publish core image` step ended while BuildKit was still installing Ubuntu packages for amd64/arm64. The `core-buildcache` manifest imported successfully. Final runner output was `context canceled` while copying `SUMMARY.md` and `pathcmd.txt`; every dependent tier was blocked. - promote run 8615 / UI #1457 at main SHA `17ab6a92f30875c021d9a5ad85b09966646539e4`: gate succeeded, then `promote-release` failed during Set up job in 5s. The runner created and ran `agentic-os:release`, then emitted `context canceled` before checkout. - `forgejo.coilysiren.me/coilyco-flight-deck/agentic-os-full:latest` remains absent. Acceptance for the next carry: retrigger the supported current-main promote/release path; if cancellation repeats, isolate and repair the runner/concurrency/resource condition rather than changing Docker build semantics or cache keys. Leave a green dev-base publication with the full manifest present, or the exact external operator action.
Author
Owner

DISPATCH-DEFERRED: attempted ward agent engineer coilyco-flight-deck/agentic-os#583 --harness codex --workflow pull-request-and-merge after recording runs 8606/8615. This director's host dispatch broker at host.docker.internal:59660 is unreachable, so Ward rejected the launch before reservation or container creation. A healthy director broker should dispatch #583; do not use direct Docker seeding from this surface.

DISPATCH-DEFERRED: attempted `ward agent engineer coilyco-flight-deck/agentic-os#583 --harness codex --workflow pull-request-and-merge` after recording runs 8606/8615. This director's host dispatch broker at `host.docker.internal:59660` is unreachable, so Ward rejected the launch before reservation or container creation. A healthy director broker should dispatch #583; do not use direct Docker seeding from this surface.
Author
Owner

Evidence from the v0.775.0-tmp pin landing in e98b4e6:

  • The agent ran the supported local dev-base-build --tier core path twice.
  • Both attempts failed before the Ward clone at docker/dev-base/core/Dockerfile:20.
  • The builder stage's single apt-get update -o Acquire::Retries=3 timed out against ports.ubuntu.com:80, then package installation had no index.
  • The parallel core stage reached the same mirror successfully and already has an outer three-attempt loop. The Ward builder stage does not.

This may explain the fast publish-core failure class tracked here. The focused pin commit does not change mirror retry behavior, and the agent made no speculative live-CI iteration.

Evidence from the `v0.775.0-tmp` pin landing in `e98b4e6`: * The agent ran the supported local `dev-base-build --tier core` path twice. * Both attempts failed before the Ward clone at `docker/dev-base/core/Dockerfile:20`. * The builder stage's single `apt-get update -o Acquire::Retries=3` timed out against `ports.ubuntu.com:80`, then package installation had no index. * The parallel core stage reached the same mirror successfully and already has an outer three-attempt loop. The Ward builder stage does not. This may explain the fast `publish-core` failure class tracked here. The focused pin commit does not change mirror retry behavior, and the agent made no speculative live-CI iteration.
Author
Owner

Live log evidence from dev-base-publish run #1496 corrects the earlier local-build inference. The plan-draft task succeeded. The publish-core task failed during runner setup before checkout. Its setup log emitted 'runs-on' key not defined in dev-base-publish/plan-draft, then started data.forgejo.org/oci/node:lts, attempted the Docker pull, and terminated with context canceled. Every build step was canceled, so run #1496 reached neither the Ward clone nor Ubuntu apt. The release revision does contain runs-on: docker on plan-draft, which makes that line a misleading runner warning rather than a missing workflow key. The live blocker needs runner/cache/network inspection for the default node image pull. No speculative workflow push was made.

Live log evidence from dev-base-publish run #1496 corrects the earlier local-build inference. The plan-draft task succeeded. The publish-core task failed during runner setup before checkout. Its setup log emitted `'runs-on' key not defined in dev-base-publish/plan-draft`, then started `data.forgejo.org/oci/node:lts`, attempted the Docker pull, and terminated with `context canceled`. Every build step was canceled, so run #1496 reached neither the Ward clone nor Ubuntu apt. The release revision does contain `runs-on: docker` on plan-draft, which makes that line a misleading runner warning rather than a missing workflow key. The live blocker needs runner/cache/network inspection for the default node image pull. No speculative workflow push was made.
Author
Owner

Resolved with repository and runner fixes, followed by a full live publication.

  • agentic-os commit 31d462a moves WARD_CONFIG_REF_COMMIT stamping after the expensive cacheable tool layers and adds bounded retry to the hidden Ward builder apt install. Focused tests and pre-commit passed. The native changed-ref rebuild took 6.36s, and a fresh builder importing the registry cache took 38.89s with expensive layers cached.
  • dev-base run 1506 proved the cache repair in CI: core passed in 3m08s. Kubernetes then identified the remaining context canceled failures as runner evictions because docker-lib exceeded its 12 GiB EmptyDir limit.
  • infrastructure commit d699de7 raises each runner's private Docker scratch to 24 GiB with a 12 GiB request. Commit df1025a restores the missing manifest document boundaries that had blocked Flux reconciliation. Flux applied both commits and rolled all runners. The live DinD limit is 24 GiB, and no new eviction occurred after rollout.
  • dev-base retry run 1507 passed every tier. Go and ops, the tiers evicted in the prior run, passed in 3m20s and 6m32s. Agent passed in 5m33s and full passed in 5m56s.
  • Release retry run 1509 passed every retag and the final release job after run 1508 encountered one transient setup download reset.
  • The published full image is the unprefixed tag in the consolidated package, not the retired agentic-os-full package. agentic-os:latest now resolves to the same multi-architecture digest as agentic-os:draft-31d462ad3aa43e5f1b134c557670dbb28a684577, with amd64 and arm64 manifests present.

The durable cross-pod cache is the registry buildcache. Each DinD daemon keeps private, disposable scratch. A shared /var/lib/docker PVC is neither required nor safe across runner daemons.

Resolved with repository and runner fixes, followed by a full live publication. * agentic-os commit `31d462a` moves `WARD_CONFIG_REF_COMMIT` stamping after the expensive cacheable tool layers and adds bounded retry to the hidden Ward builder apt install. Focused tests and pre-commit passed. The native changed-ref rebuild took 6.36s, and a fresh builder importing the registry cache took 38.89s with expensive layers cached. * dev-base run 1506 proved the cache repair in CI: core passed in 3m08s. Kubernetes then identified the remaining `context canceled` failures as runner evictions because `docker-lib` exceeded its 12 GiB EmptyDir limit. * infrastructure commit `d699de7` raises each runner's private Docker scratch to 24 GiB with a 12 GiB request. Commit `df1025a` restores the missing manifest document boundaries that had blocked Flux reconciliation. Flux applied both commits and rolled all runners. The live DinD limit is 24 GiB, and no new eviction occurred after rollout. * dev-base retry run 1507 passed every tier. Go and ops, the tiers evicted in the prior run, passed in 3m20s and 6m32s. Agent passed in 5m33s and full passed in 5m56s. * Release retry run 1509 passed every retag and the final release job after run 1508 encountered one transient setup download reset. * The published full image is the unprefixed tag in the consolidated package, not the retired `agentic-os-full` package. `agentic-os:latest` now resolves to the same multi-architecture digest as `agentic-os:draft-31d462ad3aa43e5f1b134c557670dbb28a684577`, with amd64 and arm64 manifests present. The durable cross-pod cache is the registry buildcache. Each DinD daemon keeps private, disposable scratch. A shared `/var/lib/docker` PVC is neither required nor safe across runner daemons.
Sign in to join this conversation.
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set

Reference
coilyco-flight-deck/agentic-os#583
No description provided.