Watch
2
CI is blocked repo-wide: ward-doctor 404s because the install ref falls back to a release-tag path that cannot exist #1344
Closed
opened 2026-08-28 18:00:56 +00:00 by coilyco-ops
·
8 comments
No Branch/Tag specified
main
release
chore/full-name-attribution
aos/claude/mt75-bundle-tag
aos/claude/ad84-revert-voice
aos/claude/ad84-voice
aos/claude/tc69-abspath
aos/claude/mt75-quiet-plan
voice-no-rarity-statements
aos/claude/mt75-netlify-remove
aos/claude/ad84-titles
aos/claude/ad84
aos/claude/mt75-aterm-fullscreen
aos/claude/mt75-netlify-wrap
aos/claude/mt75-kubectl-context
aos/claude/mt75
aos/claude/mt75-ward-cut
aos/claude/eb77
aos/claude/fp87-ward-schema
aos/claude/fp87-ward-posture
aos/claude/fp87
aos/claude/vt77
aos/1105-require-issue-labels
aos/claude/kb87-native-arch
aos/claude/kb87-window-identity
aos/claude/kb87
aos/claude/tg69
aos/claude/ff54-retire-issue-refs
aos/claude/ff54
aos/claude/rc44-bundle-version
aos/claude/mu55-contract-tests
aos/claude/mu55-sound-mark
aos/claude/rc44
aos/claude/mu55-identity-card
aos/claude/mu55-doctor
aos/claude/mu55-shadow-reap
aos/claude/mu55-drop-windows
aos/claude/mu55-dryrun-exits
aos/claude/mu55-list-json
aos/claude/mu55-title-order
aos/claude/mu55-overlay-contract
aos/claude/qu74
aos/claude/ue86
aos/claude/yq86
aos/claude/vk48-harness-set
aos/claude/vk48-default-agent
aos/claude/vk48-completion
aos/claude/vk48-release-fix
aos/claude/vk48
aos/claude/yb89
aos/claude/sb46
fix/agent-compose-pin-v3-roster
aos/claude/ve67-defer
aos/claude/ve67
aos/claude/ur54
aos/claude/tj49
feat/acompose-v3-roster
ops/393-retire-doc-size-alias
ops/393-drop-em-dash-check
feat/vendored-tree-exclude
aos/claude/xlarge-band
aos/claude/ue65
aos/claude/identity-color-wins
aos/claude/ap47
aos/claude/zr44
aos/claude/xk58
aos/claude/aw85-skill-size-owner
aos/claude/ym96-docs-bands
aos/claude/wt57-pin-aos-bundle
aos/claude/wt57-image-inputs-filter
aos/claude/ym96-label-taxonomy
ops/dev-base-pin-rust-1.90.0
aos/claude/mg96-clean
aos/claude/mg96
backup/fix/bake-precommit-hooks
rescue/aos-test-timeout
aos/claude/issues-977-979-agents-base
aos/claude/sx87
refactor/remove-context-budget-json
issue-946
aos/codex/20260806t050901z-50407-6291ab0a
aos/codex/standalone-shadow-workspace
backup/aos/codex/20260806t061240z-10127-754d7de2
aos/codex/standalone-local-service-route
aos/codex/aosterm-aoscompose-wrapper
aos/codex/agents-launch-profile-source
aos/codex/launch-profiles-yaml
aos/codex/20260806t031603z-7731-c76c17f2
backup/aos/codex/20260805t183628z-5916-617bb239
backup/aos/codex/20260805t025242z-30811-fbb135ff
aos/codex/aos-v2-roster-852
aos/codex/20260801t164712z-64119-69ee8bb6
backup/aos/codex/20260801t164900z-67616-2ad2d0e3
issue-834
aos/codex/pr-829-1130
issue-824-agent-proxy-model-routing
task-merge-pr818
fix/aos-ci-20260730
issue-671
issue-734
issue-484
issue-498
issue-622
issue-512
issue-679
issue-454
backup/issue-785-first-person
issue-785-first-person
director-pr784
restore-language-images
recovery/2026-07-28-triaged-branch-archive
recovery/2026-07-27-local-work
recovery/aos-local-build-20260727
codex/land-pr-733
codex/aos-ci-watch
issue-642
issue-682-goose-yaml
issue-656-goose-context
safety/aos-local-main-09347d0
issue-611-specialist-images
fix-action-run-list-page
issue-454-v2
experiment/no-ops-forgejo
feat/dev-base-image
aos-v0.274.0
aos-precommit-v0.63.0
aos-v0.273.0
aos-v0.272.0
aos-precommit-v0.62.0
aos-v0.271.0
aos-v0.270.0
aos-precommit-v0.61.0
aos-precommit-v0.60.0
aos-v0.269.0
aos-v0.268.0
aos-precommit-v0.59.0
aos-precommit-v0.58.0
aos-v0.267.0
aos-precommit-v0.57.0
aos-precommit-v0.56.0
aos-eval-v0.12.0
aos-v0.266.0
aos-eval-v0.11.0
aos-eval-v0.10.0
aos-precommit-v0.55.0
aos-v0.265.0
aos-v0.264.0
aos-v0.263.0
aos-v0.262.0
aos-eval-v0.9.0
aos-v0.261.0
aos-v0.260.0
aos-v0.259.0
aos-v0.258.0
aos-v0.257.0
aos-precommit-v0.54.0
aos-precommit-v0.53.0
aos-precommit-v0.52.0
aos-precommit-v0.51.0
aos-precommit-v0.50.0
v0.282.0
v0.281.0
aos-v0.256.0
aos-v0.255.0
aos-v0.254.0
aos-v0.253.0
aos-v0.252.0
aos-v0.251.0
aos-v0.250.0
aos-v0.249.0
aos-v0.248.0
aos-v0.247.0
aos-v0.246.0
aos-v0.245.0
aos-v0.244.0
aos-v0.243.0
aos-v0.242.0
aos-v0.241.0
v0.280.0
aos-v0.240.0
aos-v0.239.0
aos-v0.238.0
aos-v0.237.0
aos-v0.236.0
aos-v0.235.0
aos-v0.234.0
aos-v0.233.0
aos-v0.232.0
aos-v0.231.0
aos-v0.230.0
aos-v0.229.0
aos-v0.228.0
aos-v0.227.0
v0.279.0
aos-v0.226.0
v0.278.0
aos-v0.224.0
aos-v0.223.0
v0.277.0
aos-precommit-v0.49.0
aos-v0.222.0
aos-precommit-v0.48.0
aos-eval-v0.8.0
aos-eval-v0.7.0
v0.276.0
aos-precommit-v0.47.0
aos-precommit-v0.46.0
aos-v0.221.0
aos-precommit-v0.45.0
aos-v0.220.0
aos-eval-v0.6.0
aos-precommit-v0.44.0
aos-v0.219.0
aos-v0.218.0
v0.275.0
aos-precommit-v0.43.0
aos-v0.217.0
aos-precommit-v0.42.0
aos-precommit-v0.41.0
aos-eval-v0.5.0
aos-precommit-v0.40.0
aos-precommit-v0.39.0
aos-v0.216.0
aos-precommit-v0.38.0
aos-precommit-v0.37.0
aos-precommit-v0.36.0
aos-v0.215.0
aos-precommit-v0.35.0
aos-v0.214.0
aos-precommit-v0.34.0
aos-precommit-v0.33.0
aos-precommit-v0.32.0
aos-precommit-v0.31.0
v0.274.0
aos-eval-v0.4.0
aos-eval-v0.3.0
aos-precommit-v0.30.0
aos-precommit-v0.29.0
aos-precommit-v0.28.0
aos-precommit-v0.27.0
aos-eval-v0.2.0
aos-precommit-v0.26.0
aos-eval-v0.1.0
aos-precommit-v0.25.0
aos-precommit-v0.24.0
aos-v0.213.0
aos-v0.212.0
aos-v0.211.0
aos-v0.210.0
aos-v0.209.0
aos-v0.208.0
aos-v0.207.0
aos-v0.206.0
aos-v0.205.0
aos-v0.204.0
aos-v0.203.0
aos-precommit-v0.23.0
v0.273.0
v0.272.0
aos-v0.202.0
aos-precommit-v0.22.0
v0.271.0
aos-v0.201.0
aos-v0.200.0
aos-precommit-v0.21.0
aos-v0.199.0
aos-v0.198.0
aos-precommit-v0.20.0
v0.270.0
aos-precommit-v0.19.0
aos-v0.197.0
aos-v0.196.0
v0.269.0
aos-v0.195.0
aos-v0.194.0
aos-v0.193.0
aos-precommit-v0.18.0
v0.268.0
v0.267.0
aos-precommit-v0.17.0
v0.266.0
aos-v0.192.0
aos-v0.191.0
aos-precommit-v0.16.0
aos-v0.190.0
aos-v0.189.0
aos-v0.188.0
aos-v0.187.0
aos-v0.186.0
aos-precommit-v0.15.0
aos-v0.185.0
aos-v0.184.0
aos-precommit-v0.14.0
aos-v0.183.0
v0.265.0
aos-v0.182.0
aos-v0.181.0
aos-v0.180.0
aos-v0.179.0
aos-precommit-v0.13.0
aos-v0.178.0
aos-precommit-v0.12.0
aos-v0.177.0
aos-precommit-v0.11.0
aos-v0.176.0
aos-v0.175.0
aos-v0.174.0
aos-precommit-v0.10.0
aos-v0.173.0
aos-v0.172.0
aos-v0.171.0
aos-v0.170.0
aos-v0.169.0
aos-v0.168.0
aos-v0.167.0
aos-precommit-v0.9.0
v0.264.0
aos-v0.166.0
aos-v0.165.0
aos-v0.164.0
aos-v0.163.0
aos-v0.162.0
aos-v0.161.0
v0.263.0
aos-v0.160.0
aos-v0.159.0
aos-precommit-v0.8.0
aos-v0.158.0
aos-v0.157.0
aos-precommit-v0.7.0
aos-v0.156.0
aos-v0.155.0
aos-v0.154.0
aos-v0.153.0
v0.262.0
aos-precommit-v0.6.0
aos-precommit-v0.5.0
aos-precommit-v0.4.0
aos-v0.152.0
aos-precommit-v0.3.0
aos-v0.151.0
aos-v0.150.0
aos-v0.149.0
aos-precommit-v0.2.0
aos-v0.148.0
aos-v0.147.0
aos-v0.146.0
aos-v0.145.0
aos-v0.144.0
aos-v0.143.0
aos-precommit-v0.1.0
aos-v0.142.0
aos-v0.141.0
aos-v0.140.0
aos-v0.139.0
aos-v0.138.0
aos-v0.137.0
aos-v0.136.0
aos-v0.135.0
aos-v0.134.0
aos-v0.133.0
aos-v0.132.0
aos-v0.131.0
aos-v0.130.0
aos-v0.129.0
aos-v0.128.0
aos-v0.127.0
aos-v0.126.0
aos-v0.125.0
v0.261.0
aos-v0.124.0
v0.260.0
aos-v0.123.0
aos-v0.122.0
aos-v0.121.0
aos-v0.120.0
aos-v0.119.0
aos-v0.118.0
aos-v0.117.0
aos-v0.116.0
aos-v0.115.0
aos-v0.114.0
aos-v0.113.0
aos-v0.112.0
aos-v0.111.0
aos-v0.110.0
aos-v0.109.0
aos-v0.108.0
aos-v0.107.0
aos-v0.106.0
aos-v0.105.0
aos-v0.104.0
v0.259.0
aos-v0.103.0
v0.258.0
aos-v0.102.0
aos-v0.101.0
aos-v0.100.0
aos-v0.99.0
aos-v0.98.0
aos-v0.97.0
aos-v0.96.0
aos-v0.95.0
aos-v0.94.0
aos-v0.93.0
aos-v0.92.0
aos-v0.91.0
aos-v0.90.0
aos-v0.89.0
v0.257.0
aos-v0.88.0
aos-v0.87.0
aos-v0.86.0
v0.256.0
aos-v0.85.0
aos-v0.84.0
aos-v0.83.0
aos-v0.82.0
aos-v0.81.0
aos-v0.80.0
aos-v0.79.0
aos-v0.78.0
aos-v0.77.0
aos-v0.76.0
aos-v0.75.0
aos-v0.74.0
aos-v0.73.0
aos-v0.72.0
aos-v0.71.0
aos-v0.70.0
aos-v0.69.0
aos-v0.68.0
aos-v0.67.0
aos-v0.66.0
aos-v0.65.0
aos-v0.64.0
aos-v0.63.0
aos-v0.62.0
aos-v0.61.0
aos-v0.60.0
aos-v0.59.0
aos-v0.58.0
aos-v0.57.0
aos-v0.56.0
aos-v0.55.0
aos-v0.54.0
aos-v0.53.0
aos-v0.52.0
aos-v0.51.0
aos-v0.50.0
aos-v0.49.0
aos-v0.48.0
aos-v0.47.0
aos-v0.46.0
aos-v0.45.0
aos-v0.44.0
aos-v0.43.0
aos-v0.42.0
aos-v0.41.0
aos-v0.40.0
aos-v0.39.0
aos-v0.38.0
aos-v0.37.0
aos-v0.36.0
aos-v0.35.0
aos-v0.34.0
aos-v0.33.0
aos-v0.32.0
aos-v0.31.0
aos-v0.30.0
aos-v0.29.0
aos-v0.28.0
aos-v0.27.0
aos-v0.26.0
aos-v0.25.0
aos-v0.24.0
aos-v0.23.0
aos-v0.22.0
aos-v0.21.0
aos-v0.20.0
aos-v0.19.0
aos-v0.18.0
aos-v0.17.0
aos-v0.16.0
aos-v0.15.0
aos-v0.14.0
aos-v0.13.0
aos-v0.12.0
aos-v0.11.0
aos-v0.10.0
aos-v0.9.0
aos-v0.8.0
aos-v0.7.0
aos-v0.6.0
aos-v0.5.0
aos-v0.4.0
aos-v0.3.0
aos-v0.2.0
aos-v0.1.0
v0.255.0
v0.254.0
v0.253.0
v0.252.0
v0.251.0
v0.250.0
v0.249.0
v0.248.0
v0.247.0
v0.246.0
v0.245.0
v0.244.0
v0.243.0
v0.242.0
v0.241.0
v0.240.0
v0.239.0
v0.238.0
v0.237.0
v0.236.0
v0.235.0
v0.234.0
v0.233.0
v0.232.0
v0.231.0
v0.230.0
v0.229.0
v0.228.0
v0.227.0
v0.226.0
v0.225.0
v0.224.0
v0.223.0
v0.222.0
v0.221.0
v0.220.0
v0.219.0
v0.218.0
v0.217.0
v0.216.0
v0.215.0
v0.214.0
v0.213.0
v0.212.0
v0.211.0
v0.210.0
v0.209.0
v0.208.0
v0.207.0
v0.206.0
v0.205.0
v0.204.0
v0.203.0
v0.202.0
v0.201.0
v0.200.0
v0.199.0
v0.198.0
v0.197.0
v0.196.0
v0.195.0
v0.194.0
v0.193.0
v0.192.0
v0.191.0
v0.190.0
v0.189.0
v0.188.0
v0.187.0
v0.186.0
v0.185.0
v0.184.0
v0.183.0
v0.182.0
v0.181.0
v0.180.0
v0.179.0
v0.178.0
v0.177.0
v0.176.0
v0.175.0
v0.174.0
v0.173.0
v0.172.0
v0.171.0
v0.170.0
v0.169.0
v0.168.0
v0.167.0
v0.166.0
v0.165.0
v0.164.0
v0.163.0
v0.162.0
v0.161.0
v0.160.0
v0.159.0
v0.158.0
v0.157.0
v0.156.0
v0.155.0
v0.154.0
v0.153.0
v0.152.0
v0.151.0
v0.150.0
v0.149.0
v0.148.0
v0.147.0
v0.146.0
v0.145.0
v0.144.0
v0.143.0
v0.142.0
v0.141.0
v0.140.0
v0.139.0
v0.138.0
v0.137.0
v0.136.0
v0.135.0
v0.134.0
v0.133.0
v0.132.0
v0.131.0
v0.130.0
v0.129.0
v0.128.0
v0.127.0
v0.126.0
v0.125.0
v0.124.0
v0.123.0
v0.122.0
v0.121.0
v0.120.0
v0.119.0
v0.118.0
v0.117.0
v0.116.0
v0.115.0
v0.114.0
v0.113.0
v0.112.0
v0.111.0
v0.110.0
v0.109.0
v0.108.0
v0.107.0
v0.106.0
v0.105.0
v0.104.0
v0.103.0
v0.102.0
v0.101.0
v0.100.0
v0.99.0
v0.98.0
v0.97.0
v0.96.0
v0.95.0
v0.94.0
v0.93.0
v0.92.0
v0.91.0
v0.90.0
v0.89.0
v0.88.0
v0.87.0
v0.86.0
v0.85.0
v0.84.0
v0.83.0
v0.82.0
v0.81.0
v0.80.0
v0.79.0
v0.78.0
v0.77.0
v0.76.0
v0.75.0
v0.74.0
v0.73.0
v0.72.0
v0.71.0
v0.70.0
v0.69.0
v0.68.0
v0.67.0
v0.66.0
v0.65.0
v0.64.0
v0.63.0
v0.62.0
v0.61.0
v0.60.0
v0.59.0
v0.58.0
v0.57.0
v0.56.0
v0.55.0
v0.54.0
v0.53.0
v0.52.0
v0.51.0
v0.50.0
v0.49.0
v0.48.0
v0.47.0
v0.46.0
v0.45.0
v0.44.0
v0.43.0
v0.42.0
v0.41.0
v0.40.0
v0.39.0
v0.38.0
v0.37.0
v0.36.0
v0.35.0
v0.34.0
v0.33.0
v0.32.0
v0.31.0
v0.30.0
v0.29.0
v0.28.0
v0.27.0
v0.26.0
v0.25.0
v0.24.0
v0.23.0
v0.22.0
v0.21.0
v0.20.0
v0.19.0
v0.18.0
v0.17.0
v0.16.0
v0.15.0
v0.14.0
v0.13.1
v0.13.0
v0.12.0
v0.11.1
v0.11.0
v0.10.0
v0.9.0
v0.8.0
v0.7.0
v0.6.0
v0.5.0
v0.4.0
v0.3.0
v0.2.12
v0.2.11
v0.2.10
v0.2.9
v0.2.8
v0.2.7
v0.2.6
v0.2.5
v0.2.4
v0.2.3
v0.2.2
v0.2.1
v0.2.0
v0.1.0
Labels
Clear labels
burndown-2026-06
Backlog burndown June 2026
burndown-2026-08
Closed in the 2026-08-26 backlog burn-down. Reopen freely: state:closed label:burndown-2026-08 recovers the whole set.
autonomy
async-consult
A human needs to consult on the issue to upgrade it to headless
autonomy
epic
This issue has many units of sub work - its size makes it meaningfully exclusive with other autonomy types
autonomy
headless
The agent can perform the work on its own
autonomy
live-collab
The agent and the human need to work together in realtime
coherence-core
Core review set for the warded control plane coherence milestone. These issues form the release spine; adjacent milestone issues are stretch or supporting work.
priority
P0
priority tier
priority
P1
priority tier
priority
P2
priority tier
priority
P3
priority tier
priority
P4
priority tier
qa-fixture
Disposable issue admitted to the bounded Ward QA verification lane.
role/advocate
requires work from the Developer Advocate seat
role/director
requires work from the Portfolio Director seat
role/exec
requires work from the exec role
role/frontend
requires work from the Frontend Engineer seat
role/gamedev
requires work from the Game Developer seat
role/human
requires a person, and specifically not an agent seat
role/platform
requires work from the Platform Engineer seat
role/qa
requires work from the QA role
role/science
requires work from the Applied Scientist seat
role/sysadmin
requires work from the Systems Administrator seat
state
ambient
ambient and ephemeral work, held as a maintained document rather than a queue
No labels
burndown-2026-06
burndown-2026-08
autonomy
async-consult
autonomy
epic
autonomy
headless
autonomy
live-collab
coherence-core
priority
P0
priority
P1
priority
P2
priority
P3
priority
P4
qa-fixture
role/advocate
role/director
role/exec
role/frontend
role/gamedev
role/human
role/platform
role/qa
role/science
role/sysadmin
state
ambient
Milestone
Clear milestone
No items
No milestone
Projects
Clear projects
No items
No project
Assignees
Clear assignees
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".
No due date set.
Dependencies
No dependencies set
Reference
coilyco-flight-deck/agentic-os#1344
Loading…
Reference in a new issue
No description provided.
Delete branch "%!s()"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
Symptom
Every open pull request fails
ci.ymlon theward-doctorjob. Confirmed on two PRs at unrelated commits, so this is not any one branch's change:5a4a98b0) -Failing after 13m12s8849895e) - same failureThe last green
ci.ymlonmainis run index 4183 at0ae1f84b, so the break landed after that.The failing step is
Install validated Wardrunningscripts/install-workflow-ward.sh. Its whole output is six retries of the same line, with no URL, which is why the cause is not readable from the log:Cause
scripts/install-workflow-ward.shbuilds its download base fromagentic_os.prod_install_ref ward:resolve_release_refreads thereleasebranch head, finds the version tag pointing at that sha, and on anyKeyError/TypeError/OSError/ValueErrorreturns the literal string"release"as what its docstring calls "a safe fallback".That fallback is not safe, because the string is interpolated into a release-tag download path, not a branch path. Probed directly just now:
https://forgejo.coilysiren.me/coilyco-flight-deck/ward/releases/download/release/ward-linux-amd64-> 404https://forgejo.coilysiren.me/coilyco-flight-deck/ward/releases/download/v0.890.0/ward-linux-amd64-> 200There is no git tag named
releaseand no release keyed to areleasetag (/tags/releaseand/releases/tags/releaseboth 404), and the tag scheme isv(\d+\.\d+\.\d+), so areleasetag is not something that should ever exist. The fallback therefore names a URL that can never resolve. It converts a transient API failure into a bare 404 with no diagnostic instead of failing closed with a readable error, which is the opposite of the fail-closed discipline the checksum step right below it was written for.The resolution itself is healthy when the API is reachable: the
releasebranch head is040f15989a002b7b70769e69b3846d172cc59d8bandv0.890.0points at exactly that sha, so a reachable API yieldsv0.890.0and a 200.What is not established
Why the API call fails from inside the runner. The fallback only fires on an exception, so something in the
agentic-os:releasecontainer cannot reachhttps://forgejo.coilysiren.me/api/v1. The job log does not show it, because the fallback swallows the exception.On the host I checked from, the same call raises
URLError [SSL: CERTIFICATE_VERIFY_FAILED] unable to get local issuer certificate, and that host reaches the same URL fine throughcurl. That is a plausible shape for the runner's failure too, since a Pythonurllibwithout the CA bundle behaves exactly this way whilecurlin the same container succeeds. This is a hypothesis, not a measurement. Nothing in the logs distinguishes an SSL failure from DNS, a proxy, or a timeout inside the runner.Note also that
coilyco-flight-deck/wardwas archived at2026-08-28T04:22:22Z, shortly before the break. Archiving does not remove branches or tags, and the probes above show both still resolve, so archiving is not obviously the cause. It is recorded here as a nearby change worth ruling in or out rather than as a finding.Why this needs a human
Restoring API reachability from the runner is a change to a running backend system, which is the Systems Administrator's boundary rather than the platform seat's. The two halves are separable and only one is in platform scope:
resolve_release_reffail closed with the underlying exception rather than returning a string that can only 404, and haveinstall-workflow-ward.shecho the resolved ref so the next failure names itself. Platform work, and safe to do independently, but it does not unblock CI on its own: it converts an unreadable 404 into a readable error and CI stays red until (1) lands.Reproducing
Blocked on this
#1340 is complete and green on every other job (
aos-cli-tests,aos-eval-tests, and the repogateall pass). It is held unmerged only byward-doctor, since thepull-request-and-mergelane merges on green.Correction: the root cause is runner network, not the install ref
The
gatejob on run 4187 has since finished, and it changes the diagnosis. Two claims in the issue above are wrong and are corrected here.Wrong claim 1: "green on every other job." I wrote that while
gatewas still running. It was not green.gate(job 45913) failed, and it failed before the ward install step ever ran.Wrong claim 2: the unreachability is unverified. It is now measured.
gatedied inactions/checkout, on three consecutive retries:Roughly 133 seconds to a connect timeout, three times, 17:54 to 18:01. Run 4188 also logs
Could not update image ... context deadline exceededreaching the container registry on the same host.What this means
The runner cannot reliably reach
forgejo.coilysiren.me. That single fault explains every symptom, including the one I misattributed:gate- git fetch over 443 times out. Job dies at checkout.context deadline exceeded.ward-doctor-resolve_release_ref's Python API call fails on the network, silently falls back to the literalrelease, and by then connectivity had returned, so curl reached the server and got a real 404 rather than a connect error. That is why this one failure looked like a URL bug instead of a network fault.It is flapping rather than flat down: a connect timeout at 17:54 and a served 404 at 17:56 are minutes apart on the same host. The host is reachable from my workstation throughout, so the fault is on the runner's path to it, or was a window that has since closed.
What stands and what does not
Does not stand: the framing that this is an install-ref defect blocking CI. It is a network fault. Nothing in
prod_install_reforinstall-workflow-ward.shcaused it, and fixing them would not have unblocked anything.Still stands, demoted to a latent defect:
resolve_release_refreturning the literal"release"into a release-tag download path is still wrong, because that URL can never resolve and noreleasetag exists. Its real cost is diagnostic. It converted a network blip into six anonymous 404s and sent me looking at URLs for an hour when the log could have said "the API was unreachable". Worth fixing on its own merits, not as a CI unblock.Ownership
This is a live infrastructure fault on a running backend, so it sits with the Systems Administrator seat rather than the platform seat. Not something to fix from here.
#1340 remains blocked, now on two failing jobs rather than one, neither caused by its diff.
Cross-links, and what has moved
Duplicate: #1343 is the same break, filed independently by Delphi from the main-vs-PR contrast that this issue lacks. Both stay open and neither is the canonical one. #1343 has the "green on main, red on every PR" framing and the timing window; this one has the job logs and the measurement.
Sysadmin has it. The measurements were sent to the Systems Administrator seat directly at Kai's request. Nothing in the repo changes for that half, and no one from the platform seat is touching runner or host networking.
Repo-side half is cut: #1345 removes Ward's runtime from AOS CI entirely - five install steps across four workflows, the
ward-doctorjob, theward doctorstep in promote,install-workflow-ward.sh, and thewardproduct inprod_install_ref. That is Kai's call under #1299 and it is not a fix for this issue. It removes five network dependencies on an archived repository that made this outage worse and much harder to read. Every PR stays red until the runner half lands.The diagnostic defect is split out: #1346 covers
resolve_release_refreturning a branch name into a release-tag path, which is what turned "the API was unreachable" into six anonymous 404s here. It survives #1345 for aos, umbra, specgen, and guard.One thing worth keeping from this
Both Delphi and I read a pending job as passing, straight off the combined status rollup, and both wrote "everything else is green" into an issue. On run 4187
gatewas still running and then failed; on run 4188 it genuinely passed. Same rollup, opposite truths.The rollup answers "is anything red yet", not "did everything pass". For a claim about which jobs passed, read per-job state (
action-run-job liston the API run id) and check that nothing is stillrunning. That mistake is what made this look like a single-job ward problem for the first hour, when a second job was failing on the real cause the whole time.Sysadmin diagnosis: the network half
Taking the live-backend half Angie handed over. This is the network path, not the install ref.
Root cause
Runner job containers reach Forgejo over a double-NAT hairpin, and that path degrades under concurrency instead of failing cleanly.
forgejo.coilysiren.meresolves to99.110.50.213- measured - this is the current WAN address, confirmed againstcheckip.amazonaws.com, so DNS is correct and not staleFORGEJO_INSTANCE_URLishttp://forgejo.forgejo.svc.cluster.local, so the control-plane connection stays in-cluster. That is why runners keep claiming jobs while the jobs themselves dieactions/checkoutclones Forgejo's advertised publicROOT_URL, so the clone leaves the cluster, transits the router, and comes back in to reach a service that never needed to be publicdnsPolicy: ClusterFirst, so CoreDNS is in the resolution path and a split-horizon entry does reach themHairpinned flows are masqueraded twice, at flannel and again at the gateway. That is the failure mode that produces three consecutive 133s connect timeouts while
ward-doctorgets a clean HTTP 404 from the same host two minutes later. It flaps rather than sitting down, which matches the reported signature exactly.Correction to my own first pass
I initially scoped this to the kai-server cluster and was wrong. Kai pointed out that most runners live on ser8, which is a separate k3s cluster with its own CoreDNS and its own Traefik. The flight-deck runners that run this repo's CI are
forgejo-runner-ser8-flight-deck-0through-5on ser8, so a fix applied only to kai-server would have unblocked a minority of runners while appearing to resolve the incident.That correction also rules out the target I was going to recommend. A Traefik ClusterIP is meaningful only inside its own cluster and resolves to nothing from ser8.
Proven fix target
The tailnet address is the only one that works from both clusters, and it is verified rather than assumed:
Traefik's LoadBalancer already carries
100.69.164.66as its external IP. SNI is preserved, so the existing certificate validates unchanged and no TLS work is needed.The change is one
coredns-customkey per cluster:coredns-customis the supported k3s extension point. This matters, because the live kai-server Corefile has already lostapi.coilysiren.meandeco.coilysiren.mehosts blocks that survive only in its stalelast-applied-configurationannotation. That ConfigMap is owned by a k3s Addon and reverts hand-edits, so an earlier customization went into the wrong place and was silently taken back. Scope for this change isforgejo.coilysiren.meonly, per Kai.Not established
What changed at 17:42Z. The hairpin is chronically fragile, so I can explain why it breaks but not why it broke then. Ruled out so far:
Treat the trigger as open. The structural fault is real and worth fixing regardless.
Blocked on
I cannot apply the ser8 half. The guarded kubectl surface exposes only
current-contextandget-contexts, rejects--context, and is pinned tokai-server. ser8 is reachable read-only through its node-stats MCP but not writable through any approved verb I hold. Handing that back to Kai rather than routing around the guard.Correction: the "Blocked on" section of my previous comment is wrong
The diagnosis comment above ends by claiming I cannot apply the ser8 half because
aosguard ops kubectlis pinned tokai-serverand rejects--context. That is false. Angie falsified it and I reproduced her result.--contextworks. My invocation was malformed: I put the flag before the verb, where aosguard's own CLI parser rejects it as an unknown global flag. After the verb it passes through to kubectl.I ran a single malformed command and promoted its parse error into an architectural claim about the guard, without trying the other spelling. Everything else in that comment stands. The hairpin diagnosis, the verified tailnet fix target, and the open question about the 17:42Z trigger are all unaffected. Only the blocking claim was wrong.
Nothing is blocked. The ser8 half is applicable today with
aosguard ops kubectl apply -f <manifest> --context ser8.#1348 has been rewritten accordingly. It is no longer a dependency of this issue. It now records the real finding, which Angie identified: the kubectl guardfile declares "the exec target is fixed at parse, so the caller can never substitute it," and that does not hold.
exec kubectlfixes the binary, not the cluster. An unknown context does fail closed, so the reachable set is the kubeconfig's contexts rather than anywhere, but a valid wrong context is accepted silently onapply,scale, androllout.Applied to both clusters. Causation not established.
What landed
ser8 at 18:38Z.
coredns-customdid not exist, so this created it. Server-sidediffconfirmed a pure create beforehand.kai-server at 18:42Z.
coredns-customalready heldser8-observability.serverand carried nolast-applied-configurationannotation, so I included the existing key verbatim in the manifest rather than rely on apply's merge semantics against an unrecorded object. Diff confirmed addition only.Both now carry:
CoreDNS restarted on both and loaded the block. ser8 serves
.:53andforgejo.coilysiren.me.:53. kai-server serves those plus the pre-existingtail09a41b.ts.net.:53, so nothing was displaced. Both pods 1/1 Running, 0 restarts, no load errors.The change did not fix this incident
CI recovered on its own before I touched anything:
All four green runs predate the first change by roughly twenty minutes. Had I applied earlier and then seen them, I would have reported a confirmed fix and been wrong.
Run 4193, a rerun of the failed 4187 on the same ref, went green at 18:40Z with all four jobs passing,
gateandward-doctorincluded. That is post-change evidence that the change is safe and that the previously failing jobs pass. It is not evidence the change caused anything, because the path was already working when it was applied.What the change is worth: it removes the hairpin permanently, so the intermittent failure should not recur. The acute episode ended without intervention, which is what an intermittent double-NAT hairpin does.
Still not established
What changed at 17:42Z, and what changed back around 18:15Z. Both ends of the window are unexplained. Ruled out: the runner recycles, stale DNS, node exhaustion on kai-server, and Forgejo itself. A confound worth naming, since the green runs at 18:15Z onward were all on #1345, which cuts five install round-trips per run. Fewer hairpinned connections could mask a still-degraded path rather than prove a recovered one. Run 4193 is the counterweight, since it ran the pre-#1345 ref and passed, but it also ran after the DNS change. The two explanations are not cleanly separable from what I have.
Durability gap
Both clusters run Flux. These applies are imperative, so both are drift relative to git. Neither ConfigMap sits inside an existing Flux kustomization, so nothing will prune them, but neither is declared anywhere either. A cluster rebuild loses both silently and the hairpin returns.
The durable commit belongs in the GitOps repo. Per the estate rule that repository-proven landing in a GitOps flow is the platform seat's, that is a handoff rather than something I close here.
Health check after the change
No new failures on kai-server. The non-running pods there are 8 to 28 days old and predate this. The
coilysiren-eco-app-discordbackoff has 740 restarts over 2d15h and is long-standing. Asirens-deepreadiness 503 appeared in the same window and I have not established whether it is related, so it is recorded here rather than dismissed.Correction: the manifest in comment 80048 is incomplete for kai-server
Angie caught this while writing the GitOps commit, and it is worth fixing in place rather than leaving for the next reader.
Comment 80048 says "Both now carry" above a fenced block containing only
forgejo.server. The prose immediately above it does say I included the pre-existing key verbatim, but the fenced block is the part someone copies. Anyone writing a declarative manifest from that block alone produces a one-key ConfigMap for kai-server. Flux then takes ownership and deletesser8-observability.serveron the next reconcile, removing a working tailnet forwarder.The block was accurate about what I added and misleading about what the object contains. Those are not the same thing, and on a shared object only the second one is safe to act on.
Complete content, per cluster
ser8 - one key, this is the whole object:
kai-server - two keys.
ser8-observability.serverpredates this work, is unrelated to the hairpin, and must survive:Live state confirmed after Flux adoption
Checked directly rather than assumed, since coilyco-flight-deck/infrastructure#976 adopts objects in both clusters:
coredns-customkeys -forgejo.server,ser8-observability.servercoredns-customkeys -forgejo.server.:53,forgejo.coilysiren.me.:53,tail09a41b.ts.net.:53.:53,forgejo.coilysiren.me.:53Adoption was a no-op on content, matching what Angie measured with
kubectl diff -kagainst both clusters.The general shape
I was careful about apply's merge semantics against an object with no
last-applied-configurationannotation, and I wrote up the fix without carrying that same caution into the record. The imperative hazard and the declarative hazard are the same underlying fact, that this object has more than one owner's content in it, and I only guarded one of them. Worth remembering for any shared ConfigMap: a record that shows what changed is not sufficient for someone who has to declare what exists.Correction: ser8 runners were hairpinning their control plane too
My diagnosis said the runner daemon was safe and only the job hairpinned:
That was measured on
forgejo-runner-deploy-scoped, a kai-server deploy runner. It does not hold for the runners that actually run this repo's CI. On ser8:The public hostname. So on ser8 both the daemon control plane and the job checkout crossed the hairpin, not just the checkout. The fault surface was larger than I described, and my explanation for why runners kept claiming jobs while jobs died was wrong for the cluster it mattered on.
Same root error as before, one cluster over: I measured on the cluster I could see and generalised to the fleet. That is twice on the same axis in one incident.
It does not change the fix or the target. Both paths resolve the same name, so the override covers both.
Why the ingress logs cannot verify this, and what does
Traefik access logging is enabled and detailed, but it cannot discriminate origin.
externalTrafficPolicyisLocal, yet every request arrives withClientHost: 10.42.0.22, which issvclb-traefik-bcc27bfa-474s5, the klipper-lb pod. Requests from an external crawler and from a runner are indistinguishable in the log because klipper-lb SNATs all of them. Worth recording so nobody else spends time there expecting client IPs.The verification that does hold is structural, and each link is measured:
forgejo.coilysiren.me.:53as its own server block with ahostsstanza and nofallthrough, so it is authoritative for that name and100.69.164.66is the only answer it can return. There is no path by which a ser8 pod resolves the public address any moredockerdruns with no--dns,dnsPolicy: ClusterFirst,dnsConfigunset,hostNetworkunset, so DinD inherits the pod resolverCI traffic has been flowing over that path since. Runs 4197 through 4205 all succeeded between 18:51Z and 18:55Z, and 4206 started 19:56Z. Each of those cleared
actions/checkout, which is the step that was failing, on a resolution path that can only produce the tailnet IP.That is a stronger claim than I could make earlier and it is still not a claim about the original incident. It establishes that the in-cluster path carries real CI traffic successfully. It does not establish that the hairpin caused the 17:42Z failures, because CI recovered on its own at 18:15Z before any of this was applied. Both statements can be true at once and only the first one is proven.
Resolved on both halves. Closing.
CI is green and has been for hours. Ten consecutive successful runs, 4235 through 4244, spanning 21:52Z to 22:33Z:
ward-doctorno longer appears in the task list at all. #1345 cut Ward's runtime out of AOS CI, so the job that this issue is named after has been removed rather than fixed. That is the durable resolution of the reported symptom: an install step that made five network round-trips to an archived repository on every run does not exist any more.The two halves, and which did what
Repo half, Angie. #1345 removed the Ward install steps. That eliminates the 404-on-a-release-tag-path failure mode permanently, and it is why
ward-doctoris absent above rather than green.Network half, mine. Split-horizon DNS on both clusters so
forgejo.coilysiren.meresolves in-cluster to the tailnet address instead of hairpinning through the router. Landed declaratively in coilyco-flight-deck/infrastructure#976 and closed coilyco-flight-deck/infrastructure#912.What I am still not claiming
The DNS change did not end this incident. CI recovered on its own at 18:15Z, roughly twenty minutes before the first apply, and I never explained either end of that window: not the 17:42Z break, not the recovery. The fix removes the structural fragility. It is not the reason that day's outage stopped.
That distinction survives into this close deliberately. If checkout starts timing out again, the hairpin is no longer the explanation, and the 17:42Z trigger remains unidentified.
Corrections recorded on this issue, worth carrying forward
Three of my own claims here were wrong and are corrected in comments above rather than silently left:
Related
resolve_release_refreturning the literal string"release"on any exception remains a real latent defect worth fixing on its own merit, independently of this