Backlog burn-down 2026-08-28: what was closed and why #981
Labels
No labels
burndown-2026-06
burndown-2026-08
autonomy
async-consult
autonomy
epic
autonomy
headless
autonomy
live-collab
coherence-core
priority
P0
priority
P1
priority
P2
priority
P3
priority
P4
qa-fixture
role/advocate
role/director
role/exec
role/frontend
role/gamedev
role/human
role/platform
role/qa
role/science
role/sysadmin
state
ambient
No milestone
No project
No assignees
1 participant
Notifications
Due date
No due date set.
Dependencies
No dependencies set
Reference
coilyco-flight-deck/infrastructure#981
Loading…
Reference in a new issue
No description provided.
Delete branch "%!s()"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
Kai asked for the backlog to get as small as possible. This is the record for that sweep so no individual close has to carry the reasoning alone.
Everything closed here is tagged
burndown-2026-08.state:closed label:burndown-2026-08recovers the whole set. Reopen freely, no permission needed and no explanation owed.Starting count: 124 open issues.
Closed with verified live evidence
Each of these was checked against the running system today rather than judged from its title.
appshealth gate - nowREADY=True. Its evidence block namesforgejo-runner-deploy-atlas, so this and #834 were one fault filed twice by different sweeps that did not find each otherReleased- now all 43Bound, zero ReleasedClosed as duplicates
Kept the survivor named in each line.
Kept deliberately from the retention cluster because they are distinct problems rather than restatements: #905 (attachments, 39.9 GiB, has specific data), #624 (upload staging mechanism), #733 (backup rollout).
Closed as stale
Three or more months old, priority P3, no comments, no activity. All still plausible, none currently moving, all recoverable by label.
#46, #48, #110, #157, #184, #185, #253, #293 and #689.
#689 is a verification task for work that has since been superseded. The rest are dormant improvements rather than live faults.
Deliberately kept open, because they are failing right now
Recording these so a future burn-down does not sweep them by age. Each was verified live today.
1/1 Runningwith 0 restarts, which is exactly why nobody noticedinfrastructureandbridge-deploystill suspended, confirmed bykubectl get kustomization -Asirens-discord-opsrunner still exists, 1/1, 31dREGISTER_DOMAIN, six failures confirmed todayStanding issues, re-anchored rather than closed
The point-in-time snapshots closed into these two, so the concern survives without the backlog asserting stale numbers.
What this sweep did not do
Merge clusters that need a judgement call about which issue keeps the detail, specifically the runner-recycle group (#868, #840, #706) and the runner-scope group (#956, #955, #682, #693). Each contains real distinct content, and collapsing them would lose more than it saves.
The closes above are reversible by design. If any of them were wrong, the label recovers them as a set rather than one at a time.
Corrections to the plan above, after reading the issues instead of their titles
Two entries in the original body are wrong. Both errors are the same shape: I planned the merges from titles and priorities, and reading the bodies reversed the answer twice.
#903 was NOT closed, and should not be
The body lists it as "subsumed by #600" on the strength of its title saying the disk pressure was relieved. Reading it, #903 is the most actionable issue in the retention cluster, not a duplicate of anything:
older_than_30dat zero for every owner, accumulating at roughly 3 GB/daykeep_countper container package#903 stays open. If anything it is the survivor of that cluster rather than a casualty, and #600 is the vaguer umbrella.
#680 was closed, and it was the wrong survivor
Closed as a duplicate of #505 on the heuristic "keep the one with comments." Wrong. #505 is the original charter; #680 was the continuation, carrying the completed work (the stack already scaled to zero) and a seven-item remaining checklist.
The guarded Forgejo surface has
closebut no reopen, so I could not undo it. Recovered by migrating #680's full current-state and acceptance list into #505 (comment 80264). Nothing is lost, but the close itself was a mistake.#48 also stays open
Listed for stale-closure. Keeping it, because #903 independently records
k3s/server/dbat 8.5 GiB with a note that "this estate has had sustained event storms", which is live corroboration for the kine watchdog cascade #48 describes. Dormant is not the same as irrelevant.The lesson worth keeping
Every merge in a burn-down needs both bodies read, not both titles. Titles encode what the filer knew at filing time, and in a backlog this old that is frequently the least current thing about the issue. The two errors here cost one irreversible close and would have cost the retention cluster's best analysis.
Continuation, 2026-08-29 ~04:20Z. One correction to this record, and the two deferred clusters resolved.
Correction: the guarded surface does have reopen
This record states "the guarded Forgejo surface has
closebut no reopen, so I could not undo it."It has one.
aosguard ops forgejo issue reopen <owner> <repo> <index>, authorized bycan reopen issue, always sending the fixed body{"state": "open"}. So the #680 close was recoverable at the time and is recoverable now.Not reopening #680. Its full current state and acceptance list were migrated into #505 in comment 80264, so reopening it today would restore a duplicate rather than recover anything. The close stays wrong for the reason this record already gives, and the recovery stands. What changes is only the claim that no undo existed, which would otherwise teach the next sweep to treat every close as irreversible.
The two clusters this sweep deferred
Both were read in full and verified against the live clusters rather than judged from titles.
Runner-recycle group, #868 / #840 / #706. Not merged, and the sweep was right. These are three distinct defects that share one component: a coverage gap in the selector, a simultaneity problem in the scheduling, and a safety problem with active jobs. Collapsing them would lose two of the three.
docker-lib-forgejo-runner-build-0still Bound with zero mounts at 5.99 GB or more, andsirens-discord-opsstill unrecycled, now provable by pod age rather than by label read: 6d3h against 12h2m for every labelled sibling.Runner-scope group, #956 / #955 / #682 / #693. One close, one re-anchor, two confirmed accurate.
coilyco-bridge/deploy#608. Evidence in comment 80435.A closure error worth the same treatment as #903
#846 was closed on half its content. Its title carries two findings: "~16 GB in unmounted runner PVCs" and "three PVs stuck Released". The Released half is genuinely fixed, 43 volumes all
Bound, zeroReleased, and that is what this record verified. The unmounted-PVC half was never addressed and is still true, measured at 15.4 GB across eleven claims today.Same shape as the #903 error already recorded here, one step further on: not a title read this time, but a two-claim title where verifying one claim closed both. The lesson generalizes. A compound title needs every claim checked, not the one the sweep happened to measure.
Handled by folding the full measured inventory into #868, which already owned the stranded-PVC decision, rather than reopening #846 whose other half really is resolved. One issue owns PVC reclaim instead of two. Evidence in comment 80441.
Net effect on the count
Infrastructure went 109 to 108 open. One close, four re-anchors, one correction. The re-anchors are the point rather than a consolation: four issues that a future age-based sweep would have read as stale now carry today's evidence and say plainly which are live, which is waiting on a decision, and which one is finished.
Ambient findings parked here, per the director's instruction, plus the #846 reopen.
Portia decided the
STATE: ambient and ephemeral workdocuments do not get created yet, and asked that ambient findings from this pass land here meanwhile. This issue already exists and is already scoped to the sweep, so nothing new was filed to hold them.#846 is reopened
Recorded above that this sweep closed it on one of its three problems. Portia verified and gated the reopen, and I executed it with
aosguard ops forgejo issue reopen. Retitled to the surviving scope, with complete measurements in comment 80513, and#870warned in comment 80517 because it is the issue a hardware purchase would be priced from.Correcting the correction I posted earlier today: I wrote that the reopen verb exists, which is true, and then said #680 was "recoverable after all" without noting that two seats had independently hit the missing verb. It is missing from the Forgejo MCP surface and present on aosguard.
coilyco-bridge/deploy#395is open to addcan reopen issueto both MCP guardfiles for parity. So "close is one-way" is true for any seat on the MCP surface and false for this one, and that distinction is the thing worth carrying, not the bare existence of the verb.Ambient: eight Failed pods on kai-server, no issue filed
None of these is a live fault. Each is a terminated pod object left behind by a rollout, holding no compute and serving nothing:
Filed nowhere on purpose, under the
#482test: no owning seat wants them, and there is no done-condition a reader would check. What makes them worth writing down at all is second-order: they make every pod list read as partly broken, so the next person scanning for a real fault has eight false positives to dismiss first. That cost is real and it is not worth an issue.The two
coilysiren-eco-appentries are the pre-deploy#818rollout and are named there already.Ambient: three failed recycle Jobs, 36 to 38 days old
forgejo-runner-recycleJobs from 2026-07-21, -22 and -23 persist withfailed: 2andBackoffLimitExceeded. The last four runs of that CronJob all succeeded. Noted on#869for the same reason as above: they make the CronJob read as failing at a glance when it is not.Ambient: external-dns is restarting
external-dns-77dc66dd47-v6jbcshows 22 restarts in 6d3h, roughly 3.6 a day. Not filed, because#971already records six confirmed Route53REGISTER_DOMAINfailures and this is plausibly the same fault seen from the pod side. Stated as a candidate link rather than a diagnosis - nobody has connected them, and whoever picks up #971 should check the restart count before assuming it is only an API-level failure.