Backlog burn-down 2026-08-28: what was closed and why #981

Open
opened 2026-08-28 21:32:38 +00:00 by coilyco-ops · 3 comments
Owner

Kai asked for the backlog to get as small as possible. This is the record for that sweep so no individual close has to carry the reasoning alone.

Everything closed here is tagged burndown-2026-08. state:closed label:burndown-2026-08 recovers the whole set. Reopen freely, no permission needed and no explanation owed.

Starting count: 124 open issues.

Closed with verified live evidence

Each of these was checked against the running system today rather than judged from its title.

  • #912 - runners hairpin the public edge - fixed today, split-horizon DNS on both clusters, landed in #976
  • #834 - atlas runner crash-looping - StatefulSet no longer exists
  • #836 - Flux apps health gate - now READY=True. Its evidence block names forgejo-runner-deploy-atlas, so this and #834 were one fault filed twice by different sweeps that did not find each other
  • #824 - galaxy-gen CI red - runs 161 and 162 green. Closed as an instance, explicitly not as the #783 class
  • #966 - kai-server 82.8 percent - now 77.9 percent, roughly 45 GB reclaimed
  • #846 - three PVs stuck Released - now all 43 Bound, zero Released
  • #871 - ser8 growth attribution - the trend reversed, 54.7 to 47.3 percent

Closed as duplicates

Kept the survivor named in each line.

  • #680 - duplicate of #505, retire kai-server observability after ser8 cutover
  • #622 - duplicate of #806. Evaluate comes before adopt, and #806 is the evaluation
  • #864 - duplicate of #811, same nameserver-limit root cause. #811 is broader, covering both nodes
  • #903 - subsumed by #600, the umbrella retention issue. #903's own title records that the disk pressure motivating it was already relieved

Kept deliberately from the retention cluster because they are distinct problems rather than restatements: #905 (attachments, 39.9 GiB, has specific data), #624 (upload staging mechanism), #733 (backup rollout).

Closed as stale

Three or more months old, priority P3, no comments, no activity. All still plausible, none currently moving, all recoverable by label.

#46, #48, #110, #157, #184, #185, #253, #293 and #689.

#689 is a verification task for work that has since been superseded. The rest are dormant improvements rather than live faults.

Deliberately kept open, because they are failing right now

Recording these so a future burn-down does not sweep them by age. Each was verified live today.

  • #765 - Tailscale operator OAuth. Failing continuously for 22 days, last seen 21:07Z today. Raised P3 to P2. The pod is 1/1 Running with 0 restarts, which is exactly why nobody noticed
  • #837 - Flux infrastructure and bridge-deploy still suspended, confirmed by kubectl get kustomization -A
  • #771 - the untracked sirens-discord-ops runner still exists, 1/1, 31d
  • #869 - the recycle cron timezone bug, confirmed and worse than filed: kai-server fires 16:15Z while ser8 fires 09:15Z, so the two clusters recycle seven hours apart
  • #971 - Route53 REGISTER_DOMAIN, six failures confirmed today

Standing issues, re-anchored rather than closed

  • #870 - kai-server capacity, now carries today's 77.9 percent
  • #916 - ser8 capacity, now carries today's 47.3 percent

The point-in-time snapshots closed into these two, so the concern survives without the backlog asserting stale numbers.

What this sweep did not do

Merge clusters that need a judgement call about which issue keeps the detail, specifically the runner-recycle group (#868, #840, #706) and the runner-scope group (#956, #955, #682, #693). Each contains real distinct content, and collapsing them would lose more than it saves.

The closes above are reversible by design. If any of them were wrong, the label recovers them as a set rather than one at a time.

Kai asked for the backlog to get as small as possible. This is the record for that sweep so no individual close has to carry the reasoning alone. **Everything closed here is tagged `burndown-2026-08`.** `state:closed label:burndown-2026-08` recovers the whole set. Reopen freely, no permission needed and no explanation owed. Starting count: **124 open issues.** ## Closed with verified live evidence Each of these was checked against the running system today rather than judged from its title. * **#912** - runners hairpin the public edge - fixed today, split-horizon DNS on both clusters, landed in #976 * **#834** - atlas runner crash-looping - StatefulSet no longer exists * **#836** - Flux `apps` health gate - now `READY=True`. Its evidence block names `forgejo-runner-deploy-atlas`, so this and #834 were one fault filed twice by different sweeps that did not find each other * **#824** - galaxy-gen CI red - runs 161 and 162 green. Closed as an instance, explicitly not as the #783 class * **#966** - kai-server 82.8 percent - now **77.9 percent**, roughly 45 GB reclaimed * **#846** - three PVs stuck `Released` - now **all 43 `Bound`**, zero Released * **#871** - ser8 growth attribution - the trend **reversed**, 54.7 to **47.3 percent** ## Closed as duplicates Kept the survivor named in each line. * **#680** - duplicate of **#505**, retire kai-server observability after ser8 cutover * **#622** - duplicate of **#806**. Evaluate comes before adopt, and #806 is the evaluation * **#864** - duplicate of **#811**, same nameserver-limit root cause. #811 is broader, covering both nodes * **#903** - subsumed by **#600**, the umbrella retention issue. #903's own title records that the disk pressure motivating it was already relieved Kept deliberately from the retention cluster because they are distinct problems rather than restatements: **#905** (attachments, 39.9 GiB, has specific data), **#624** (upload staging mechanism), **#733** (backup rollout). ## Closed as stale Three or more months old, priority P3, no comments, no activity. All still plausible, none currently moving, all recoverable by label. **#46, #48, #110, #157, #184, #185, #253, #293** and **#689**. #689 is a verification task for work that has since been superseded. The rest are dormant improvements rather than live faults. ## Deliberately kept open, because they are failing right now Recording these so a future burn-down does not sweep them by age. Each was verified live today. * **#765** - Tailscale operator OAuth. Failing continuously for 22 days, last seen 21:07Z today. **Raised P3 to P2.** The pod is `1/1 Running` with 0 restarts, which is exactly why nobody noticed * **#837** - Flux `infrastructure` and `bridge-deploy` still suspended, confirmed by `kubectl get kustomization -A` * **#771** - the untracked `sirens-discord-ops` runner still exists, 1/1, 31d * **#869** - the recycle cron timezone bug, confirmed and **worse than filed**: kai-server fires 16:15Z while ser8 fires 09:15Z, so the two clusters recycle seven hours apart * **#971** - Route53 `REGISTER_DOMAIN`, six failures confirmed today ## Standing issues, re-anchored rather than closed * **#870** - kai-server capacity, now carries today's 77.9 percent * **#916** - ser8 capacity, now carries today's 47.3 percent The point-in-time snapshots closed into these two, so the concern survives without the backlog asserting stale numbers. ## What this sweep did not do Merge clusters that need a judgement call about which issue keeps the detail, specifically the runner-recycle group (#868, #840, #706) and the runner-scope group (#956, #955, #682, #693). Each contains real distinct content, and collapsing them would lose more than it saves. The closes above are reversible by design. If any of them were wrong, the label recovers them as a set rather than one at a time.
Author
Owner

Corrections to the plan above, after reading the issues instead of their titles

Two entries in the original body are wrong. Both errors are the same shape: I planned the merges from titles and priorities, and reading the bodies reversed the answer twice.

#903 was NOT closed, and should not be

The body lists it as "subsumed by #600" on the strength of its title saying the disk pressure was relieved. Reading it, #903 is the most actionable issue in the retention cluster, not a duplicate of anything:

  • Complete measured attribution of 133 GiB of Forgejo application data, packages at 93.3 GB and attachments at 39.4 GiB
  • The decisive finding that zero blobs are unreferenced, so garbage collection can reclaim nothing and retention is the only available lever. That single fact is why the problem cannot self-heal
  • 2,039 package versions with older_than_30d at zero for every owner, accumulating at roughly 3 GB/day
  • A four-stage remediation plan where stage 1 is config-only and needs one decision, the keep_count per container package
  • It reframes agentic-os#1060, showing the dev-base publish is blocked by 695 prior versions of the image it is trying to push

#903 stays open. If anything it is the survivor of that cluster rather than a casualty, and #600 is the vaguer umbrella.

#680 was closed, and it was the wrong survivor

Closed as a duplicate of #505 on the heuristic "keep the one with comments." Wrong. #505 is the original charter; #680 was the continuation, carrying the completed work (the stack already scaled to zero) and a seven-item remaining checklist.

The guarded Forgejo surface has close but no reopen, so I could not undo it. Recovered by migrating #680's full current-state and acceptance list into #505 (comment 80264). Nothing is lost, but the close itself was a mistake.

#48 also stays open

Listed for stale-closure. Keeping it, because #903 independently records k3s/server/db at 8.5 GiB with a note that "this estate has had sustained event storms", which is live corroboration for the kine watchdog cascade #48 describes. Dormant is not the same as irrelevant.

The lesson worth keeping

Every merge in a burn-down needs both bodies read, not both titles. Titles encode what the filer knew at filing time, and in a backlog this old that is frequently the least current thing about the issue. The two errors here cost one irreversible close and would have cost the retention cluster's best analysis.

## Corrections to the plan above, after reading the issues instead of their titles Two entries in the original body are wrong. Both errors are the same shape: I planned the merges from titles and priorities, and reading the bodies reversed the answer twice. ### #903 was NOT closed, and should not be The body lists it as "subsumed by #600" on the strength of its title saying the disk pressure was relieved. Reading it, **#903 is the most actionable issue in the retention cluster**, not a duplicate of anything: * Complete measured attribution of 133 GiB of Forgejo application data, packages at 93.3 GB and attachments at 39.4 GiB * The decisive finding that **zero blobs are unreferenced**, so garbage collection can reclaim nothing and retention is the only available lever. That single fact is why the problem cannot self-heal * 2,039 package versions with `older_than_30d` at zero for every owner, accumulating at roughly 3 GB/day * A four-stage remediation plan where stage 1 is config-only and needs one decision, the `keep_count` per container package * It reframes agentic-os#1060, showing the dev-base publish is blocked by 695 prior versions of the image it is trying to push **#903 stays open.** If anything it is the survivor of that cluster rather than a casualty, and #600 is the vaguer umbrella. ### #680 was closed, and it was the wrong survivor Closed as a duplicate of #505 on the heuristic "keep the one with comments." Wrong. #505 is the original charter; **#680 was the continuation**, carrying the completed work (the stack already scaled to zero) and a seven-item remaining checklist. The guarded Forgejo surface has `close` but **no reopen**, so I could not undo it. Recovered by migrating #680's full current-state and acceptance list into #505 (comment 80264). Nothing is lost, but the close itself was a mistake. ### #48 also stays open Listed for stale-closure. Keeping it, because #903 independently records `k3s/server/db` at 8.5 GiB with a note that "this estate has had sustained event storms", which is live corroboration for the kine watchdog cascade #48 describes. Dormant is not the same as irrelevant. ### The lesson worth keeping Every merge in a burn-down needs both bodies read, not both titles. Titles encode what the filer knew at filing time, and in a backlog this old that is frequently the least current thing about the issue. The two errors here cost one irreversible close and would have cost the retention cluster's best analysis.
Author
Owner

Continuation, 2026-08-29 ~04:20Z. One correction to this record, and the two deferred clusters resolved.

Correction: the guarded surface does have reopen

This record states "the guarded Forgejo surface has close but no reopen, so I could not undo it."

It has one. aosguard ops forgejo issue reopen <owner> <repo> <index>, authorized by can reopen issue, always sending the fixed body {"state": "open"}. So the #680 close was recoverable at the time and is recoverable now.

Not reopening #680. Its full current state and acceptance list were migrated into #505 in comment 80264, so reopening it today would restore a duplicate rather than recover anything. The close stays wrong for the reason this record already gives, and the recovery stands. What changes is only the claim that no undo existed, which would otherwise teach the next sweep to treat every close as irreversible.

The two clusters this sweep deferred

Both were read in full and verified against the live clusters rather than judged from titles.

Runner-recycle group, #868 / #840 / #706. Not merged, and the sweep was right. These are three distinct defects that share one component: a coverage gap in the selector, a simultaneity problem in the scheduling, and a safety problem with active jobs. Collapsing them would lose two of the three.

  • #868 both halves still live at ten days. docker-lib-forgejo-runner-build-0 still Bound with zero mounts at 5.99 GB or more, and sirens-discord-ops still unrecycled, now provable by pod age rather than by label read: 6d3h against 12h2m for every labelled sibling.
  • #840 still live and wider than filed. The herd is now two herds, twelve on kai-server and eleven on ser8, because the general pools moved under #693.
  • #706 untouched here. It needs the 09:10Z actor identified, which is a log question rather than a state question.

Runner-scope group, #956 / #955 / #682 / #693. One close, one re-anchor, two confirmed accurate.

  • #955 closed. Every repository it named as unable to see a deploy runner now has an exact-repository runner live. The two exceptions are resolved rather than outstanding: atlas is gone with its repository, reddit-mcp was retired by coilyco-bridge/deploy#608. Evidence in comment 80435.
  • #693 re-anchored, not closed. The target topology is live. Org-scoped general pools and the image builder on ser8, only exact-repository deploy runners plus the tap writer on kai-server. Its done-condition is now a decision rather than a migration, specifically rollout step 7's retention call on 9.21 GB of scaled-zero PVCs. Evidence in comment 80439.
  • #682 and #956 both still accurate. The target is a four-pod idle floor behind an autoscaling controller. Observed today: 25 permanently warm runner pods across the two clusters and no controller. Neither is stale and neither should be swept. They are sequential rather than duplicate, since #956's own preconditions require #682 landed first.

A closure error worth the same treatment as #903

#846 was closed on half its content. Its title carries two findings: "~16 GB in unmounted runner PVCs" and "three PVs stuck Released". The Released half is genuinely fixed, 43 volumes all Bound, zero Released, and that is what this record verified. The unmounted-PVC half was never addressed and is still true, measured at 15.4 GB across eleven claims today.

Same shape as the #903 error already recorded here, one step further on: not a title read this time, but a two-claim title where verifying one claim closed both. The lesson generalizes. A compound title needs every claim checked, not the one the sweep happened to measure.

Handled by folding the full measured inventory into #868, which already owned the stranded-PVC decision, rather than reopening #846 whose other half really is resolved. One issue owns PVC reclaim instead of two. Evidence in comment 80441.

Net effect on the count

Infrastructure went 109 to 108 open. One close, four re-anchors, one correction. The re-anchors are the point rather than a consolation: four issues that a future age-based sweep would have read as stale now carry today's evidence and say plainly which are live, which is waiting on a decision, and which one is finished.

## Continuation, 2026-08-29 ~04:20Z. One correction to this record, and the two deferred clusters resolved. ### Correction: the guarded surface does have reopen This record states "the guarded Forgejo surface has `close` but **no reopen**, so I could not undo it." **It has one.** `aosguard ops forgejo issue reopen <owner> <repo> <index>`, authorized by `can reopen issue`, always sending the fixed body `{"state": "open"}`. So the #680 close was recoverable at the time and is recoverable now. **Not reopening #680.** Its full current state and acceptance list were migrated into #505 in comment 80264, so reopening it today would restore a duplicate rather than recover anything. The close stays wrong for the reason this record already gives, and the recovery stands. What changes is only the claim that no undo existed, which would otherwise teach the next sweep to treat every close as irreversible. ### The two clusters this sweep deferred Both were read in full and verified against the live clusters rather than judged from titles. **Runner-recycle group, #868 / #840 / #706. Not merged, and the sweep was right.** These are three distinct defects that share one component: a coverage gap in the selector, a simultaneity problem in the scheduling, and a safety problem with active jobs. Collapsing them would lose two of the three. * **#868** both halves still live at ten days. `docker-lib-forgejo-runner-build-0` still Bound with zero mounts at 5.99 GB or more, and `sirens-discord-ops` still unrecycled, now provable by pod age rather than by label read: 6d3h against 12h2m for every labelled sibling. * **#840** still live and **wider than filed**. The herd is now two herds, twelve on kai-server and eleven on ser8, because the general pools moved under #693. * **#706** untouched here. It needs the 09:10Z actor identified, which is a log question rather than a state question. **Runner-scope group, #956 / #955 / #682 / #693. One close, one re-anchor, two confirmed accurate.** * **#955 closed.** Every repository it named as unable to see a deploy runner now has an exact-repository runner live. The two exceptions are resolved rather than outstanding: atlas is gone with its repository, reddit-mcp was retired by `coilyco-bridge/deploy#608`. Evidence in comment 80435. * **#693 re-anchored, not closed.** The target topology is live. Org-scoped general pools and the image builder on ser8, only exact-repository deploy runners plus the tap writer on kai-server. **Its done-condition is now a decision rather than a migration**, specifically rollout step 7's retention call on 9.21 GB of scaled-zero PVCs. Evidence in comment 80439. * **#682 and #956 both still accurate.** The target is a four-pod idle floor behind an autoscaling controller. Observed today: **25 permanently warm runner pods across the two clusters and no controller.** Neither is stale and neither should be swept. They are sequential rather than duplicate, since #956's own preconditions require #682 landed first. ### A closure error worth the same treatment as #903 **#846 was closed on half its content.** Its title carries two findings: "~16 GB in unmounted runner PVCs" and "three PVs stuck Released". The Released half is genuinely fixed, 43 volumes all `Bound`, zero `Released`, and that is what this record verified. **The unmounted-PVC half was never addressed and is still true**, measured at 15.4 GB across eleven claims today. Same shape as the #903 error already recorded here, one step further on: not a title read this time, but a **two-claim title where verifying one claim closed both**. The lesson generalizes. A compound title needs every claim checked, not the one the sweep happened to measure. Handled by folding the full measured inventory into **#868**, which already owned the stranded-PVC decision, rather than reopening #846 whose other half really is resolved. One issue owns PVC reclaim instead of two. Evidence in comment 80441. ### Net effect on the count Infrastructure went 109 to 108 open. **One close, four re-anchors, one correction.** The re-anchors are the point rather than a consolation: four issues that a future age-based sweep would have read as stale now carry today's evidence and say plainly which are live, which is waiting on a decision, and which one is finished.
Author
Owner

Ambient findings parked here, per the director's instruction, plus the #846 reopen.

Portia decided the STATE: ambient and ephemeral work documents do not get created yet, and asked that ambient findings from this pass land here meanwhile. This issue already exists and is already scoped to the sweep, so nothing new was filed to hold them.

#846 is reopened

Recorded above that this sweep closed it on one of its three problems. Portia verified and gated the reopen, and I executed it with aosguard ops forgejo issue reopen. Retitled to the surviving scope, with complete measurements in comment 80513, and #870 warned in comment 80517 because it is the issue a hardware purchase would be priced from.

Correcting the correction I posted earlier today: I wrote that the reopen verb exists, which is true, and then said #680 was "recoverable after all" without noting that two seats had independently hit the missing verb. It is missing from the Forgejo MCP surface and present on aosguard. coilyco-bridge/deploy#395 is open to add can reopen issue to both MCP guardfiles for parity. So "close is one-way" is true for any seat on the MCP surface and false for this one, and that distinction is the thing worth carrying, not the bare existence of the verb.

Ambient: eight Failed pods on kai-server, no issue filed

None of these is a live fault. Each is a terminated pod object left behind by a rollout, holding no compute and serving nothing:

registry-65969b794c-wsv2w                    Failed   28d19h
steam-mcp-oauth2-proxy-7f7cfb699c-4stmd      Failed   28d19h
bluesky-mcp-oauth2-proxy-6667c9b7d8-r5h5n    Failed   28d19h
galaxy-gen-55dd49b4c8-2t48p                  Failed   14d6h
playwright-mcp-84b68d6b9b-ldcfn              Failed   13d20h
node-stats-mcp-db79b675d-6v6qr               Failed    8d18h
coilysiren-eco-app-app-7dcb7d8c8b-fpvgm      Failed    8d19h
coilysiren-eco-app-discord-58d94c74cf-4kr85  Failed    8d19h

Filed nowhere on purpose, under the #482 test: no owning seat wants them, and there is no done-condition a reader would check. What makes them worth writing down at all is second-order: they make every pod list read as partly broken, so the next person scanning for a real fault has eight false positives to dismiss first. That cost is real and it is not worth an issue.

The two coilysiren-eco-app entries are the pre-deploy#818 rollout and are named there already.

Ambient: three failed recycle Jobs, 36 to 38 days old

forgejo-runner-recycle Jobs from 2026-07-21, -22 and -23 persist with failed: 2 and BackoffLimitExceeded. The last four runs of that CronJob all succeeded. Noted on #869 for the same reason as above: they make the CronJob read as failing at a glance when it is not.

Ambient: external-dns is restarting

external-dns-77dc66dd47-v6jbc shows 22 restarts in 6d3h, roughly 3.6 a day. Not filed, because #971 already records six confirmed Route53 REGISTER_DOMAIN failures and this is plausibly the same fault seen from the pod side. Stated as a candidate link rather than a diagnosis - nobody has connected them, and whoever picks up #971 should check the restart count before assuming it is only an API-level failure.

## Ambient findings parked here, per the director's instruction, plus the #846 reopen. Portia decided the `STATE: ambient and ephemeral work` documents do not get created yet, and asked that ambient findings from this pass land here meanwhile. This issue already exists and is already scoped to the sweep, so nothing new was filed to hold them. ### #846 is reopened Recorded above that this sweep closed it on one of its three problems. Portia verified and gated the reopen, and I executed it with `aosguard ops forgejo issue reopen`. Retitled to the surviving scope, with complete measurements in comment 80513, and `#870` warned in comment 80517 because it is the issue a hardware purchase would be priced from. **Correcting the correction I posted earlier today:** I wrote that the reopen verb exists, which is true, and then said #680 was "recoverable after all" without noting that two seats had independently hit the missing verb. It is missing from the **Forgejo MCP** surface and present on **aosguard**. `coilyco-bridge/deploy#395` is open to add `can reopen issue` to both MCP guardfiles for parity. So "close is one-way" is true for any seat on the MCP surface and false for this one, and that distinction is the thing worth carrying, not the bare existence of the verb. ### Ambient: eight Failed pods on kai-server, no issue filed None of these is a live fault. Each is a terminated pod object left behind by a rollout, holding no compute and serving nothing: ``` registry-65969b794c-wsv2w Failed 28d19h steam-mcp-oauth2-proxy-7f7cfb699c-4stmd Failed 28d19h bluesky-mcp-oauth2-proxy-6667c9b7d8-r5h5n Failed 28d19h galaxy-gen-55dd49b4c8-2t48p Failed 14d6h playwright-mcp-84b68d6b9b-ldcfn Failed 13d20h node-stats-mcp-db79b675d-6v6qr Failed 8d18h coilysiren-eco-app-app-7dcb7d8c8b-fpvgm Failed 8d19h coilysiren-eco-app-discord-58d94c74cf-4kr85 Failed 8d19h ``` **Filed nowhere on purpose**, under the `#482` test: no owning seat wants them, and there is no done-condition a reader would check. What makes them worth writing down at all is second-order: **they make every pod list read as partly broken**, so the next person scanning for a real fault has eight false positives to dismiss first. That cost is real and it is not worth an issue. The two `coilysiren-eco-app` entries are the pre-`deploy#818` rollout and are named there already. ### Ambient: three failed recycle Jobs, 36 to 38 days old `forgejo-runner-recycle` Jobs from 2026-07-21, -22 and -23 persist with `failed: 2` and `BackoffLimitExceeded`. The last four runs of that CronJob all succeeded. Noted on `#869` for the same reason as above: they make the CronJob read as failing at a glance when it is not. ### Ambient: external-dns is restarting `external-dns-77dc66dd47-v6jbc` shows **22 restarts in 6d3h**, roughly 3.6 a day. Not filed, because `#971` already records six confirmed Route53 `REGISTER_DOMAIN` failures and this is plausibly the same fault seen from the pod side. **Stated as a candidate link rather than a diagnosis** - nobody has connected them, and whoever picks up #971 should check the restart count before assuming it is only an API-level failure.
Sign in to join this conversation.
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set

Reference
coilyco-flight-deck/infrastructure#981
No description provided.