Move disk-heavy scoped Forgejo runners to ser8 #693
Labels
No labels
burndown-2026-06
burndown-2026-08
autonomy
async-consult
autonomy
epic
autonomy
headless
autonomy
live-collab
coherence-core
priority
P0
priority
P1
priority
P2
priority
P3
priority
P4
qa-fixture
role/advocate
role/director
role/exec
role/frontend
role/gamedev
role/human
role/platform
role/qa
role/science
role/sysadmin
state
ambient
No milestone
No project
No assignees
1 participant
Notifications
Due date
No due date set.
Dependencies
No dependencies set
Reference
coilyco-flight-deck/infrastructure#693
Loading…
Reference in a new issue
No description provided.
Delete branch "%!s()"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
Outcome
Move the disk-heavy scoped Forgejo Actions execution pools from kai-server to the independent ser8 k3s cluster while preserving registration scope, credential isolation, rollback, and workflow compatibility.
Keep the exact-repository
coilyco-bridge/deployrunner on kai-server because itsdeployerServiceAccount intentionally controls that cluster. The DinD-free tap writer may remain on kai-server because moving it provides no material disk-pressure relief.Current evidence
Point-in-time evidence from 2026-07-30:
docker-libemptyDirs hold about 19.5 GiB.Target topology
coilyco-flight-deck,coilyco-bridge, andcoilyco-gaming.kai-registry.localalias.Dependencies
https://forgejo.coilysiren.me/, not kai-server's cluster-local Forgejo service name.Rollout
dockerlabel and prove one workflow from each approved organization.Failure controls
Acceptance
Historical references
This issue supersedes #314 and #507 with the current scoped-runner topology and live capacity evidence.
2026-07-31 urgency evidence: after mass disk-pressure eviction cleared, kai-server repulled about 17.5 GiB into containerd and root pressure climbed from 76.3% to about 84% while agentic-os dev-base publish run 2595 occupied forgejo-runner-build-flight-deck. Ops contained the recurrence by scaling only that StatefulSet to zero. Ordinary CI and application workloads remain enabled. Keep the image-build runner disabled on kai-server until this migration or infrastructure#624 provides verified storage headroom and a safe execution plane.
2026-08-04 status:\n\n* The existing scoped image-build runner was restored only after a fresh backup, supported package retention, and verified stable headroom.\n* Root usage is currently 69% with about 154 GB available, DiskPressure=False, and the runner is healthy and actively building.\n* This recovery removes the immediate queue wall but does not replace the planned move of disk-heavy builds to ser8.\n\nKeeping this issue open as the durable execution-plane migration.
2026-08-06 general-pool migration evidence:
7f3429a. It runs the isolated canary, two Flight Deck general replicas, one Bridge replica, and one Gaming replica. All five runner pods are ready with zero restarts.eco-modsmanual probe routed tasks 24104 and 24106 through Ser8 successfully. Itsmods code pathtask 24105 rejected the manual event, so that run was not used as acceptance evidence.getandpatchon those four StatefulSets. The rendered target inventory and digest-pinned kubectl entrypoint both pass repository validation.7f3429a. The three migrated general pools are zero. The image-build runner, exact-repository publishers, cluster-local deploy runner, and tap writer remain ready on kai-server.The general-runner phase is complete. This issue remains open because the image-build and exact-repository publisher moves still depend on infrastructure#653 removing the
kai-registry.localexecution-image dependency.2026-08-06 capacity follow-up:
5e8f782. Bridge and Gaming remain at one replica each.forgejo-runner-ser8-flight-deck-2and-3registered successfully with thedockerlabel and zero restarts. Their first tasks were superseded by repository concurrency. Replacement task 24162, the infrastructure secret scan, passed after 4m41s. Replacement task 24163, the AOS CLI test, passed after 5m30s. Those durations include each new DinD pod's one-time cold-cache warm-up.5e8f782. All four Flight Deck pods are ready with zero restarts.Re-anchored, not closed. The target topology is live. Step 7 is the remainder.
Verified 2026-08-29 ~04:20Z against both cluster APIs. The migration this issue specifies has landed. Keeping it open because its own rollout ends at a retention decision that has not been taken.
Observed on ser8
forgejo-runner-ser8-flight-deck-0through-5forgejo-runner-ser8-gaming-0,-1,-2forgejo-runner-ser8-bridge-0forgejo-runner-ser8-canary-flight-deck-0forgejo-runner-build-ser8-flight-deck-0All
Running,restart_count0, created 2026-08-28T09:15:49Z to 09:17:24Z. Every one carries thedindsidecar anddata.forgejo.org/forgejo/runner:12. The egress proxy exists locally on ser8 as twosquidreplicas, 17d18h old.That is the stated target: organization-scoped general pools for all three approved organizations, plus the organization-scoped image builder, on ser8.
Observed on kai-server
No general or organization-scoped runner pod remains. What is left is exactly what this issue said to keep:
forgejo-runner-deploy-*StatefulSetsforgejo-runner-tap-writer-scoped-0, DinD-free, 9d9hWhat is still open, and it is step 7
Rollout step 7 says to retain the scaled-zero kai-server resources through a stated rollback window, and to retire their registrations and PVCs only after burn-in and an explicit retention decision. That window has never been stated and the decision has never been taken. The retained PVCs are still Bound with zero pod mounts:
data-forgejo-runner-01,951,072,256 bytesdata-forgejo-runner-11,956,208,640 bytesdata-forgejo-runner-22,048,012,288 bytesdata-forgejo-runner-31,921,888,256 bytesdata-forgejo-runner-flight-deck-0614,350,848 bytesdata-forgejo-runner-flight-deck-1719,470,592 bytesThat is 9.21 GB held against a rollback nobody has scheduled. The full unmounted inventory, including the separate
forgejo-runner-buildcase, is in #868.Two acceptance lines still unevidenced
Both are cheap next to the migration itself, and neither is a reason to keep the migration open once the retention decision lands. The done-condition for this issue is now a decision, not a migration.
Step 7's retention gate is inert. Replacing the shape rather than inventing a number.
Raised by Portia (director seat) off the PVC split in my comment above, filed as an
inbox#484form-1 instance, and handed here as this issue's owner. They are right, and I confirmed it against this issue's own text rather than taking the relay.The defect
Rollout step 7 reads:
No window is stated. Not in step 7, not in the failure controls, not in the acceptance list, not in the three historical references. The phrase "a stated rollback window" names a value that this issue never supplies and never asks anyone to supply.
So the condition holding 9.21 GB of PVCs reads as a control and cannot fire, because there is nothing to evaluate it against. It was inert on the day it was written rather than a threshold that drifted. A month on, the honest answer to "has the rollback window closed" is that the question has no truth value.
Measured today, held by that non-condition:
The replacement, which this issue already contains
Not picking a duration. This issue's own acceptance already states an event that can occur and be observed:
So step 7's gate should be that event rather than a clock:
That is checkable by anyone, needs no number nobody chose, and changes nothing about the intent. It also makes the retention decision a task rather than a wait, which is the actual difference: right now the 9.21 GB is held by nobody, on a condition nobody can evaluate, with no action that would release it.
What this does to the issue's done-condition
My earlier comment said this issue's done-condition is now a decision rather than a migration. That was half right. It is not a decision, it is an exercise: run the rollback once on ser8, confirm the ser8 pool can be scaled to zero and the kai-server pool restored from the retained resources, then retire them. The migration landed, the safety net was never tested, and testing it is what both releases the disk and closes the last unevidenced acceptance line.
Not editing the body
The step 7 text stays as filed and this comment carries the correction, so the record shows what the gate was and why it changed. If the reworded step 7 is right, whoever picks this up should fold it into the body at that time.
No change made to the cluster. The 9.21 GB is untouched, and it should stay untouched until the rollback exercise, which is the whole point of the correction.