ser8 read-only watcher: alert on the lock-class crashloop signature on kai-server #315
Labels
No labels
burndown-2026-06
burndown-2026-08
autonomy
async-consult
autonomy
epic
autonomy
headless
autonomy
live-collab
coherence-core
priority
P0
priority
P1
priority
P2
priority
P3
priority
P4
qa-fixture
role/advocate
role/director
role/exec
role/frontend
role/gamedev
role/human
role/platform
role/qa
role/science
role/sysadmin
state
ambient
No milestone
No project
No assignees
1 participant
Notifications
Due date
No due date set.
Dependencies
No dependencies set
Reference
coilyco-flight-deck/infrastructure#315
Loading…
Reference in a new issue
No description provided.
Delete branch "%!s()"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
Mitigation 2 from #311. The o11y outage ran 45 days because the alerting stack was the casualty. ser8 holds a read-only kubeconfig for kai-server and sits in a different failure domain, so it watches for the deterministic signature: recent k3s restart plus CrashLoopBackOff plus a lock-class log error (flock.lock, boltdb open, address already in use, Another server instance). The alert carries the diagnosis and reap commands, not a generic pod-unhealthy. Interim shape: a small poller until the full o11y stack on ser8 takes over the job.
Backlog burndown 2026-06-17: closing low-priority (P3/P4) to bring the open count to a manageable level. Nothing lost — reopen if this resurfaces. Batch tag:
burndown-2026-06.