investigate: recurring probe timeouts across stateful services #592
Labels
No labels
burndown-2026-06
autonomy
async-consult
autonomy
epic
autonomy
headless
autonomy
live-collab
coherence-core
priority
P0
priority
P1
priority
P2
priority
P3
priority
P4
qa-fixture
role/ai
role/creator
role/design
role/director
role/engineer
role/exec
role/human
role/ops
role/qa
No milestone
No project
No assignees
1 participant
Notifications
Due date
No due date set.
Dependencies
No dependencies set
Reference
coilyco-flight-deck/infrastructure#592
Loading…
Reference in a new issue
No description provided.
Delete branch "%!s()"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
Problem
The node has recorded recurring readiness and liveness probe timeouts for about 12 days across unrelated stateful workloads, including Forgejo database, GlitchTip Valkey, SigNoz Zookeeper, Eco GNOME database, and Open WebUI.
The affected workloads are currently Ready after the 2026-07-22 recovery, so this issue tracks chronic latency rather than a current outage. The shared node and overlapping DiskPressure history suggest host I/O or runtime latency should be measured before anyone widens probe thresholds.
Scope
Acceptance
Related to the node-pressure evidence in #560.