ser8: root filesystem climbing ~10% -> ~54% over 30 days - attribute the growth #871

Closed
opened 2026-08-19 11:42:34 +00:00 by coilyco-ops · 2 comments
Owner

Filed by Claude (ops seat), from the same read-only disk-pressure pass on 2026-08-19. Nothing changed. Surfaced while pulling the kai-server disk trend (filed alongside this as the kai-server capacity issue); ser8 has the opposite shape and is worth its own look.

Signal

node_stats.filesystem.utilization{host.name=ser8}, daily avg, 2026-07-25 to 2026-08-19:

10 11 11 11 12 12 12 13 13 13 | 22 34 39 41 40 40 41 42 43 48 50 50 49 50 53 54

Unlike kai-server (flat, oscillating), ser8 climbs monotonically: flat at ~10-13% through 2026-08-03, then a step change starting 2026-08-04 (13% -> ~40% over three days), then a steady grind to ~54% by today. No self-correction.

Reading

Not urgent - 54% with plenty of headroom. The concern is the slope and the step: something got added around 2026-08-04 (~27 points of ser8's disk) and keeps growing. A monotonic climb with no oscillation reads as accumulation, not churn.

Candidates to check

ser8 is the warm-standby / DR node (Beelink) and runs Agent Proxy, Node Stats MCP, SigNoz MCP, and the ser8 SigNoz observability data plane. The prime suspect for steady growth is SigNoz retention (metrics/logs/traces accumulating), with the Agent Proxy trajectory ledger and any DR/backup landing as secondary candidates. The step around 2026-08-04 suggests a discrete enablement - a new data source, a retention bump, or a service that started writing.

Action

  • Attribute the growth with the node-stats-ser8 breakdown (var-lib and the SigNoz data dirs).
  • Determine whether it is expected retention growth or unbounded accumulation.
  • If unbounded, set a retention or size bound on the owning service.

Acceptance

  • The growing consumer on ser8 named and measured.
  • A retention or size bound set if the growth is unbounded, or a recorded note that the growth is expected and the ceiling it plateaus at.
Filed by Claude (ops seat), from the same read-only disk-pressure pass on 2026-08-19. Nothing changed. Surfaced while pulling the kai-server disk trend (filed alongside this as the kai-server capacity issue); ser8 has the **opposite shape** and is worth its own look. ## Signal `node_stats.filesystem.utilization{host.name=ser8}`, daily avg, 2026-07-25 to 2026-08-19: 10 11 11 11 12 12 12 13 13 13 | 22 34 39 41 40 40 41 42 43 48 50 50 49 50 53 54 Unlike kai-server (flat, oscillating), ser8 **climbs monotonically**: flat at ~10-13% through 2026-08-03, then a **step change starting 2026-08-04** (13% -> ~40% over three days), then a steady grind to ~54% by today. No self-correction. ## Reading Not urgent - 54% with plenty of headroom. The concern is the **slope and the step**: something got added around 2026-08-04 (~27 points of ser8's disk) and keeps growing. A monotonic climb with no oscillation reads as accumulation, not churn. ## Candidates to check ser8 is the warm-standby / DR node (Beelink) and runs Agent Proxy, Node Stats MCP, SigNoz MCP, and the **ser8 SigNoz observability data plane**. The prime suspect for steady growth is SigNoz retention (metrics/logs/traces accumulating), with the Agent Proxy trajectory ledger and any DR/backup landing as secondary candidates. The step around 2026-08-04 suggests a discrete enablement - a new data source, a retention bump, or a service that started writing. ## Action - Attribute the growth with the `node-stats-ser8` breakdown (var-lib and the SigNoz data dirs). - Determine whether it is expected retention growth or unbounded accumulation. - If unbounded, set a retention or size bound on the owning service. ## Acceptance - The growing consumer on ser8 named and measured. - A retention or size bound set if the growth is unbounded, or a recorded note that the growth is expected and the ceiling it plateaus at.
Author
Owner

Live-state check while triaging the disk cluster of issues. The numbers this was filed against have moved:

kai-server  /      480 GB   316 GB used   140 GB free   70.86%   status: ok
ser8        /      980 GB   421 GB used   518 GB free   44.90%

kai-server's DiskPressure condition has been False since 2026-08-23T00:35:35Z. Neither host is near a threshold.

Not closing this. A capacity-planning or attribution issue is not resolved by the current reading being comfortable, and the question it asks stays valid. But it should be worked from these numbers rather than the ones in its title, and it is not urgent.

One input that changed: the three orphaned local-path PVs stuck Released for 82 days are now deleted (#863, #901). They held no data on disk, so this reclaimed nothing, but they no longer distort a volume inventory.

Live-state check while triaging the disk cluster of issues. The numbers this was filed against have moved: ``` kai-server / 480 GB 316 GB used 140 GB free 70.86% status: ok ser8 / 980 GB 421 GB used 518 GB free 44.90% ``` kai-server's `DiskPressure` condition has been `False` since `2026-08-23T00:35:35Z`. Neither host is near a threshold. Not closing this. A capacity-planning or attribution issue is not resolved by the current reading being comfortable, and the question it asks stays valid. But it should be worked from these numbers rather than the ones in its title, and it is not urgent. One input that changed: the three orphaned local-path PVs stuck `Released` for 82 days are now deleted (#863, #901). They held no data on disk, so this reclaimed nothing, but they no longer distort a volume inventory.
Author
Owner

The growth reversed. Closing into #916.

Measured on ser8 2026-08-28: 47.3 percent used, 444.8 GB of 980.7 GB, 494.9 GB free.

This issue asks to attribute a climb from ~10 percent to ~54 percent over 30 days. The filesystem is now below where the climb was reported, so the trend this tracks is not currently happening.

Closing because a growth-attribution issue whose growth has gone away cannot be worked as written. The attribution question is real but needs a live trend to chase, and there isn't one right now.

#916 is the standing ser8 capacity issue and carries today's numbers, so the concern has somewhere to live.

Worth saying plainly

I did not attribute the reversal any more than this issue attributed the climb. An unexplained drop is exactly as weak a signal as an unexplained rise, and if it climbs again the right move is a fresh measurement with a real time series rather than reopening this.

Related and still open: #808, the ser8 gaming runner evicted for exceeding its 24Gi docker-lib EmptyDir, which is a plausible contributor to both directions and bounded per workload rather than per node.

## The growth reversed. Closing into #916. Measured on ser8 2026-08-28: **47.3 percent** used, 444.8 GB of 980.7 GB, 494.9 GB free. This issue asks to attribute a climb from ~10 percent to ~54 percent over 30 days. The filesystem is now **below** where the climb was reported, so the trend this tracks is not currently happening. Closing because a growth-attribution issue whose growth has gone away cannot be worked as written. The attribution question is real but needs a live trend to chase, and there isn't one right now. **#916 is the standing ser8 capacity issue** and carries today's numbers, so the concern has somewhere to live. ### Worth saying plainly I did not attribute the reversal any more than this issue attributed the climb. An unexplained drop is exactly as weak a signal as an unexplained rise, and if it climbs again the right move is a fresh measurement with a real time series rather than reopening this. Related and still open: #808, the ser8 gaming runner evicted for exceeding its 24Gi docker-lib EmptyDir, which is a plausible contributor to both directions and bounded per workload rather than per node.
coilyco-ops 2026-08-28 21:15:52 +00:00
Sign in to join this conversation.
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set

Reference
coilyco-flight-deck/infrastructure#871
No description provided.