Watch
2
design: per-driver context-management spec + role-aware three-tier (proactive/immediate/peripheral) context probe #373
Closed
opened 2026-06-26 07:21:17 +00:00 by coilysiren
·
5 comments
No Branch/Tag specified
main
release
docs/readme-agents-surface
fix/pr-repair-verb-lookup
aos/claude/aw85-docs-bands
aos/claude/aw85-declare-band
aos/claude/bk79-autonomy-label-scope
fix/gofmt-runner
chore/umbra-rename
aos/claude/mg96-fm
claude/agents-temp-clone-note
aos/claude/qa57
remove-format-exec-gate-refusal
fix/ward-1649-ci-fixture
fix/detached-ci-exec
issue-1626-generic-agent-broker
issue-1177
issue-1160
issue-1501
ward-salvage/ward-12486bc7
ward-salvage/ward-babfa2ba
ward-salvage/ward-a832df03
ward-salvage/ward-85c795b2
ward-salvage/ward-b87f8859
issue-1484
ward-salvage/ward-86b72dcf
ward-salvage/ward-b7b26d1d
recovery/2026-07-28-triaged-branch-archive
recovery/2026-07-27-local-work
issue-1584
issue-1571
issue-1524
ward-salvage/ward-3800c2d1
ward-salvage/ward-70175da3
ward-salvage/ward-bcbfef78
issue-1298
issue-737-signoz-deferred
ward-salvage/ward-c6aa5da9
ward-salvage/ward-94dd7346
ward-salvage/ward-58aa3fe3
ward-salvage/ward-ff6c509f
ward-salvage/ward-e80f0460
ward-salvage/ward-3e2e4760
ward-salvage/ward-649addd7
ward-salvage/ward-890e4d29
ward-salvage/ward-645b750b
ward-salvage/ward-bf851a72
ward-salvage/ward-98c8652f
ward-salvage/ward-bb10b620
ward-salvage/ward-1ee9be1c
ward-salvage/ward-d18595f6
ward-salvage/ward-5f914692
ward-salvage/ward-6ca05cbd
ward-salvage/ward-14f676cf
ward-salvage/ward-de811c20
ward-salvage/ward-eba3e824
ward-salvage/ward-16eb4ad0
ward-salvage/ward-c58e9c43
ward-salvage/ward-7487270b
v0.890.0
v0.889.0
v0.888.0
v0.887.0
v0.886.0
v0.885.0
v0.884.0
v0.883.0
v0.882.0
v0.881.0
v0.880.0
v0.879.0
v0.878.0
v0.877.0
v0.876.0
v0.875.0
v0.874.0
v0.873.0
v0.872.0
v0.871.0
v0.870.0
v0.869.0
v0.868.0
v0.867.0
v0.866.0
v0.865.0
v0.864.0
v0.863.0
v0.862.0
v0.861.0
v0.860.0
v0.859.0
v0.858.0
v0.857.0
v0.856.0
v0.855.0
v0.854.0
v0.853.0
v0.852.0
v0.851.0
v0.850.0
v0.849.0
v0.848.0
v0.847.0
v0.846.0
v0.845.0
v0.844.0
v0.843.0
v0.842.0
v0.841.0
v0.840.0
v0.839.0
v0.838.0
v0.837.0
v0.836.0
v0.835.0
v0.834.0
v0.833.0
v0.832.0
v0.830.0
v0.831.0
v0.829.0
v0.828.0
v0.827.0
v0.826.0
v0.825.0
v0.824.0
v0.823.0
v0.822.0
v0.821.0
v0.820.0
v0.819.0
v0.818.0
v0.817.0
v0.816.0
v0.815.0
v0.814.0
v0.813.0
v0.812.0
v0.811.0
v0.810.0
v0.809.0
v0.808.0
v0.807.0
v0.806.0
v0.805.0
v0.804.0
v0.803.0
v0.802.0
v0.801.0
v0.800.0
v0.799.0
v0.798.0
v0.797.0
v0.796.0
v0.795.0
v0.794.0
v0.793.0
v0.792.0
v0.791.0
v0.790.0
v0.789.0
v0.788.0
v0.787.0
v0.786.0
v0.785.0
v0.784.0
v0.783.0
v0.782.0
v0.781.0
v0.780.0
v0.779.0
v0.778.0
v0.777.0
v0.775.0-tmp
v0.776.0
v0.775.0
v0.774.0
v0.773.0
v0.772.0
v0.771.0
v0.770.0
v0.769.0
v0.768.0
v0.767.0
v0.766.0
v0.765.0
v0.764.0
v0.763.0
v0.762.0
v0.761.0
v0.760.0
v0.759.0
v0.758.0
v0.757.0
v0.756.0
v0.755.0
v0.754.0
v0.753.0
v0.752.0
v0.751.0
v0.750.0
v0.749.0
v0.748.0
v0.747.0
v0.746.0
v0.745.0
v0.744.0
v0.743.0
v0.742.0
v0.741.0
v0.740.0
v0.739.0
v0.738.0
v0.737.0
v0.736.0
v0.735.0
v0.734.0
v0.733.0
v0.732.0
v0.731.0
v0.730.0
v0.729.0
v0.728.0
v0.727.0
v0.726.0
v0.725.0
v0.724.0
v0.723.0
v0.722.0
v0.721.0
v0.720.0
v0.719.0
v0.718.0
v0.717.0
v0.716.0
v0.715.0
v0.714.0
v0.713.0
v0.712.0
v0.711.0
v0.710.0
v0.709.0
v0.708.0
v0.707.0
v0.706.0
v0.705.0
v0.704.0
v0.703.0
v0.702.0
v0.701.0
v0.700.0
v0.699.0
v0.698.0
v0.697.0
v0.696.0
v0.695.0
v0.694.0
v0.693.0
v0.692.0
v0.691.0
v0.690.0
v0.689.0
v0.688.0
v0.687.0
v0.686.0
v0.685.0
v0.684.0
v0.683.0
v0.682.0
v0.681.0
v0.680.0
v0.679.0
v0.678.0
v0.677.0
v0.676.0
v0.675.0
v0.674.0
v0.673.0
v0.672.0
v0.671.0
v0.670.0
v0.669.0
v0.668.0
v0.667.0
v0.666.0
v0.665.0
v0.664.0
v0.663.0
v0.662.0
v0.661.0
v0.660.0
v0.659.0
v0.658.0
v0.657.0
v0.656.0
v0.655.0
v0.654.0
v0.653.0
v0.652.0
v0.651.0
v0.650.0
v0.649.0
v0.648.0
v0.647.0
v0.646.0
v0.645.0
v0.644.0
v0.643.0
v0.642.0
v0.641.0
v0.640.0
v0.639.0
v0.638.0
v0.637.0
v0.636.0
v0.635.0
v0.634.0
v0.633.0
v0.632.0
v0.631.0
v0.630.0
v0.629.0
v0.628.0
v0.627.0
v0.626.0
v0.625.0
v0.624.0
v0.623.0
v0.622.0
v0.621.0
v0.620.0
v0.619.0
v0.618.0
v0.617.0
v0.616.0
v0.615.0
v0.614.0
v0.613.0
v0.612.0
v0.611.0
v0.610.0
v0.609.0
v0.608.0
v0.607.0
v0.606.0
v0.605.0
v0.604.0
v0.603.0
v0.602.0
v0.601.0
v0.600.0
v0.599.0
v0.598.0
v0.597.0
v0.596.0
v0.595.0
v0.594.0
v0.593.0
v0.592.0
v0.591.0
v0.590.0
v0.589.0
v0.588.0
v0.587.0
v0.586.0
v0.585.0
v0.584.0
v0.583.0
v0.582.0
v0.581.0
v0.580.0
v0.579.0
v0.578.0
v0.577.0
v0.576.0
v0.575.0
v0.574.0
v0.573.0
v0.572.0
v0.571.0
v0.570.0
v0.569.0
v0.568.0
v0.567.0
v0.566.0
v0.565.0
v0.564.0
v0.563.0
v0.562.0
v0.561.0
v0.560.0
v0.559.0
v0.558.0
v0.557.0
v0.556.0
v0.555.0
v0.554.0
v0.553.0
v0.552.0
v0.551.0
v0.550.0
v0.549.0
v0.548.0
v0.547.0
v0.546.0
v0.545.0
v0.544.0
v0.543.0
v0.542.0
v0.541.0
v0.540.0
v0.539.0
v0.538.0
v0.537.0
v0.536.0
v0.535.0
v0.534.0
v0.533.0
v0.532.0
v0.531.0
v0.530.0
v0.529.0
v0.528.0
v0.527.0
v0.526.0
v0.525.0
v0.524.0
v0.523.0
v0.522.0
v0.521.0
v0.520.0
v0.519.0
v0.518.0
v0.517.0
v0.516.0
v0.515.0
v0.514.0
v0.513.0
v0.512.0
v0.511.0
v0.510.0
v0.509.0
v0.508.0
v0.507.0
v0.506.0
v0.505.0
v0.504.0
v0.503.0
v0.502.0
v0.501.0
v0.500.0
v0.499.0
v0.498.0
v0.497.0
v0.496.0
v0.495.0
v0.494.0
v0.493.0
v0.492.0
v0.491.0
v0.490.0
v0.489.0
v0.488.0
v0.487.0
v0.486.0
v0.485.0
v0.484.0
v0.483.0
v0.482.0
v0.481.0
v0.480.0
v0.479.0
v0.478.0
v0.477.0
v0.476.0
v0.475.0
v0.474.0
v0.473.0
v0.472.0
v0.471.0
v0.470.0
v0.469.0
v0.468.0
v0.467.0
v0.466.0
v0.465.0
v0.464.0
v0.463.0
v0.462.0
v0.461.0
v0.460.0
v0.459.0
v0.458.0
v0.457.0
v0.456.0
v0.455.0
v0.454.0
v0.453.0
v0.452.0
v0.451.0
v0.450.0
v0.449.0
v0.448.0
v0.447.0
v0.446.0
v0.445.0
v0.444.0
v0.443.0
v0.442.0
v0.441.0
v0.440.0
v0.439.0
v0.438.0
v0.437.0
v0.436.0
v0.435.0
v0.434.0
v0.433.0
v0.432.0
v0.431.0
v0.430.0
v0.429.0
v0.428.0
v0.427.0
v0.426.0
v0.425.0
v0.424.0
v0.423.0
v0.422.0
v0.421.0
v0.420.0
v0.419.0
v0.418.0
v0.417.0
v0.416.0
v0.415.0
v0.414.0
v0.413.0
v0.412.0
v0.411.0
v0.410.0
v0.409.0
v0.408.0
v0.407.0
v0.406.0
v0.405.0
v0.404.0
v0.403.0
v0.402.0
v0.401.0
v0.400.0
v0.399.0
v0.398.0
v0.397.0
v0.396.0
v0.395.0
v0.394.0
v0.393.0
v0.392.0
v0.391.0
v0.390.0
v0.389.0
v0.388.0
v0.387.0
v0.386.0
v0.385.0
v0.384.0
v0.383.0
v0.382.0
v0.381.0
v0.380.0
v0.379.0
v0.378.0
v0.377.0
v0.376.0
v0.375.0
v0.374.0
v0.373.0
v0.372.0
v0.371.0
v0.370.0
v0.369.0
v0.368.0
v0.367.0
v0.366.0
v0.365.0
v0.364.0
v0.363.0
v0.362.0
v0.361.0
v0.360.0
v0.359.0
v0.358.0
v0.357.0
v0.356.0
v0.355.0
v0.354.0
v0.353.0
v0.352.0
v0.351.0
v0.350.0
v0.349.0
v0.348.0
v0.347.0
v0.346.0
v0.345.0
v0.344.0
v0.343.0
v0.342.0
v0.341.0
v0.340.0
v0.339.0
v0.338.0
v0.337.0
v0.336.0
v0.335.0
v0.334.0
v0.333.0
v0.332.0
v0.331.0
v0.330.0
v0.329.0
v0.328.0
v0.327.0
v0.326.0
v0.325.0
v0.324.0
v0.323.0
v0.322.0
v0.321.0
v0.320.0
v0.319.0
v0.318.0
v0.317.0
v0.316.0
v0.315.0
v0.314.0
v0.313.0
v0.312.0
v0.311.0
v0.310.0
v0.309.0
v0.308.0
v0.307.0
v0.306.0
v0.305.0
v0.304.0
v0.303.0
v0.302.0
v0.301.0
v0.300.0
v0.299.0
v0.298.0
v0.297.0
v0.296.0
v0.295.0
v0.294.0
v0.293.0
v0.292.0
v0.291.0
v0.290.0
v0.289.0
v0.288.0
v0.287.0
v0.286.0
v0.285.0
v0.284.0
v0.283.0
v0.282.0
v0.281.0
v0.280.0
v0.279.0
v0.278.0
v0.277.0
v0.276.0
v0.275.0
v0.274.0
v0.273.0
v0.272.0
v0.271.0
v0.270.0
v0.269.0
v0.268.0
v0.267.0
v0.266.0
v0.265.0
v0.264.0
v0.263.0
v0.262.0
v0.261.0
v0.260.0
v0.259.0
v0.258.0
v0.257.0
v0.256.0
v0.255.0
v0.254.0
v0.253.0
v0.252.0
v0.251.0
v0.250.0
v0.249.0
v0.248.0
v0.247.0
v0.246.0
v0.245.0
v0.244.0
v0.243.0
v0.242.0
v0.241.0
v0.240.0
v0.239.0
v0.238.0
v0.237.0
v0.236.0
v0.235.0
v0.234.0
v0.233.0
v0.232.0
v0.231.0
v0.230.0
v0.229.0
v0.228.0
v0.227.0
v0.226.0
v0.225.0
v0.224.0
v0.223.0
v0.222.0
v0.221.0
v0.220.0
v0.219.0
v0.218.0
v0.217.0
v0.216.0
v0.215.0
v0.214.0
v0.213.0
v0.212.0
v0.211.0
v0.210.0
v0.209.0
v0.208.0
v0.207.0
v0.206.0
v0.205.0
v0.204.0
v0.203.0
v0.202.0
v0.201.0
v0.200.0
v0.199.0
v0.198.0
v0.197.0
v0.196.0
v0.195.0
v0.194.0
v0.193.0
v0.192.0
v0.191.0
v0.190.0
v0.189.0
v0.188.0
v0.187.0
v0.186.0
v0.185.0
v0.184.0
v0.183.0
v0.182.0
v0.181.0
v0.180.0
v0.179.0
v0.178.0
v0.177.0
v0.176.0
v0.175.0
v0.174.0
v0.173.0
v0.172.0
v0.171.0
v0.170.0
v0.169.0
v0.168.0
v0.167.0
v0.166.0
v0.165.0
v0.164.0
v0.163.0
v0.162.0
v0.161.0
v0.160.0
v0.159.0
v0.158.0
v0.157.0
v0.156.0
v0.155.0
v0.154.0
v0.153.0
v0.152.0
v0.151.0
v0.150.0
v0.149.0
v0.148.0
v0.147.0
v0.146.0
v0.145.0
v0.144.0
v0.143.0
v0.142.0
v0.141.0
v0.140.0
v0.139.0
v0.138.0
v0.137.0
v0.136.0
v0.135.0
v0.134.0
v0.133.0
v0.132.0
v0.131.0
v0.130.0
v0.129.0
v0.128.0
v0.127.0
v0.126.0
v0.125.0
v0.124.0
v0.123.0
v0.122.0
v0.121.0
v0.120.0
v0.119.0
v0.118.0
v0.117.0
v0.116.0
v0.115.0
v0.114.0
v0.113.0
v0.112.0
v0.111.0
v0.110.0
v0.109.0
v0.108.0
v0.107.0
v0.106.0
v0.105.0
v0.104.0
v0.103.0
v0.102.0
v0.101.0
v0.100.0
v0.99.0
v0.98.0
v0.97.0
v0.96.0
v0.95.0
v0.94.0
v0.93.0
v0.92.0
v0.91.0
v0.90.0
v0.89.0
v0.88.0
v0.87.0
v0.86.0
v0.85.0
v0.84.0
v0.83.0
v0.82.0
v0.81.0
v0.80.0
v0.79.0
v0.78.0
v0.77.0
v0.76.0
v0.75.0
v0.74.0
v0.73.0
v0.72.0
v0.71.0
v0.70.0
v0.69.0
v0.68.0
v0.67.0
v0.66.0
v0.65.0
v0.64.0
v0.63.0
v0.62.0
v0.61.0
v0.60.0
v0.59.0
v0.58.0
v0.57.0
v0.56.0
v0.55.0
v0.54.0
v0.53.0
v0.52.0
v0.51.0
v0.50.0
v0.49.0
v0.48.0
v0.47.0
v0.46.0
v0.45.0
v0.44.0
v0.43.0
v0.42.0
v0.41.0
v0.40.0
v0.39.0
v0.38.0
v0.37.0
v0.36.0
v0.35.0
v0.34.0
v0.33.0
v0.32.0
v0.31.0
v0.30.0
v0.29.0
v0.28.0
v0.27.0
v0.26.0
v0.25.0
v0.24.0
v0.23.0
v0.22.0
v0.21.0
v0.20.0
v0.19.0
v0.18.0
v0.17.0
v0.16.0
v0.15.0
v0.14.0
v0.13.0
v0.12.0
v0.11.0
v0.10.0
v0.9.0
v0.8.0
v0.7.0
v0.6.0
v0.5.8
v0.5.7
v0.5.6
v0.5.5
v0.5.4
v0.5.3
v0.5.2
v0.5.1
v0.5.0
v0.4.0
v0.3.0
v0.2.2
v0.2.1
v0.2.0
v0.1.3
v0.1.2
v0.1.1
v0.1.0
v0.0.18
v0.0.17
v0.0.16
v0.0.15
v0.0.14
v0.0.13
v0.0.12
v0.0.11
v0.0.10
v0.0.9
v0.0.8
v0.0.7
v0.0.6
v0.0.5
v0.0.4
v0.0.3
v0.0.2
v0.0.1
Labels
Clear labels
burndown-2026-06
Backlog burndown June 2026
pressure-test
Cold-read release pressure-test findings and coordination
sunday-sprint
Burn-down by Sunday 2026-06-07
autonomy
async-consult
A human needs to consult on the issue to upgrade it to headless
autonomy
epic
This issue has many units of sub work - its size makes it meaningfully exclusive with other autonomy types
autonomy
headless
The agent can perform the work on its own
autonomy
live-collab
The agent and the human need to work together in realtime
coherence-core
Core review set for the warded control plane coherence milestone. These issues form the release spine; adjacent milestone issues are stretch or supporting work.
priority
P0
priority tier
priority
P1
priority tier
priority
P2
priority tier
priority
P3
priority tier
priority
P4
priority tier
qa-fixture
Disposable issue admitted to the bounded Ward QA verification lane.
role/advocate
requires work from the Developer Advocate seat
role/director
requires work from the Portfolio Director seat
role/exec
requires work from the exec role
role/frontend
requires work from the Frontend Engineer seat
role/gamedev
requires work from the Game Developer seat
role/human
requires a person, and specifically not an agent seat
role/platform
requires work from the Platform Engineer seat
role/qa
requires work from the QA role
role/science
requires work from the Applied Scientist seat
role/sysadmin
requires work from the Systems Administrator seat
state
ambient
ambient and ephemeral work, held as a maintained document rather than a queue
No labels
burndown-2026-06
pressure-test
sunday-sprint
autonomy
async-consult
autonomy
epic
autonomy
headless
autonomy
live-collab
coherence-core
priority
P0
priority
P1
priority
P2
priority
P3
priority
P4
qa-fixture
role/advocate
role/director
role/exec
role/frontend
role/gamedev
role/human
role/platform
role/qa
role/science
role/sysadmin
state
ambient
Milestone
Clear milestone
No items
No milestone
Projects
Clear projects
No items
No project
Assignees
Clear assignees
No assignees
2 participants
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".
No due date set.
Dependencies
No dependencies set
Reference
coilyco-flight-deck/ward#373
Loading…
Reference in a new issue
No description provided.
Delete branch "%!s()"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
Motivation
Low-capability drivers (the qwen trivial tier, any small local model, a degraded
remote) fail not only when context is missing but when it is too much.
Three tiers of context load a driver, and each drowns a weak model differently:
(composed AGENTS/CLAUDE doctrine, every mounted skill's frontmatter, the
statusline, the issue-corpus index, the front-load seed block).
/workspace/<name>,one tool call away. A 3k-file repo is a real grep/attention surface even though
it is not in the prompt.
/substrate/<name>reference repos (8 today)the agent can grep without leaving its box. Cheap to provide, not free to a
small model that now has a much larger haystack to pick wrongly from.
The thesis: a low-capability warded driver is fundamentally incapable of its
task if any tier is too heavy, so each tier must be curated per driver. We want a
programmatic, agent-runnable probe that spins (or composes) a warded role and
reports all three tiers with numbers, plus a per-driver context-management
spec that declares and enforces the budget. This issue is design-first:
propose the probe architecture and the spec schema, do not lock a schema yet.
What already exists (do not rebuild)
check_context_budget(agentic-os,ward context-budget,agentic_os/pre_commit/check_context_budget.py,docs/context-budget.md).Measures the proactive axis only - composed doc + skill frontmatter + an
mcp server-count note - per harness (claude/codex/opencode), static,
host-side, chars/4 proxy, with
--checkfor CI and per-harness budgets.agent-adapters.yamlper-agentcontextLevel: 2|1|0(full/scoped/minimal),the
WARD_CONTEXT_LEVELladder the entrypoint composes against(
docs/agent-adapter-manifest.md). This is the only per-driver contextknob today: one integer, not a spec.
preclone-repos.txt, 8 repos) - the peripheral tier.context commands like goose ask use." This issue is the per-role, three-tier
generalization of that seed.
The gap
check_context_budgetmeasures claude vs codexvs opencode. It does not model the warded roles (engineer / advisor /
director / architect), which differ in
WARD_CONTEXT_LEVELcomposition andin live container-only additions (the container doctrine overlay, the architect
read-only section, the issue-corpus index ward#367, the statusline ward#294,
the front-load seed block). The user explicitly wants engineer / consultant
(advisor) / director measured as they actually land in the container.
(working-dir clone: file count, bytes, token-proxy) or peripherally-available
(substrate: repo count, total grep surface) tiers, which the thesis says matter
most for weak drivers.
contextLevelis one int, not a spec. There is no per-driver declarationof a budget across the three tiers, and no enforcement.
Proposed shape (design these, do not lock yet)
container and report: proactive (extend/reuse
check_context_budget's doc+skillaccounting plus the container-only additions), immediately-available (walk
/workspace/<name>: tracked-file count, bytes, token-proxy), peripherally-available (walk
/substrate/<name>: per-repo and total). Programmatic,--check-able, agent-runnable. Spin-real-container vs static-composition is anopen design fork - state the trade and pick.
agent-adapters.yamlfrom a single
contextLevelint to a per-tier budget block (proactive /immediate / peripheral), so each driver declares its ceiling and the probe's
--checkenforces it. The first schema instance sets the pattern everythingcopies - propose it, surface it for Kai to shape, do not autonomously commit a
locked schema.
authored in agentic-os (extend
check_context_budgetwith the working-dir +substrate tiers, measurement-only, reversible). The role-spin probe + the
spec schema live in ward, which owns roles, containers, substrate, and
WARD_CONTEXT_LEVEL. Cross-file an agentic-os companion for the measurement leg.Acceptance (design-first)
docs/<feature>.md) covering: the probearchitecture (real-spin vs compose), the per-tier accounting for each role, the
proposed per-driver spec schema, and the agentic-os/ward split.
static, where
--checkenforcement lives) rather than silently resolving them.agent-adapters.yamluntil Kai blesses it.Non-goals
Filed from a read-only architect session (design walk on per-driver context
curation). Generalizes agentic-os#269; builds on
docs/context-budget.md.🔎 ward agent advisor
ward agent advisor --driver clauderan a one-shot standard research pass on this question:Now I have enough grounding. The key structural fact I confirmed:
contextLevelis keyed per-driver (claude=2, codex=1, qwen=0, goose=2), while the actually-composed context is per-role (the entrypoint prependsAGENTS.container.md, conditionally appends doctrine by level, and injects the read-only block for architect). That driver-vs-role asymmetry is the crux the schema has to resolve. Writing the proposal.Design: role-aware three-tier context probe + per-driver context-management spec
This is a design walk, not an implementation. It proposes the probe architecture, a candidate spec schema, the real-spin vs static-compose trade, and the agentic-os/ward authoring split. It deliberately leaves the schema unlocked and surfaces the open forks at the end for Kai to shape.
The crux nobody has named yet: driver-keyed knob, role-keyed load
The single most important finding from the current wiring:
contextLevelis keyed per-driver. Inagent-adapters.yamlit is one int per harness (claude=2, codex=1, qwen=0, goose=2). It is the driver's declared appetite.entrypoint.sh:compose_contextalways prependsAGENTS.container.md, then appends host doctrine gated on the level, then injects the read-only architect block whenWARD_READONLY=1. So the same driver carries materially different proactive bytes inengineervsarchitect, before you even count the container-only additions (issue-corpus index ward#367, statusline ward#294, front-load seed).So the load the thesis cares about is a (role x driver) cell, but the only ceiling we can declare today is a driver column. Any spec has to decide how a per-driver ceiling is checked against a per-role load. That decision is the schema, and it is fork #1 below. Everything else is plumbing.
The three tiers, as they actually land in the container
~/.claude/CLAUDE.md(for goose, the mirrored.goosehints) plus every mounted skill's frontmatter plus the statusline plus the issue-corpus index plus the front-load seed. This is the only tier that is literally bytes-in-the-prompt, so it is the only one where a token-proxy is a true "loaded" number.check_context_budgetalready measures the doc+skill+mcp slice of this per harness, statically, host-side. The gap is the container-only additions (theAGENTS.container.mdoverlay, the read-only block, ward#367, ward#294, the seed) which only exist once the container composes./workspace/<target>clone plus any--repo-granted working copies. Not in the prompt, one tool call away. The honest unit here is haystack size: tracked-file count and total bytes, with a token-proxy as a secondary number. A 3k-file repo is a real attention surface for a weak driver even at zero prompt cost./substrate/<name>references. The manifest is 8 repos today, minus the target when the target is also on the manifest (the entrypoint skips re-warming it). Same unit as immediate (repo count, total grep surface), one notch cheaper to provide and one notch easier for a small model to pick wrongly from.A clean way to think about the units: proactive is measured in loaded tokens (hard, prompt-resident), immediate and peripheral are measured in reachable surface (file count + bytes, token-proxy advisory). Conflating them into one "token" number is a trap, because 600k tokens of greppable substrate is not remotely the same load as 6k tokens of prompt. The spec should keep the units distinct.
Probe architecture
A single agent-runnable entrypoint, proposed surface
ward context-probe --role <r> --driver <d> [--check], emitting a structured report per tier:The report is the deliverable;
--checkcompares each tier against the spec's ceiling for that driver and exits non-zero on breach. The accounting primitives (count tracked files, sum bytes, chars/4 token-proxy, parse skill frontmatter) are the same onescheck_context_budgetalready uses for the proactive doc - the probe reuses them and points them at three roots instead of one.The fork that gates everything else: real-spin vs static-compose
compose_context's logic host-side and walk a clone, the waycheck_context_budgetalready models the proactive doc without a container. Fast, no docker, CI/pre-commit friendly, deterministic. Cost: it must re-implement the bash composition and every container-only addition, so it drifts the momententrypoint.shor ward#367/ward#294 changes, and it cannot truthfully measure immediate/peripheral without actually cloning the workspace and warming substrate (at which point it is not really "static" anymore).ward context-probeboots the actual container through the real entrypoint withWARD_CONTEXT_LEVELandWARD_READONLYset for the role, then a probe step measures the real~/.claude/CLAUDE.md, the real/workspace, the real/substrate. Truthful by construction - it measures the artifact, not a model of it, so it captures every container-only addition for free and never drifts. Cost: needs the docker socket, the gitcache/network for substrate warm, is slower, and is non-deterministic across runs as substrate repos grow.Recommendation: hybrid, with real-spin as the source of truth. ward already owns container spin, so the real-spin probe is cheap for ward to build and is the only thing that measures immediate/peripheral honestly. Keep the existing fast static path in
check_context_budgetfor the proactive tier only, as the pre-commit-speed gate that catches doctrine bloat on every push. Run the full three-tier real-spin probe on a docker-capable CI lane or cron, not in pre-commit. Stating the trade plainly: two measurement paths means the static proactive number can disagree with the real one (it will under-count by exactly the container-only additions), so the static path must be documented as a lower-bound smoke check, not the authority. Whether Kai accepts two paths or forces one is fork #2.Candidate spec schema (NOT for commit - shape this, Kai)
Extend each
agent-adapters.yamlentry with an optionalcontextblock. KeepcontextLevelexactly as-is so nothing breaks and the composition ladder is untouched - the new block is a ceiling the probe enforces, not a new composition input.Design intent baked into this shape, each a thing to accept or reject:
(role, driver)cell against this single column and flags any role whose composed load exceeds it. This keeps the schema flat and matches wherecontextLevelalready lives. The alternative - a full(role x driver)matrix - is more precise but multiplies the surface 4x and is the heavier first-pattern to copy.proactivein tokens (real, loaded),immediate/peripheralinfilesand/or token-proxy (reachable). Mixing them into one number is the trap above.context:block is unbudgeted (today's behavior), so adoption is incremental and claude/goose can stay unbudgeted while qwen gets a tight one first.I am deliberately not committing this. The issue is right that the first schema instance sets the pattern everything copies, and the per-driver-vs-matrix choice (fork #1) changes the shape fundamentally.
Authoring-vs-rollout split (per aos doctrine)
check_context_budget/agentic_os/pre_commit/check_context_budget.pyso the tier-accounting functions (tracked-file count, byte sum, token-proxy, frontmatter parse) are reusable against an arbitrary path, not just the composed doc. This is measurement-only, reversible, host-side, and stays in the catalog where validators live. It does not learn about ward's roles or containers. A companion issue cross-files this leg into agentic-os.compose_context, the/workspaceclone, the substrate warm, andWARD_CONTEXT_LEVEL. So the role-spin probe (ward context-probe), the schema inagent-adapters.yaml, and the--checkenforcement live here. ward calls the aos primitives and points them at the three real roots.This is the same split the manifest comment already describes: aos publishes the validator, ward embeds and orchestrates. The probe is "how to measure a tier" (aos) wired to "what the tiers are and what budget to enforce" (ward).
One scope flag that falls out of the peripheral tier: the thesis says each tier must be curated per driver, but
/substrateis one sharedpreclone-repos.txtfor all drivers today. A peripheral budget number only measures and warns. Actually curating peripheral per-driver means a driver-aware substrate manifest (a qwen run warms 3 repos, a claude run warms 8), which is a bigger change than this issue's measurement scope and should be its own ticket. This design covers measure-and-enforce-a-ceiling, not driver-filtered substrate seeding.Open forks for Kai (resolve before any commit)
(role x driver)matrix vs per-driver with per-role overrides. And the unit for immediate/peripheral: file-count, token-proxy, or grep-surface bytes - and are those tiers hard--checkfailures or warn-only while only proactive blocks.--checkgates. Fast static proactive in pre-commit, full three-tier real-spin in a docker-capable CI/cron lane, or both. And the sharper question: does a budget breach only fail CI, or does the entrypoint refuse to run a driver in a role whose composed load exceeds that driver's ceiling at runtime.contextLevel. Does the budget block coexist withcontextLevel(level composes, budget enforces - proposed) or eventually subsume it (derive the int from the budget).preclone-repos.txt) in scope, or does this issue stop at measure-and-budget and curation gets its own ticket (recommended).Acceptance check against the issue
This proposal gives the probe architecture (real-spin authoritative, static proactive smoke check), the per-tier accounting for each role as it actually lands in the container, a candidate per-driver schema presented unlocked, and the aos/ward split with a companion cross-file named. The forks above are surfaced for Kai rather than silently resolved, and no schema is committed to
agent-adapters.yaml.Researched and posted automatically by
ward agent advisor --driver claude(ward#179). This is one-shot research, not a carried change - verify before acting on it.— Claude (she/her), via
ward agentCanonical for #372 (merged). Flipped headless: agent drafts the per-driver context-management spec + three-tier probe design as a doc PR; Kai reviews there. Recorded by Claude Code (Fable) during the 2026-07-01 ward launch triage session with Kai.
🔒 Reserved by
ward agent --driver claude— containerengineer-claude-ward-373on hostKAI-DESKTOP-TOWERis carrying this issue (reserved 2026-07-03T06:24:44Z). Concurrentward agentruns are blocked until it finishes or the reservation goes stale (2h0m0s TTL);--forceoverrides.— Claude (she/her), via
ward agentDesign proposal: role-aware three-tier context probe + per-driver spec
Landed the design as
docs/context-probe.md(kept under the repo's 80-line/4000-char doc cap). This comment carries the fuller accounting the doc points back to. No schema is committed - the fleet KDL is untouched; the block below is proposed for you to shape.The three tiers, precisely
compose_contextinentrypoint.sh:AGENTS.container.mdalways, then perWARD_CONTEXT_LEVEL(L2 = targetCLAUDE.md+AGENTS.md, L1 =AGENTS.md, L0 = nothing), then role overlays (read-only block, front-load seed, statusline, issue-corpus index), plus mounted-skill frontmatter. Measure: chars/4, reusingcheck_context_budget's doc+skill accounting and adding the container-only overlays that never exist in the host-side load point./workspace/<name>:git ls-fileswalk -> tracked-file count, bytes, token proxy. Tracked-only so build/vendor trees do not inflate it./substrate/<name>frompreclone-repos.txt(8 today): same walk per-repo + a total.Per-role accounting
Tiers 2/3 are role-invariant (same clone, same substrate). Roles diverge in the proactive tier only:
The driver axis (claude L2 / codex L1 / opencode L0) then scales the doctrine portion.
Probe architecture (recommendation, one of the open forks)
Static-compose by default +
--spinverifier. The proactive tier is deterministic from inputs ward owns (embedded doctrine, the level ladder, the overlay strings viareadOnlyContextBlock/agentSeedPrompt), so a host-side probe reproduces it with no Docker. Tiers 2/3 are agit ls-fileswalk.--spinboots one real container per (role, driver) cell and diffs measured vs. predicted, catching composition drift. Reuse ward's own compose functions so drift is a test failure, not a silent wrong number.Proposed per-driver spec (NOT committed)
The issue names
agent-adapters.yaml, but that mirror was retired in ward#419 - the live source isward-kdl.fleet.kdl(dialect 2). Candidate additive block:Absent block = unbudgeted (backward-compatible).
Authoring split + companion
check_context_budgetwith the working-dir + substrate walkers, measurement-only).WARD_CONTEXT_LEVEL).Open forks for you (not silently resolved)
peripheralgates the manifest, per-driver vs per-(driver, role). First instance sets the pattern.--spin.--checklives - ward verb, aos hook, or split (aos cannot reach up into ward's role model).WARD-OUTCOME: done - landed docs/context-probe.md (three-tier role-aware probe + proposed per-driver context-budget block, schema unlocked) and cross-filed the measurement leg as agentic-os#319.
Design-first felt like the right call here - the real work was reading how context actually composes (the
compose_contextladder, the read-only overlay, the front-load seed) so the tier accounting matched what lands in a container rather than what the docs claim. Two things fought back. First, the issue namesagent-adapters.yamlas the spec home, but that mirror was retired in ward#419 - the live source is the fleet KDL, so I proposed the block there and flagged the staleness. Second, and more annoyingly, the repo-wide code-comments hook was already red on main from the recent Windows-guard commits, which blocked my commit without--no-verify. I trimmed those blocks - then discovered mid-rebase that #168 had independently trimmed the exact same lines, so I dropped my redundant commit and kept just the doc.Confidence is high on the design being a faithful, useful map and on the tier boundaries. I deliberately did not lock the schema - the KDL block, real-spin-vs-static, and where
--checkenforcement lives are all named as forks for you rather than resolved. The one rough edge worth a look: the 4000-char doc cap forced the durable doc to be a summary with the full accounting in the issue thread, which is a slightly awkward split if the thread ages out of attention. If the design survives your read, the natural follow-ups are the two implementation issues (the aos walker in #319, the ward probe verb) once you settle the forks.