Watch
2
[epic] Stop the warded-workflow full stop: landability gate + convergence circuit-breaker + merge queue #1069
Closed
opened 2026-07-10 18:12:52 +00:00 by coilyco-ops
·
5 comments
No Branch/Tag specified
main
release
docs/readme-agents-surface
fix/pr-repair-verb-lookup
aos/claude/aw85-docs-bands
aos/claude/aw85-declare-band
aos/claude/bk79-autonomy-label-scope
fix/gofmt-runner
chore/umbra-rename
aos/claude/mg96-fm
claude/agents-temp-clone-note
aos/claude/qa57
remove-format-exec-gate-refusal
fix/ward-1649-ci-fixture
fix/detached-ci-exec
issue-1626-generic-agent-broker
issue-1177
issue-1160
issue-1501
ward-salvage/ward-12486bc7
ward-salvage/ward-babfa2ba
ward-salvage/ward-a832df03
ward-salvage/ward-85c795b2
ward-salvage/ward-b87f8859
issue-1484
ward-salvage/ward-86b72dcf
ward-salvage/ward-b7b26d1d
recovery/2026-07-28-triaged-branch-archive
recovery/2026-07-27-local-work
issue-1584
issue-1571
issue-1524
ward-salvage/ward-3800c2d1
ward-salvage/ward-70175da3
ward-salvage/ward-bcbfef78
issue-1298
issue-737-signoz-deferred
ward-salvage/ward-c6aa5da9
ward-salvage/ward-94dd7346
ward-salvage/ward-58aa3fe3
ward-salvage/ward-ff6c509f
ward-salvage/ward-e80f0460
ward-salvage/ward-3e2e4760
ward-salvage/ward-649addd7
ward-salvage/ward-890e4d29
ward-salvage/ward-645b750b
ward-salvage/ward-bf851a72
ward-salvage/ward-98c8652f
ward-salvage/ward-bb10b620
ward-salvage/ward-1ee9be1c
ward-salvage/ward-d18595f6
ward-salvage/ward-5f914692
ward-salvage/ward-6ca05cbd
ward-salvage/ward-14f676cf
ward-salvage/ward-de811c20
ward-salvage/ward-eba3e824
ward-salvage/ward-16eb4ad0
ward-salvage/ward-c58e9c43
ward-salvage/ward-7487270b
v0.890.0
v0.889.0
v0.888.0
v0.887.0
v0.886.0
v0.885.0
v0.884.0
v0.883.0
v0.882.0
v0.881.0
v0.880.0
v0.879.0
v0.878.0
v0.877.0
v0.876.0
v0.875.0
v0.874.0
v0.873.0
v0.872.0
v0.871.0
v0.870.0
v0.869.0
v0.868.0
v0.867.0
v0.866.0
v0.865.0
v0.864.0
v0.863.0
v0.862.0
v0.861.0
v0.860.0
v0.859.0
v0.858.0
v0.857.0
v0.856.0
v0.855.0
v0.854.0
v0.853.0
v0.852.0
v0.851.0
v0.850.0
v0.849.0
v0.848.0
v0.847.0
v0.846.0
v0.845.0
v0.844.0
v0.843.0
v0.842.0
v0.841.0
v0.840.0
v0.839.0
v0.838.0
v0.837.0
v0.836.0
v0.835.0
v0.834.0
v0.833.0
v0.832.0
v0.830.0
v0.831.0
v0.829.0
v0.828.0
v0.827.0
v0.826.0
v0.825.0
v0.824.0
v0.823.0
v0.822.0
v0.821.0
v0.820.0
v0.819.0
v0.818.0
v0.817.0
v0.816.0
v0.815.0
v0.814.0
v0.813.0
v0.812.0
v0.811.0
v0.810.0
v0.809.0
v0.808.0
v0.807.0
v0.806.0
v0.805.0
v0.804.0
v0.803.0
v0.802.0
v0.801.0
v0.800.0
v0.799.0
v0.798.0
v0.797.0
v0.796.0
v0.795.0
v0.794.0
v0.793.0
v0.792.0
v0.791.0
v0.790.0
v0.789.0
v0.788.0
v0.787.0
v0.786.0
v0.785.0
v0.784.0
v0.783.0
v0.782.0
v0.781.0
v0.780.0
v0.779.0
v0.778.0
v0.777.0
v0.775.0-tmp
v0.776.0
v0.775.0
v0.774.0
v0.773.0
v0.772.0
v0.771.0
v0.770.0
v0.769.0
v0.768.0
v0.767.0
v0.766.0
v0.765.0
v0.764.0
v0.763.0
v0.762.0
v0.761.0
v0.760.0
v0.759.0
v0.758.0
v0.757.0
v0.756.0
v0.755.0
v0.754.0
v0.753.0
v0.752.0
v0.751.0
v0.750.0
v0.749.0
v0.748.0
v0.747.0
v0.746.0
v0.745.0
v0.744.0
v0.743.0
v0.742.0
v0.741.0
v0.740.0
v0.739.0
v0.738.0
v0.737.0
v0.736.0
v0.735.0
v0.734.0
v0.733.0
v0.732.0
v0.731.0
v0.730.0
v0.729.0
v0.728.0
v0.727.0
v0.726.0
v0.725.0
v0.724.0
v0.723.0
v0.722.0
v0.721.0
v0.720.0
v0.719.0
v0.718.0
v0.717.0
v0.716.0
v0.715.0
v0.714.0
v0.713.0
v0.712.0
v0.711.0
v0.710.0
v0.709.0
v0.708.0
v0.707.0
v0.706.0
v0.705.0
v0.704.0
v0.703.0
v0.702.0
v0.701.0
v0.700.0
v0.699.0
v0.698.0
v0.697.0
v0.696.0
v0.695.0
v0.694.0
v0.693.0
v0.692.0
v0.691.0
v0.690.0
v0.689.0
v0.688.0
v0.687.0
v0.686.0
v0.685.0
v0.684.0
v0.683.0
v0.682.0
v0.681.0
v0.680.0
v0.679.0
v0.678.0
v0.677.0
v0.676.0
v0.675.0
v0.674.0
v0.673.0
v0.672.0
v0.671.0
v0.670.0
v0.669.0
v0.668.0
v0.667.0
v0.666.0
v0.665.0
v0.664.0
v0.663.0
v0.662.0
v0.661.0
v0.660.0
v0.659.0
v0.658.0
v0.657.0
v0.656.0
v0.655.0
v0.654.0
v0.653.0
v0.652.0
v0.651.0
v0.650.0
v0.649.0
v0.648.0
v0.647.0
v0.646.0
v0.645.0
v0.644.0
v0.643.0
v0.642.0
v0.641.0
v0.640.0
v0.639.0
v0.638.0
v0.637.0
v0.636.0
v0.635.0
v0.634.0
v0.633.0
v0.632.0
v0.631.0
v0.630.0
v0.629.0
v0.628.0
v0.627.0
v0.626.0
v0.625.0
v0.624.0
v0.623.0
v0.622.0
v0.621.0
v0.620.0
v0.619.0
v0.618.0
v0.617.0
v0.616.0
v0.615.0
v0.614.0
v0.613.0
v0.612.0
v0.611.0
v0.610.0
v0.609.0
v0.608.0
v0.607.0
v0.606.0
v0.605.0
v0.604.0
v0.603.0
v0.602.0
v0.601.0
v0.600.0
v0.599.0
v0.598.0
v0.597.0
v0.596.0
v0.595.0
v0.594.0
v0.593.0
v0.592.0
v0.591.0
v0.590.0
v0.589.0
v0.588.0
v0.587.0
v0.586.0
v0.585.0
v0.584.0
v0.583.0
v0.582.0
v0.581.0
v0.580.0
v0.579.0
v0.578.0
v0.577.0
v0.576.0
v0.575.0
v0.574.0
v0.573.0
v0.572.0
v0.571.0
v0.570.0
v0.569.0
v0.568.0
v0.567.0
v0.566.0
v0.565.0
v0.564.0
v0.563.0
v0.562.0
v0.561.0
v0.560.0
v0.559.0
v0.558.0
v0.557.0
v0.556.0
v0.555.0
v0.554.0
v0.553.0
v0.552.0
v0.551.0
v0.550.0
v0.549.0
v0.548.0
v0.547.0
v0.546.0
v0.545.0
v0.544.0
v0.543.0
v0.542.0
v0.541.0
v0.540.0
v0.539.0
v0.538.0
v0.537.0
v0.536.0
v0.535.0
v0.534.0
v0.533.0
v0.532.0
v0.531.0
v0.530.0
v0.529.0
v0.528.0
v0.527.0
v0.526.0
v0.525.0
v0.524.0
v0.523.0
v0.522.0
v0.521.0
v0.520.0
v0.519.0
v0.518.0
v0.517.0
v0.516.0
v0.515.0
v0.514.0
v0.513.0
v0.512.0
v0.511.0
v0.510.0
v0.509.0
v0.508.0
v0.507.0
v0.506.0
v0.505.0
v0.504.0
v0.503.0
v0.502.0
v0.501.0
v0.500.0
v0.499.0
v0.498.0
v0.497.0
v0.496.0
v0.495.0
v0.494.0
v0.493.0
v0.492.0
v0.491.0
v0.490.0
v0.489.0
v0.488.0
v0.487.0
v0.486.0
v0.485.0
v0.484.0
v0.483.0
v0.482.0
v0.481.0
v0.480.0
v0.479.0
v0.478.0
v0.477.0
v0.476.0
v0.475.0
v0.474.0
v0.473.0
v0.472.0
v0.471.0
v0.470.0
v0.469.0
v0.468.0
v0.467.0
v0.466.0
v0.465.0
v0.464.0
v0.463.0
v0.462.0
v0.461.0
v0.460.0
v0.459.0
v0.458.0
v0.457.0
v0.456.0
v0.455.0
v0.454.0
v0.453.0
v0.452.0
v0.451.0
v0.450.0
v0.449.0
v0.448.0
v0.447.0
v0.446.0
v0.445.0
v0.444.0
v0.443.0
v0.442.0
v0.441.0
v0.440.0
v0.439.0
v0.438.0
v0.437.0
v0.436.0
v0.435.0
v0.434.0
v0.433.0
v0.432.0
v0.431.0
v0.430.0
v0.429.0
v0.428.0
v0.427.0
v0.426.0
v0.425.0
v0.424.0
v0.423.0
v0.422.0
v0.421.0
v0.420.0
v0.419.0
v0.418.0
v0.417.0
v0.416.0
v0.415.0
v0.414.0
v0.413.0
v0.412.0
v0.411.0
v0.410.0
v0.409.0
v0.408.0
v0.407.0
v0.406.0
v0.405.0
v0.404.0
v0.403.0
v0.402.0
v0.401.0
v0.400.0
v0.399.0
v0.398.0
v0.397.0
v0.396.0
v0.395.0
v0.394.0
v0.393.0
v0.392.0
v0.391.0
v0.390.0
v0.389.0
v0.388.0
v0.387.0
v0.386.0
v0.385.0
v0.384.0
v0.383.0
v0.382.0
v0.381.0
v0.380.0
v0.379.0
v0.378.0
v0.377.0
v0.376.0
v0.375.0
v0.374.0
v0.373.0
v0.372.0
v0.371.0
v0.370.0
v0.369.0
v0.368.0
v0.367.0
v0.366.0
v0.365.0
v0.364.0
v0.363.0
v0.362.0
v0.361.0
v0.360.0
v0.359.0
v0.358.0
v0.357.0
v0.356.0
v0.355.0
v0.354.0
v0.353.0
v0.352.0
v0.351.0
v0.350.0
v0.349.0
v0.348.0
v0.347.0
v0.346.0
v0.345.0
v0.344.0
v0.343.0
v0.342.0
v0.341.0
v0.340.0
v0.339.0
v0.338.0
v0.337.0
v0.336.0
v0.335.0
v0.334.0
v0.333.0
v0.332.0
v0.331.0
v0.330.0
v0.329.0
v0.328.0
v0.327.0
v0.326.0
v0.325.0
v0.324.0
v0.323.0
v0.322.0
v0.321.0
v0.320.0
v0.319.0
v0.318.0
v0.317.0
v0.316.0
v0.315.0
v0.314.0
v0.313.0
v0.312.0
v0.311.0
v0.310.0
v0.309.0
v0.308.0
v0.307.0
v0.306.0
v0.305.0
v0.304.0
v0.303.0
v0.302.0
v0.301.0
v0.300.0
v0.299.0
v0.298.0
v0.297.0
v0.296.0
v0.295.0
v0.294.0
v0.293.0
v0.292.0
v0.291.0
v0.290.0
v0.289.0
v0.288.0
v0.287.0
v0.286.0
v0.285.0
v0.284.0
v0.283.0
v0.282.0
v0.281.0
v0.280.0
v0.279.0
v0.278.0
v0.277.0
v0.276.0
v0.275.0
v0.274.0
v0.273.0
v0.272.0
v0.271.0
v0.270.0
v0.269.0
v0.268.0
v0.267.0
v0.266.0
v0.265.0
v0.264.0
v0.263.0
v0.262.0
v0.261.0
v0.260.0
v0.259.0
v0.258.0
v0.257.0
v0.256.0
v0.255.0
v0.254.0
v0.253.0
v0.252.0
v0.251.0
v0.250.0
v0.249.0
v0.248.0
v0.247.0
v0.246.0
v0.245.0
v0.244.0
v0.243.0
v0.242.0
v0.241.0
v0.240.0
v0.239.0
v0.238.0
v0.237.0
v0.236.0
v0.235.0
v0.234.0
v0.233.0
v0.232.0
v0.231.0
v0.230.0
v0.229.0
v0.228.0
v0.227.0
v0.226.0
v0.225.0
v0.224.0
v0.223.0
v0.222.0
v0.221.0
v0.220.0
v0.219.0
v0.218.0
v0.217.0
v0.216.0
v0.215.0
v0.214.0
v0.213.0
v0.212.0
v0.211.0
v0.210.0
v0.209.0
v0.208.0
v0.207.0
v0.206.0
v0.205.0
v0.204.0
v0.203.0
v0.202.0
v0.201.0
v0.200.0
v0.199.0
v0.198.0
v0.197.0
v0.196.0
v0.195.0
v0.194.0
v0.193.0
v0.192.0
v0.191.0
v0.190.0
v0.189.0
v0.188.0
v0.187.0
v0.186.0
v0.185.0
v0.184.0
v0.183.0
v0.182.0
v0.181.0
v0.180.0
v0.179.0
v0.178.0
v0.177.0
v0.176.0
v0.175.0
v0.174.0
v0.173.0
v0.172.0
v0.171.0
v0.170.0
v0.169.0
v0.168.0
v0.167.0
v0.166.0
v0.165.0
v0.164.0
v0.163.0
v0.162.0
v0.161.0
v0.160.0
v0.159.0
v0.158.0
v0.157.0
v0.156.0
v0.155.0
v0.154.0
v0.153.0
v0.152.0
v0.151.0
v0.150.0
v0.149.0
v0.148.0
v0.147.0
v0.146.0
v0.145.0
v0.144.0
v0.143.0
v0.142.0
v0.141.0
v0.140.0
v0.139.0
v0.138.0
v0.137.0
v0.136.0
v0.135.0
v0.134.0
v0.133.0
v0.132.0
v0.131.0
v0.130.0
v0.129.0
v0.128.0
v0.127.0
v0.126.0
v0.125.0
v0.124.0
v0.123.0
v0.122.0
v0.121.0
v0.120.0
v0.119.0
v0.118.0
v0.117.0
v0.116.0
v0.115.0
v0.114.0
v0.113.0
v0.112.0
v0.111.0
v0.110.0
v0.109.0
v0.108.0
v0.107.0
v0.106.0
v0.105.0
v0.104.0
v0.103.0
v0.102.0
v0.101.0
v0.100.0
v0.99.0
v0.98.0
v0.97.0
v0.96.0
v0.95.0
v0.94.0
v0.93.0
v0.92.0
v0.91.0
v0.90.0
v0.89.0
v0.88.0
v0.87.0
v0.86.0
v0.85.0
v0.84.0
v0.83.0
v0.82.0
v0.81.0
v0.80.0
v0.79.0
v0.78.0
v0.77.0
v0.76.0
v0.75.0
v0.74.0
v0.73.0
v0.72.0
v0.71.0
v0.70.0
v0.69.0
v0.68.0
v0.67.0
v0.66.0
v0.65.0
v0.64.0
v0.63.0
v0.62.0
v0.61.0
v0.60.0
v0.59.0
v0.58.0
v0.57.0
v0.56.0
v0.55.0
v0.54.0
v0.53.0
v0.52.0
v0.51.0
v0.50.0
v0.49.0
v0.48.0
v0.47.0
v0.46.0
v0.45.0
v0.44.0
v0.43.0
v0.42.0
v0.41.0
v0.40.0
v0.39.0
v0.38.0
v0.37.0
v0.36.0
v0.35.0
v0.34.0
v0.33.0
v0.32.0
v0.31.0
v0.30.0
v0.29.0
v0.28.0
v0.27.0
v0.26.0
v0.25.0
v0.24.0
v0.23.0
v0.22.0
v0.21.0
v0.20.0
v0.19.0
v0.18.0
v0.17.0
v0.16.0
v0.15.0
v0.14.0
v0.13.0
v0.12.0
v0.11.0
v0.10.0
v0.9.0
v0.8.0
v0.7.0
v0.6.0
v0.5.8
v0.5.7
v0.5.6
v0.5.5
v0.5.4
v0.5.3
v0.5.2
v0.5.1
v0.5.0
v0.4.0
v0.3.0
v0.2.2
v0.2.1
v0.2.0
v0.1.3
v0.1.2
v0.1.1
v0.1.0
v0.0.18
v0.0.17
v0.0.16
v0.0.15
v0.0.14
v0.0.13
v0.0.12
v0.0.11
v0.0.10
v0.0.9
v0.0.8
v0.0.7
v0.0.6
v0.0.5
v0.0.4
v0.0.3
v0.0.2
v0.0.1
Labels
Clear labels
burndown-2026-06
Backlog burndown June 2026
pressure-test
Cold-read release pressure-test findings and coordination
sunday-sprint
Burn-down by Sunday 2026-06-07
autonomy
async-consult
A human needs to consult on the issue to upgrade it to headless
autonomy
epic
This issue has many units of sub work - its size makes it meaningfully exclusive with other autonomy types
autonomy
headless
The agent can perform the work on its own
autonomy
live-collab
The agent and the human need to work together in realtime
coherence-core
Core review set for the warded control plane coherence milestone. These issues form the release spine; adjacent milestone issues are stretch or supporting work.
priority
P0
priority tier
priority
P1
priority tier
priority
P2
priority tier
priority
P3
priority tier
priority
P4
priority tier
qa-fixture
Disposable issue admitted to the bounded Ward QA verification lane.
role/advocate
requires work from the Developer Advocate seat
role/director
requires work from the Portfolio Director seat
role/exec
requires work from the exec role
role/frontend
requires work from the Frontend Engineer seat
role/gamedev
requires work from the Game Developer seat
role/human
requires a person, and specifically not an agent seat
role/platform
requires work from the Platform Engineer seat
role/qa
requires work from the QA role
role/science
requires work from the Applied Scientist seat
role/sysadmin
requires work from the Systems Administrator seat
state
ambient
ambient and ephemeral work, held as a maintained document rather than a queue
No labels
burndown-2026-06
pressure-test
sunday-sprint
autonomy
async-consult
autonomy
epic
autonomy
headless
autonomy
live-collab
coherence-core
priority
P0
priority
P1
priority
P2
priority
P3
priority
P4
qa-fixture
role/advocate
role/director
role/exec
role/frontend
role/gamedev
role/human
role/platform
role/qa
role/science
role/sysadmin
state
ambient
Milestone
Clear milestone
No items
No milestone
Projects
Clear projects
No items
No project
Assignees
Clear assignees
No assignees
1 participant
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".
No due date set.
Dependencies
No dependencies set
Reference
coilyco-flight-deck/ward#1069
Loading…
Reference in a new issue
No description provided.
Delete branch "%!s()"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
[epic] Stop the warded-workflow full stop: detect the jam, stop churning
The second full-stop layer this week is the warded workflow itself. When landing is blocked - a keystone issue breaks a repo's test suite (ward#1009 red the whole ward suite), or the merge queue jams on base-branch conflicts - the workflow does not detect it. It churns: it keeps dispatching engineers whose work cannot merge, and re-files "replace unmergeable PR #X" issues in a self-perpetuating loop, burning all 12 engineer slots producing PRs that never land. It looks like full activity and lands nothing. A full stop that actively spins is worse than one that sits: it wastes engineer-hours and pollutes the backlog with replace-PR churn.
Evidence (this week)
Systemic work (children of this epic)
Folded reactive issues (symptoms of this epic)
Acceptance
Context
Filed at Kai's direction as the workflow-layer twin of the infrastructure post-bounce resilience epic. Both are "one keystone failure halts the layer and nothing detects the jam." The infra epic self-recovers a bounced fleet; this one keeps the workflow from spinning against an un-landable state. Refs: ward#1009 (the keystone that jammed the suite), ward#1063 / agentic-os#445 (merge-queue), ward#1016/#1025/#1008/#1024/#1013/#996/#990/#993.
Root cause of the pileup this epic addresses (2026-07-10): a Codex director in
/goal(slash-goal) mode hammered the goal criteria until it spun off the rails, mass-dispatching engineers that opened ~30 branches with no backpressure. Many hit the red #1009 suite and took the forbidden shortcut of restoring the deleted.ward/ward-kdlbundle, so the PRs were mutually-conflicting and unmergeable. Kai's directive: a ward broker + pre-flight gate (BOTH layers) would have stopped it. Concrete children: ward#1073 (open-PR backpressure gate, default 6, broker + pre-flight) and aos#450 (codex binary + director model upgrade). Cleanup closed ward#1053/#1054/#1025 (specgen-restores) + #1043 (empty).🔎 ward agent advisor
ward agent advisor --harness clauderan a one-shot standard research pass on this question:Advisory: implementation plan for the warded-workflow full-stop epic (ward#1069)
This is single-repo work in ward - all three components land in
cmd/ward/'s director/dispatch code, so no cross-repo fan-out. aos#445 and ward#1063 were point-fixes against the same forgejo merge machinery ward already owns; the merge-queue work item here generalizes them into ward's director, it does not spawn a parallel effort in agentic-os. Below is a grounded plan keyed to the real code paths, sequenced into landable PRs.Where the churn actually happens today (the code, not the symptom)
The dispatch tick is
directorDispatchTick(cmd/ward/agent_director_heartbeat.go:200). It computes one global slot budgetavail := cfg.maxParallel - inflight, pullspicks := backlogQueuedPicks(entries)(agent_director.go:1128, ranked tier -> repo -> num across the whole scope), probesbe.probeForgeHealth, asks the LLMbe.decide(...), and dispatches each pick viabe.dispatch.Three gaps produce the full-stop-that-spins:
forgeHealth(agent_director_heartbeat.go:27), which is coarse forge liveness (can we read the top issue), not "is this repo'smainsuite green." So when ward#1009 red the whole suite, every pick still looked dispatchable.RedispatchAttempts+reconcileNoOutcome(agent_director.go:1499, capredispatchAttemptCap = 3) bound only pre-launch death re-queues. A run that launches, produces an unmergeable PR, and parksblockedcarries no failed-landing counter, so the same root cause re-queues forever and the replace-PR issues pile up.runDirectorMergeRepo(agent_director_merge.go:83) loops open PRs and merges each eligible one in a single pass without re-checking mergeability against the new main after each merge, and never rebases the losers - so a moving main strands the rest as conflicts (ward#1063 / aos#445).Component 1 - Landability gate on dispatch
Goal: before dispatching into repo X, confirm X can land -
main's required status contexts are green and the repo is not merge-jammed. If jammed, drop X's picks this tick and escalate the keystone once.New file
cmd/ward/agent_director_landability.go:directorRepoLandable(ctx, cl *forgejoClient, repo string) (landable bool, reason string, keystone int). Reuse the exact machinerydirectorMergeStatusGate(agent_director_merge.go:234) already uses, but pointed atmaininstead of a PR head:cl.getBranch(ctx, owner, repo, "main")for required contexts, resolve main's tip SHA,cl.getCommitCombinedStatus(ctx, owner, repo, sha), thenbuildDirectorMergeStatusSummary(...). Red required context -> not landable. Both client methods already exist (forgejo_ops.go:514,:549) - no new REST surface needed.map[string]landabilitybuilt at the top of the tick) so a 12-issue backlog does not fan into 12 duplicatemainreads.Wire-in: in
directorDispatchTick, afterpicks := backlogQueuedPicks(entries), partition picks byp.repoand filter out any pick whose repo is not landable. AddlandableRepos(ctx, repos) map[string]landabilityto thedirectorBackendinterface (agent_director_heartbeat.go:60) so the live backend does the forge reads and tests inject a fake - same seam pattern asprobeForgeHealth. When a repo is dropped, post the keystone escalation exactly once (guard on a ledger flag, see Component 2's dedup) rather than every tick.Keystone identification: cheapest correct version is "repo main is red -> hold all dispatch into that repo and comment on the highest-ranked open issue labelled the keystone / the repo's tracking issue." A fuller version parses which failing context maps to which fix issue; ship the coarse hold first (it already satisfies the acceptance criterion "stops dispatching into a red-main repo") and refine keystone attribution in a follow-up.
Component 2 - Convergence / circuit-breaker
Goal: after K failed landings on the same issue lineage, stop re-queueing it and stop generating replace-PR issues; surface the keystone instead.
Extend
backlogEntry(agent_director.go:84) withLandFailures intandLineage string(the root issue the replace-PR chain descends from), mirroring the existingRedispatchAttemptsfield and its yaml treatment.Increment point: in the reconcile loop (
agent_director.go:~1480, alongsidereconcileNoOutcome), when an entry resolves toblocked/failedwith an unmergeable / base-conflict root cause - classify it with the logic already indirectorMergeConflictReason/directorMergeConflictReasonFromComments(agent_director_merge.go:130) - bumpLandFailures. WhenLandFailures >= landFailureCap(proposelandFailureCap = 2, matching the acceptance line "no replace-PR issue filed more than once"), transition the entry to a new terminal-ish stateland-blockedthatbacklogQueuedPicksexcludes (add the case there), exactly howorphaned-needs-redispatchalready parks out of the pick set.Suppress the re-file: the "replace unmergeable PR #X" issues are currently created by the engineer/director agent (LLM behavior), not by deterministic ward code - so the breaker's deterministic job is (a) park the lineage out of the pick set (above) and (b) gate any auto-file. The clean hook is the reap/salvage reopen path (
container_reap.go:690notifySalvage,:1101reportUnlandedExtraRepos, both callreopenIssue): before reopening/refiling for a lineage, count prior cycles on that lineage (from the ledgerLandFailures) and, past the cap, post one keystone-escalation comment instead of reopening. Lineage identity: parse theCloses #root/ "replace unmergeable PR #N" reference chain (directorLinkedIssueNumber,agent_director_merge.go:455, already extracts the linked number) and thread the root forward asLineage.This reuses the shape of the ward#595 bounded-retry design, just keyed on landing outcome rather than pre-launch death.
Component 3 - Serialized merge queue
Goal: engineers on one repo land one-at-a-time against a stable base; a moving main auto-rebases the queue rather than stranding conflicts.
Rework
runDirectorMergeRepo(agent_director_merge.go:83) from "merge every eligible PR in one pass" to a serialized queue step:directorMergeEligibility), order them (oldest-linked-issue first, or existing rank).mergePullRequestWithHead(already head-pinned,forgejo_ops.go:446).updatePullRequestBranch(ctx, owner, repo, number)inforgejo_ops.gocalling Forgejo'sPOST /repos/{o}/{r}/pulls/{n}/update(the rebase-onto-base endpoint). It does not exist yet; every other method it composes with does.pr.Mergeableon the next tick so the queue drains one clean merge at a time.This converts the ward#1063 / aos#445 "repair a jammed queue by hand" into a standing invariant: serialize + auto-rebase, so PRs never all conflict against a moving main. Because dispatch already forwards a PR ref into engineers (per recent
issue-913work), pairing serialized merge with rebase-onto-main also stabilizes the base each new engineer branches from.Capacity protection (folds ward#1016 / ward#1025) - cheap, do it first
The root cause named in the thread (Codex
/goaldirector mass-dispatching ~30 branches with no backpressure) is the same global-budget gap.directorDispatchTick'savailis a single scope-wide number with no per-repo ceiling. Add aperRepoMaxtobacklogConfig(agent_director.go:176, default ~ half ofmaxParallel, min 1) and, when selectingchosen, cap how many picks share onep.repo. This directly satisfies "one jammed repo cannot consume all engineer slots" and is a small, independently-landable PR. It composes with ward#1073's open-PR backpressure gate (default 6): ward#1073 caps outstanding PRs per repo at the broker + pre-flight layer, this caps concurrent engineers per repo at the director layer - the two are complementary rails, not duplicates, so land ward#1073 and this per-repo slot cap without stepping on each other.Suggested PR sequence (each independently landable, ordered by unblock value)
backlogConfig.perRepoMax+ selection filter indirectorDispatchTick). Smallest, highest immediate churn-relief, no new forge reads. Ties off ward#1016/#1025.agent_director_landability.go+directorBackend.landableReposseam + dispatch filter). Depends on nothing; reusesgetBranch/getCommitCombinedStatus. Delivers the headline acceptance criterion ("red main -> stop dispatching").backlogEntry.LandFailures/Lineage, increment in reconcile,land-blockedstate, refile guard innotifySalvage). Depends on B's landability classification for a clean keystone signal but can ship in parallel keyed on merge-conflict reason.updatePullRequestBranchclient method + serializedrunDirectorMergeRepo. Depends on nothing structurally; generalizes ward#1063 / aos#445. Land last since it is the largest behavioral change to the merge path.Each PR is small enough to carry its own test file next to the existing
agent_director_*_test.go/agent_dispatch_*_test.gosuites, and each satisfies one line of the epic's Acceptance block on its own.Open questions to confirm before implementation
landFailureCap = 2(park after two failed landings) reads as the safe default - confirm the exact count.ward agent director merge(heartbeat-driven) and not a new standaloneward agent queuecommand - the former reuses all existing eligibility/status-gate code.Recommendation: proceed with PR-A immediately (it is the direct antidote to the ~30-branch stampede in the thread's root-cause note and is nearly free), then B/C/D in sequence. All four are ward-only; file them as four children of this epic.
Researched and posted automatically by
ward agent advisor --harness claude(ward#179). This is one-shot research, not a carried change - verify before acting on it.— Claude (she/her), via
ward agentCleanup note from the director surface on 2026-07-10:
The dispatch lane repeated the failure mode this epic is meant to stop. The surface found seven recent ward engineers that exited together with code 2, one still-running
agentic-os#445engineer stuck in conflict churn, and several no-PR issue branches more than 100 commits behindmainthat kept being resumed instead of forcing a clean start or pausing the issue.Cleanup performed:
agentic-os#445engineer so it could not race cleanup.engineer-*containers from the host Docker view.issue-980,issue-1073,issue-930,issue-786,issue-1033,issue-1064,issue-1096,issue-1087,issue-1084,issue-1045,issue-1085,issue-1083; agentic-osissue-462,issue-455; deployissue-130,issue-134.issue-1007,issue-1011,issue-1006,issue-1000,issue-998,issue-988.issue-133, which duplicated the open salvage PR branchward-salvage/deploy-5e3b3535.coilyco-flight-deck/agentic-os#436.Source behavior still missing: the director should fail closed before dispatch when an issue has a stale no-PR branch, repeated recent failed runs, or an open PR backlog above the configured threshold. It should mark/report the issue as needing branch cleanup or human review instead of launching another engineer against the same stale state.
Evidence for this epic from the burn-down session: it ran exactly the merge-queue shape this epic describes, by hand. Strict one-at-a-time landing, merge onto the moving main per slot, verify with ward build plus ward test plus pre-commit, push, confirm registration, delete the branch. That loop cleared a ten PR jam that had hard-blocked dispatch, with every landed push producing its release tag (v0.641.0 through v0.649.0). Two design notes from the run: conflicts reshuffled after every land, so per-slot re-resolution is mandatory and any precomputed conflict analysis goes stale immediately, and Forgejo never marked push-landed PRs merged, so the queue needs its own close-with-provenance step (ward#1170). The convergence circuit-breaker matters too - the jam only existed because the backpressure gate refused the repair work that would have relieved it (ward#1073).
Closing as superseded by #1620. The autonomous burndown and redispatch loop that produced convergence churn is being removed rather than expanded with another scheduler and merge queue.