Strategy seats answer one altitude above the question #355
Labels
No labels
burndown-2026-08
autonomy
async-consult
autonomy
epic
autonomy
headless
autonomy
live-collab
coherence-core
priority
P0
priority
P1
priority
P2
priority
P3
priority
P4
qa-fixture
role/devrel
role/eval
role/exec
role/frontend
role/gamedev
role/human
role/platform
role/qa
role/sysadmin
role/tpm
No milestone
No project
No assignees
1 participant
Notifications
Due date
No due date set.
Dependencies
No dependencies set
Reference
coilyco-flight-deck/agent-compose#355
Loading…
Reference in a new issue
No description provided.
Delete branch "%!s()"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
Axis D from #351. Two instances, both on strategy-shaped seats, both within 29 hours of each other.
What happened
c34b11762026-08-22T22:47:29Z, eval seat. The human opens with "strategic correction" and points out that the agent had been sourcing content for a talk four months out from material that will not survive four months. The agent was working at the wrong time horizon for the artifact.fca5aa4e2026-08-23T02:02:06Z, tpm seat. The agent delivered an abstract framing of what a presentation's claims must survive. The human calls it wrong altitude and wrong subject matter, then supplies three concrete section sketches naming specific repos and specific demos. The gap between what was delivered and what was wanted is roughly one level of abstraction.The shape
In both the agent produced the layer above the one asked for: durability criteria instead of content, claim structure instead of sections. Neither output was wrong, and both were unusable for the request in front of them.
Notable that both land on seats whose charters are about deciding and measuring rather than building. A charter written in terms of what a seat is responsible for may be read as an instruction about what altitude to answer at, which would make this a charter-register problem rather than a personality one.
That reading is inference. Two instances on adjacent seats in one 29-hour window is equally consistent with one bad stretch, and the corpus cannot currently separate those.
What would resolve it
First, more instances, because n=2 in one window does not distinguish a systematic register problem from a coincidence. This is the axis most likely to dissolve on a second round.
If it holds, the question to answer is whether a charter that describes responsibility is being read as a directive about output altitude, and whether the fix belongs in the role bodies or in the personality meld.
Acceptance condition for round 2: either three or more instances appearing across separate weeks and more than two seats, which promotes it to a real axis, or a second round finding nothing further, which closes it.
Limits
Two instances, adjacent seats, one window, both in presentation-preparation work rather than in ordinary task work. The weakest axis in the round and flagged as such. Method and corpus in #351.
Joins the pre-board edit set
Kai decided on #357 that this issue lands before the 91-case board is graded,
alongside #352 and #353. Those two land as one edit. This one is separate,
because altitude is role prose rather than boundary text and folding three axes
into one diff would make it unreviewable.
Why it made the pre-board set, measured at
4eac6ab:tpm-per-outwardandeval-per-empiricalplus three role-fit cases per seat that sit close to the same behavior. Grading those against unedited prose would measure the gap this issue already documented rather than the fix.What the board will and will not tell you about this
Worth stating plainly so the result is not over-read.
The board has no altitude case.
evalkit.matrixderives boundary, role-fit,and personality tiers from the roster, and altitude is none of those. The closest
instruments are
tpm-per-outwardandtpm-fit-eval, and neither is a directtest.
So the acceptance condition for this issue cannot be a board outcome, unlike
#352. If the fix needs to be verifiable, it needs either a new tier or an
explicit decision that transcript mining is its verification, which is #351's
loop and depends on #350's join key existing.
I have not picked between those, because the answer changes what gets built.
Flagging it rather than letting the issue land with no way to tell whether it
worked.
Recorded on #357.
Half landed in
e672939. The other half needs a decision.The eval instance is addressed. That charter already required
"representative measured evidence" and left
representativeundefined on bothaxes this issue names. It now reads:
role-evalgoes 295 to 329 words against its 400 cap. Go suite, smoke, andpre-commit all green.
The tpm instance is blocked, and the blocker is measured rather than assumed
role-tpmsits at exactly 400 words against a 400-word cap. Counted with theGo rule in
roleSkillBodyWordCount, which strips the#heading line, so thisis the number the validator sees rather than a
wc -wapproximation. Every otherseat has room:
frontend 122,eval 105before the edit,sysadmin 97,platform 88,gamedev 33,devrel 19. tpm has 0.So the clause cannot be added. Something has to come out first, and about half
that body is the merge, close, and revert gate landed today in
69cdc5bande1939e3. Choosing what to trade away is a charter judgment, and it is not theeval seat's to make, particularly against text a different seat wrote hours ago.
Options, for Kai
role-tpmand add the clause. Cheapest, and it costs whatever gets cut. The gate paragraph is the only block big enough to yield 30 words, and it is the newest text in the file.maxRoleSkillBodyWordsabove 400. One constant ininternal/person/person.go:32. It is a real budget with a real reason, and #360 argues composed bodies are already too large rather than too small.I have not picked. Option 3 is the most interesting and the most expensive, and
its cost is not mine to weigh.
The verifiability problem stands
Repeating from above so it is not lost under the landing: the board has no
altitude tier.
evalkit.matrixderives boundary, role-fit, and personalityonly. So unlike
#352, this
issue's fix has no board outcome that confirms it worked.
If option 3 is taken, that changes: a grounded personality case on tpm is a
direct test, and this issue becomes verifiable as a side effect. Worth weighing
alongside its cost.
Recorded on #357.