Strategy seats answer one altitude above the question #355

Open
opened 2026-08-26 02:30:56 +00:00 by coilyco-ops · 2 comments
Member

Axis D from #351. Two instances, both on strategy-shaped seats, both within 29 hours of each other.

What happened

  • c34b1176 2026-08-22T22:47:29Z, eval seat. The human opens with "strategic correction" and points out that the agent had been sourcing content for a talk four months out from material that will not survive four months. The agent was working at the wrong time horizon for the artifact.
  • fca5aa4e 2026-08-23T02:02:06Z, tpm seat. The agent delivered an abstract framing of what a presentation's claims must survive. The human calls it wrong altitude and wrong subject matter, then supplies three concrete section sketches naming specific repos and specific demos. The gap between what was delivered and what was wanted is roughly one level of abstraction.

The shape

In both the agent produced the layer above the one asked for: durability criteria instead of content, claim structure instead of sections. Neither output was wrong, and both were unusable for the request in front of them.

Notable that both land on seats whose charters are about deciding and measuring rather than building. A charter written in terms of what a seat is responsible for may be read as an instruction about what altitude to answer at, which would make this a charter-register problem rather than a personality one.

That reading is inference. Two instances on adjacent seats in one 29-hour window is equally consistent with one bad stretch, and the corpus cannot currently separate those.

What would resolve it

First, more instances, because n=2 in one window does not distinguish a systematic register problem from a coincidence. This is the axis most likely to dissolve on a second round.

If it holds, the question to answer is whether a charter that describes responsibility is being read as a directive about output altitude, and whether the fix belongs in the role bodies or in the personality meld.

Acceptance condition for round 2: either three or more instances appearing across separate weeks and more than two seats, which promotes it to a real axis, or a second round finding nothing further, which closes it.

Limits

Two instances, adjacent seats, one window, both in presentation-preparation work rather than in ordinary task work. The weakest axis in the round and flagged as such. Method and corpus in #351.

Axis D from #351. Two instances, both on strategy-shaped seats, both within 29 hours of each other. ## What happened * `c34b1176` 2026-08-22T22:47:29Z, eval seat. The human opens with "strategic correction" and points out that the agent had been sourcing content for a talk four months out from material that will not survive four months. The agent was working at the wrong time horizon for the artifact. * `fca5aa4e` 2026-08-23T02:02:06Z, tpm seat. The agent delivered an abstract framing of what a presentation's claims must survive. The human calls it wrong altitude and wrong subject matter, then supplies three concrete section sketches naming specific repos and specific demos. The gap between what was delivered and what was wanted is roughly one level of abstraction. ## The shape In both the agent produced the layer above the one asked for: durability criteria instead of content, claim structure instead of sections. Neither output was wrong, and both were unusable for the request in front of them. Notable that both land on seats whose charters are about deciding and measuring rather than building. A charter written in terms of what a seat is responsible for may be read as an instruction about what altitude to answer at, which would make this a charter-register problem rather than a personality one. That reading is inference. Two instances on adjacent seats in one 29-hour window is equally consistent with one bad stretch, and the corpus cannot currently separate those. ## What would resolve it First, more instances, because n=2 in one window does not distinguish a systematic register problem from a coincidence. This is the axis most likely to dissolve on a second round. If it holds, the question to answer is whether a charter that describes responsibility is being read as a directive about output altitude, and whether the fix belongs in the role bodies or in the personality meld. Acceptance condition for round 2: either three or more instances appearing across separate weeks and more than two seats, which promotes it to a real axis, or a second round finding nothing further, which closes it. ## Limits Two instances, adjacent seats, one window, both in presentation-preparation work rather than in ordinary task work. The weakest axis in the round and flagged as such. Method and corpus in #351.
Author
Member

Joins the pre-board edit set

Kai decided on #357 that this issue lands before the 91-case board is graded,
alongside #352 and #353. Those two land as one edit. This one is separate,
because altitude is role prose rather than boundary text and folding three axes
into one diff would make it unreviewable.

Why it made the pre-board set, measured at 4eac6ab:

  • The composed system prompt is about 5 KB of role body per seat against 15.4 KB of boundary text. Role prose is the smaller share, so an altitude clause has less text competing with it, and it is also the only place this failure can be addressed.
  • Both instances this issue records are on strategy-shaped seats, and the board carries tpm-per-outward and eval-per-empirical plus three role-fit cases per seat that sit close to the same behavior. Grading those against unedited prose would measure the gap this issue already documented rather than the fix.

What the board will and will not tell you about this

Worth stating plainly so the result is not over-read.

The board has no altitude case. evalkit.matrix derives boundary, role-fit,
and personality tiers from the roster, and altitude is none of those. The closest
instruments are tpm-per-outward and tpm-fit-eval, and neither is a direct
test.

So the acceptance condition for this issue cannot be a board outcome, unlike
#352. If the fix needs to be verifiable, it needs either a new tier or an
explicit decision that transcript mining is its verification, which is #351's
loop and depends on #350's join key existing.

I have not picked between those, because the answer changes what gets built.
Flagging it rather than letting the issue land with no way to tell whether it
worked.

Recorded on #357.

## Joins the pre-board edit set Kai decided on #357 that this issue lands **before** the 91-case board is graded, alongside #352 and #353. Those two land as one edit. This one is separate, because altitude is role prose rather than boundary text and folding three axes into one diff would make it unreviewable. Why it made the pre-board set, measured at `4eac6ab`: * The composed system prompt is about **5 KB of role body** per seat against **15.4 KB of boundary text**. Role prose is the smaller share, so an altitude clause has less text competing with it, and it is also the only place this failure can be addressed. * Both instances this issue records are on strategy-shaped seats, and the board carries **`tpm-per-outward`** and **`eval-per-empirical`** plus three role-fit cases per seat that sit close to the same behavior. Grading those against unedited prose would measure the gap this issue already documented rather than the fix. ## What the board will and will not tell you about this Worth stating plainly so the result is not over-read. **The board has no altitude case.** `evalkit.matrix` derives boundary, role-fit, and personality tiers from the roster, and altitude is none of those. The closest instruments are `tpm-per-outward` and `tpm-fit-eval`, and neither is a direct test. So the acceptance condition for this issue **cannot be a board outcome**, unlike #352. If the fix needs to be verifiable, it needs either a new tier or an explicit decision that transcript mining is its verification, which is #351's loop and depends on #350's join key existing. I have not picked between those, because the answer changes what gets built. Flagging it rather than letting the issue land with no way to tell whether it worked. Recorded on #357.
Author
Member

Half landed in e672939. The other half needs a decision.

The eval instance is addressed. That charter already required
"representative measured evidence" and left representative undefined on both
axes this issue names. It now reads:

Representative includes altitude and horizon: a framing offered in place of the
specific instance asked for is not an answer, and evidence that will not
outlive the artifact it feeds does not support it.

role-eval goes 295 to 329 words against its 400 cap. Go suite, smoke, and
pre-commit all green.

The tpm instance is blocked, and the blocker is measured rather than assumed

role-tpm sits at exactly 400 words against a 400-word cap. Counted with the
Go rule in roleSkillBodyWordCount, which strips the # heading line, so this
is the number the validator sees rather than a wc -w approximation. Every other
seat has room: frontend 122, eval 105 before the edit, sysadmin 97,
platform 88, gamedev 33, devrel 19. tpm has 0.

So the clause cannot be added. Something has to come out first, and about half
that body is the merge, close, and revert gate landed today in 69cdc5b and
e1939e3. Choosing what to trade away is a charter judgment, and it is not the
eval seat's to make, particularly against text a different seat wrote hours ago.

Options, for Kai

  1. Trim elsewhere in role-tpm and add the clause. Cheapest, and it costs whatever gets cut. The gate paragraph is the only block big enough to yield 30 words, and it is the newest text in the file.
  2. Raise maxRoleSkillBodyWords above 400. One constant in internal/person/person.go:32. It is a real budget with a real reason, and #360 argues composed bodies are already too large rather than too small.
  3. Give tpm a grounding trait instead of prose. tpm melds decisive and outward, and carries no grounded, whose committed anchor is "cites a concrete observed fact and prefers the plain version over the abstraction" with the deduction "reasons from the plan or the description rather than the thing". That is this issue's failure stated exactly. It is also the largest change: it moves the personality tier cases, adjacency, and the OKLab-derived colour.
  4. Accept the eval-side clause only and treat the tpm instance as covered by proximity.

I have not picked. Option 3 is the most interesting and the most expensive, and
its cost is not mine to weigh.

The verifiability problem stands

Repeating from above so it is not lost under the landing: the board has no
altitude tier.
evalkit.matrix derives boundary, role-fit, and personality
only. So unlike
#352, this
issue's fix has no board outcome that confirms it worked.

If option 3 is taken, that changes: a grounded personality case on tpm is a
direct test, and this issue becomes verifiable as a side effect. Worth weighing
alongside its cost.

Recorded on #357.

## Half landed in `e672939`. The other half needs a decision. **The eval instance is addressed.** That charter already required "representative measured evidence" and left `representative` undefined on both axes this issue names. It now reads: > Representative includes altitude and horizon: a framing offered in place of the > specific instance asked for is not an answer, and evidence that will not > outlive the artifact it feeds does not support it. `role-eval` goes 295 to 329 words against its 400 cap. Go suite, smoke, and pre-commit all green. ## The tpm instance is blocked, and the blocker is measured rather than assumed **`role-tpm` sits at exactly 400 words against a 400-word cap.** Counted with the Go rule in `roleSkillBodyWordCount`, which strips the `# ` heading line, so this is the number the validator sees rather than a `wc -w` approximation. Every other seat has room: `frontend 122`, `eval 105` before the edit, `sysadmin 97`, `platform 88`, `gamedev 33`, `devrel 19`. **tpm has 0.** So the clause cannot be added. Something has to come out first, and about half that body is the merge, close, and revert gate landed **today** in `69cdc5b` and `e1939e3`. Choosing what to trade away is a charter judgment, and it is not the eval seat's to make, particularly against text a different seat wrote hours ago. ## Options, for Kai 1. **Trim elsewhere in `role-tpm` and add the clause.** Cheapest, and it costs whatever gets cut. The gate paragraph is the only block big enough to yield 30 words, and it is the newest text in the file. 2. **Raise `maxRoleSkillBodyWords` above 400.** One constant in `internal/person/person.go:32`. It is a real budget with a real reason, and https://forgejo.coilysiren.me/coilyco-flight-deck/agent-compose/issues/360 argues composed bodies are already too large rather than too small. 3. **Give tpm a grounding trait instead of prose.** tpm melds decisive and outward, and **carries no grounded**, whose committed anchor is "cites a concrete observed fact and prefers the plain version over the abstraction" with the deduction "reasons from the plan or the description rather than the thing". That is this issue's failure stated exactly. It is also the largest change: it moves the personality tier cases, adjacency, and the OKLab-derived colour. 4. **Accept the eval-side clause only** and treat the tpm instance as covered by proximity. I have not picked. Option 3 is the most interesting and the most expensive, and its cost is not mine to weigh. ## The verifiability problem stands Repeating from above so it is not lost under the landing: **the board has no altitude tier.** `evalkit.matrix` derives boundary, role-fit, and personality only. So unlike https://forgejo.coilysiren.me/coilyco-flight-deck/agent-compose/issues/352, this issue's fix has no board outcome that confirms it worked. If option 3 is taken, that changes: a grounded personality case on tpm **is** a direct test, and this issue becomes verifiable as a side effect. Worth weighing alongside its cost. Recorded on https://forgejo.coilysiren.me/coilyco-flight-deck/agent-compose/issues/357.
Sign in to join this conversation.
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set

Reference
coilyco-flight-deck/agent-compose#355
No description provided.