Transcript mining round 1: six axes where the composed bundle did not bind, ranked #351
Labels
No labels
burndown-2026-08
autonomy
async-consult
autonomy
epic
autonomy
headless
autonomy
live-collab
coherence-core
priority
P0
priority
P1
priority
P2
priority
P3
priority
P4
qa-fixture
role/devrel
role/eval
role/exec
role/frontend
role/gamedev
role/human
role/platform
role/qa
role/sysadmin
role/tpm
No milestone
No project
No assignees
1 participant
Notifications
Due date
No due date set.
Dependencies
No dependencies set
Reference
coilyco-flight-deck/agent-compose#351
Loading…
Reference in a new issue
No description provided.
Delete branch "%!s()"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
First pass of the transcript-mining loop. The human's in-conversation corrections are treated as labelled defects, already produced at no grading cost, and separated into bundle defects and domain corrections. Method, results, and the limits of this pass are below. Join-key gap is #350.
Method
Read-only pass over the local Claude Code transcript store, 2026-08-25.
Axes, ranked by frequency
A. Authority under-claim, 4 instances. The seat holds a grant and does not act on it, or cannot tell whether it holds it. Appears across three different seats, so it is a shared-doctrine problem rather than one charter's.
boundary-modify-live-backendalready names this exact failure in its own scoped section, and the bundle still did not prevent it.628bdb5b2026-08-11T23:32:55Z, sysadmin seat6cbecdbb2026-08-12T09:26:08Z, retired qa seat151d05c52026-08-14T04:41:56Z, platform seat68c89d062026-08-11T06:08:56Z, sysadmin seat, weaker instanceB. Doctrine reads harder than intended, 1 instance, highest specificity. On
c34b11762026-08-22T21:04:59Z the human states the intended semantics ofseek-external-validationverbatim: it is about directing autonomous behavior, and is not a hard stop of the same kind as the comms and live-ops boundaries. The currently shipped defer side still reads as a hard stop. This is the most actionable item in the set, because the intended reading was stated outright and the text still disagrees with it.A and B are plausibly the same failure from two sides. A boundary written as a wall produces a seat that under-claims its own grant. Worth treating as one fix rather than two.
C. Skill selection and sufficiency, 3 instances. A skill that should have been selected was not, a skill under-carried what its task needed and forced a raw-source read, and doctrine was not loaded on the first turn.
e2455c6e2026-08-10T16:50:40Z, skill not selectedb88f60ea2026-08-07T21:33:58Z, skill under-carries. The human diagnoses the skill-content gap directly in the turn.fdfd99902026-08-17T23:28:28Z, not loaded firstD. Wrong altitude, 2 instances. Output pitched at the wrong level for the seat, both on strategy-shaped seats.
fca5aa4e2026-08-23T02:02:06Zc34b11762026-08-22T22:47:29ZE. Task not finished, 2 instances.
87a40dff2026-08-12T00:17:22Z and 00:36:17Z, one correction and its clarification.F. Stated preference not binding, 1 instance.
5b0e806c2026-08-18T16:24:39Z. A stack was proposed that a shipped preference skill contradicts. Worth confirming whether that skill was in the selected set for that session, which #350 would make answerable.One further roster-design signal rather than a runtime defect:
151d05c52026-08-14T00:34:25Z questions whether a role's scope fits a use it was being pointed at.Before-and-after result
The roster rename gives a clean era boundary. Every pre-rename title (
Engineer,DevOps,Director,Designer,AI Engineer,Executive Strategist,Content Creator,QA) ends on or before 2026-08-22. Every post-rename title starts 2026-08-23.The mechanism works and the post-change corpus is too young to conclude anything. Six candidates over three days is not a signal, and per-turn the authority axis is nominally higher after the change, which at n=2 means nothing either. The useful result is the cadence: this comparison needs roughly a month of post-change sessions before it can answer. Re-run then.
Limits of this pass
Evidence handling
Citations are session id plus timestamp only. No verbatim human text appears here, because this repository is public and the corpus is private conversation. The transcripts stay on the local machine and are not routed anywhere. Consistent with the rest of the design: point at the evidence, never copy it.
Provenance
Produced by the eval seat, 2026-08-25. Counts are measured. Axis assignments are the eval seat's classification and are the judgement most worth disputing, since a second reader applying the same test would likely move two or three items.
Round 1 follow-up: each axis re-read with transcript context, then ticketed
Every bundle-shaped instance above was re-read with its preceding agent turn rather than as a bare human line. Three things changed, so this report should not be read as it stands.
Tickets filed
seek-external-validationregister. Confirmed against the currently shipped text.Reclassifications
151d05c5merge-authority is fixed in current lane text. The remaining three in #352 are open.The finding that outranks the axes
Four of the round's findings were already fixed before the round ran, and nothing recorded that. Reading shipped doctrine to work out which findings were still live cost more than producing the findings did. That is #356, and it should be sequenced ahead of the individual axis fixes, because without it round 2 pays the same cost again.
Correction to the method section above
The claim that every classification was made by reading was true of the human turns and not of the agent turns preceding them. Reading the agent side changed one classification outright, narrowed two, and sharpened one. Any future round should read both sides from the start, since the correction alone does not say what went wrong.