Eval prep: the measured gap on the seven-seat board, and every eval, prose, and content ticket folded into one ordering #357
Labels
No labels
burndown-2026-08
autonomy
async-consult
autonomy
epic
autonomy
headless
autonomy
live-collab
coherence-core
priority
P0
priority
P1
priority
P2
priority
P3
priority
P4
qa-fixture
role/devrel
role/eval
role/exec
role/frontend
role/gamedev
role/human
role/platform
role/qa
role/sysadmin
role/tpm
No milestone
No project
No assignees
1 participant
Notifications
Due date
No due date set.
Dependencies
No dependencies set
Reference
coilyco-flight-deck/agent-compose#357
Loading…
Reference in a new issue
No description provided.
Delete branch "%!s()"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
Filed by Lucia (eval seat), 2026-08-25. Prep record for the first seven-seat
board run. Nothing here is a grade. It is the measured gap, plus every open
ticket that is eval, prose, or content shaped folded into one ordering, so the
run does not start against text that is about to move.
Measured at
4eac6abin a detached worktree. Commands and outputs below are thething rather than a description of it.
The gap, measured
just evalkit-matrixreturns91 challenges: 56 boundary, 21 role-fit, 14 personality, andper role: devrel 13, eval 13, frontend 13, gamedev 13, platform 13, sysadmin 13, tpm 13. Perfectly even, which the nine-seat roster never was.challenges.yamlcarries the sysadmin slice only:sysadmin-sec-{in,out},sysadmin-mlb-{in,out},sysadmin-bfs-{in,out},sysadmin-sev-{in,out},sysadmin-fit-{within,platform,tpm},sysadmin-per-{protective,grounded}. 78 cases have no prompt.evaluations/holds no v3 record. The newest graded pair isevaluations/pilot/ops-board-2026-08-12and-regraded, both nine-seat era, retired by the rename inae06747.evaluations/reflow-v3/attributes.yamlcarrieseval 4, frontend 4, gamedev 4, devrel 3, platform 3, sysadmin 3, tpm 3. Plus 4 owner pairs, which need no target prose, that is the 28 pairs behind the 56 boundary cases.The instrument is green
Both legs were run rather than assumed.
just evalkit-checkpasses: ruff clean, 9 files formatted, mypy strict clean over 9 source files,25 passed in 2.48s.just evalkit-smokereaches Agent Proxy athttp://ser8:8080/v1, lists seven models, and the declared subjectevaluation/deepseek-v4-proanswers withusage: prompt 96, completion 21.reasoning words: 14.aos-eval, version 0.1.1resolves.So the run leg is not what is blocking. Authoring and grading are.
Fold-in
Four buckets. The bucket decides when a ticket touches the run, not how
important it is.
Bucket 1: the subject under test. These move the text the board measures
Grading before these land buys evidence about text that is about to change. Each
one edits role, boundary, or personality prose, which is the Developer Platform
Engineer's to write. Eval owns the acceptance condition and hands the edit over.
seek-external-validationreads as a hard stop - highest confidence in the set, because Kai stated the intended semantics verbatim and the shipped defer side still disagrees with ite81a96cnow records composed body size in the manifest, which is the instrument this needs. Re-measure against the seven seats before it can bind#351records that axes A and B are plausibly one failure seen from two sides:a boundary written as a wall produces a seat that under-claims its own grant.
Worth fixing as one edit rather than four.
Bucket 2: board inputs. Findings that should become cases or join keys
Bucket 3: instrument and methodology
docs/evaluation.mdand the tree. Its one open clause is "re-earn the whole board", which is #318. Rescope it to that or close itCLAUDE.mdcontamination class that the plain-call execution model made structurally impossible. Eight packs for eight retired seats. Close it against #262 and #318aos-eval exportis one way, which removes the round trip step 4 called the hard requirement. What is missing is the explicit recorded decision with its revisit triggerBucket 4: outside the freeze
Content and platform work that does not change what the board measures. Listed
so the exclusion is recorded rather than silent.
ef0c205just smokered on main - partly addressed byeb29a6dThe remaining open issues (#329 and its children #330 through #341, plus #345,
#347, #348, #349, #276, #342, #343) are the inversion program and platform work.
They are not eval, prose, or content shaped and are out of scope here.
The ordering question this raises
Buckets 1 and 3 pull against each other. Landing the doctrine edits first means
the board measures the roster we intend to ship, and loses the before-and-after
that would show the edits worked. Grading first buys that baseline and costs a
second pass over 91 cases by hand.
Not deciding this myself. It is a genuine fork, the wrong branch is expensive
to undo in human grading hours, and #351's axis classification is my own seat's
work, which I do not independently accept as an evaluation contract.
Provenance
Roster and counts measured at
4eac6ab, 2026-08-25, in a detached worktree, withno mutation to any checkout. Ticket classification is the eval seat's and is the
judgement most worth disputing.
Decisions, from Kai, 2026-08-25
coilysiren/inbox#434only. The audience-as-grader pairs come out of the same authoring pass. #324 stays separate, since a shoot plan is different work.Correction to bucket 1: #275 is not stale, and its defect got worse
I wrote above that #275's numbers were pre-v3 and needed re-measuring before the
ticket could bind. I then measured it, and that was the wrong call. The
ticket understates the problem now.
The artifact #275 cites,
services/sirens-echo/rendered/sirens-deep-bundle.txt,no longer exists. The current render is
coilyco-bridge/deploy/services/sirens-echo/rendered/roles-bundle.txt, whichcarries
composed body bytesper role directly. Measured there:librarian5,947director7,604exec10,359qa10,580design11,502ai11,891ops13,380creator48,274engineer145,593That is a 24.5x spread, against the 7x #275 reported.
engineeris 3x thenext largest and carries 57 skills and 0 boundaries, including every
coding-*umbrella, everypersonal-preference-*, the wholewriting-*set,and the ops and qa tooling sets. It is the role with the most reach and the only
core role with nothing bounding it.
Two caveats I am not going to paper over:
sirens-echo#1147and#1155already track. These are nine retired seats. The spread is real and currently deployed, and it is not a measurement of the seven-seat roster.engineerhaving zero boundaries may beboundary-omitworking as designed, dropping defer-side boundaries whose owning seat that deployment lacks.opskeeps 3 andaikeeps 2, so the omission is partial rather than total. I have not readdocs/boundary-omission.mdclosely enough to say which it is, and that read is the platform seat's.The v3 comparison, for what it is worth, is a different and narrower
artifact:
just evalkit-promptscomposes thecompileddelivery at frontiertier, which inlines role, boundaries, personalities, and the invariant but no
ordinary skills. Across the seven seats that is
frontend23,545 todevrel24,741 bytes, a 1.05x spread. Flat, and about 24 KB each. That is not
evidence the 24.5x defect is fixed, because ordinary skills are exactly what
produces the outlier and this delivery excludes them.
The board data survives
ae06747intactae06747retitles four roles and renames three seats: eval becomes Evie (she)and Evaluation Engineer, sysadmin becomes Vera (she), tpm becomes Portia (they),
platform displays as Agentic Platform Engineer, frontend as Design Engineer, and
tpm as Portfolio Director.
I checked whether that invalidates authored cases. It does not. Neither
challenges.yamlnorevaluations/reflow-v3/attributes.yamlcontains a seatname or a display title. Both are keyed on slugs, and the commit message
confirms slugs hold. The 13 authored sysadmin prompts carry over unchanged.
Live evidence for #350, from this session
Worth recording because it happened rather than because it was predicted. This
prep was produced by a session whose identity card reads Lucia (she), Agent
Evaluation Engineer, which is the pre-
ae06747seat. Main ships Evie (she),Evaluation Engineer. I only found that out by diffing the composed delivery
against the commit, and nothing in my own transcript would have told me which
bundle I was running.
That is precisely the join-key gap #350 describes, observed from the inside. The
signature on this issue's body names a seat that main has already retired, and I
am leaving it as filed rather than rewriting it, since it is the artifact.
State change: the board is authored, and the ordering is settled
All 91 cases now have prompts. Landed in
19d5abf. Detail on #318, and thecount that mattered in this issue's first section is now
91 authored, 91 derived, 0 unauthored.Second correction, this one to my own claim above
The body of this issue says neither
challenges.yamlnorattributes.yamlcarries a seat name or display title, and I used that to conclude
ae06747invalidated nothing.
Half wrong.
attributes.yamlis clean.challenges.yamlcarries displaytitles in its targets, three before this change and seven after, because a target
has to name the seat receiving the handoff. My check grepped personal names
across both files and titles across only one, so the conclusion outran the
evidence.
The conclusion survives by luck rather than by the reasoning I gave: the sysadmin
slice had already been updated to the post-
ae06747titles, so the file wascurrent. A retitle does invalidate authored targets, and that is now recorded
in the file header rather than left to be rediscovered.
Ordering, decided
Kai's second decision, after the measurement below: land #352, #353, and #355,
then grade once, before the #329 inversion. #355 joins the pre-board edit set.
The reasoning worth keeping, since it splits the tickets differently than this
issue's buckets did.
Boundary text is 64% of the composed system prompt. Measured across all seven
seats from
just evalkit-prompts: every seat carries about 15.4 KB of boundarybodies against about 5 KB of role body, so the boundary block is 3x the role
charter and near-identical seat to seat. #352 lands in the shared frame of that
block, and 56 of the 91 cases test it. Grading before it lands would grade the
majority of what the subject reads, knowing it is about to change.
The #329 program is the opposite case, and the board should predate it. I
read the epic end to end. Its entire completion list is mechanism: no
.kdlremains,
internal/personcarries no semantic logic, the compositor installsfrom PyPI, a roster change goes live without a rebuild. Nothing in it changes a
word of doctrine.
That inverts the argument for holding. #333 verifies the port "with the Go tests
as the oracle", and unit tests can prove the compositor computes the same meld
and the same OKLab centroid. They cannot prove the composed text still produces
the same behavior in a model. A graded board is that oracle, and it only works
if it predates the inversion. #339 deletes the
composecommand thatscripts/eval-prompts.shshells out to today, so this is not hypothetical.Two bucket-1 corrections from the same pass
compileddelivery the subject receives carries role instructions, identity card, invariant, boundaries, and personalities, and no ordinary skills. Skill selection is invisible to the eval, so #354 belongs in bucket 4 rather than bucket 1.just smokeon88f87dcpasses all ten stages, including the two assertions the ticket names.eb29a6dfixed it. The ticket is stale and is the platform seat's to close.Where this now sits