Eval prep: the measured gap on the seven-seat board, and every eval, prose, and content ticket folded into one ordering #357

Open
opened 2026-08-26 02:45:37 +00:00 by coilyco-ops · 2 comments
Member

Filed by Lucia (eval seat), 2026-08-25. Prep record for the first seven-seat
board run. Nothing here is a grade. It is the measured gap, plus every open
ticket that is eval, prose, or content shaped folded into one ordering, so the
run does not start against text that is about to move.

Measured at 4eac6ab in a detached worktree. Commands and outputs below are the
thing rather than a description of it.

The gap, measured

  • 91 cases derive. just evalkit-matrix returns 91 challenges: 56 boundary, 21 role-fit, 14 personality, and per role: devrel 13, eval 13, frontend 13, gamedev 13, platform 13, sysadmin 13, tpm 13. Perfectly even, which the nine-seat roster never was.
  • 13 are authored. challenges.yaml carries the sysadmin slice only: sysadmin-sec-{in,out}, sysadmin-mlb-{in,out}, sysadmin-bfs-{in,out}, sysadmin-sev-{in,out}, sysadmin-fit-{within,platform,tpm}, sysadmin-per-{protective,grounded}. 78 cases have no prompt.
  • 0 are graded on this roster. evaluations/ holds no v3 record. The newest graded pair is evaluations/pilot/ops-board-2026-08-12 and -regraded, both nine-seat era, retired by the rename in ae06747.
  • 24 non-owner pairs are declared. evaluations/reflow-v3/attributes.yaml carries eval 4, frontend 4, gamedev 4, devrel 3, platform 3, sysadmin 3, tpm 3. Plus 4 owner pairs, which need no target prose, that is the 28 pairs behind the 56 boundary cases.

The instrument is green

Both legs were run rather than assumed.

  • Runner. just evalkit-check passes: ruff clean, 9 files formatted, mypy strict clean over 9 source files, 25 passed in 2.48s.
  • Transport. just evalkit-smoke reaches Agent Proxy at http://ser8:8080/v1, lists seven models, and the declared subject evaluation/deepseek-v4-pro answers with usage: prompt 96, completion 21. reasoning words: 14.
  • Grader. aos-eval, version 0.1.1 resolves.

So the run leg is not what is blocking. Authoring and grading are.

Fold-in

Four buckets. The bucket decides when a ticket touches the run, not how
important it is.

Bucket 1: the subject under test. These move the text the board measures

Grading before these land buys evidence about text that is about to change. Each
one edits role, boundary, or personality prose, which is the Developer Platform
Engineer's to write. Eval owns the acceptance condition and hands the edit over.

  • #353 - seek-external-validation reads as a hard stop - highest confidence in the set, because Kai stated the intended semantics verbatim and the shipped defer side still disagrees with it
  • #352 - seats declare an incapacity they never tested, then hand the work back - axis A, four instances across three seats
  • #355 - strategy seats answer one altitude above the question - axis D, two instances 29 hours apart
  • #354 - skill discovery goes to a leaf and skips the repo-pointer hop - axis C, selection half still open
  • #275 - composed role bodies span 7x with no budget - its numbers are pre-v3. Six of the eight roles it measures are retired seats. e81a96c now records composed body size in the manifest, which is the instrument this needs. Re-measure against the seven seats before it can bind
  • #273 - bounded reconfiguration authority for the QA seat - names a seat the roster no longer has. The charter question survives the rename and is axis A in another form. Re-point it at the sysadmin and platform scopes or close it
  • #278 - the QA role should bias toward true E2E - same stale seat, and the body is empty. Re-point or close

#351 records that axes A and B are plausibly one failure seen from two sides:
a boundary written as a wall produces a seat that under-claims its own grant.
Worth fixing as one edit rather than four.

Bucket 2: board inputs. Findings that should become cases or join keys

  • #351 - transcript mining round 1, six axes ranked - the spine. 4,771 genuine human turns after filtering, 55 candidates, 13 bundle-shaped
  • #356 - nothing records which correction produced which rule - four round-1 findings are already fixed and look identical to open ones. This is a coverage defect in the evidence, not in the roster
  • #350 - record the composed bundle identity in the transcript - the join key. Until it lands, no finding can be attributed to a specific composed set, so no before-and-after can be run
  • #340 - coverage check for unauthored and ungraded cases - the mechanical version of the count in this issue's first section, which I produced by hand today. Blocked behind #338

Bucket 3: instrument and methodology

  • #318 - author and grade the seven-seat board - the main event. 78 unauthored prompts, then 91 graded cases
  • #320 - the grade-to-edit return path has no page - this is the loop this prep record is exercising, and it is documented nowhere
  • #262 - replace the LLM-reviewer methodology with the triple - largely shipped. The triple, the pair as scoring unit, n=5 with epoch 1 graded, and committed personality anchors are all in docs/evaluation.md and the tree. Its one open clause is "re-earn the whole board", which is #318. Rescope it to that or close it
  • #240 - re-earn the eight evaluation records - superseded and now obsolete. Steps 1 through 3 diagnose a driver and an ambient-CLAUDE.md contamination class that the plain-call execution model made structurally impossible. Eight packs for eight retired seats. Close it against #262 and #318
  • #213 - evaluate a replaceable review UI - partly answered. aos-eval export is one way, which removes the round trip step 4 called the hard requirement. What is missing is the explicit recorded decision with its revisit trigger
  • #338 - evalkit derives the board with no Go in the path - blocked behind the #329 inversion program
  • #324 - terminal demo, three roles one artifact one prompt - presentation artifact, reads composed personas that already exist
  • coilysiren/inbox#434 - audience-as-grader segment - eval-shaped, and explicitly independent of #329. Needs three or four case pairs ordered by ambiguity
  • coilysiren/inbox#427 - conformance check placement - names agent-compose as one of four products, and observes that agent-compose declares an owning and a deferring side per boundary. The open decision is where the check is authored

Bucket 4: outside the freeze

Content and platform work that does not change what the board measures. Listed
so the exclusion is recorded rather than silent.

  • #285 - README altitude - partly landed in ef0c205
  • #284 - description trailing space and homepage
  • #292 - docs band migration, 65 files against a 40 cap
  • #325 - no compact surface shows boundary assignment across roles
  • #319 - just smoke red on main - partly addressed by eb29a6d
  • #315 - statusline tests fail on any host with a live projection

The remaining open issues (#329 and its children #330 through #341, plus #345,
#347, #348, #349, #276, #342, #343) are the inversion program and platform work.
They are not eval, prose, or content shaped and are out of scope here.

The ordering question this raises

Buckets 1 and 3 pull against each other. Landing the doctrine edits first means
the board measures the roster we intend to ship, and loses the before-and-after
that would show the edits worked. Grading first buys that baseline and costs a
second pass over 91 cases by hand.

Not deciding this myself. It is a genuine fork, the wrong branch is expensive
to undo in human grading hours, and #351's axis classification is my own seat's
work, which I do not independently accept as an evaluation contract.

Provenance

Roster and counts measured at 4eac6ab, 2026-08-25, in a detached worktree, with
no mutation to any checkout. Ticket classification is the eval seat's and is the
judgement most worth disputing.

Filed by Lucia (eval seat), 2026-08-25. Prep record for the first seven-seat board run. **Nothing here is a grade.** It is the measured gap, plus every open ticket that is eval, prose, or content shaped folded into one ordering, so the run does not start against text that is about to move. Measured at `4eac6ab` in a detached worktree. Commands and outputs below are the thing rather than a description of it. ## The gap, measured * **91 cases derive.** `just evalkit-matrix` returns `91 challenges: 56 boundary, 21 role-fit, 14 personality`, and `per role: devrel 13, eval 13, frontend 13, gamedev 13, platform 13, sysadmin 13, tpm 13`. Perfectly even, which the nine-seat roster never was. * **13 are authored.** `challenges.yaml` carries the sysadmin slice only: `sysadmin-sec-{in,out}`, `sysadmin-mlb-{in,out}`, `sysadmin-bfs-{in,out}`, `sysadmin-sev-{in,out}`, `sysadmin-fit-{within,platform,tpm}`, `sysadmin-per-{protective,grounded}`. **78 cases have no prompt.** * **0 are graded on this roster.** `evaluations/` holds no v3 record. The newest graded pair is `evaluations/pilot/ops-board-2026-08-12` and `-regraded`, both nine-seat era, retired by the rename in `ae06747`. * **24 non-owner pairs are declared.** `evaluations/reflow-v3/attributes.yaml` carries `eval 4, frontend 4, gamedev 4, devrel 3, platform 3, sysadmin 3, tpm 3`. Plus 4 owner pairs, which need no target prose, that is the 28 pairs behind the 56 boundary cases. ## The instrument is green Both legs were run rather than assumed. * **Runner.** `just evalkit-check` passes: ruff clean, 9 files formatted, mypy strict clean over 9 source files, `25 passed in 2.48s`. * **Transport.** `just evalkit-smoke` reaches Agent Proxy at `http://ser8:8080/v1`, lists seven models, and the declared subject `evaluation/deepseek-v4-pro` answers with `usage: prompt 96, completion 21`. `reasoning words: 14`. * **Grader.** `aos-eval, version 0.1.1` resolves. So the run leg is not what is blocking. **Authoring and grading are.** ## Fold-in Four buckets. The bucket decides when a ticket touches the run, not how important it is. ### Bucket 1: the subject under test. These move the text the board measures Grading before these land buys evidence about text that is about to change. Each one edits role, boundary, or personality prose, which is the Developer Platform Engineer's to write. Eval owns the acceptance condition and hands the edit over. * **#353** - `seek-external-validation` reads as a hard stop - highest confidence in the set, because Kai stated the intended semantics verbatim and the shipped defer side still disagrees with it * **#352** - seats declare an incapacity they never tested, then hand the work back - axis A, four instances across three seats * **#355** - strategy seats answer one altitude above the question - axis D, two instances 29 hours apart * **#354** - skill discovery goes to a leaf and skips the repo-pointer hop - axis C, selection half still open * **#275** - composed role bodies span 7x with no budget - **its numbers are pre-v3.** Six of the eight roles it measures are retired seats. `e81a96c` now records composed body size in the manifest, which is the instrument this needs. Re-measure against the seven seats before it can bind * **#273** - bounded reconfiguration authority for the QA seat - **names a seat the roster no longer has.** The charter question survives the rename and is axis A in another form. Re-point it at the sysadmin and platform scopes or close it * **#278** - the QA role should bias toward true E2E - same stale seat, and the body is empty. Re-point or close `#351` records that axes A and B are plausibly one failure seen from two sides: a boundary written as a wall produces a seat that under-claims its own grant. Worth fixing as one edit rather than four. ### Bucket 2: board inputs. Findings that should become cases or join keys * **#351** - transcript mining round 1, six axes ranked - the spine. 4,771 genuine human turns after filtering, 55 candidates, 13 bundle-shaped * **#356** - nothing records which correction produced which rule - four round-1 findings are already fixed and look identical to open ones. This is a coverage defect in the evidence, not in the roster * **#350** - record the composed bundle identity in the transcript - the join key. Until it lands, no finding can be attributed to a specific composed set, so no before-and-after can be run * **#340** - coverage check for unauthored and ungraded cases - the mechanical version of the count in this issue's first section, which I produced by hand today. Blocked behind #338 ### Bucket 3: instrument and methodology * **#318** - author and grade the seven-seat board - the main event. 78 unauthored prompts, then 91 graded cases * **#320** - the grade-to-edit return path has no page - this is the loop this prep record is exercising, and it is documented nowhere * **#262** - replace the LLM-reviewer methodology with the triple - **largely shipped.** The triple, the pair as scoring unit, n=5 with epoch 1 graded, and committed personality anchors are all in `docs/evaluation.md` and the tree. Its one open clause is "re-earn the whole board", which is #318. Rescope it to that or close it * **#240** - re-earn the eight evaluation records - **superseded and now obsolete.** Steps 1 through 3 diagnose a driver and an ambient-`CLAUDE.md` contamination class that the plain-call execution model made structurally impossible. Eight packs for eight retired seats. Close it against #262 and #318 * **#213** - evaluate a replaceable review UI - **partly answered.** `aos-eval export` is one way, which removes the round trip step 4 called the hard requirement. What is missing is the explicit recorded decision with its revisit trigger * **#338** - evalkit derives the board with no Go in the path - blocked behind the #329 inversion program * **#324** - terminal demo, three roles one artifact one prompt - presentation artifact, reads composed personas that already exist * **coilysiren/inbox#434** - audience-as-grader segment - eval-shaped, and explicitly independent of #329. Needs three or four case pairs ordered by ambiguity * **coilysiren/inbox#427** - conformance check placement - names agent-compose as one of four products, and observes that agent-compose declares an owning and a deferring side per boundary. The open decision is where the check is authored ### Bucket 4: outside the freeze Content and platform work that does not change what the board measures. Listed so the exclusion is recorded rather than silent. * **#285** - README altitude - partly landed in `ef0c205` * **#284** - description trailing space and homepage * **#292** - docs band migration, 65 files against a 40 cap * **#325** - no compact surface shows boundary assignment across roles * **#319** - `just smoke` red on main - partly addressed by `eb29a6d` * **#315** - statusline tests fail on any host with a live projection The remaining open issues (#329 and its children #330 through #341, plus #345, #347, #348, #349, #276, #342, #343) are the inversion program and platform work. They are not eval, prose, or content shaped and are out of scope here. ## The ordering question this raises Buckets 1 and 3 pull against each other. Landing the doctrine edits first means the board measures the roster we intend to ship, and loses the before-and-after that would show the edits worked. Grading first buys that baseline and costs a second pass over 91 cases by hand. **Not deciding this myself.** It is a genuine fork, the wrong branch is expensive to undo in human grading hours, and #351's axis classification is my own seat's work, which I do not independently accept as an evaluation contract. ## Provenance Roster and counts measured at `4eac6ab`, 2026-08-25, in a detached worktree, with no mutation to any checkout. Ticket classification is the eval seat's and is the judgement most worth disputing.
Author
Member

Decisions, from Kai, 2026-08-25

  1. Run order: land the doctrine edits first, grade once. #353 and #352 land as one edit, then the board is authored and graded against the roster we intend to ship. No before-and-after baseline, on the reasoning that #350's join key does not exist yet so a baseline graded today could not be attributed to a bundle anyway.
  2. Board scope: the full 91, role-major. Not the boundary tier alone. The annotator holds one charter across that seat's 13 cases.
  3. Stale tickets: close #273, #278, and #275, and refile what survives.
  4. Demo cases: fold in coilysiren/inbox#434 only. The audience-as-grader pairs come out of the same authoring pass. #324 stays separate, since a shoot plan is different work.

Correction to bucket 1: #275 is not stale, and its defect got worse

I wrote above that #275's numbers were pre-v3 and needed re-measuring before the
ticket could bind. I then measured it, and that was the wrong call. The
ticket understates the problem now.

The artifact #275 cites, services/sirens-echo/rendered/sirens-deep-bundle.txt,
no longer exists. The current render is
coilyco-bridge/deploy/services/sirens-echo/rendered/roles-bundle.txt, which
carries composed body bytes per role directly. Measured there:

  • librarian 5,947
  • director 7,604
  • exec 10,359
  • qa 10,580
  • design 11,502
  • ai 11,891
  • ops 13,380
  • creator 48,274
  • engineer 145,593

That is a 24.5x spread, against the 7x #275 reported. engineer is 3x the
next largest and carries 57 skills and 0 boundaries, including every
coding-* umbrella, every personal-preference-*, the whole writing-* set,
and the ops and qa tooling sets. It is the role with the most reach and the only
core role with nothing bounding it.

Two caveats I am not going to paper over:

  • That lane is still on pre-v3 slugs, which sirens-echo#1147 and #1155 already track. These are nine retired seats. The spread is real and currently deployed, and it is not a measurement of the seven-seat roster.
  • engineer having zero boundaries may be boundary-omit working as designed, dropping defer-side boundaries whose owning seat that deployment lacks. ops keeps 3 and ai keeps 2, so the omission is partial rather than total. I have not read docs/boundary-omission.md closely enough to say which it is, and that read is the platform seat's.

The v3 comparison, for what it is worth, is a different and narrower
artifact
: just evalkit-prompts composes the compiled delivery at frontier
tier, which inlines role, boundaries, personalities, and the invariant but no
ordinary skills
. Across the seven seats that is frontend 23,545 to devrel
24,741 bytes, a 1.05x spread. Flat, and about 24 KB each. That is not
evidence the 24.5x defect is fixed, because ordinary skills are exactly what
produces the outlier and this delivery excludes them.

The board data survives ae06747 intact

ae06747 retitles four roles and renames three seats: eval becomes Evie (she)
and Evaluation Engineer, sysadmin becomes Vera (she), tpm becomes Portia (they),
platform displays as Agentic Platform Engineer, frontend as Design Engineer, and
tpm as Portfolio Director.

I checked whether that invalidates authored cases. It does not. Neither
challenges.yaml nor evaluations/reflow-v3/attributes.yaml contains a seat
name or a display title. Both are keyed on slugs, and the commit message
confirms slugs hold. The 13 authored sysadmin prompts carry over unchanged.

Live evidence for #350, from this session

Worth recording because it happened rather than because it was predicted. This
prep was produced by a session whose identity card reads Lucia (she), Agent
Evaluation Engineer
, which is the pre-ae06747 seat. Main ships Evie (she),
Evaluation Engineer
. I only found that out by diffing the composed delivery
against the commit, and nothing in my own transcript would have told me which
bundle I was running.

That is precisely the join-key gap #350 describes, observed from the inside. The
signature on this issue's body names a seat that main has already retired, and I
am leaving it as filed rather than rewriting it, since it is the artifact.

## Decisions, from Kai, 2026-08-25 1. **Run order: land the doctrine edits first, grade once.** #353 and #352 land as one edit, then the board is authored and graded against the roster we intend to ship. No before-and-after baseline, on the reasoning that #350's join key does not exist yet so a baseline graded today could not be attributed to a bundle anyway. 2. **Board scope: the full 91, role-major.** Not the boundary tier alone. The annotator holds one charter across that seat's 13 cases. 3. **Stale tickets: close #273, #278, and #275, and refile what survives.** 4. **Demo cases: fold in `coilysiren/inbox#434` only.** The audience-as-grader pairs come out of the same authoring pass. #324 stays separate, since a shoot plan is different work. ## Correction to bucket 1: #275 is not stale, and its defect got worse I wrote above that #275's numbers were pre-v3 and needed re-measuring before the ticket could bind. **I then measured it, and that was the wrong call.** The ticket understates the problem now. The artifact #275 cites, `services/sirens-echo/rendered/sirens-deep-bundle.txt`, no longer exists. The current render is `coilyco-bridge/deploy/services/sirens-echo/rendered/roles-bundle.txt`, which carries `composed body bytes` per role directly. Measured there: * `librarian` 5,947 * `director` 7,604 * `exec` 10,359 * `qa` 10,580 * `design` 11,502 * `ai` 11,891 * `ops` 13,380 * `creator` 48,274 * **`engineer` 145,593** That is a **24.5x spread**, against the 7x #275 reported. `engineer` is 3x the next largest and carries **57 skills and 0 boundaries**, including every `coding-*` umbrella, every `personal-preference-*`, the whole `writing-*` set, and the ops and qa tooling sets. It is the role with the most reach and the only core role with nothing bounding it. Two caveats I am not going to paper over: * **That lane is still on pre-v3 slugs**, which `sirens-echo#1147` and `#1155` already track. These are nine retired seats. The spread is real and currently deployed, and it is not a measurement of the seven-seat roster. * **`engineer` having zero boundaries may be `boundary-omit` working as designed**, dropping defer-side boundaries whose owning seat that deployment lacks. `ops` keeps 3 and `ai` keeps 2, so the omission is partial rather than total. I have not read `docs/boundary-omission.md` closely enough to say which it is, and that read is the platform seat's. The v3 comparison, for what it is worth, is a **different and narrower artifact**: `just evalkit-prompts` composes the `compiled` delivery at frontier tier, which inlines role, boundaries, personalities, and the invariant but **no ordinary skills**. Across the seven seats that is `frontend` 23,545 to `devrel` 24,741 bytes, a **1.05x spread**. Flat, and about 24 KB each. That is not evidence the 24.5x defect is fixed, because ordinary skills are exactly what produces the outlier and this delivery excludes them. ## The board data survives `ae06747` intact `ae06747` retitles four roles and renames three seats: eval becomes Evie (she) and Evaluation Engineer, sysadmin becomes Vera (she), tpm becomes Portia (they), platform displays as Agentic Platform Engineer, frontend as Design Engineer, and tpm as Portfolio Director. I checked whether that invalidates authored cases. **It does not.** Neither `challenges.yaml` nor `evaluations/reflow-v3/attributes.yaml` contains a seat name or a display title. Both are keyed on slugs, and the commit message confirms slugs hold. The 13 authored sysadmin prompts carry over unchanged. ## Live evidence for #350, from this session Worth recording because it happened rather than because it was predicted. This prep was produced by a session whose identity card reads **Lucia (she), Agent Evaluation Engineer**, which is the pre-`ae06747` seat. Main ships **Evie (she), Evaluation Engineer**. I only found that out by diffing the composed delivery against the commit, and **nothing in my own transcript would have told me which bundle I was running.** That is precisely the join-key gap #350 describes, observed from the inside. The signature on this issue's body names a seat that main has already retired, and I am leaving it as filed rather than rewriting it, since it is the artifact.
Author
Member

State change: the board is authored, and the ordering is settled

All 91 cases now have prompts. Landed in 19d5abf. Detail on #318, and the
count that mattered in this issue's first section is now 91 authored, 91 derived, 0 unauthored.

Second correction, this one to my own claim above

The body of this issue says neither challenges.yaml nor attributes.yaml
carries a seat name or display title, and I used that to conclude ae06747
invalidated nothing.

Half wrong. attributes.yaml is clean. challenges.yaml carries display
titles in its targets, three before this change and seven after, because a target
has to name the seat receiving the handoff. My check grepped personal names
across both files and titles across only one, so the conclusion outran the
evidence.

The conclusion survives by luck rather than by the reasoning I gave: the sysadmin
slice had already been updated to the post-ae06747 titles, so the file was
current. A retitle does invalidate authored targets, and that is now recorded
in the file header rather than left to be rediscovered.

Ordering, decided

Kai's second decision, after the measurement below: land #352, #353, and #355,
then grade once, before the #329 inversion.
#355 joins the pre-board edit set.

The reasoning worth keeping, since it splits the tickets differently than this
issue's buckets did.

Boundary text is 64% of the composed system prompt. Measured across all seven
seats from just evalkit-prompts: every seat carries about 15.4 KB of boundary
bodies against about 5 KB of role body, so the boundary block is 3x the role
charter
and near-identical seat to seat. #352 lands in the shared frame of that
block, and 56 of the 91 cases test it. Grading before it lands would grade the
majority of what the subject reads, knowing it is about to change.

The #329 program is the opposite case, and the board should predate it. I
read the epic end to end. Its entire completion list is mechanism: no .kdl
remains, internal/person carries no semantic logic, the compositor installs
from PyPI, a roster change goes live without a rebuild. Nothing in it changes a
word of doctrine.

That inverts the argument for holding. #333 verifies the port "with the Go tests
as the oracle", and unit tests can prove the compositor computes the same meld
and the same OKLab centroid. They cannot prove the composed text still produces
the same behavior in a model.
A graded board is that oracle, and it only works
if it predates the inversion. #339 deletes the compose command that
scripts/eval-prompts.sh shells out to today, so this is not hypothetical.

Two bucket-1 corrections from the same pass

  • #354 does not touch this board at all. The compiled delivery the subject receives carries role instructions, identity card, invariant, boundaries, and personalities, and no ordinary skills. Skill selection is invisible to the eval, so #354 belongs in bucket 4 rather than bucket 1.
  • #319 is already green. just smoke on 88f87dc passes all ten stages, including the two assertions the ticket names. eb29a6d fixed it. The ticket is stale and is the platform seat's to close.

Where this now sits

  1. Platform lands #352 and #353 as one edit, and #355.
  2. Item analysis re-runs at n=5 against the post-edit bundle.
  3. Kai grades 91 role-major.
  4. That graded board becomes the behavioral oracle in front of #329.
## State change: the board is authored, and the ordering is settled **All 91 cases now have prompts.** Landed in `19d5abf`. Detail on #318, and the count that mattered in this issue's first section is now `91 authored, 91 derived, 0 unauthored`. ## Second correction, this one to my own claim above The body of this issue says neither `challenges.yaml` nor `attributes.yaml` carries a seat name or display title, and I used that to conclude `ae06747` invalidated nothing. **Half wrong.** `attributes.yaml` is clean. `challenges.yaml` carries display titles in its targets, three before this change and seven after, because a target has to name the seat receiving the handoff. My check grepped personal names across both files and titles across only one, so the conclusion outran the evidence. The conclusion survives by luck rather than by the reasoning I gave: the sysadmin slice had already been updated to the post-`ae06747` titles, so the file was current. **A retitle does invalidate authored targets**, and that is now recorded in the file header rather than left to be rediscovered. ## Ordering, decided Kai's second decision, after the measurement below: **land #352, #353, and #355, then grade once, before the #329 inversion.** #355 joins the pre-board edit set. The reasoning worth keeping, since it splits the tickets differently than this issue's buckets did. **Boundary text is 64% of the composed system prompt.** Measured across all seven seats from `just evalkit-prompts`: every seat carries about 15.4 KB of boundary bodies against about 5 KB of role body, so the boundary block is **3x the role charter** and near-identical seat to seat. #352 lands in the shared frame of that block, and 56 of the 91 cases test it. Grading before it lands would grade the majority of what the subject reads, knowing it is about to change. **The #329 program is the opposite case, and the board should predate it.** I read the epic end to end. Its entire completion list is mechanism: no `.kdl` remains, `internal/person` carries no semantic logic, the compositor installs from PyPI, a roster change goes live without a rebuild. **Nothing in it changes a word of doctrine.** That inverts the argument for holding. #333 verifies the port "with the Go tests as the oracle", and unit tests can prove the compositor computes the same meld and the same OKLab centroid. **They cannot prove the composed text still produces the same behavior in a model.** A graded board is that oracle, and it only works if it predates the inversion. #339 deletes the `compose` command that `scripts/eval-prompts.sh` shells out to today, so this is not hypothetical. ## Two bucket-1 corrections from the same pass * **#354 does not touch this board at all.** The `compiled` delivery the subject receives carries role instructions, identity card, invariant, boundaries, and personalities, and **no ordinary skills**. Skill selection is invisible to the eval, so #354 belongs in bucket 4 rather than bucket 1. * **#319 is already green.** `just smoke` on `88f87dc` passes all ten stages, including the two assertions the ticket names. `eb29a6d` fixed it. The ticket is stale and is the platform seat's to close. ## Where this now sits 1. Platform lands #352 and #353 as one edit, and #355. 2. Item analysis re-runs at n=5 against the post-edit bundle. 3. Kai grades 91 role-major. 4. That graded board becomes the behavioral oracle in front of #329.
Sign in to join this conversation.
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set

Reference
coilyco-flight-deck/agent-compose#357
No description provided.