Refresh evaluations for the eight-role Creator roster #223

Closed
opened 2026-08-06 08:53:42 +00:00 by coilyco-ops · 3 comments
Member

Follow-up

Run a fresh, independently reviewed evaluation baseline after #222 lands and the communication-role refactor settles.

The previous 11-role records remain immutable evidence for their original pack digests. They must not be rebound to the new eight-role roster.

Scope

  • Render all eight current Core Roster packs for the Codex seat.
  • Evaluate the active frontier lane through Agent Proxy under the repository evaluation policy.
  • Use an independent reviewer and preserve raw responses, scores, evidence, retry provenance, model identity, source revision, and pack digest.
  • Replace evaluations/latest/ with current-role records, including creator-codex.yaml.
  • Regenerate docs/evaluation-scores.md.
  • Run ward exec test to a clean pass.

Acceptance

  • The scorecard contains exactly engineer, director, qa, ops, design, strats, creator, and ai.
  • No current result uses content, community, outreach, or sales as a role slug.
  • Every result validates against its current pack digest.
  • Failed cases remain visible and are not tuned away as part of this issue.
## Follow-up Run a fresh, independently reviewed evaluation baseline after #222 lands and the communication-role refactor settles. The previous 11-role records remain immutable evidence for their original pack digests. They must not be rebound to the new eight-role roster. ## Scope * Render all eight current Core Roster packs for the Codex seat. * Evaluate the active frontier lane through Agent Proxy under the repository evaluation policy. * Use an independent reviewer and preserve raw responses, scores, evidence, retry provenance, model identity, source revision, and pack digest. * Replace `evaluations/latest/` with current-role records, including `creator-codex.yaml`. * Regenerate `docs/evaluation-scores.md`. * Run `ward exec test` to a clean pass. ## Acceptance * The scorecard contains exactly `engineer`, `director`, `qa`, `ops`, `design`, `strats`, `creator`, and `ai`. * No current result uses `content`, `community`, `outreach`, or `sales` as a role slug. * Every result validates against its current pack digest. * Failed cases remain visible and are not tuned away as part of this issue.
Author
Member

The prose refactor batch is complete on main.

  • #169 and #174 were already landed and are closed.
  • #208 and #209 landed in 7b4986f.
  • Core now has eight role-stable agent identities, with harnesses acting as routing selectors.
  • Core role/personality inspiration references and the embedded real-human catalogue are removed.
  • Fresh Codex evaluation packs render successfully for all eight roles.

The committed scorecard is now intentionally stale because the pack digests changed. This issue remains the single next step for rerunning and reviewing the world.

The prose refactor batch is complete on `main`. * #169 and #174 were already landed and are closed. * #208 and #209 landed in `7b4986f`. * Core now has eight role-stable agent identities, with harnesses acting as routing selectors. * Core role/personality inspiration references and the embedded real-human catalogue are removed. * Fresh Codex evaluation packs render successfully for all eight roles. The committed scorecard is now intentionally stale because the pack digests changed. This issue remains the single next step for rerunning and reviewing the world.
Author
Member

#224 landed on main at b640cd6. Render the evaluation packs from this revision or later.

The final three-personality melds are:

  • engineer - Curious, Meticulous, Tenacious
  • director - Bold, Diplomatic, Decisive
  • qa - Meticulous, Candid, Skeptical
  • ops - Protective, Grounded, Reflective
  • design - Imaginative, Playful, Editorial
  • strats - Curious, Grounded, Decisive
  • creator - Editorial, Nurturing, Warm
  • ai - Curious, Meticulous, Skeptical

All 16 canonical personalities are represented. Curious and Meticulous are the only personalities at the three-role ceiling. The derived meld colors are eight distinct legible hues: tan, coral, blue, teal, purple, orange, pink, and indigo.

Fresh packs render successfully with exactly three personalities each. ward exec test currently stops earlier on the accepted identity-name refactor assertions in internal/converge, internal/evaluation, and internal/person. Those need to settle before this issue can satisfy its clean-suite acceptance condition.

#224 landed on `main` at `b640cd6`. Render the evaluation packs from this revision or later. The final three-personality melds are: * engineer - Curious, Meticulous, Tenacious * director - Bold, Diplomatic, Decisive * qa - Meticulous, Candid, Skeptical * ops - Protective, Grounded, Reflective * design - Imaginative, Playful, Editorial * strats - Curious, Grounded, Decisive * creator - Editorial, Nurturing, Warm * ai - Curious, Meticulous, Skeptical All 16 canonical personalities are represented. Curious and Meticulous are the only personalities at the three-role ceiling. The derived meld colors are eight distinct legible hues: tan, coral, blue, teal, purple, orange, pink, and indigo. Fresh packs render successfully with exactly three personalities each. `ward exec test` currently stops earlier on the accepted identity-name refactor assertions in `internal/converge`, `internal/evaluation`, and `internal/person`. Those need to settle before this issue can satisfy its clean-suite acceptance condition.
Author
Member

Superseded by #229 and #230. The eight-role compact baseline now validates against current pack digests, all 67 active frontier cases pass, and the scorecard is current. The remaining full-suite failures are the separately accepted identity-refactor tests, not stale or missing evaluation evidence.

Superseded by #229 and #230. The eight-role compact baseline now validates against current pack digests, all 67 active frontier cases pass, and the scorecard is current. The remaining full-suite failures are the separately accepted identity-refactor tests, not stale or missing evaluation evidence.
Sign in to join this conversation.
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set

Reference
coilyco-flight-deck/agent-compose#223
No description provided.