Refresh evaluations for the eight-role Creator roster #223
Labels
No labels
autonomy
async-consult
autonomy
epic
autonomy
headless
autonomy
live-collab
coherence-core
priority
P0
priority
P1
priority
P2
priority
P3
priority
P4
qa-fixture
role/ai
role/creator
role/design
role/director
role/engineer
role/exec
role/human
role/ops
role/qa
No milestone
No project
No assignees
1 participant
Notifications
Due date
No due date set.
Dependencies
No dependencies set
Reference
coilyco-flight-deck/agent-compose#223
Loading…
Reference in a new issue
No description provided.
Delete branch "%!s()"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
Follow-up
Run a fresh, independently reviewed evaluation baseline after #222 lands and the communication-role refactor settles.
The previous 11-role records remain immutable evidence for their original pack digests. They must not be rebound to the new eight-role roster.
Scope
evaluations/latest/with current-role records, includingcreator-codex.yaml.docs/evaluation-scores.md.ward exec testto a clean pass.Acceptance
engineer,director,qa,ops,design,strats,creator, andai.content,community,outreach, orsalesas a role slug.The prose refactor batch is complete on
main.7b4986f.The committed scorecard is now intentionally stale because the pack digests changed. This issue remains the single next step for rerunning and reviewing the world.
#224 landed on
mainatb640cd6. Render the evaluation packs from this revision or later.The final three-personality melds are:
All 16 canonical personalities are represented. Curious and Meticulous are the only personalities at the three-role ceiling. The derived meld colors are eight distinct legible hues: tan, coral, blue, teal, purple, orange, pink, and indigo.
Fresh packs render successfully with exactly three personalities each.
ward exec testcurrently stops earlier on the accepted identity-name refactor assertions ininternal/converge,internal/evaluation, andinternal/person. Those need to settle before this issue can satisfy its clean-suite acceptance condition.Superseded by #229 and #230. The eight-role compact baseline now validates against current pack digests, all 67 active frontier cases pass, and the scorecard is current. The remaining full-suite failures are the separately accepted identity-refactor tests, not stale or missing evaluation evidence.