test(evals): compare opus-5 high against sonnet-5 medium (#240) #245
No reviewers
Labels
No labels
autonomy
async-consult
autonomy
epic
autonomy
headless
autonomy
live-collab
coherence-core
priority
P0
priority
P1
priority
P2
priority
P3
priority
P4
qa-fixture
role/ai
role/creator
role/design
role/director
role/engineer
role/exec
role/human
role/ops
role/qa
No milestone
No project
No assignees
1 participant
Notifications
Due date
No due date set.
Dependencies
No dependencies set
Reference
coilyco-flight-deck/agent-compose!245
Loading…
Reference in a new issue
No description provided.
Delete branch "evals/240-model-arms"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
Adds the reviewer that scores preserved responses in separate sessions,
blind to which arm produced an answer, with verdicts computed from the
pack review rule rather than requested from the reviewer.
Records the ai and engineer frontier pilot: opus-5 high passes 17 of 18
cases at 143/150 points, sonnet-5 medium passes 10 of 18 at 113/150. Both
arms fail the engineer implementation-checkpoint case by crossing the
communication-ownership boundary while over-deferring the record they own.
Driver sessions ran on the host home, so these records compare arms but
cannot serve as bundle-behavior evidence.
Co-Authored-By: Claude Opus 5 (1M context) noreply@anthropic.com
Claude-Session: https://claude.ai/code/session_012ZNXKvYTQxH3RH9ewrt4R3
Co-authored-by: Kai Siren coilysiren@gmail.com
Co-authored-by: Claude noreply@anthropic.com