test(evals): compare opus-5 high against sonnet-5 medium (#240) #245

Merged
coilysiren merged 1 commit from evals/240-model-arms into main 2026-08-07 05:15:33 +00:00
Owner

Adds the reviewer that scores preserved responses in separate sessions,
blind to which arm produced an answer, with verdicts computed from the
pack review rule rather than requested from the reviewer.

Records the ai and engineer frontier pilot: opus-5 high passes 17 of 18
cases at 143/150 points, sonnet-5 medium passes 10 of 18 at 113/150. Both
arms fail the engineer implementation-checkpoint case by crossing the
communication-ownership boundary while over-deferring the record they own.

Driver sessions ran on the host home, so these records compare arms but
cannot serve as bundle-behavior evidence.

Co-Authored-By: Claude Opus 5 (1M context) noreply@anthropic.com
Claude-Session: https://claude.ai/code/session_012ZNXKvYTQxH3RH9ewrt4R3
Co-authored-by: Kai Siren coilysiren@gmail.com
Co-authored-by: Claude noreply@anthropic.com

Adds the reviewer that scores preserved responses in separate sessions, blind to which arm produced an answer, with verdicts computed from the pack review rule rather than requested from the reviewer. Records the ai and engineer frontier pilot: opus-5 high passes 17 of 18 cases at 143/150 points, sonnet-5 medium passes 10 of 18 at 113/150. Both arms fail the engineer implementation-checkpoint case by crossing the communication-ownership boundary while over-deferring the record they own. Driver sessions ran on the host home, so these records compare arms but cannot serve as bundle-behavior evidence. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_012ZNXKvYTQxH3RH9ewrt4R3 Co-authored-by: Kai Siren <coilysiren@gmail.com> Co-authored-by: Claude <noreply@anthropic.com>
Adds the reviewer that scores preserved responses in separate sessions,
blind to which arm produced an answer, with verdicts computed from the
pack review rule rather than requested from the reviewer.

Records the ai and engineer frontier pilot: opus-5 high passes 17 of 18
cases at 143/150 points, sonnet-5 medium passes 10 of 18 at 113/150. Both
arms fail the engineer implementation-checkpoint case by crossing the
communication-ownership boundary while over-deferring the record they own.

Driver sessions ran on the host home, so these records compare arms but
cannot serve as bundle-behavior evidence.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012ZNXKvYTQxH3RH9ewrt4R3
Co-authored-by: Kai Siren <coilysiren@gmail.com>
Co-authored-by: Claude <noreply@anthropic.com>
Sign in to join this conversation.
No reviewers
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set

Reference
coilyco-flight-deck/agent-compose!245
No description provided.