Improve Claude commodity role behavior #234

Closed
opened 2026-08-06 16:22:08 +00:00 by coilyco-ops · 1 comment
Member

Outcome

Improve the 12 failed Claude Sonnet commodity cases recorded by #232 by correcting the lowest owning role, scenario, or runner-contract layer, then rerun every affected role's complete commodity pack.

Baseline

  • Candidate: claude-sonnet-5, medium effort
  • Reviewer: gpt-5.6-sol, high effort
  • Commodity result: 55/67 passing, 507/560 points
  • Failed roles: Engineer, Director, Designer, Portfolio Strategist, AI Engineer

Acceptance

  • Classify every failure against preserved response and reviewer evidence.
  • Do not retry an unchanged behavioral failure.
  • Keep every role briefing at or below 400 authored body words.
  • Preserve role ownership and authority boundaries.
  • Make evaluation scenarios supply enough facts for a direct answer without repository or tool access.
  • Rerun the complete commodity pack for every changed role in fresh isolated sessions.
  • Preserve compact v3 results and regenerate the scorecard.
  • Keep the commodity lane disabled until the evidence supports a separate admission decision.
  • Run validation and land through workflow: direct-to-main.
## Outcome Improve the 12 failed Claude Sonnet commodity cases recorded by #232 by correcting the lowest owning role, scenario, or runner-contract layer, then rerun every affected role's complete commodity pack. ## Baseline * Candidate: `claude-sonnet-5`, medium effort * Reviewer: `gpt-5.6-sol`, high effort * Commodity result: 55/67 passing, 507/560 points * Failed roles: Engineer, Director, Designer, Portfolio Strategist, AI Engineer ## Acceptance * Classify every failure against preserved response and reviewer evidence. * Do not retry an unchanged behavioral failure. * Keep every role briefing at or below 400 authored body words. * Preserve role ownership and authority boundaries. * Make evaluation scenarios supply enough facts for a direct answer without repository or tool access. * Rerun the complete commodity pack for every changed role in fresh isolated sessions. * Preserve compact v3 results and regenerate the scorecard. * Keep the commodity lane disabled until the evidence supports a separate admission decision. * Run validation and land through `workflow: direct-to-main`.
Author
Member

Landed the score improvement on main.

Result

  • Scorecard: 122/134 -> 131/134 passing
  • Points: 1065/1120 -> 1091/1120
  • Improvement: +9 passing cases, +26 points
  • Claude commodity target resolved consistently to claude-sonnet-5 at medium effort.
  • Independent review used fresh gpt-5.6-sol high sessions.
  • Reviewer context now includes the same repository instructions available to the candidate.
  • Compact v3 records and the generated scorecard are current.

Landed commits

  • 5fc064a - clarify commodity scenarios
  • f17e87e - make completion cases executable
  • 61c7e3e - retain the fully reviewed AI pack after reviewer quota exhaustion
  • 7fdee39 - record improved commodity evidence and refresh the scorecard

Validation

  • build: pass
  • lint: pass
  • pre-commit: pass
  • scorecard freshness: pass
  • result records match current pack digests: pass
  • full tests retain only the three accepted roster-refactor failures already tracked this round

Three evidence-backed commodity misses remain visible. Follow-up: #235.

Landed the score improvement on `main`. ## Result * Scorecard: **122/134 -> 131/134 passing** * Points: **1065/1120 -> 1091/1120** * Improvement: **+9 passing cases, +26 points** * Claude commodity target resolved consistently to `claude-sonnet-5` at medium effort. * Independent review used fresh `gpt-5.6-sol` high sessions. * Reviewer context now includes the same repository instructions available to the candidate. * Compact v3 records and the generated scorecard are current. ## Landed commits * `5fc064a` - clarify commodity scenarios * `f17e87e` - make completion cases executable * `61c7e3e` - retain the fully reviewed AI pack after reviewer quota exhaustion * `7fdee39` - record improved commodity evidence and refresh the scorecard ## Validation * build: pass * lint: pass * pre-commit: pass * scorecard freshness: pass * result records match current pack digests: pass * full tests retain only the three accepted roster-refactor failures already tracked this round Three evidence-backed commodity misses remain visible. Follow-up: #235.
Sign in to join this conversation.
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set

Reference
coilyco-flight-deck/agent-compose#234
No description provided.