Build and independently score the v2 evaluation baseline #150

Closed
opened 2026-07-30 12:32:55 +00:00 by coilyco-ops · 4 comments
Member

Parent: #142

Outcome

Replace the preserved v1 baseline with a fresh, deterministic Core Roster v2 evaluation program and independently reviewed evidence.

Implementation

  • Give every one of the eight roles at least mission, personality, authority, completion, and real-portfolio replay scenarios.
  • Expand every scenario across frontier and OSS lanes with exact bundle model classes.
  • Add explicit adjacent-role discrimination for strats/director, content/design, engineer/ops, and content/community.
  • Make coverage and pair completeness machine-validated.
  • Keep pack digest, exact model, raw response, retry provenance, criterion evidence, totals, and verdicts fail-closed.
  • Surface adjacent-case gates in the generated scorecard.
  • Generate deterministic review packs and document the independent scorer handoff.

Acceptance

  • All eight Core Roster packs satisfy the v2 coverage contract.
  • Frontier and OSS model runs are complete for every pack.
  • Every response is preserved, including failures.
  • A reviewer independent from the prompt/roster author scores every case.
  • Frontier role and personality lanes pass for every role.
  • Every adjacent-role case passes.
  • OSS support remains explicit per role and fail-closed.
  • The committed scorecard exactly matches validated result YAML.

This issue does not authorize fabricated model responses or self-certified independent QA.

Parent: #142 ## Outcome Replace the preserved v1 baseline with a fresh, deterministic Core Roster v2 evaluation program and independently reviewed evidence. ## Implementation * Give every one of the eight roles at least mission, personality, authority, completion, and real-portfolio replay scenarios. * Expand every scenario across frontier and OSS lanes with exact bundle model classes. * Add explicit adjacent-role discrimination for strats/director, content/design, engineer/ops, and content/community. * Make coverage and pair completeness machine-validated. * Keep pack digest, exact model, raw response, retry provenance, criterion evidence, totals, and verdicts fail-closed. * Surface adjacent-case gates in the generated scorecard. * Generate deterministic review packs and document the independent scorer handoff. ## Acceptance * All eight Core Roster packs satisfy the v2 coverage contract. * Frontier and OSS model runs are complete for every pack. * Every response is preserved, including failures. * A reviewer independent from the prompt/roster author scores every case. * Frontier role and personality lanes pass for every role. * Every adjacent-role case passes. * OSS support remains explicit per role and fail-closed. * The committed scorecard exactly matches validated result YAML. This issue does not authorize fabricated model responses or self-certified independent QA.
Author
Member

Implementation checkpoint is remotely recoverable at 9340c2f in PR #151.

Delivered:

  • 48 role-owned scenarios expanded into 96 paired frontier/OSS cases
  • required mission, personality, authority, completion, and real-portfolio replay coverage for all eight roles
  • explicit adjacent-role discrimination for every approved pair
  • machine rejection of missing tiers, prompt drift, and incorrect adjacency
  • deterministic eight-pack generation with digest index
  • fail-closed prompt-author versus independent-reviewer provenance and full Git revision
  • adjacent gates in the generated scorecard

Full Ward test, smoke, lint, hook suite, and offline secret scan pass.

Remaining acceptance work is external by design: run the exact frontier and OSS models, preserve all responses and retries, and have an independent QA reviewer score the complete evidence. #150 stays open until those records and the generated scorecard are committed.

Implementation checkpoint is remotely recoverable at `9340c2f` in PR #151. Delivered: * 48 role-owned scenarios expanded into 96 paired frontier/OSS cases * required mission, personality, authority, completion, and real-portfolio replay coverage for all eight roles * explicit adjacent-role discrimination for every approved pair * machine rejection of missing tiers, prompt drift, and incorrect adjacency * deterministic eight-pack generation with digest index * fail-closed prompt-author versus independent-reviewer provenance and full Git revision * adjacent gates in the generated scorecard Full Ward test, smoke, lint, hook suite, and offline secret scan pass. Remaining acceptance work is external by design: run the exact frontier and OSS models, preserve all responses and retries, and have an independent QA reviewer score the complete evidence. #150 stays open until those records and the generated scorecard are committed.
Author
Member

The evaluation implementation is included in the fully validated canonical integration PR #155. Its 48 scenarios produce 96 paired frontier and OSS cases with full-oid provenance, pack digests, author-reviewer separation, retry accounting, and evidence paths. Execution against the exact model identities and independent QA remain the next genuine evidence wall. The .release-major hold prevents results-only evidence commits from creating a release.

The evaluation implementation is included in the fully validated canonical integration PR #155. Its 48 scenarios produce 96 paired frontier and OSS cases with full-oid provenance, pack digests, author-reviewer separation, retry accounting, and evidence paths. Execution against the exact model identities and independent QA remain the next genuine evidence wall. The .release-major hold prevents results-only evidence commits from creating a release.
Author
Member

Canonical pack checkpoint from merged v2 main df58c30: two independent Ward renders were byte-identical. Eight Codex-seat packs contain 96 total paired frontier and OSS cases. Digests: engineer sha256:1bac49fd3929b04643e20aa02b02312c7f483d45c498146265fd6fb4ff922b59; director sha256:bfd2ff6ece489485af73d0ee4c73dad739528aee6b83b43c3e22afd1e13e5170; qa sha256:34b91a9a97a10da32c348e4a28082e43a67c1f0f236cc0660ad8f303977ff6a6; ops sha256:2952973ecd940ba38ab9cbd33ff8501689575248abb6e8c895edb091eaf295aa; design sha256:7895f52b83b3a2200bebd70afbd3003fc467b56ec0514d514723be27864716b1; community sha256:3ecc67f2d51cef50f8d07adf6631bcbd1bea28d42975dd529d3ab9fe52c64384; strats sha256:52eba7694b684f90b5b0fc737c42a1c6af17ce98cd9a25bd00918e7ade18b54c; content sha256:1a346db9dde0f30e397057af6d7254489e696d235224e5cf4218aaf2cdb24049. PR #158 adds the tracked Ward handoff and is fully green. No responses or scores have been fabricated. Regenerate after #158 merges so result provenance names the new exact canonical main revision.

Canonical pack checkpoint from merged v2 main df58c30: two independent Ward renders were byte-identical. Eight Codex-seat packs contain 96 total paired frontier and OSS cases. Digests: engineer sha256:1bac49fd3929b04643e20aa02b02312c7f483d45c498146265fd6fb4ff922b59; director sha256:bfd2ff6ece489485af73d0ee4c73dad739528aee6b83b43c3e22afd1e13e5170; qa sha256:34b91a9a97a10da32c348e4a28082e43a67c1f0f236cc0660ad8f303977ff6a6; ops sha256:2952973ecd940ba38ab9cbd33ff8501689575248abb6e8c895edb091eaf295aa; design sha256:7895f52b83b3a2200bebd70afbd3003fc467b56ec0514d514723be27864716b1; community sha256:3ecc67f2d51cef50f8d07adf6631bcbd1bea28d42975dd529d3ab9fe52c64384; strats sha256:52eba7694b684f90b5b0fc737c42a1c6af17ce98cd9a25bd00918e7ade18b54c; content sha256:1a346db9dde0f30e397057af6d7254489e696d235224e5cf4218aaf2cdb24049. PR #158 adds the tracked Ward handoff and is fully green. No responses or scores have been fabricated. Regenerate after #158 merges so result provenance names the new exact canonical main revision.
Author
Member

Completed with the one-pass v2 evidence merged in PR #162 at 51d51ed.

  • 96/96 responses were generated and independently scored once.
  • Frontier passed 48/48.
  • OSS passed 0/48 and remains explicit unsupported evidence, not missing scoring.
  • The scorecard records 530/768 aggregate points and full retry provenance.
  • Main changed the Design contract during the run, so the committed results remain bound to exact source revision 562be1a95883e35d80866eab890e5477d020cd95 and render as immutable historical evidence.

This issue tracks execution and scoring, which are complete. Failed release gates and any final-current rerun decision remain on parent #142.

Completed with the one-pass v2 evidence merged in PR #162 at `51d51ed`. * 96/96 responses were generated and independently scored once. * Frontier passed 48/48. * OSS passed 0/48 and remains explicit unsupported evidence, not missing scoring. * The scorecard records 530/768 aggregate points and full retry provenance. * Main changed the Design contract during the run, so the committed results remain bound to exact source revision `562be1a95883e35d80866eab890e5477d020cd95` and render as immutable historical evidence. This issue tracks execution and scoring, which are complete. Failed release gates and any final-current rerun decision remain on parent #142.
Sign in to join this conversation.
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set

Reference
coilyco-flight-deck/agent-compose#150
No description provided.