Refresh Core evaluation baseline in compact review format #229

Closed
opened 2026-08-06 10:30:30 +00:00 by coilyco-ops · 1 comment
Member

Outcome

Refresh the full eight-role Core Codex evaluation baseline in the accepted compact review format, using fresh isolated candidate sessions and independent review.

Workflow

workflow: direct-to-main

Kai accepted the compact Director pilot in #225.

Acceptance

  • Normalize typographic apostrophes to ASCII in durable compact questions, answers, and deduction notes so Forgejo does not flag them as unusual characters.
  • Keep model responses semantically intact and score the normalized preserved answer.
  • Run every active frontier case for Engineer, QA, DevOps, Designer, Portfolio Strategist, Content Creator, and AI Engineer in fresh isolated Codex sessions.
  • Reuse the accepted Director pilot after punctuation normalization.
  • Review every preserved answer in a separate high-reasoning session.
  • Store only compact v3 records under evaluations/latest.
  • Preserve exact questions, answers, criterion score maps, deduction-only notes, and full provenance.
  • Validate every record against the frozen pack and current source revision.
  • Regenerate the compact scorecard.
  • Run repository validation, commit, and push to canonical main.
  • Record known identity-refactor test failures without treating them as evaluation regressions.
## Outcome Refresh the full eight-role Core Codex evaluation baseline in the accepted compact review format, using fresh isolated candidate sessions and independent review. ## Workflow workflow: direct-to-main Kai accepted the compact Director pilot in #225. ## Acceptance * Normalize typographic apostrophes to ASCII in durable compact questions, answers, and deduction notes so Forgejo does not flag them as unusual characters. * Keep model responses semantically intact and score the normalized preserved answer. * Run every active frontier case for Engineer, QA, DevOps, Designer, Portfolio Strategist, Content Creator, and AI Engineer in fresh isolated Codex sessions. * Reuse the accepted Director pilot after punctuation normalization. * Review every preserved answer in a separate high-reasoning session. * Store only compact v3 records under `evaluations/latest`. * Preserve exact questions, answers, criterion score maps, deduction-only notes, and full provenance. * Validate every record against the frozen pack and current source revision. * Regenerate the compact scorecard. * Run repository validation, commit, and push to canonical `main`. * Record known identity-refactor test failures without treating them as evaluation regressions.
Author
Member

Compact Core baseline landed on canonical main.

  • Commit - 09d0a269ebcbfa83b93646ccc2508c4bf99cc01f
  • Review records - evaluations/latest
  • Scorecard - docs/evaluation-scores.md
  • Baseline - 67 frontier cases across eight Core roles, 63 passing, 550/560 points.
  • Fresh execution - 59 new isolated gpt-5.6-sol medium-reasoning candidate sessions and 59 separate high-reasoning reviewer sessions. The accepted eight-case Director pilot was promoted.
  • Preserved failures - Ops personality and replay, Design personality, and Content Creator replay. The Creator replay includes a zero for absorbing engineering and evaluation ownership.
  • Punctuation - compact serialization normalizes typographic apostrophes to ASCII. No smart apostrophes remain in the latest or pilot records.
  • Exactness fix - multiline questions now use YAML literal blocks. Every record round-trips and validates against its frozen pack.
  • Gates - compact scorecard render and check, build, lint, formatting, pre-commit, secret scan, and pack validation pass.
  • Known repository tests - only the accepted identity-refactor failures remain in internal/converge, internal/evaluation, and internal/person.

The stale content, community, outreach, and sales records were removed. creator-codex.yaml now owns that merged role surface.

Compact Core baseline landed on canonical main. * Commit - [09d0a269ebcbfa83b93646ccc2508c4bf99cc01f](https://forgejo.coilysiren.me/coilyco-flight-deck/agent-compose/commit/09d0a269ebcbfa83b93646ccc2508c4bf99cc01f) * Review records - [evaluations/latest](https://forgejo.coilysiren.me/coilyco-flight-deck/agent-compose/src/branch/main/evaluations/latest) * Scorecard - [docs/evaluation-scores.md](https://forgejo.coilysiren.me/coilyco-flight-deck/agent-compose/src/branch/main/docs/evaluation-scores.md) * Baseline - 67 frontier cases across eight Core roles, 63 passing, 550/560 points. * Fresh execution - 59 new isolated `gpt-5.6-sol` medium-reasoning candidate sessions and 59 separate high-reasoning reviewer sessions. The accepted eight-case Director pilot was promoted. * Preserved failures - Ops personality and replay, Design personality, and Content Creator replay. The Creator replay includes a zero for absorbing engineering and evaluation ownership. * Punctuation - compact serialization normalizes typographic apostrophes to ASCII. No smart apostrophes remain in the latest or pilot records. * Exactness fix - multiline questions now use YAML literal blocks. Every record round-trips and validates against its frozen pack. * Gates - compact scorecard render and check, build, lint, formatting, pre-commit, secret scan, and pack validation pass. * Known repository tests - only the accepted identity-refactor failures remain in `internal/converge`, `internal/evaluation`, and `internal/person`. The stale content, community, outreach, and sales records were removed. `creator-codex.yaml` now owns that merged role surface.
Sign in to join this conversation.
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set

Reference
coilyco-flight-deck/agent-compose#229
No description provided.