Refresh portfolio-native evaluation baselines for every previously scored role #109
Labels
No labels
autonomy
async-consult
autonomy
epic
autonomy
headless
autonomy
live-collab
coherence-core
priority
P0
priority
P1
priority
P2
priority
P3
priority
P4
qa-fixture
role/ai
role/creator
role/design
role/director
role/engineer
role/exec
role/human
role/ops
role/qa
No milestone
No project
No assignees
1 participant
Notifications
Due date
No due date set.
Dependencies
No dependencies set
Reference
coilyco-flight-deck/agent-compose#109
Loading…
Reference in a new issue
No description provided.
Delete branch "%!s()"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
Outcome
Refresh every stale portfolio-native evaluation baseline from one immutable
canonical revision after the profile, library, role-skill, seat, copy-contract,
and evaluation-matrix work has landed.
Independent QA owns rubric scoring. Kai does not manually score every response.
QA escalates only an ambiguous criterion, contradictory evidence, or a policy
decision the rubric cannot resolve.
Run contract: #115
Format dependency: #125
Exact baseline set
Refresh these eleven Codex records:
advisor-codex.yamlceo-codex.yamlcustomer-success-codex.yamldesigner-codex.yamldirector-codex.yamlengineer-codex.yamlops-codex.yamlpm-codex.yamlqa-codex.yamlsales-codex.yamlsocial-codex.yamlLeave
community-codex.yamlandcommunity-discord.yamlunchanged unlessresult-format migration requires a mechanical provenance update. Their existing
responses and scores remain accepted evidence.
Frontier cases use
gpt-5.6-sol. OSS cases useqwen3:4b. The run executes allfour cases for every listed role. CEO's OSS cases remain evidence for its
compatibility gate even while production composition keeps CEO frontier-only.
Result format
#125 owns
agent-compose.evaluation-result.v2, canonical pack digesting, and v1compatibility. This issue consumes that released format and does not define a
second result schema.
All eleven newly written records must validate against the exact immutable
source revision and pack digest.
Run protocol
submitted verbatim.
totals and verdicts from the unchanged rubric.
provenance records the retry.
role boundary to preserve an earlier verdict.
Done condition
issue, QA reviewer, evaluation date, retry provenance, and pack digest are
present.
ward exec testandward exec smokepass.mainand closes this issue.Landed in signed commit
fca7468. The refresh records all 11 requested Codex roles at one immutable source revision with exact v2 pack digests, explicit retry provenance, frozen raw responses, and independent QA evidence for every criterion. QA scored 19/44 cases passing: gpt-5.6-sol 19/22 and qwen3:4b 0/22. The current qwen3:4b outputs exposed reasoning, truncated, and sometimes invented facts, so the baselines preserve those failures. ward exec test, the complete pre-commit suite, and ward exec smoke pass.