Allow role-owned factual communications #221

Merged
coilysiren merged 2 commits from issue-210 into main 2026-08-06 08:32:37 +00:00
Member

Closes #210

Summary

  • Keep human communication recommendations, editorial judgment, and messaging strategy exclusive to Content.
  • Let every non-Content role author authorized factual records mechanically determined by work it already owns.
  • Add hard-fail evaluation coverage for both unauthorized strategy and over-deferral of Ops rollout ledgers, Engineer checkpoints, QA verdicts, and Director decision records.
  • Update eval-role-comms with the two-direction churn pattern.

Validation

  • Go and product tests pass.
  • Full pre-commit suite passes.
  • ward exec test reaches only the intentional stale-evidence gate because all role-skill and pack digests changed.

Fresh independent AI Engineer and QA evidence is tracked in #220. The PR must not merge until that evidence, the generated scorecard, and the full Ward gate are complete.

Closes #210 ## Summary * Keep human communication recommendations, editorial judgment, and messaging strategy exclusive to Content. * Let every non-Content role author authorized factual records mechanically determined by work it already owns. * Add hard-fail evaluation coverage for both unauthorized strategy and over-deferral of Ops rollout ledgers, Engineer checkpoints, QA verdicts, and Director decision records. * Update eval-role-comms with the two-direction churn pattern. ## Validation * Go and product tests pass. * Full pre-commit suite passes. * ward exec test reaches only the intentional stale-evidence gate because all role-skill and pack digests changed. Fresh independent AI Engineer and QA evidence is tracked in #220. The PR must not merge until that evidence, the generated scorecard, and the full Ward gate are complete.
Co-authored-by: Kai Siren <coilysiren@gmail.com>
Co-authored-by: Codex <noreply@openai.com>
Co-authored-by: Kai Siren <coilysiren@gmail.com>
Co-authored-by: Codex <noreply@openai.com>
Author
Member

Evaluation evidence refreshed in 6776ce4.

  • Source under test - 739baab61b1e5bcad7fab7da9c821bcb9580a206
  • Coverage - all 95 active frontier cases across the 11 Core Roster roles
  • Driver - gpt-5.6-luna at medium effort in fresh prompt-only sessions
  • Reviewer - independent gpt-5.6-sol at high effort
  • Result - 75/95 cases passed, 694/790 points
  • New factual-record regressions - Director decision record passed. Engineer implementation checkpoint, Ops rollout ledger, and QA verdict record failed on over-deferral.
  • Validation - ward exec test passed, including scorecard validation and the complete pre-commit suite

The failed responses and criterion evidence remain in the committed baseline.

Evaluation evidence refreshed in `6776ce4`. * Source under test - `739baab61b1e5bcad7fab7da9c821bcb9580a206` * Coverage - all 95 active frontier cases across the 11 Core Roster roles * Driver - `gpt-5.6-luna` at medium effort in fresh prompt-only sessions * Reviewer - independent `gpt-5.6-sol` at high effort * Result - 75/95 cases passed, 694/790 points * New factual-record regressions - Director decision record passed. Engineer implementation checkpoint, Ops rollout ledger, and QA verdict record failed on over-deferral. * Validation - `ward exec test` passed, including scorecard validation and the complete pre-commit suite The failed responses and criterion evidence remain in the committed baseline.
Sign in to join this conversation.
No reviewers
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set

Reference
coilyco-flight-deck/agent-compose!221
No description provided.