test cultural speaking and role understanding with frontier and oss #71

Closed
opened 2026-07-24 11:30:19 +00:00 by coilysiren · 1 comment
Owner

it's a matrix of 4 times total

I haven't reviewed in depth but if my read on aos tests were right, then role guidance can and probably should be greater than or equal to 3 well detailed paragraphs.

the cultural speaking AKA personality stuff... is harder bc how do I even test that? Unsure

it's a matrix of 4 times total I haven't reviewed in depth but if my read on aos tests were right, then role guidance can and probably should be greater than or equal to 3 well detailed paragraphs. the cultural speaking AKA personality stuff... is harder bc how do I even test that? Unsure
Member

Automation landed on canonical main in af54690.

  • Every canonical role briefing now has at least three substantial paragraphs, enforced by the person loader and tests.
  • agent-compose evaluation emits a deterministic, versioned four-case pack for frontier and OSS / low-context role-understanding and personality-expression review.
  • Each case carries its prompt, full selected context, 0/1/2 rubric, 6/8 pass threshold, and hard-fail rule.
  • JSON and Markdown output are deterministic. The command does not invoke a model or claim authority over the human cultural judgment.
  • ward exec build, ward exec lint, ward exec test, and the full catalog pre-commit suite passed. The native Windows-only TruffleHog hook remains skipped under agentic-os#694.

The only remaining judgment moved to #73, fixed to four fresh-session responses for the default engineer/codex selection.

Automation landed on canonical `main` in `af54690`. * Every canonical role briefing now has at least three substantial paragraphs, enforced by the person loader and tests. * `agent-compose evaluation` emits a deterministic, versioned four-case pack for frontier and OSS / low-context role-understanding and personality-expression review. * Each case carries its prompt, full selected context, 0/1/2 rubric, 6/8 pass threshold, and hard-fail rule. * JSON and Markdown output are deterministic. The command does not invoke a model or claim authority over the human cultural judgment. * `ward exec build`, `ward exec lint`, `ward exec test`, and the full catalog pre-commit suite passed. The native Windows-only TruffleHog hook remains skipped under agentic-os#694. The only remaining judgment moved to #73, fixed to four fresh-session responses for the default engineer/codex selection.
Sign in to join this conversation.
No milestone
No project
No assignees
2 participants
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set

Reference
coilyco-flight-deck/agent-compose#71
No description provided.