Measure fine-grained local composed-role instruction adherence #653

Open
opened 2026-07-23 14:55:36 +00:00 by coilyco-ops · 0 comments
Member

Context

AOS #652 proved that the selected local model receives all ten composed roles, identifies each role, and answers a real role-specific question. The answers also exposed finer instruction-following drift that the role-routing marker intentionally does not measure. Several responses used em dashes, and one QA response used a prose table despite the universal voice rules.

This does not block role selection or composed-skill delivery. It is a separate model-quality and instruction-density question.

Scope

  • Add a repeatable evaluation layer for universal instruction adherence after role selection is proven.
  • Separate projection failures, role-selection failures, and model instruction-following failures in results.
  • Cover the universal prose rules that can be checked deterministically without grading answer quality.
  • Preserve the real role-specific questions and bounded one-role-at-a-time execution from AOS #652.
  • Avoid turning one stochastic answer into a model-wide verdict. Define the sample count and reporting threshold.

Acceptance

  • The probe can report role confirmation separately from universal-style adherence.
  • Tests cover case and display-name normalization plus each deterministic style rule selected for enforcement.
  • Documentation states which findings fail a run, which remain advisory, and how repeated samples are interpreted.
  • No credential, backend address, or tracked transcript enters the repository.

Dependency and boundary

  • Context campaign: AOS #655
  • Run after AOS #656 establishes the structural payload and after the first reduction pass is remeasured.
  • A style or adherence failure must not be attributed to role composition until the structural report proves the expected doctrine and selected skill metadata reached that lane.
  • This ticket remains a behavioral evaluation. It does not own AGENTS.md clipping, skill placement, MCP mounting, or harness adapter measurement.
## Context AOS #652 proved that the selected local model receives all ten composed roles, identifies each role, and answers a real role-specific question. The answers also exposed finer instruction-following drift that the role-routing marker intentionally does not measure. Several responses used em dashes, and one QA response used a prose table despite the universal voice rules. This does not block role selection or composed-skill delivery. It is a separate model-quality and instruction-density question. ## Scope * Add a repeatable evaluation layer for universal instruction adherence after role selection is proven. * Separate projection failures, role-selection failures, and model instruction-following failures in results. * Cover the universal prose rules that can be checked deterministically without grading answer quality. * Preserve the real role-specific questions and bounded one-role-at-a-time execution from AOS #652. * Avoid turning one stochastic answer into a model-wide verdict. Define the sample count and reporting threshold. ## Acceptance * The probe can report role confirmation separately from universal-style adherence. * Tests cover case and display-name normalization plus each deterministic style rule selected for enforcement. * Documentation states which findings fail a run, which remain advisory, and how repeated samples are interpreted. * No credential, backend address, or tracked transcript enters the repository. ## Dependency and boundary * Context campaign: AOS #655 * Run after AOS #656 establishes the structural payload and after the first reduction pass is remeasured. * A style or adherence failure must not be attributed to role composition until the structural report proves the expected doctrine and selected skill metadata reached that lane. * This ticket remains a behavioral evaluation. It does not own AGENTS.md clipping, skill placement, MCP mounting, or harness adapter measurement.
Sign in to join this conversation.
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set

Reference
coilyco-flight-deck/agentic-os#653
No description provided.