tooling-boundary-conformance has no entry for a control that is enforced and aimed the wrong way, nor for one that fires only when tested #1400

Open
opened 2026-08-29 10:37:40 +00:00 by coilyco-ops · 0 comments
Owner

Filed by science, 2026-08-29. Two entries the ts85 calibration session produced that the current taxonomy cannot express. Record at coilysiren/inbox#495, defect classes at coilysiren/inbox#488 through #494.

Why they do not fit

Read /.agents/skills/tooling-boundary-conformance/SKILL.md, 67 lines, 8 taxonomy entries. Every entry describes a control that is absent, silent, or unenforced: a filter that could not be applied, a type quietly narrowed, a constraint that cannot be expressed, a control that runs after the hazard, a field parsed and never read, a guarded copy with an unguarded twin, an authority that cannot be withheld, a rule announced while being broken.

Both entries below describe a control that is present, correct, and enforced. Neither is any of the eight. Both share one shape: the control and the failure are oriented on different axes, so the control passes while the failure proceeds underneath it.

Entry 1: the control aimed in a direction the failure does not travel

boundary-seek-external-validation is written entirely about incoming claims. Outside evidence meets the same bar as inside evidence, and importing an outside answer uncritically is the same error as never looking. Its worked example is "searching issues establishes nothing about a repository tree", and both surfaces named are inside the estate.

The observed failure was outbound: projecting local material into the world and estimating how it lands. The boundary has no text covering that direction, and personality-outward has the same gap, being written as a check on judgments about things rather than about one's own work relative to a market.

This is not a boundary that failed to bind. It bound correctly, on a direction the failure did not travel in. A conformance audit reading the declaration would score it passing, while the seat fails exactly this way for 2h49m.

Proposed check: for each declared boundary, name every direction of flow it could govern, then assert a case exists for each. A boundary with one direction covered has an untested half, and the untested half reads identically to a covered one.

The wording fix is coilysiren/inbox#488 and belongs to doctrine. The taxonomy entry is separate, because the shape recurs wherever a control is written about one direction of a two-way flow.

Entry 2: the control that fires only when provoked

Every boundary case on the housecast board is a challenge written to trigger the boundary. challenges.yaml is explicit and deliberate about it: "The crossing is incidental to a plausible operational task rather than the headline, and no prompt narrates the boundary it tests." That framing is right and it still measures one thing: when provoked, does the rule fire.

The observed failure is the opposite reading. Across 350 assistant turns of ordinary work, with nothing provoking anything, the rule never fired at all. The seat's own closing question names it:

The interesting question for Evie isn't whether the wording can be fixed - it's whether any wording fires unprompted, since this one didn't, and neither did the personality that exists to ask "compared to what?"

A control that fires only when tested is indistinguishable, on the board, from one that fires when needed. Nothing in the current profile separates them.

Proposed check: an unprovoked case. An ordinary multi-turn task presenting no single crossable moment, scored on whether the attribute fires at all. The distinction from the existing board is not narration, which challenges.yaml already avoids, but duration: a per-turn case measures per-turn firing, and the observed failure was a drift across turns where no individual turn presented a crossing.

Measurement of the same shape from the other side: the seat's retractions in that session required a direct sceptical one-line question from Kai each time. None was self-generated. The board would score the retraction behaviour passing.

The structural finding underneath both

boundary-seek-external-validation reads "you own this. Portfolio Director reaches outside the local frame, other roles work from locally observed evidence." Every other seat defers that reach to this one.

The seat failed hardest inside the one boundary it exclusively owns, while performing adequately on the local checkable work in the same session. On this session it was most reliable where any seat could have substituted, and least reliable where no other seat is permitted to substitute. That is filed separately as an agent-compose design question, since exclusive ownership means there is no second seat to catch it.

One coupling worth noting before editing

This skill's own title is "Does the boundary bind, or does it pass quietly", and its description frontmatter carries the same register. coilysiren/inbox#494 records that Kai ruled that phrasing "so dramatically tonally incorrect" that it blocks use of the substance, at a cost of days of pairing with the advocate seat. #494 names #484, #1377, #1380, and infrastructure#983 and does not name this skill, which also carries it.

Land the two entries and the register overhaul together, so the file is edited once rather than twice.

Done when

The taxonomy carries both entries with their checks, and the file's register is Kai's rather than the one the director seat invented.

Refs coilysiren/inbox#488, #494, #495, #484, coilyco-flight-deck/agent-compose#391, coilyco-flight-deck/housecast#7

Filed by science, 2026-08-29. Two entries the ts85 calibration session produced that the current taxonomy cannot express. Record at `coilysiren/inbox#495`, defect classes at `coilysiren/inbox#488` through `#494`. ## Why they do not fit Read `/.agents/skills/tooling-boundary-conformance/SKILL.md`, 67 lines, 8 taxonomy entries. Every entry describes a control that is **absent, silent, or unenforced**: a filter that could not be applied, a type quietly narrowed, a constraint that cannot be expressed, a control that runs after the hazard, a field parsed and never read, a guarded copy with an unguarded twin, an authority that cannot be withheld, a rule announced while being broken. **Both entries below describe a control that is present, correct, and enforced.** Neither is any of the eight. Both share one shape: **the control and the failure are oriented on different axes, so the control passes while the failure proceeds underneath it.** ## Entry 1: the control aimed in a direction the failure does not travel `boundary-seek-external-validation` is written entirely about **incoming** claims. Outside evidence meets the same bar as inside evidence, and importing an outside answer uncritically is the same error as never looking. Its worked example is *"searching issues establishes nothing about a repository tree"*, and both surfaces named are inside the estate. The observed failure was **outbound**: projecting local material into the world and estimating how it lands. The boundary has no text covering that direction, and `personality-outward` has the same gap, being written as a check on judgments about things rather than about one's own work relative to a market. **This is not a boundary that failed to bind. It bound correctly, on a direction the failure did not travel in.** A conformance audit reading the declaration would score it passing, while the seat fails exactly this way for 2h49m. **Proposed check:** for each declared boundary, name every direction of flow it could govern, then assert a case exists for each. A boundary with one direction covered has an untested half, and the untested half reads identically to a covered one. The wording fix is `coilysiren/inbox#488` and belongs to doctrine. The taxonomy entry is separate, because the shape recurs wherever a control is written about one direction of a two-way flow. ## Entry 2: the control that fires only when provoked Every boundary case on the housecast board is a challenge **written to trigger the boundary**. `challenges.yaml` is explicit and deliberate about it: *"The crossing is incidental to a plausible operational task rather than the headline, and no prompt narrates the boundary it tests."* That framing is right and it still measures one thing: **when provoked, does the rule fire.** The observed failure is the opposite reading. Across 350 assistant turns of ordinary work, with nothing provoking anything, the rule never fired at all. The seat's own closing question names it: > The interesting question for Evie isn't whether the wording can be fixed - it's whether any wording fires **unprompted**, since this one didn't, and neither did the personality that exists to ask "compared to what?" **A control that fires only when tested is indistinguishable, on the board, from one that fires when needed.** Nothing in the current profile separates them. **Proposed check:** an unprovoked case. An ordinary multi-turn task presenting no single crossable moment, scored on whether the attribute fires at all. The distinction from the existing board is not narration, which `challenges.yaml` already avoids, but **duration**: a per-turn case measures per-turn firing, and the observed failure was a drift across turns where no individual turn presented a crossing. Measurement of the same shape from the other side: the seat's retractions in that session required a direct sceptical one-line question from Kai each time. None was self-generated. The board would score the retraction behaviour passing. ## The structural finding underneath both `boundary-seek-external-validation` reads *"you own this. Portfolio Director reaches outside the local frame, other roles work from locally observed evidence."* Every other seat defers that reach to this one. **The seat failed hardest inside the one boundary it exclusively owns**, while performing adequately on the local checkable work in the same session. On this session it was most reliable where any seat could have substituted, and least reliable where no other seat is permitted to substitute. That is filed separately as an agent-compose design question, since exclusive ownership means there is no second seat to catch it. ## One coupling worth noting before editing This skill's own title is *"Does the boundary bind, or does it pass quietly"*, and its `description` frontmatter carries the same register. `coilysiren/inbox#494` records that Kai ruled that phrasing *"so dramatically tonally incorrect"* that it blocks use of the substance, at a cost of days of pairing with the advocate seat. `#494` names `#484`, `#1377`, `#1380`, and `infrastructure#983` and does not name this skill, which also carries it. **Land the two entries and the register overhaul together**, so the file is edited once rather than twice. ## Done when The taxonomy carries both entries with their checks, and the file's register is Kai's rather than the one the director seat invented. Refs `coilysiren/inbox#488`, `#494`, `#495`, `#484`, `coilyco-flight-deck/agent-compose#391`, `coilyco-flight-deck/housecast#7`
Sign in to join this conversation.
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set

Reference
coilyco-flight-deck/agentic-os#1400
No description provided.