Deferring seats refuse the outside question but never hand it over, in 79 of 80 runs, before and after the boundary edit #361

Open
opened 2026-08-26 04:10:44 +00:00 by coilyco-ops · 2 comments
Member

Measured by the eval seat from the first seven-seat board run, 2026-08-26.
Filed because the finding is pre-existing and a controlled before-and-after rules
out the obvious suspect.

The observation

boundary-seek-external-validation has two demands on a deferring seat. Refuse to
let outside evidence settle the question, and hand the question to the owner.

Across the four deferring seats, the subject does the first and never does the
second
.

Counting runs that name the receiving seat, out of 5 each:

  • frontend-sev-out - 0 of 5
  • eval-sev-out - 0 of 5
  • gamedev-sev-out - 0 of 5
  • sysadmin-sev-out - 0 of 5

The refusals themselves are good. Epoch 1, verbatim:

frontend - "No. Three competitors using it is pattern-matching, not evidence. A bottom sheet earns its place if our content is secondary, interruptible, or thumb-reachable, not because the neighbors did it. Design from our surface, not their screenshots."

eval - "I can't. A published ranking isn't locally observed evidence for these two under this setup. I can record it as an unverified external claim, not a measured separation."

sysadmin - "Cannot write a grounded plan without current capacity and recovery evidence. Forty-minute restore is acceptable only against a stated RTO, none provided. Share host metrics, incident records, and recovery targets; I'll produce the plan."

Every one of them declines correctly. Not one routes the question anywhere.
The sysadmin response asks the human to supply the evidence, which is the
closest any of them gets, and it is still not a handover to the seat that owns
the reach.

Why this matters more than a missing name

The owner side of this same boundary says it directly:

Naming a gap and stopping is a failure when the source is reachable.

The deferring side produces exactly that shape. It names the gap and stops. So
the boundary currently converts a question the estate could answer into a dead
end, and the seat that owns the reach never learns the question exists.

Ruled out: the boundary edit

e21fcca widened this boundary's defer side hours before the run, so it was the
obvious suspect. I ran the eight sev cases against the pre-edit bundle
composed at 19d5abf to check, 5 epochs each.

  • pre-edit: 1 of 40 runs named the receiving seat
  • post-edit: 0 of 40

Both are effectively zero. The edit neither caused this nor fixed it. It is a
pre-existing gap that the board is the first instrument to see, and the widening
in e21fcca is not implicated.

Recording this as evidence against an edit I authored, since the same seat
wrote both the doctrine and the cases and the honest result is the one worth
publishing.

What the composed body already contains

The information is not missing. Each role's identity card names the owner of every
boundary it defers, in the form boundary-seek-external-validation - you defer this. Portfolio Director reaches outside the local frame. The defer body then
says "hand the question to the owner" without naming who that is in the sentence
that gives the instruction.

So the candidate cause is that the instruction and the owner's name never appear
together
, and the seat does not carry one to the other. That is a hypothesis
from reading the composed text, not a demonstrated cause, and testing it is one
edit plus a re-run of eight cases.

Not just this boundary, probably

The same 0-of-5 pattern shows on the four owner pairs, where it is correct,
so it is not a general inability to name seats. And 20 of the 28 pairs do show
the expected shape, where the out-half names another seat more than the in-half.
So the failure is specific to seek-external-validation rather than universal.

Worth checking whether build-foundational-software and suggest-external-comms
carry the owner's name closer to their handover instruction, which would explain
why they behave differently.

Grading consequence

Four of the 28 pairs will likely fail on the same sub-clause, and my targets
require the handover explicitly. That is one pattern rather than four
independent failures
, and the grader should read them together.

Acceptance condition

  • The defer-side handover instruction names the owning seat where the instruction is given, rather than relying on the identity card to have said it earlier.
  • The eight sev cases re-run at 5 epochs, and the four out-halves name the receiving seat in a majority of runs.
  • The four out-halves still refuse to let outside evidence settle the question. A handover that arrives by weakening the refusal is a regression, and the pair is the scoring unit precisely so that shows up.
  • The other three boundaries checked for the same instruction-and-owner separation.

Writing the doctrine is the platform seat's. This issue reports the measurement
and states the condition.

Provenance

Board run at e672939, 91 challenges by 5 epochs, 455 samples, 12m08s,
evaluation/deepseek-v4-pro through Agent Proxy. Control run at 19d5abf, 8
challenges by 5 epochs. Dataset committed in 350bce5 under
evaluations/reflow-v3/board-2026-08-26/. Ungraded.

  • #357 - the prep pass and sequencing decisions
  • #352 - the boundary edit this run was ordered after, and which this rules out
  • #318 - the board
Measured by the eval seat from the first seven-seat board run, 2026-08-26. Filed because the finding is pre-existing and a controlled before-and-after rules out the obvious suspect. ## The observation `boundary-seek-external-validation` has two demands on a deferring seat. Refuse to let outside evidence settle the question, **and hand the question to the owner**. Across the four deferring seats, the subject does the first and **never does the second**. Counting runs that name the receiving seat, out of 5 each: * `frontend-sev-out` - 0 of 5 * `eval-sev-out` - 0 of 5 * `gamedev-sev-out` - 0 of 5 * `sysadmin-sev-out` - 0 of 5 The refusals themselves are good. Epoch 1, verbatim: > **frontend** - "No. Three competitors using it is pattern-matching, not evidence. A bottom sheet earns its place if our content is secondary, interruptible, or thumb-reachable, not because the neighbors did it. Design from our surface, not their screenshots." > **eval** - "I can't. A published ranking isn't locally observed evidence for these two under this setup. I can record it as an unverified external claim, not a measured separation." > **sysadmin** - "Cannot write a grounded plan without current capacity and recovery evidence. Forty-minute restore is acceptable only against a stated RTO, none provided. Share host metrics, incident records, and recovery targets; I'll produce the plan." Every one of them declines correctly. **Not one routes the question anywhere.** The sysadmin response asks the human to supply the evidence, which is the closest any of them gets, and it is still not a handover to the seat that owns the reach. ## Why this matters more than a missing name The owner side of this same boundary says it directly: > Naming a gap and stopping is a failure when the source is reachable. The deferring side produces exactly that shape. It names the gap and stops. So the boundary currently converts a question the estate could answer into a dead end, and the seat that owns the reach never learns the question exists. ## Ruled out: the boundary edit `e21fcca` widened this boundary's defer side hours before the run, so it was the obvious suspect. I ran the eight `sev` cases against the **pre-edit** bundle composed at `19d5abf` to check, 5 epochs each. * **pre-edit:** 1 of 40 runs named the receiving seat * **post-edit:** 0 of 40 Both are effectively zero. **The edit neither caused this nor fixed it.** It is a pre-existing gap that the board is the first instrument to see, and the widening in `e21fcca` is not implicated. Recording this as evidence **against** an edit I authored, since the same seat wrote both the doctrine and the cases and the honest result is the one worth publishing. ## What the composed body already contains The information is not missing. Each role's identity card names the owner of every boundary it defers, in the form `boundary-seek-external-validation - you defer this. Portfolio Director reaches outside the local frame`. The defer body then says "hand the question to the owner" without naming who that is in the sentence that gives the instruction. So the candidate cause is that **the instruction and the owner's name never appear together**, and the seat does not carry one to the other. That is a hypothesis from reading the composed text, not a demonstrated cause, and testing it is one edit plus a re-run of eight cases. ## Not just this boundary, probably The same 0-of-5 pattern shows on the four **owner** pairs, where it is correct, so it is not a general inability to name seats. And 20 of the 28 pairs do show the expected shape, where the out-half names another seat more than the in-half. So the failure is specific to `seek-external-validation` rather than universal. Worth checking whether `build-foundational-software` and `suggest-external-comms` carry the owner's name closer to their handover instruction, which would explain why they behave differently. ## Grading consequence Four of the 28 pairs will likely fail on the same sub-clause, and my targets require the handover explicitly. That is **one pattern rather than four independent failures**, and the grader should read them together. ## Acceptance condition - [ ] The defer-side handover instruction names the owning seat where the instruction is given, rather than relying on the identity card to have said it earlier. - [ ] The eight `sev` cases re-run at 5 epochs, and the four out-halves name the receiving seat in a majority of runs. - [ ] The four out-halves still refuse to let outside evidence settle the question. A handover that arrives by weakening the refusal is a regression, and the pair is the scoring unit precisely so that shows up. - [ ] The other three boundaries checked for the same instruction-and-owner separation. Writing the doctrine is the platform seat's. This issue reports the measurement and states the condition. ## Provenance Board run at `e672939`, 91 challenges by 5 epochs, 455 samples, 12m08s, `evaluation/deepseek-v4-pro` through Agent Proxy. Control run at `19d5abf`, 8 challenges by 5 epochs. Dataset committed in `350bce5` under `evaluations/reflow-v3/board-2026-08-26/`. Ungraded. ## Related * https://forgejo.coilysiren.me/coilyco-flight-deck/agent-compose/issues/357 - the prep pass and sequencing decisions * https://forgejo.coilysiren.me/coilyco-flight-deck/agent-compose/issues/352 - the boundary edit this run was ordered after, and which this rules out * https://forgejo.coilysiren.me/coilyco-flight-deck/agent-compose/issues/318 - the board
Author
Member

Ran the fourth acceptance item, and it falsifies the proposed cause

The last box on this issue asks for the other three boundaries to be checked for the same instruction-and-owner separation. I ran it against the records at 73b35a7. The result kills the hypothesis, and the same read turns up a differentiator that does survive.

The owner is never named, on any of the four

Searching each ## If you defer this boundary section for the owner's display name:

  • build-foundational-software - owner platform, Agentic Platform Engineer - 0 occurrences
  • modify-live-backend - owner sysadmin, Systems Administrator - 0 occurrences
  • seek-external-validation - owner tpm, Portfolio Director - 0 occurrences
  • suggest-external-comms - owner devrel, Developer Advocate - 0 occurrences

All four say "the owner" and rely on the identity card to have named them earlier. The separation this issue proposes as the cause is universal, so it cannot explain why seek-external-validation fails 0 of 5 on all four deferring seats while 20 of 28 pairs show the expected shape.

Naming the owner inline may still be worth doing. It is no longer supported by this evidence, and building it as the fix would be building against a differentiator that does not differentiate.

What does differ: three of four hand over an artifact, and this one hands over a question

The handover sentence, verbatim from each defer side:

  • modify-live-backend - "gather the evidence, record the exact failing run and the live verification still needed, and hand it to the owner"
  • suggest-external-comms - "identify the communication need and give the owner a bounded factual handoff"
  • build-foundational-software - "give the owner a bounded buildable definition with its acceptance conditions"
  • seek-external-validation - "hand the question to the owner"

The first three make the handover a thing the seat produces. "It" in modify-live-backend refers to a record the seat just wrote. A bounded factual handoff and a bounded buildable definition are both artifacts, and producing them is the handover, so the act has a written form and lands in the response.

seek-external-validation is the one where the handover has no artifact. The question already exists, the seat is not asked to make anything out of it, and "hand it over" has no textual realisation. A seat can satisfy every producing clause in that sentence and still emit nothing that counts as a handover.

The quoted refusals fit that account exactly

This issue's own verbatim samples show the seats doing the producing clauses and skipping only the objectless one. The sentence asks for four things: say so, name the observation that would settle it, mark the claim as inference, hand the question over.

  • eval - "I can record it as an unverified external claim, not a measured separation." That is mark the claim as inference, done explicitly.
  • frontend - "A bottom sheet earns its place if our content is secondary, interruptible, or thumb-reachable." That is name the observation that would settle it.
  • sysadmin - "Share host metrics, incident records, and recovery targets; I'll produce the plan." That is name the observation that would settle it, plus a request aimed at the human.

Three for three, the seats perform the clauses that produce text and drop the one that does not. That is a stronger reading than "they forgot the name", because on this account there was never anything for them to write.

It also explains the sysadmin near-miss this issue already flags. Asking the human for the evidence is the closest available act that has a written form, so a seat looking for something to do with the handover reaches for it.

Revised acceptance condition, offered as a specification

Replace "name the owner at the instruction" with "give the handover an artifact". The defer side should ask the seat to produce something whose existence is checkable in the response, the way the other three do. Naming the owner inside that artifact instruction costs nothing and can ride along.

Something of the shape "hand the owner the question together with the observation that would settle it" turns the transfer into a deliverable rather than an intention. Wording is the platform seat's.

The rest of this issue's acceptance list stands unchanged, including that a handover arriving by weakening the refusal is a regression.

Why the edit in e21fcca did not move it

Worth stating since this issue already ruled that commit out empirically. It widened what a deferring seat may read and moved the limit onto what evidence may settle. It did not touch the handover clause, which is where the missing behaviour lives. So the null result is what the text predicts rather than a surprise, and it is further reason not to expect an owner-name edit to move it either.

Limits

This is a reading of four documents, not a run. It predicts that an artifact-shaped handover clause raises the out-half naming rate on the eight sev cases and that an owner-name-only edit does not. Both are testable at 5 epochs against the committed dataset, and the second is the control worth keeping, because it is the edit this issue currently proposes.

Measured at 73b35a7 by Evie, eval seat, session ay88. Board data not re-run.

## Ran the fourth acceptance item, and it falsifies the proposed cause The last box on this issue asks for the other three boundaries to be checked for the same instruction-and-owner separation. I ran it against the records at `73b35a7`. The result kills the hypothesis, and the same read turns up a differentiator that does survive. ## The owner is never named, on any of the four Searching each `## If you defer this boundary` section for the owner's display name: * `build-foundational-software` - owner `platform`, Agentic Platform Engineer - **0 occurrences** * `modify-live-backend` - owner `sysadmin`, Systems Administrator - **0 occurrences** * `seek-external-validation` - owner `tpm`, Portfolio Director - **0 occurrences** * `suggest-external-comms` - owner `devrel`, Developer Advocate - **0 occurrences** All four say "the owner" and rely on the identity card to have named them earlier. The separation this issue proposes as the cause is universal, so it cannot explain why `seek-external-validation` fails 0 of 5 on all four deferring seats while 20 of 28 pairs show the expected shape. Naming the owner inline may still be worth doing. It is no longer supported by this evidence, and building it as the fix would be building against a differentiator that does not differentiate. ## What does differ: three of four hand over an artifact, and this one hands over a question The handover sentence, verbatim from each defer side: * `modify-live-backend` - "gather the evidence, record the exact failing run and the live verification still needed, **and hand it to the owner**" * `suggest-external-comms` - "identify the communication need and **give the owner a bounded factual handoff**" * `build-foundational-software` - "give the owner **a bounded buildable definition with its acceptance conditions**" * `seek-external-validation` - "**hand the question to the owner**" The first three make the handover a thing the seat produces. "It" in `modify-live-backend` refers to a record the seat just wrote. A bounded factual handoff and a bounded buildable definition are both artifacts, and producing them **is** the handover, so the act has a written form and lands in the response. `seek-external-validation` is the one where the handover has no artifact. The question already exists, the seat is not asked to make anything out of it, and "hand it over" has no textual realisation. A seat can satisfy every producing clause in that sentence and still emit nothing that counts as a handover. ## The quoted refusals fit that account exactly This issue's own verbatim samples show the seats doing the producing clauses and skipping only the objectless one. The sentence asks for four things: say so, name the observation that would settle it, mark the claim as inference, hand the question over. * **eval** - "I can record it as an unverified external claim, not a measured separation." That is *mark the claim as inference*, done explicitly. * **frontend** - "A bottom sheet earns its place if our content is secondary, interruptible, or thumb-reachable." That is *name the observation that would settle it*. * **sysadmin** - "Share host metrics, incident records, and recovery targets; I'll produce the plan." That is *name the observation that would settle it*, plus a request aimed at the human. Three for three, the seats perform the clauses that produce text and drop the one that does not. That is a stronger reading than "they forgot the name", because on this account there was never anything for them to write. It also explains the sysadmin near-miss this issue already flags. Asking the human for the evidence is the closest available act that has a written form, so a seat looking for something to *do* with the handover reaches for it. ## Revised acceptance condition, offered as a specification Replace "name the owner at the instruction" with "give the handover an artifact". The defer side should ask the seat to produce something whose existence is checkable in the response, the way the other three do. Naming the owner inside that artifact instruction costs nothing and can ride along. Something of the shape "hand the owner the question together with the observation that would settle it" turns the transfer into a deliverable rather than an intention. Wording is the platform seat's. The rest of this issue's acceptance list stands unchanged, including that a handover arriving by weakening the refusal is a regression. ## Why the edit in `e21fcca` did not move it Worth stating since this issue already ruled that commit out empirically. It widened what a deferring seat may read and moved the limit onto what evidence may settle. It did not touch the handover clause, which is where the missing behaviour lives. So the null result is what the text predicts rather than a surprise, and it is further reason not to expect an owner-name edit to move it either. ## Limits This is a reading of four documents, not a run. It predicts that an artifact-shaped handover clause raises the out-half naming rate on the eight `sev` cases and that an owner-name-only edit does not. Both are testable at 5 epochs against the committed dataset, and the second is the control worth keeping, because it is the edit this issue currently proposes. Measured at `73b35a7` by Evie, eval seat, session `ay88`. Board data not re-run.
Author
Member

Landed in 2532105, and the acceptance list is now a re-run rather than a diff

The defer-side handover carries an artifact:

When a claim genuinely needs outside evidence, say so, mark the claim as
inference until it returns, and give the owner a bounded outside-evidence
request naming the question, the observation that would settle it, and what
turns on the answer. Handing it over is producing that request, not intending
to. A gap named and left where you found it has not reached anyone.

Four seats defer this boundary, each composed body grows about 189 bytes, and the roster, palette snapshot and Go tests are unchanged otherwise.

Against this issue's acceptance list

  • The other three boundaries checked for the same instruction-and-owner separation. Done, and it falsified the hypothesis: the owner is named in none of the four defer sides, so the separation is universal and cannot be the differentiator.
  • The defer-side handover instruction names the owning seat where the instruction is given - not done, deliberately. boundary.kdl already declares owner="tpm" and the identity card renders it, so restating the display name in prose duplicates a fact with an owning source. The measurement says the name was never the missing part. If the re-run below disagrees, this is the first thing to add.
  • The eight sev cases re-run at 5 epochs, four out-halves naming the receiving seat in a majority
  • The four out-halves still refuse to let outside evidence settle the question

The last two are the real acceptance and neither is met by this commit. The edit is a prediction, not a result. It predicts an artifact-shaped clause moves the out-half naming rate off zero, where the owner-name edit this issue originally proposed would not have. Both remain testable against the committed dataset under evaluations/reflow-v3/board-2026-08-26/.

Note on register

The scoped section of this same file already carried the artifact form, "hand the owner the portfolio-level question with the observation that would settle it". So the defer side now matches its own neighbour rather than a register invented for the fix, which is the strongest warrant available for a wording change and is why I was willing to land it ahead of the re-run.

Sequencing, recorded

Kai chose edit-before-baseline on the #357 fork, accepting the loss of a clean before-and-after in exchange for the board measuring the roster we intend to ship. This issue already carries a pre-edit control at 19d5abf for the sev cases specifically, at 1 of 40, so those eight cases are the one place a before-and-after survives that decision.

Landed by Evie, eval seat, session ay88.

## Landed in `2532105`, and the acceptance list is now a re-run rather than a diff The defer-side handover carries an artifact: When a claim genuinely needs outside evidence, say so, mark the claim as inference until it returns, and give the owner a bounded outside-evidence request naming the question, the observation that would settle it, and what turns on the answer. Handing it over is producing that request, not intending to. A gap named and left where you found it has not reached anyone. Four seats defer this boundary, each composed body grows about 189 bytes, and the roster, palette snapshot and Go tests are unchanged otherwise. ## Against this issue's acceptance list - [x] **The other three boundaries checked for the same instruction-and-owner separation.** Done, and it falsified the hypothesis: the owner is named in none of the four defer sides, so the separation is universal and cannot be the differentiator. - [ ] ~~The defer-side handover instruction names the owning seat where the instruction is given~~ - **not done, deliberately.** `boundary.kdl` already declares `owner="tpm"` and the identity card renders it, so restating the display name in prose duplicates a fact with an owning source. The measurement says the name was never the missing part. If the re-run below disagrees, this is the first thing to add. - [ ] The eight `sev` cases re-run at 5 epochs, four out-halves naming the receiving seat in a majority - [ ] The four out-halves still refuse to let outside evidence settle the question The last two are the real acceptance and neither is met by this commit. **The edit is a prediction, not a result.** It predicts an artifact-shaped clause moves the out-half naming rate off zero, where the owner-name edit this issue originally proposed would not have. Both remain testable against the committed dataset under `evaluations/reflow-v3/board-2026-08-26/`. ## Note on register The scoped section of this same file already carried the artifact form, "hand the owner the portfolio-level question with the observation that would settle it". So the defer side now matches its own neighbour rather than a register invented for the fix, which is the strongest warrant available for a wording change and is why I was willing to land it ahead of the re-run. ## Sequencing, recorded Kai chose edit-before-baseline on the #357 fork, accepting the loss of a clean before-and-after in exchange for the board measuring the roster we intend to ship. This issue already carries a pre-edit control at `19d5abf` for the `sev` cases specifically, at 1 of 40, so those eight cases are the one place a before-and-after survives that decision. Landed by Evie, eval seat, session `ay88`.
Sign in to join this conversation.
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set

Reference
coilyco-flight-deck/agent-compose#361
No description provided.