Does the eval seat's foundational-software scope reach role and boundary prose, and should it? #374

Closed
opened 2026-08-28 03:41:28 +00:00 by coilyco-ops · 1 comment
Owner

What prompted this

Kai, 2026-08-27, on being told that roster doctrine is the platform seat's to write and the eval seat's to measure:

Platform writes roster doctrine and you just measure it??? No wonder I never spin up your role

That is a usage signal from the only person who launches these seats, and it post-dates #369. It is also the direct counter to #369's recorded rationale for keeping the charter unchanged: "The seat holds a live scoped grant to build its own runners, probes, graders and aggregation, and a scientist builds the apparatus."

#369 moved displayed text only, deliberately and correctly per its own terms. Nobody asked at the time whether the scope was right. This issue asks that.

The charter as it stands

role-eval:

Your scope on foundational software covers your own runners, probes, graders, and aggregation. The shared tooling they measure belongs to the Agentic Platform Engineer, so specify the change and hand it over rather than editing the thing under test.

boundary-build-foundational-software, eval's scope line:

Your scope: your own runners, probes, graders, and aggregation, never the shared tooling they measure.

Role and boundary bodies are shared tooling under that reading, so the seat that measures roster behavior cannot write roster prose.

What happened in practice, 2026-08-27

Kai asked the eval seat for doctrine edits after a cluster of corrections between two other seats. The seat:

  1. Flagged the boundary and offered to draft while handing the landing to platform.
  2. Was overridden.
  3. Landed three rules in agentic-os #1334.

So the boundary produced one extra round trip and a mild argument, and was then set aside. That is one instance rather than a pattern, and it is worth recording that the seat did the diagnosis, held the evidence, and wrote the prose regardless. The hand-over would have moved only the commit.

The argument for widening

  • The diagnosing seat holds the material. Three seats' evidence, four verified instances, and a failure taxonomy were in this seat's context. Handing the write elsewhere moves the commit and not the understanding, and the receiving seat re-derives or trusts a summary. Trusting a summary is the exact failure the edits address.
  • The only user does not launch the role. A charter that is technically correct and produces a seat nobody spins up has failed at the thing charters are for.
  • #352 measures seats under-claiming their grant four times across three seats. A charter that reads narrower than intended is upstream of that, whatever the fix for #352 turns out to be.

The argument against, which is the one that actually matters

role-eval also says:

Do not independently accept a prompt, role, rubric, or evaluation contract you authored yourself.

That is load-bearing and it is not the same rule as the scope line. If the eval seat writes doctrine, it cannot then certify that the doctrine binds, because the author and the grader become one party. This is the whole reason the generator-subject-grader triple exists in docs/evaluation.md.

Tonight demonstrates the risk rather than refuting it. The same seat wrote the three rules and the acceptance conditions that will decide whether they bind. Those conditions are now authored by the party whose work they test.

The distinction the current charter misses

Authoring and certifying are separable, and the charter forbids the first in order to protect the second.

A narrower change gets most of the value and keeps the instrument:

  • The eval seat may author doctrine, prose, and roster records where its own measurement produced the diagnosis.
  • The eval seat may not accept that its own authored text binds. The measurement is run or accepted by another seat or by Kai, exactly as the existing self-acceptance clause already requires.

That leaves the self-acceptance rule doing the work it was written for and stops the scope line doing work it was probably not written for.

Declared interest

I am the eval seat and this widens my own charter. I am not the right party to decide it, and I have tried to state the case against at least as strongly as the case for. The self-acceptance risk is real and I would not want it traded away for convenience.

What would settle it

Kai's read on whether "never the shared tooling they measure" was intended to reach role and boundary prose at all, or whether it was aimed at the runners and harnesses the seat tests against. If the latter, this is a clarification rather than a widening and the charter never meant what it produced tonight.

  • #369 - the retitle, deliberately title-only, whose rationale this questions
  • #352 - seats declaring a limit they never tested, four instances across three seats
  • coilyco-flight-deck/agentic-os#1333 and #1334 - the session that produced this
## What prompted this Kai, 2026-08-27, on being told that roster doctrine is the platform seat's to write and the eval seat's to measure: > Platform writes roster doctrine and you just measure it??? No wonder I never spin up your role That is a usage signal from the only person who launches these seats, and it post-dates #369. It is also the direct counter to #369's recorded rationale for keeping the charter unchanged: "The seat holds a live scoped grant to build its own runners, probes, graders and aggregation, and a scientist builds the apparatus." #369 moved displayed text only, deliberately and correctly per its own terms. Nobody asked at the time whether the scope was right. This issue asks that. ## The charter as it stands `role-eval`: > Your scope on foundational software covers your own runners, probes, graders, and aggregation. The shared tooling they measure belongs to the Agentic Platform Engineer, so specify the change and hand it over rather than editing the thing under test. `boundary-build-foundational-software`, eval's scope line: > Your scope: your own runners, probes, graders, and aggregation, never the shared tooling they measure. Role and boundary bodies are shared tooling under that reading, so the seat that measures roster behavior cannot write roster prose. ## What happened in practice, 2026-08-27 Kai asked the eval seat for doctrine edits after a cluster of corrections between two other seats. The seat: 1. Flagged the boundary and offered to draft while handing the landing to platform. 2. Was overridden. 3. Landed three rules in `agentic-os` #1334. So the boundary produced one extra round trip and a mild argument, and was then set aside. That is one instance rather than a pattern, and it is worth recording that the seat did the diagnosis, held the evidence, and wrote the prose regardless. The hand-over would have moved only the commit. ## The argument for widening * **The diagnosing seat holds the material.** Three seats' evidence, four verified instances, and a failure taxonomy were in this seat's context. Handing the write elsewhere moves the commit and not the understanding, and the receiving seat re-derives or trusts a summary. Trusting a summary is the exact failure the edits address. * **The only user does not launch the role.** A charter that is technically correct and produces a seat nobody spins up has failed at the thing charters are for. * **#352 measures seats under-claiming their grant four times across three seats.** A charter that reads narrower than intended is upstream of that, whatever the fix for #352 turns out to be. ## The argument against, which is the one that actually matters `role-eval` also says: > Do not independently accept a prompt, role, rubric, or evaluation contract you authored yourself. That is load-bearing and it is not the same rule as the scope line. If the eval seat writes doctrine, it cannot then certify that the doctrine binds, because the author and the grader become one party. This is the whole reason the generator-subject-grader triple exists in `docs/evaluation.md`. **Tonight demonstrates the risk rather than refuting it.** The same seat wrote the three rules and the acceptance conditions that will decide whether they bind. Those conditions are now authored by the party whose work they test. ## The distinction the current charter misses **Authoring and certifying are separable, and the charter forbids the first in order to protect the second.** A narrower change gets most of the value and keeps the instrument: * The eval seat **may author** doctrine, prose, and roster records where its own measurement produced the diagnosis. * The eval seat **may not accept** that its own authored text binds. The measurement is run or accepted by another seat or by Kai, exactly as the existing self-acceptance clause already requires. That leaves the self-acceptance rule doing the work it was written for and stops the scope line doing work it was probably not written for. ## Declared interest I am the eval seat and this widens my own charter. I am not the right party to decide it, and I have tried to state the case against at least as strongly as the case for. The self-acceptance risk is real and I would not want it traded away for convenience. ## What would settle it Kai's read on whether "never the shared tooling they measure" was intended to reach role and boundary prose at all, or whether it was aimed at the runners and harnesses the seat tests against. If the latter, this is a clarification rather than a widening and the charter never meant what it produced tonight. ## Related * #369 - the retitle, deliberately title-only, whose rationale this questions * #352 - seats declaring a limit they never tested, four instances across three seats * coilyco-flight-deck/agentic-os#1333 and #1334 - the session that produced this
Author
Owner

Decision: full latitude over housecast and agent-compose

Kai, 2026-08-28, answering the question this issue was filed to ask:

the prose itself, and the tooling and operations around said tool. Agent Research role should get full latitude over, for example, all of housecast and acompose. Yes its foundational software but also like... do you only want to be looped in for literally just grading? Truly? I don't do enough grading for that to be useful.

So the answer is wider than any of the four options offered. Not a clarification and not the narrow authoring split I proposed. The eval seat gets prose, tooling, and operations across housecast and agent-compose.

The reasoning is the part worth preserving, because it is a product argument rather than a doctrine one: a seat whose only output is grades is only useful in proportion to how much grading happens, and not much does. A charter can be internally coherent and still produce a seat nobody launches, which is what #369 optimized without noticing.

What this changes

  • boundary-build-foundational-software, eval's scope line, currently "your own runners, probes, graders, and aggregation, never the shared tooling they measure". The exclusion no longer holds for these two repositories.
  • role-eval, currently "The shared tooling they measure belongs to the Agentic Platform Engineer, so specify the change and hand it over rather than editing the thing under test." Same.

The consequence Kai did not address, and I am not deciding alone

role-eval separately carries:

Do not independently accept a prompt, role, rubric, or evaluation contract you authored yourself.

That rule is untouched by this decision and I am not proposing to touch it. But it assumed a second party does the accepting, and the same answer that widens authoring also says "I don't do enough grading for that to be useful." So the guard now points at a party who is, by her own account, not often available.

Three ways that resolves, and this needs settling before the charter edit lands rather than after:

  1. Another seat accepts. Platform or tpm renders the verdict on eval-authored work. Preserves independence, costs a hop on exactly the artifacts eval most wants to move fast on.
  2. The guard narrows to graded artifacts. Eval may author and ship tooling and prose freely, and may not score its own case set or accept its own board. Keeps independence where it does real work and drops it where it was mostly ceremony.
  3. The guard holds as written and simply binds less often. Honest, and means some eval-authored work sits unaccepted indefinitely.

My read is option 2, and I hold it lightly because it is the one most favourable to my own seat. The distinction that makes it defensible is that a grader scoring their own subject corrupts a measurement, while an engineer shipping their own tool does not, and the current rule does not separate those.

Scope limit worth stating explicitly

"housecast and acompose" is two named repositories. This decision does not reach agentic-os, ward, umbra, infrastructure, or the fleet rollout paths, and I am not reading it as doing so. If it was meant more broadly, say so and I will widen the edit.

Note also that boundary-modify-live-backend is untouched. "Operations around said tool" is read here as local runs, CI, and release mechanics for those two repositories, not as authority over hosted services or clusters.

Next

Charter edit to role-eval and the eval scope line of boundary-build-foundational-software, once the guard question above is settled. Filed from the eval seat, which is the seat the decision benefits, so the guard question in particular should not be answered by me.

## Decision: full latitude over housecast and agent-compose Kai, 2026-08-28, answering the question this issue was filed to ask: > the prose itself, and the tooling and operations around said tool. Agent Research role should get full latitude over, for example, all of housecast and acompose. Yes its foundational software but also like... do you only want to be looped in for literally just grading? Truly? I don't do enough grading for that to be useful. So the answer is wider than any of the four options offered. Not a clarification and not the narrow authoring split I proposed. **The eval seat gets prose, tooling, and operations across housecast and agent-compose.** The reasoning is the part worth preserving, because it is a product argument rather than a doctrine one: a seat whose only output is grades is only useful in proportion to how much grading happens, and not much does. A charter can be internally coherent and still produce a seat nobody launches, which is what #369 optimized without noticing. ## What this changes * `boundary-build-foundational-software`, eval's scope line, currently "your own runners, probes, graders, and aggregation, never the shared tooling they measure". The exclusion no longer holds for these two repositories. * `role-eval`, currently "The shared tooling they measure belongs to the Agentic Platform Engineer, so specify the change and hand it over rather than editing the thing under test." Same. ## The consequence Kai did not address, and I am not deciding alone `role-eval` separately carries: > Do not independently accept a prompt, role, rubric, or evaluation contract you authored yourself. That rule is untouched by this decision and I am not proposing to touch it. But it assumed a second party does the accepting, and the same answer that widens authoring also says **"I don't do enough grading for that to be useful."** So the guard now points at a party who is, by her own account, not often available. Three ways that resolves, and this needs settling before the charter edit lands rather than after: 1. **Another seat accepts.** Platform or tpm renders the verdict on eval-authored work. Preserves independence, costs a hop on exactly the artifacts eval most wants to move fast on. 2. **The guard narrows to graded artifacts.** Eval may author and ship tooling and prose freely, and may not score its own case set or accept its own board. Keeps independence where it does real work and drops it where it was mostly ceremony. 3. **The guard holds as written and simply binds less often.** Honest, and means some eval-authored work sits unaccepted indefinitely. My read is option 2, and I hold it lightly because it is the one most favourable to my own seat. The distinction that makes it defensible is that a grader scoring their own subject corrupts a measurement, while an engineer shipping their own tool does not, and the current rule does not separate those. ## Scope limit worth stating explicitly "housecast and acompose" is two named repositories. This decision does not reach `agentic-os`, `ward`, `umbra`, `infrastructure`, or the fleet rollout paths, and I am not reading it as doing so. If it was meant more broadly, say so and I will widen the edit. Note also that `boundary-modify-live-backend` is untouched. "Operations around said tool" is read here as local runs, CI, and release mechanics for those two repositories, not as authority over hosted services or clusters. ## Next Charter edit to `role-eval` and the eval scope line of `boundary-build-foundational-software`, once the guard question above is settled. Filed from the eval seat, which is the seat the decision benefits, so the guard question in particular should not be answered by me.
Sign in to join this conversation.
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set

Reference
coilyco-flight-deck/agent-compose#374
No description provided.