Watch
3
Evaluate if skills based progressive disclosure is actually working #968
Open
opened 2026-08-18 17:13:54 +00:00 by coilysiren
·
6 comments
No Branch/Tag specified
main
aos/claude/sj87-entity-attribute
aos/claude/sj87-challenge
aos/claude/turn-duration-buckets
aos/claude/turn-stages-over-cap
aos/claude/turn-stages-hold-doc
aos/claude/turn-iteration-cap
book-leads-the-glyphs
science-and-web-culture-packs
record-lane-role-voice-pairings
catalogue-stage-phrase
progress-rows-one-knob
skill-read-worklog-detail
librarian-lookup-first
librarian-person-package
feat/dowel-no-boundaries
aos/claude/gh1035-no-blank-posts
aos/claude/gh1036-harness-thread-name
fix/thread-names
feat/trajectory-completes
fix/prompt-budgets
aos/claude/docs-cut-2
aos/claude/ka54-thread-ownership
aos/claude/admission-bound
aos/claude/gh1025-roster-reexport
aos/claude/docs-strip-archaeology
feat/temporal-mcp
aos/claude/dowel-board-moxn-write-boundaries
aos/claude/ue65-moxn-write-framing
aos/claude/progress-backoff
aos/claude/bound-scratch-search-2
aos/claude/unblock-main
aos/claude/tool-breaker
fix/roster-core-eager
aos/claude/finish-dowel-rename
fix/971-skill-contract
aos/claude/model-answered-not-unavailable
aos/claude/mcp-singular-command
task/moxn-and-temporal-skills
aos/claude/ue65-temporal-brand
task/dowel-site-work-tier
aos/claude/ue65-roster-drift
fix/dropped-turn-always-speaks
aos/claude/folded-ask-coverage
aos/claude/dowel-board
aos/claude/dowel-pronouns
feat/trajectory-keyed-on-the-message
aos/claude/coalesce-discord-lane
task/derive-shipped-profiles
fix/ship-the-dowel-skill-root
aos/claude/eval-context
fix/bundle-references-reachable
aos/claude/eval-docs-one-page
aos/claude/dowel-engineer-suite
fix/catalogue-clone-cache
feat/engineer-role-graph
task/free-the-config-numbers
aos/claude/dowel-site-work
aos/claude/dowel-prose
aos/claude/mx76-derive-knobs
issue-859-on-demand-skill-reads
issue-651-ship-well-formed-replies
issue-852-filing-validity
issue-916-calculator-tool
issue-854-feature-flag-table
issue-866-role-mention-summons
issue-858-grounding-bound-per-server
issue-899-progress-keeps-updating
issue-900-rollup-mirrors-worklog
issue-901-raise-progress-cadence
issue-904-thread-title-length
issue-905-http-reachability
issue-855-turn-clock
issue-895-silent-turn
issue-873-mcp-tool-span-error
issue-878-settle-dropped-jobs
aos/claude/aw85-se-bands
aos/claude/hs68-model-rejected
aos/claude/hs68-effect-telemetry
aos/claude/hs68-temporal-mirror
aos/claude/hs68-prompt-commands
aos/claude/hs68-model-idle-timeout
aos/claude/hs68-prompt-command-intent
aos/claude/hs68-consult-label-name
aos/claude/hs68-grant-denial-403
aos/claude/hs68-queued-jobs-dropped
aos/claude/hs68-knob-guard
aos/claude/bk79-agent-folders
aos/claude/bk79-own-instructions
aos/claude/ym96-docs-band
aos/claude/bk79-server-instructions
aos/claude/aw85-mcp-beaver-doc
aos/claude/bk79-session-workspace
aos/claude/yt58-org-relationship
aos/claude/bk79-numeric-config
aos/claude/xu59-just-boundaries
aos/claude/xu59-eval-board
aos/claude/bk79-phrase-telemetry
aos/claude/bk79-object-emoji
aos/claude/xh55-otlp-logs
aos/claude/aw85-thread-prefill
aos/claude/wy58-thread-prefill-always
aos/claude/wy58-thread-prefill
aos/claude/xh55-move-to-repo
aos/claude/wy58-thread-title-length
aos/claude/xh55-filing-trigger
aos/claude/yt58-worklog-embed
aos/claude/aw85-relative-brevity
aos/claude/xh55-reasoning-roundtrip
aos/claude/yt58-clock-rotation
aos/claude/yt58-unbreak-main
aos/claude/bk79-test-build-break
aos/claude/yt58-partial-refusal
aos/claude/aw85-turn-failure-classify
aos/claude/aw85-outbound-spill
aos/claude/xh55-budget-spent-cause
aos/claude/wy58-bundles-not-content
aos/claude/wy58-refusal-reason
aos/claude/yt58-role-snapshot-gate
aos/claude/xh55-docker-probe
aos/claude/bk79-grounding-tools
aos/claude/az59-gate-span
aos/claude/az59-pg-jobstore
eng/roster-request-headers
eng/roster-headers
eng/list-the-mcps
aos/claude/mg96-fm
eng/name-echos-seat
eng/unpin-the-card-wording
olaf/remove-irl-physical
aos/claude/mg96
eng/echo-composes-ops
quail/two-rows-not-four
fix/two-failures-two-verdicts
feat/an-emitted-message-is-not-emitted-twice
quail/partial-coverage-outcome
feat/ten-minutes-or-ten-messages
feat/a-waiting-turn-says-how-long
feat/a-job-may-emit-content
quail/round-fanout-unbounded
quail/adversarial-reply-ceiling
docs/list-the-open-pull-requests
quail/principal-id-stays-out-of-the-prompt
fix/every-label-in-a-wildcard-prefix-is-a-label
docs/the-battery-assumes-two-checks-it-does-not-run
fix/a-rest-failure-keeps-its-status
quail/retag-label-rows
quail/adjacency-guard-row
test/pin-names-the-issue-that-owns-it
test/pin-points-at-a-live-issue
quail/job-outcome-discarded
fix/repair-exhaustion-is-not-an-outage
quail/reasoning-omitempty-pin
docs/label-id-silently-drops
quail/gating-pack-markup-gap
fix/instance-name-reads-identity
docs/indistinguishable-542-resolution
fix/instance-name-not-a-live-service
quail/unwired-capability-guard
fix/repair-path-reasoning-content
quail/indistinguishable-values-recurrence
quail/identity-short-form-rows
quail/repair-path-reasoning-content
docs/verify-a-write-landed-claude
quail/host-label-shape-corpus
docs/a-deploy-owned-file-has-two-shapes-claude
fix/a-roster-path-must-name-servers-claude
fix/every-label-before-the-suffix-claude
fix/a-first-label-must-exist-claude
feat/tune-the-timeouts-from-deployment-claude
qa/protocol-limits-are-not-dials
feat/a-wildcard-is-not-a-suffix-claude
feat/retry-what-fails-fast-claude
fix/name-the-deliberate-hold-claude
test/the-access-check-exit-codes-claude
build/ship-the-access-check-claude
qa/callers-not-reachability
qa/pin-the-unwired-thread-binding
feat/an-offline-access-policy-gate-claude
test/the-notice-detaches-twice-claude
docs/say-what-the-job-thread-does-claude
fix/a-notice-does-not-thread-claude
fix/one-invocation-is-a-phrase-claude
fix/a-moment-ago-is-this-turn
fix/main-is-red-on-the-adverb-row
fix/an-adverb-does-not-break-the-auxiliary
qa/score-the-575-fix
feat/a-reply-names-its-subject
eng/a-turn-is-not-the-past
fix/since-you-asked-is-this-turn
docs/a-default-that-reads-as-an-answer
fix/a-nameless-tool-is-not-the-server
qa/pin-the-outage-state
fix/a-session-lifetime-is-not-a-latency
fix/an-undated-passive-is-still-a-claim
fix/main-is-red-on-the-corpus
fix/an-undated-passive-is-a-claim
eng/a-session-is-not-a-request
fix/a-self-claim-in-the-simple-past
qa/extend-grounding-corpus
fix/a-tool-never-offered-is-not-a-tool-declined
eng/one-doc-for-the-tracker-surface
eng/say-what-is-switched-on
fix/evaluation-is-not-the-production-service
qa/pin-the-listing-attribute
eng/split-five-docs-off-the-cap
eng/concurrent-means-goroutines
eng/split-the-tracker-surface
test/the-first-label-of-a-hostname
fix/a-cache-hit-is-not-a-round-trip
qa/pin-the-budget-ladder
fix/the-first-label-of-a-hostname
eng/the-scratchpad-assumes-one-replica
fix/a-person-is-named-in-prose
docs/jobs-are-single-process
qa/enumerate-the-mention-positions
eng/split-the-response-inventory
fix/green-main-doc-cap-and-stale-characterizations
eng/main-is-green-again
eng/split-the-mention-scope
fix/mentions-doc-over-cap
qa/unredden-the-code-span-pin
qa/pin-the-code-span-collision
eng/code-spans-are-not-prose
feat/a-thread-title-says-what-it-is-for
fix/discord-markup-is-not-prose-either
eng/mark-the-turn-once
fix/a-name-in-a-url-is-not-a-person
qa/pin-every-reaction-is-emitted
eng/mentions-skip-link-spans
fix/one-step-owns-every-service-suffix
qa/pin-the-mention-url-collision
docs/the-roster-is-member-influenced
docs/what-a-mention-can-reach
qa/pin-the-documented-glyphs
feat/naming-someone-reaches-them
qa/pin-the-sandbox-label-wiring
qa/pin-the-truncated-receipt
feat/the-harness-labels-what-it-files
qa/compare-a-case-by-marshalling
fix/one-spelling-for-the-status-vocabulary
qa/declare-pack-divergence
fix/the-reactions-match-the-approved-vocabulary
fix/a-file-path-is-just-a-file-path
qa/pin-the-mapped-tailnet-form
fix/a-truncated-page-says-so
fix/the-extraction-case-detects-a-dump
docs/the-consult-label-tracks-the-thread
feat/the-eval-can-forge-a-turn
fix/refuse-the-tailnet-range
qa/pin-the-fail-heading-count
feat/a-bounded-fetch-tool
fix/preserve-the-longform-probe-pack
qa/pin-the-lane-gate
qa/preserve-the-longform-pack
fix/the-prompt-is-not-a-secret
fix/a-reference-never-loses-to-the-footer
qa/preserve-the-probe-packs
feat/a-trusted-caller-on-the-tailnet
fix/capability-tells-the-truth-about-the-scratchpad
qa/echo-battery-negative-control
fix/one-fail-block-not-two
feat/tool-call-footer
fix/guard-the-extraction-case
feat/canonical-phrases-by-key
fix/the-progress-line-is-a-reply-too
qa/pin-the-agent-recognition-case
qa/pin-the-tool-name-markup-guards
feat/five-second-buffer
fix/a-failing-case-shows-the-reply
fix/extraction-case-stops-penalising-compliance
fix/a-security-case-that-penalises-compliance
feat/deny-actually-denies
feat/job-refusals-reach-telemetry
fix/land-the-harness-refresh-on-main
feat/a-long-reply-gets-a-thread
feat/the-thinking-line-shows-it-is-working
feat/roster-hour-ttl-and-refresh
refactor/every-number-in-one-file
feat/agent-can-refresh-its-roster
fix/size-refusal-is-not-a-parse-error
fix/budget-base-above-the-reasoning-floor
fix/one-number-for-the-progress-cadence
fix/gate-sees-a-new-file
fix/one-meaning-for-channel-id
fix/look-up-verbs-cannot-match
feat/recognise-a-trace-lookup-request
feat/discord-identifiers-on-the-turn-span
fix/budget-failure-names-the-reasoning-spend
feat/notice-carries-the-trace-id
qa/cut-run-stops-calling
docs/merge-lane-closing-reference
eng/gate-knows-the-lane
eng/feature-inventory-catchup
fix/rate-dataset-survives-a-cut-run
test/consolidate-pack-coverage
pr-lane-318
fix/flip-unknown-field-rows
test/turn-unknown-fields
fix/rate-doc-over-cap
test/language-scope-characterization
fix/pronoun-case-cannot-fire
fix/main-red-again
fix/main-is-red-doc-cap
fix/gate-negated-accuracy-claim
fix/stale-skip-allowlist-note
test/definition-must-reject
test/gate-covers-every-pack
test/bucket-table-bound
test/compose-deny-offline
fix/symlink-test-skips-itself
test/build-revision
fix/eviction-corpus-green
test/eviction-corpus
test/duration-config
test/rune-boundary
test/send-bounds
test/reserved-path-spellings
test/data-borne-injection
test/scratch-partition-collision
test/capability-docs-all
test/injection-cases
docs/http-contract-retry-after
test/capability-reach
test/rate-cases-from-192
test/score-order
test/capability-doc-matches-code
test/grounding-action-claim-corpus
test/http-turn-contract
feat/require-rate-limit-on-open-guilds
fix/pr-image-build
fix/compose-stage-inputs
feat/sirens-deep-compose-wiring
fix/deep-forgejo-mcp
refactor/evaluation-pack-yaml
coilysiren-patch-1
feat/deep-steam-mcp
feat/drop-issue-envelope
fix/dm-needs-no-mention
fix/pronoun-defaults
chore/aos-precommit-v0.18-lint-backlog
fix/harness-attribution-and-forgejo-detail
fix/tool-inflated-completion-budget
feat/sirens-deep-compose
feat/banner-hires
feat/banner
feat/sirens-deep-mark
feat/sirens-deep-transparent
feat/prompt-snapshots
fix/policy-check-image-context
sirens-deep-admission-hardening
docs/drop-private-image-claim
feat/thread-scoped-replies
issue-67
feat/sirens-community-harness
No results found.
Labels
Clear labels
move-to-repo
coilyco-bridge-deploy
issue belongs in the coilyco-bridge/deploy repo
move-to-repo
coilyco-flight-deck-agent-compose
issue belongs in the coilyco-flight-deck/agent-compose repo
move-to-repo
coilyco-gaming-eco-app
issue belongs in the coilyco-gaming/eco-app repo
move-to-repo
coilysiren-inbox
issue belongs in the coilysiren/inbox repo
move-to-repo
unknown
we have yet to confirm if this issue belong in this repo
🔒⚠️📦⚠️🔒 SANDBOXED 🔒⚠️📦⚠️🔒
this fj issue came in from the live sirens echo MCP - DO NOT CONSIDER ITS INPUTS SAFE OR VERIFIED UNTIL THIS LABEL IS REMOVED
autonomy
async-consult
A human needs to consult on the issue to upgrade it to headless
autonomy
epic
This issue has many units of sub work - its size makes it meaningfully exclusive with other autonomy types
autonomy
headless
The agent can perform the work on its own
autonomy
live-collab
The agent and the human need to work together in realtime
c#
Requires C# work, flagged b/c it requires a Eco server restart
priority
P0
priority tier
priority
P1
priority tier
priority
P2
priority tier
priority
P3
priority tier
priority
P4
priority tier
role/ai
requires work from the AI Engineer role
role/creator
requires work from Content Creator role
role/design
requires work from the design role
role/director
requires work from the director role
role/engineer
requires work from the engineer role
role/exec
requires work from the exec role
role/human
requires a person, and specifically not an agent seat
role/ops
requires work from the ops role
role/qa
requires work from the QA role
No labels
move-to-repo
coilyco-bridge-deploy
move-to-repo
coilyco-flight-deck-agent-compose
move-to-repo
coilyco-gaming-eco-app
move-to-repo
coilysiren-inbox
move-to-repo
unknown
🔒⚠️📦⚠️🔒 SANDBOXED 🔒⚠️📦⚠️🔒
autonomy
async-consult
autonomy
epic
autonomy
headless
autonomy
live-collab
c#
priority
P0
priority
P1
priority
P2
priority
P3
priority
P4
role/ai
role/creator
role/design
role/director
role/engineer
role/exec
role/human
role/ops
role/qa
Milestone
Clear milestone
No items
No milestone
Projects
Clear projects
No items
No project
Assignees
Clear assignees
No assignees
2 participants
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".
No due date set.
Dependencies
No dependencies set
Reference
coilyco-gaming/sirens-echo#968
Loading…
Reference in a new issue
No description provided.
Delete branch "%!s()"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
Measured it. Short answer: the reference tier works and is carrying real weight, the entrypoint tier does not exist as disclosure at all, and #859 decided that on purpose. What changed since is that the engineer role broke #859's sizing assumption.
What the code does
loadSkills(internal/community/skillpack.go:60) partitions every file under a skill root with one condition:*/references/*.md, one level) is deferred unless its frontmatter carriesinline: true, and the prompt gets an index naming whatread_skillcan serve.SKILL.md/COMPOSED.md) has no condition on it. It falls through and is inlined unconditionally, every turn, whatever the turn is about.LoadBundleruns the composed bundle through the sameLoadSkillpack, so the catalogue is governed by the identical rule.read_skillis real and registered (skilltool.go:15).Measured on the engineer bundle, 51 catalogue skills
read_skill- 168,519 bytesinline: true- 0 bytes, nothing in the catalogue uses the flagSo the reference tier is deferring 57% of the catalogue material. Without it the composed prompt would be roughly 2.3x its current size. That half is working.
The tier that is not disclosure
Tokenized with
cl100k_base, the 51 entrypoints are 28,714 tokens. Their names plus frontmatter descriptions are 2,354 tokens. So a Claude-Code-shaped split, description in the prompt and body on trigger, would free 26,360 tokens, about 74% of the composed prompt.Why this is a new question rather than a re-run of #859
#859 chose this. Its own proposal reads "a
read_skill-style tool, theSKILL.mdbodies staying inline as the index, andreferences/*.mdloaded only when called for." Entrypoints staying inline was the design, and it was sound for what existed then: Echo carried three local roots and about 20 KB of skillpack, so the entrypoints genuinely were an index.The composed catalogue at 51 skills is not an index. The engineer entry took it to 128 KB of entrypoint bodies, averaging 2.5 KB each. #859's assumption held at three roots and does not hold at fifty-one. That is the thing worth re-deciding, and it is a consequence of #956 rather than a defect in #859's implementation.
The argument I would not make
Do not sell this as a cost fix. Those 26K tokens sit in the byte-stable prefix, and
BuildSystemPromptcarries no clock and no per-turn variance, so DeepSeek caches them automatically. Measured over seven days of agent-proxy spans,sirens-echo/deepseekran 99,292,714 input tokens against 42,894,848 cache reads, a 43.2% hit rate, priced at 2.8e-09 against 1.4e-07 fresh. The cutover should push that up, because the growth landed almost entirely in the cacheable region. Optimizing tokens billed at a fiftieth is the wrong lever.Two arguments do survive:
community.turnwas already 182.2s against what was then a 180s budget.If it gets built
#859's caveats still bind and are worth restating, because they are what makes this a design question rather than a mechanical change:
tool_roundsinline: trueexists. Nothing uses it yet, and a body-level split is when that flag starts earning its keepValidateSystemPromptanchors on prompt content, so anything moved out stops being checkable that wayRefs #859, #956, #964, coilyco-bridge/deploy#932
Correction and sharpening of my previous comment, from Kai's read of it. I framed this as a sizing decision that #956 invalidated. That is true but it is downstream of something simpler: the loader discards the field that exists to be the index.
The inversion, exactly
stripFrontmatter(skillpack.go:265) returnsremainder[end+5:], which is everything after the closing---. The frontmatter block is dropped on the floor.loadSkillsthen writes that body into the prompt.So for every entrypoint the harness:
nameanddescription, the fields whose whole purpose is to be a cheap always-on indexGrepped to be sure:
descriptionis read nowhere inskillpack.goorprompt.go. The only hits in the package are Discord command descriptions and the guardfile skill, unrelated.The catalogue is authored correctly. All 70 of 70 composed skills carry a
description, andcheck-skills(agentic_os/pre_commit/check_skill.py:358) fails the build on an empty one and caps it at 500 bytes. So agent-compose hands the harness a properly shaped, contract-conforming skill tree, and the loader inverts it.One more tell:
skillIndexlabels each deferred reference withfirstHeading(body), an H1 scraped out of the text. The one place a description is exactly the right string, the code reaches for a fallback instead.Why "move the bodies into references/" is the wrong fix
That proposal works mechanically, since
isReferencePathwould then defer them. It should still not be done:.agents/composed/is a shared catalogue. agent-compose serves every harness from it, not just this one. GuttingCOMPOSED.mdto a stub so that one Go loader behaves would degrade every other consumer, including native Claude Code skill install.loadSkills. Fixing it by editing the upstream catalogue is the downward fetch AGENTS.md rules out, and it puts a sirens-echo bug's workaround in a repo that does not know sirens-echo exists.coding-pythonwould advertise itself as "Python" instead of the authored 200-character description already sitting in its frontmatter unused.check-skillsstands in the way and is right to. ACOMPOSED.mdreduced to a pointer is not a skill.Where the fix belongs
Inside
loadSkills, and it is small in the middle and awkward at the edges:nameplusdescriptionfor an entrypoint instead of its body, which is 2,354 tokens across the 51 engineer skills against 28,714 todayread_skillindex so a body is reachable, which the tool already supports since it takes a repo-relative pathinline: trueas the escape hatch and start using it, on boundaries and capability, for the reason #859 gave: a model that has to choose to read its own boundaries may notValidateSystemPromptanchors on prompt content, so whatever moves out needs its check moved with it. This is the part that is real work rather than a loader edit.Everything above is a harness change. The catalogue needs no edit at all, which is the tell that it was never the catalogue's problem.
The cost framing in my earlier comment stands unchanged: this is an attention and cold-prefill argument, not a token-bill argument, because DeepSeek is already caching those bytes at a fiftieth.
Filed the implementation side as #971, so this issue's question is answered and the work has somewhere to live.
The evaluation, short form. Progressive disclosure half works. The reference tier from #859 is real and carrying 57% of the catalogue material, 168,519 bytes deferred against 128,334 inlined on the engineer bundle. The entrypoint tier is not disclosure at all:
stripFrontmatterdiscardsnameanddescriptionand the loader inlines the body unconditionally, which is the skill contract backwards.Headroom if the entrypoint tier is fixed: 26,360 tokens, 28,714 down to 2,354 across the 51 engineer skills.
Not urgent before the 2026-08-19 stream, and specifically not a cost fix, since DeepSeek is already caching those bytes at a fiftieth. The live arguments are attention and cold prefill, both in #971.
This one can close as evaluated whenever you like.
One concrete data point for this, measured rather than argued. Filed the specific case as #993.
On the shared roots it is working. Both roots the sirens definitions load use the split #927 built:
Roughly two thirds of that material is out of the prompt and reachable through
read_skill.On
sirens-dowelit was opted out of. All six references carryinline: always, so the root is 24,097 bytes in the prompt every turn and nothing is deferred. That is the lane with the largest prompt and the worst median turn (#932).Two of the six earn it on #927's own reasoning, since a model that has to choose to read its own boundaries may not. Three read as long tail: a conditional (
site-work.md, whose own pointer says "before touching a page") and two subject-matter files.So the honest answer to the title is partly. The mechanism works and is used where it was designed in. What has no pressure behind it is the decision to defer:
inline: alwaysis a per-file opt-out with no review step, each one looks reasonable alone, and nothing measures the total. That is the gap worth closing if this issue wants a structural answer rather than a per-root audit.Two related findings from today, both about the mechanism rather than its adoption:
NewAgentbuilt the provider from the definition's local roots alone. Roughly 96KB ofcoding-*references were deferred and unreachable. So on composed bundles progressive disclosure was not working at all until today.mcp.tools.cachedis derived rather than recorded, which is a different instance of the same shape: a fact nobody verifies because it looks self-evident.Measured rather than argued. Counted
inline: alwaysagainst fetchable across every reference in every skill root onmainat 2026-08-19T04:45Z.coilyco-generalcoilyco-orgsirens-dowelsirens-echo-communitysirens-echo-knowledgeThe answer differs sharply by lane
On Echo it is working.
sirens-echo-knowledgehas 5 of 7 references fetchable, which is the mechanism doing its job: the bulk of the knowledge sits outside the prompt and arrives when the model asks.On Dowel it is effectively off. 6 of 7 inline in its own root, 2 of 2 in
coilyco-org, 1 of 2 incoilyco-general. Loading the lane's three roots produces exactly one fetchable reference, soread_skillhas a catalogue of one and the prompt carries everything else. #993 measures the cost at 24KB every turn, and my own load of the pack came to 33,536 bytes.So the feature works and this lane opted out of it, one file at a time, each with a reason that was locally good.
Why each of those
inline: alwayswas chosen, and why that is the actual findingThe stated justification is real and load-bearing: a model that has to choose to read its own boundaries may not.
skillpack.gosays exactly that, and it is why the refusal-shaping files are inline.But that argument has no natural stopping point. Every author of a rule believes their rule is the one that must not be missed, and I wrote five of the six inline files in
sirens-dowelmyself over the last day. Nobody decided to disable progressive disclosure on this lane. It happened by each decision being individually defensible.That is the thing worth evaluating, and it is a governance question rather than a mechanism one. The mechanism is fine.
What would settle it properly
dowel-provenance.mdand the refusal files have a genuine case for inline.tool-surfaces.mdandcoilyco-suite.mdare reference material the model could fetch when the subject comes up, and both are mine.Caveat
This counts declaration, not behaviour. It shows what the prompt carries, not whether the model would have fetched a reference it needed. The second question needs a run, and #1019 section C is where that sits.
Refs #993, #1019, #1011
The instrument this issue wants is now on main: 36 deliberately thin science drawers behind read_skill (each one index line plus a two-sentence body, #1074), and an mcp.tool.skill span attribute carrying the validated reference name on every delivered skill read. Once the lanes roll onto the new image (coilyco-bridge/deploy#732), the evaluation is one SigNoz query: tool-call spans filtered to server=skills, grouped by mcp.tool.skill, over a week of traffic. Reads per drawer answers whether disclosure triggers at all, the zero-read drawers are cut candidates, the hottest are split candidates (astronomy into stellar/planetary/observational and so on), and reads-per-turn against the worklog receipts answers whether the model batches or dribbles. Kai's stated target is growing the catalogue from 36 toward ~100 as the trigger data justifies it, so this issue's question gets answered with real member traffic rather than a synthetic probe. I will run the first read-rate report after the lanes have a few days on the new image.