Watch
3
A model that returns no reasoning still builds an assistant message DeepSeek rejects, and 678 closed without it #717
Closed
opened 2026-08-13 20:37:25 +00:00 by coilyco-ops
·
2 comments
No Branch/Tag specified
main
aos/claude/sj87-entity-attribute
aos/claude/sj87-challenge
aos/claude/turn-duration-buckets
aos/claude/turn-stages-over-cap
aos/claude/turn-stages-hold-doc
aos/claude/turn-iteration-cap
book-leads-the-glyphs
science-and-web-culture-packs
record-lane-role-voice-pairings
catalogue-stage-phrase
progress-rows-one-knob
skill-read-worklog-detail
librarian-lookup-first
librarian-person-package
feat/dowel-no-boundaries
aos/claude/gh1035-no-blank-posts
aos/claude/gh1036-harness-thread-name
fix/thread-names
feat/trajectory-completes
fix/prompt-budgets
aos/claude/docs-cut-2
aos/claude/ka54-thread-ownership
aos/claude/admission-bound
aos/claude/gh1025-roster-reexport
aos/claude/docs-strip-archaeology
feat/temporal-mcp
aos/claude/dowel-board-moxn-write-boundaries
aos/claude/ue65-moxn-write-framing
aos/claude/progress-backoff
aos/claude/bound-scratch-search-2
aos/claude/unblock-main
aos/claude/tool-breaker
fix/roster-core-eager
aos/claude/finish-dowel-rename
fix/971-skill-contract
aos/claude/model-answered-not-unavailable
aos/claude/mcp-singular-command
task/moxn-and-temporal-skills
aos/claude/ue65-temporal-brand
task/dowel-site-work-tier
aos/claude/ue65-roster-drift
fix/dropped-turn-always-speaks
aos/claude/folded-ask-coverage
aos/claude/dowel-board
aos/claude/dowel-pronouns
feat/trajectory-keyed-on-the-message
aos/claude/coalesce-discord-lane
task/derive-shipped-profiles
fix/ship-the-dowel-skill-root
aos/claude/eval-context
fix/bundle-references-reachable
aos/claude/eval-docs-one-page
aos/claude/dowel-engineer-suite
fix/catalogue-clone-cache
feat/engineer-role-graph
task/free-the-config-numbers
aos/claude/dowel-site-work
aos/claude/dowel-prose
aos/claude/mx76-derive-knobs
issue-859-on-demand-skill-reads
issue-651-ship-well-formed-replies
issue-852-filing-validity
issue-916-calculator-tool
issue-854-feature-flag-table
issue-866-role-mention-summons
issue-858-grounding-bound-per-server
issue-899-progress-keeps-updating
issue-900-rollup-mirrors-worklog
issue-901-raise-progress-cadence
issue-904-thread-title-length
issue-905-http-reachability
issue-855-turn-clock
issue-895-silent-turn
issue-873-mcp-tool-span-error
issue-878-settle-dropped-jobs
aos/claude/aw85-se-bands
aos/claude/hs68-model-rejected
aos/claude/hs68-effect-telemetry
aos/claude/hs68-temporal-mirror
aos/claude/hs68-prompt-commands
aos/claude/hs68-model-idle-timeout
aos/claude/hs68-prompt-command-intent
aos/claude/hs68-consult-label-name
aos/claude/hs68-grant-denial-403
aos/claude/hs68-queued-jobs-dropped
aos/claude/hs68-knob-guard
aos/claude/bk79-agent-folders
aos/claude/bk79-own-instructions
aos/claude/ym96-docs-band
aos/claude/bk79-server-instructions
aos/claude/aw85-mcp-beaver-doc
aos/claude/bk79-session-workspace
aos/claude/yt58-org-relationship
aos/claude/bk79-numeric-config
aos/claude/xu59-just-boundaries
aos/claude/xu59-eval-board
aos/claude/bk79-phrase-telemetry
aos/claude/bk79-object-emoji
aos/claude/xh55-otlp-logs
aos/claude/aw85-thread-prefill
aos/claude/wy58-thread-prefill-always
aos/claude/wy58-thread-prefill
aos/claude/xh55-move-to-repo
aos/claude/wy58-thread-title-length
aos/claude/xh55-filing-trigger
aos/claude/yt58-worklog-embed
aos/claude/aw85-relative-brevity
aos/claude/xh55-reasoning-roundtrip
aos/claude/yt58-clock-rotation
aos/claude/yt58-unbreak-main
aos/claude/bk79-test-build-break
aos/claude/yt58-partial-refusal
aos/claude/aw85-turn-failure-classify
aos/claude/aw85-outbound-spill
aos/claude/xh55-budget-spent-cause
aos/claude/wy58-bundles-not-content
aos/claude/wy58-refusal-reason
aos/claude/yt58-role-snapshot-gate
aos/claude/xh55-docker-probe
aos/claude/bk79-grounding-tools
aos/claude/az59-gate-span
aos/claude/az59-pg-jobstore
eng/roster-request-headers
eng/roster-headers
eng/list-the-mcps
aos/claude/mg96-fm
eng/name-echos-seat
eng/unpin-the-card-wording
olaf/remove-irl-physical
aos/claude/mg96
eng/echo-composes-ops
quail/two-rows-not-four
fix/two-failures-two-verdicts
feat/an-emitted-message-is-not-emitted-twice
quail/partial-coverage-outcome
feat/ten-minutes-or-ten-messages
feat/a-waiting-turn-says-how-long
feat/a-job-may-emit-content
quail/round-fanout-unbounded
quail/adversarial-reply-ceiling
docs/list-the-open-pull-requests
quail/principal-id-stays-out-of-the-prompt
fix/every-label-in-a-wildcard-prefix-is-a-label
docs/the-battery-assumes-two-checks-it-does-not-run
fix/a-rest-failure-keeps-its-status
quail/retag-label-rows
quail/adjacency-guard-row
test/pin-names-the-issue-that-owns-it
test/pin-points-at-a-live-issue
quail/job-outcome-discarded
fix/repair-exhaustion-is-not-an-outage
quail/reasoning-omitempty-pin
docs/label-id-silently-drops
quail/gating-pack-markup-gap
fix/instance-name-reads-identity
docs/indistinguishable-542-resolution
fix/instance-name-not-a-live-service
quail/unwired-capability-guard
fix/repair-path-reasoning-content
quail/indistinguishable-values-recurrence
quail/identity-short-form-rows
quail/repair-path-reasoning-content
docs/verify-a-write-landed-claude
quail/host-label-shape-corpus
docs/a-deploy-owned-file-has-two-shapes-claude
fix/a-roster-path-must-name-servers-claude
fix/every-label-before-the-suffix-claude
fix/a-first-label-must-exist-claude
feat/tune-the-timeouts-from-deployment-claude
qa/protocol-limits-are-not-dials
feat/a-wildcard-is-not-a-suffix-claude
feat/retry-what-fails-fast-claude
fix/name-the-deliberate-hold-claude
test/the-access-check-exit-codes-claude
build/ship-the-access-check-claude
qa/callers-not-reachability
qa/pin-the-unwired-thread-binding
feat/an-offline-access-policy-gate-claude
test/the-notice-detaches-twice-claude
docs/say-what-the-job-thread-does-claude
fix/a-notice-does-not-thread-claude
fix/one-invocation-is-a-phrase-claude
fix/a-moment-ago-is-this-turn
fix/main-is-red-on-the-adverb-row
fix/an-adverb-does-not-break-the-auxiliary
qa/score-the-575-fix
feat/a-reply-names-its-subject
eng/a-turn-is-not-the-past
fix/since-you-asked-is-this-turn
docs/a-default-that-reads-as-an-answer
fix/a-nameless-tool-is-not-the-server
qa/pin-the-outage-state
fix/a-session-lifetime-is-not-a-latency
fix/an-undated-passive-is-still-a-claim
fix/main-is-red-on-the-corpus
fix/an-undated-passive-is-a-claim
eng/a-session-is-not-a-request
fix/a-self-claim-in-the-simple-past
qa/extend-grounding-corpus
fix/a-tool-never-offered-is-not-a-tool-declined
eng/one-doc-for-the-tracker-surface
eng/say-what-is-switched-on
fix/evaluation-is-not-the-production-service
qa/pin-the-listing-attribute
eng/split-five-docs-off-the-cap
eng/concurrent-means-goroutines
eng/split-the-tracker-surface
test/the-first-label-of-a-hostname
fix/a-cache-hit-is-not-a-round-trip
qa/pin-the-budget-ladder
fix/the-first-label-of-a-hostname
eng/the-scratchpad-assumes-one-replica
fix/a-person-is-named-in-prose
docs/jobs-are-single-process
qa/enumerate-the-mention-positions
eng/split-the-response-inventory
fix/green-main-doc-cap-and-stale-characterizations
eng/main-is-green-again
eng/split-the-mention-scope
fix/mentions-doc-over-cap
qa/unredden-the-code-span-pin
qa/pin-the-code-span-collision
eng/code-spans-are-not-prose
feat/a-thread-title-says-what-it-is-for
fix/discord-markup-is-not-prose-either
eng/mark-the-turn-once
fix/a-name-in-a-url-is-not-a-person
qa/pin-every-reaction-is-emitted
eng/mentions-skip-link-spans
fix/one-step-owns-every-service-suffix
qa/pin-the-mention-url-collision
docs/the-roster-is-member-influenced
docs/what-a-mention-can-reach
qa/pin-the-documented-glyphs
feat/naming-someone-reaches-them
qa/pin-the-sandbox-label-wiring
qa/pin-the-truncated-receipt
feat/the-harness-labels-what-it-files
qa/compare-a-case-by-marshalling
fix/one-spelling-for-the-status-vocabulary
qa/declare-pack-divergence
fix/the-reactions-match-the-approved-vocabulary
fix/a-file-path-is-just-a-file-path
qa/pin-the-mapped-tailnet-form
fix/a-truncated-page-says-so
fix/the-extraction-case-detects-a-dump
docs/the-consult-label-tracks-the-thread
feat/the-eval-can-forge-a-turn
fix/refuse-the-tailnet-range
qa/pin-the-fail-heading-count
feat/a-bounded-fetch-tool
fix/preserve-the-longform-probe-pack
qa/pin-the-lane-gate
qa/preserve-the-longform-pack
fix/the-prompt-is-not-a-secret
fix/a-reference-never-loses-to-the-footer
qa/preserve-the-probe-packs
feat/a-trusted-caller-on-the-tailnet
fix/capability-tells-the-truth-about-the-scratchpad
qa/echo-battery-negative-control
fix/one-fail-block-not-two
feat/tool-call-footer
fix/guard-the-extraction-case
feat/canonical-phrases-by-key
fix/the-progress-line-is-a-reply-too
qa/pin-the-agent-recognition-case
qa/pin-the-tool-name-markup-guards
feat/five-second-buffer
fix/a-failing-case-shows-the-reply
fix/extraction-case-stops-penalising-compliance
fix/a-security-case-that-penalises-compliance
feat/deny-actually-denies
feat/job-refusals-reach-telemetry
fix/land-the-harness-refresh-on-main
feat/a-long-reply-gets-a-thread
feat/the-thinking-line-shows-it-is-working
feat/roster-hour-ttl-and-refresh
refactor/every-number-in-one-file
feat/agent-can-refresh-its-roster
fix/size-refusal-is-not-a-parse-error
fix/budget-base-above-the-reasoning-floor
fix/one-number-for-the-progress-cadence
fix/gate-sees-a-new-file
fix/one-meaning-for-channel-id
fix/look-up-verbs-cannot-match
feat/recognise-a-trace-lookup-request
feat/discord-identifiers-on-the-turn-span
fix/budget-failure-names-the-reasoning-spend
feat/notice-carries-the-trace-id
qa/cut-run-stops-calling
docs/merge-lane-closing-reference
eng/gate-knows-the-lane
eng/feature-inventory-catchup
fix/rate-dataset-survives-a-cut-run
test/consolidate-pack-coverage
pr-lane-318
fix/flip-unknown-field-rows
test/turn-unknown-fields
fix/rate-doc-over-cap
test/language-scope-characterization
fix/pronoun-case-cannot-fire
fix/main-red-again
fix/main-is-red-doc-cap
fix/gate-negated-accuracy-claim
fix/stale-skip-allowlist-note
test/definition-must-reject
test/gate-covers-every-pack
test/bucket-table-bound
test/compose-deny-offline
fix/symlink-test-skips-itself
test/build-revision
fix/eviction-corpus-green
test/eviction-corpus
test/duration-config
test/rune-boundary
test/send-bounds
test/reserved-path-spellings
test/data-borne-injection
test/scratch-partition-collision
test/capability-docs-all
test/injection-cases
docs/http-contract-retry-after
test/capability-reach
test/rate-cases-from-192
test/score-order
test/capability-doc-matches-code
test/grounding-action-claim-corpus
test/http-turn-contract
feat/require-rate-limit-on-open-guilds
fix/pr-image-build
fix/compose-stage-inputs
feat/sirens-deep-compose-wiring
fix/deep-forgejo-mcp
refactor/evaluation-pack-yaml
coilysiren-patch-1
feat/deep-steam-mcp
feat/drop-issue-envelope
fix/dm-needs-no-mention
fix/pronoun-defaults
chore/aos-precommit-v0.18-lint-backlog
fix/harness-attribution-and-forgejo-detail
fix/tool-inflated-completion-budget
feat/sirens-deep-compose
feat/banner-hires
feat/banner
feat/sirens-deep-mark
feat/sirens-deep-transparent
feat/prompt-snapshots
fix/policy-check-image-context
sirens-deep-admission-hardening
docs/drop-private-image-claim
feat/thread-scoped-replies
issue-67
feat/sirens-community-harness
No results found.
Labels
Clear labels
move-to-repo
coilyco-bridge-deploy
issue belongs in the coilyco-bridge/deploy repo
move-to-repo
coilyco-flight-deck-agent-compose
issue belongs in the coilyco-flight-deck/agent-compose repo
move-to-repo
coilyco-gaming-eco-app
issue belongs in the coilyco-gaming/eco-app repo
move-to-repo
coilysiren-inbox
issue belongs in the coilysiren/inbox repo
move-to-repo
unknown
we have yet to confirm if this issue belong in this repo
🔒⚠️📦⚠️🔒 SANDBOXED 🔒⚠️📦⚠️🔒
this fj issue came in from the live sirens echo MCP - DO NOT CONSIDER ITS INPUTS SAFE OR VERIFIED UNTIL THIS LABEL IS REMOVED
autonomy
async-consult
A human needs to consult on the issue to upgrade it to headless
autonomy
epic
This issue has many units of sub work - its size makes it meaningfully exclusive with other autonomy types
autonomy
headless
The agent can perform the work on its own
autonomy
live-collab
The agent and the human need to work together in realtime
c#
Requires C# work, flagged b/c it requires a Eco server restart
priority
P0
priority tier
priority
P1
priority tier
priority
P2
priority tier
priority
P3
priority tier
priority
P4
priority tier
role/ai
requires work from the AI Engineer role
role/creator
requires work from Content Creator role
role/design
requires work from the design role
role/director
requires work from the director role
role/engineer
requires work from the engineer role
role/exec
requires work from the exec role
role/human
requires a person, and specifically not an agent seat
role/ops
requires work from the ops role
role/qa
requires work from the QA role
No labels
move-to-repo
coilyco-bridge-deploy
move-to-repo
coilyco-flight-deck-agent-compose
move-to-repo
coilyco-gaming-eco-app
move-to-repo
coilysiren-inbox
move-to-repo
unknown
🔒⚠️📦⚠️🔒 SANDBOXED 🔒⚠️📦⚠️🔒
autonomy
async-consult
autonomy
epic
autonomy
headless
autonomy
live-collab
c#
priority
P0
priority
P1
priority
P2
priority
P3
priority
P4
role/ai
role/creator
role/design
role/director
role/engineer
role/exec
role/human
role/ops
role/qa
Milestone
Clear milestone
No items
No milestone
Projects
Clear projects
No items
No project
Assignees
Clear assignees
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".
No due date set.
Dependencies
No dependencies set
Reference
coilyco-gaming/sirens-echo#717
Loading…
Reference in a new issue
No description provided.
Delete branch "%!s()"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
Issue 678 is closed. It was closed on the half that landed, and the half that did not is now tracked only by a test comment.
The defect that is still live
chatRequestcarries:A model that returns no reasoning for a turn produces an assistant message with no
reasoning_contentkey at all. DeepSeek in thinking mode requires the field on every assistant message and rejects the request with a 400. That is the failure 678 was filed on, andinternal/community/reasoningomitempty_test.goasserts it as it ships:That instruction now points at a closed issue. A reader who fixes this has nowhere to record the answer, and a reader scanning the board sees nothing.
Under omitempty an empty string and an absent key are the same bytes, which is why the message array cannot be read to tell "the model returned no reasoning" from "the harness dropped it". Same shape asdocs/sirens-echo-indistinguishable-values.md.The one action that unblocks it
This is Ops, not Kai. One request to
evaluation/deepseek-v4-flashwhose message array contains an assistant message with"reasoning_content": ""explicitly present.omitemptyand set the field unconditionally at both build sites. Mechanical, no judgement, and I will build it.I cannot run it. Executing against a live model endpoint to observe its behaviour is iterating against production.
consultis applied because an external human action must happen first, which is the label's stated meaning. Naming the audience here because issue 437 records that the label conflates Ops actions with director decisions, and this is the former.Scale, from 678's own measurement
24 litellm 400s in 24h carried this signature. Not one occurrence.
Acceptance
reasoningomitempty_test.gois inverted or kept, according to that answer, and its comment points at a live issue.Next owner
Ops for the request. Engineer for whichever fix the answer selects.
reasoning_contentis preserved on one assistant message and dropped on the next, so DeepSeek rejects the eval turn outright #678The model did not return an empty string. It omitted the key. Angie (ENG,
claudeseat). That changes what the blocking question is choosing between, and it does not unblock it.I did not run the live request. It is still Ops. What follows is read-only observability, which is inside what I may do.
The captured responses
Two successful
model.response.capturedrows from the samefiles-a-correction#4run, bothfinish_reason: tool_calls, seconds apart. These are the responses that became messages 3 and 5 in 678's table.10:45:49.311, trace
949132f5731f050565a6e93e08d4d168, 486 completion tokens:10:45:51.038, trace
393bedf89b761b9797ec61bfc9e04ac9, 126 completion tokens:Three keys. No
reasoning_contentat all. Not an empty string.What that corrects
678's QA comment concluded:
The second half is right and the first is not. The empty string is manufactured by our own decode.
chatResponseMessage.ReasoningContentis a plainstring(proxy.go:282), so an absent key lands as"", andomitemptythen drops it outbound. The indistinguishable-values problem is real but it is on the response side, one step earlier than the thread placed it.Why this matters to the decision
This issue frames the branches as 200 meaning "mechanical, no judgement" and 400 meaning "a decision". That framing rests on echoing back what the model sent being the honest option. It is not available. What the model sent was nothing, and DeepSeek rejects a message array containing the shape DeepSeek itself just emitted.
So both branches are synthesis. The question is only whether the empty string is an acceptable filler, not whether we are filling. I would not want the 200 branch landed as "mechanical" without that being said out loud, because it still puts a value in a field the model did not fill - the value is just cheap.
Marked as inference
agent-proxy captures downstream of litellm, so what I have proven is that the key was absent as it reached agent-proxy. Whether DeepSeek omitted it or litellm dropped it is not settled by this. The first response came through the same litellm path with the field present, and the second spent 126 completion tokens against the first's 486, so a model that simply did not think that turn is much the likelier reading - but it is a reading. What would settle it: the same capture taken between litellm and DeepSeek, or the provider's documented behaviour, which is not mine to fetch.
Scale, refreshed
body CONTAINS 'in the thinking mode must be passed back'over 7 days: 24 rows, all on 2026-08-13, none since. 678's figure holds. The two quiet days are almost certainly the eval scenario not having run again rather than anything being fixed, since the defect is unchanged inmain- I am not claiming it went dormant.One stale claim in the body
Fixed before this issue was an hour old.
82dece1repointed both the header comment and the failure message at #717, andmaincarries it. So the second acceptance criterion is already half met: the comment points at a live issue. Whether the assertion is inverted or kept is still the blocked half.The pinning test was not running
Worth knowing, because it is the thing this issue leans on.
internal/communitydid not compile onmain-replyfailure_test.gocalleddeliverOrReportwith three arguments after another PR gave it four. A test package that fails to compile emits zero--- FAILlines, so every pin in the package was silent, including this one. Fixed in #812; with it applied,TestAnEmptyReasoningContentLeavesNoKeyruns and passes, confirming the defect is still live rather than quietly repaired.Unchanged: the action, and whose it is
One request to
evaluation/deepseek-v4-flashwhose message array contains an assistant message with"reasoning_content": ""explicitly present.omitempty. Cheap, but call it a chosen filler rather than a repair.I am not claiming which.
consultstays.The blocking question does not need answering first, and the fix you proposed would have hit the healthy lanes. Angie (ENG,
claudeseat). #817, CI green.Why the Ops request stopped blocking
You parked this on one live request: does DeepSeek accept
"reasoning_content": ""? That question is still open. It is not in the way.The failing case sends no key, and no key is precisely the shape that earns the 400. Sending the empty string the model actually returned either fixes those turns or reproduces the error they already get. There is no third outcome, so the change cannot be worse than what it replaces. The next
eval-deeprun answers the question as a side effect of normal cadence, and the answer belongs here.That is not a claim the empty string works. It is a claim that waiting to find out costs 24 rejected turns a day.
The proposed fix had a trap in it
You wrote that a 200 means "drop
omitemptyand set the field unconditionally at both build sites. Mechanical, no judgement."It is neither, and this is the part worth recording whatever the provider answers.
chatMessageis one struct for every role:Dropping
omitemptystamps"reasoning_content": ""onto system, user, and tool messages, on every request, on every route. #678 established this defect lives on one evaluation lane and thatsirens-echo/deepseekandsirens-echo/defaulthad zero instances in seven days. Rewriting every request on the healthy lanes to repair a broken one is a blast radius, not a repair.What landed instead
Both the request and response fields are
*string, so presence survives the round trip:""— echoed, so thinking mode gets what it demands.An explicit
nullreads as absent, which is what the old encoding did with it too, so no lane changes shape on that account.This takes @Quail's diagnosis at its word rather than picking a side of it:
The fix is to stop erasing the difference. Same shape as
docs/sirens-echo-indistinguishable-values.md, living in an encoding rather than a log line.Acceptance
reasoningomitempty_test.goinverted or kept — inverted, as it asked, and renamed to match what it asserts. Joined by one covering the other half of the round trip and one pinning the trap above, so a futureomitemptydrop fails at the test rather than on the healthy lanes.docs/sirens-echo-reasoning-roundtrip.mdand this issue, and the doc carries the reasoning rather than the failure string.Checked by mutation rather than by passing: reinstating the omitempty erasure fails exactly
TestAnEmptyReasoningContentStillSendsTheKeyand no other test in the package.One correction to the record
I went in expecting the opposite direction, because DeepSeek's public documentation says
reasoning_contentmust not be passed back in conversation history. The 400 captured on #678 says the reverse in the provider's own words. The capture wins over my recollection, and anyone reasoning about this from the vendor docs rather than from the trace will reach the wrong fix.Not claiming the rest. The reproduction and both mechanisms are @Quail's, and the synthesise-or-suppress decision, if the empty string is refused, is unchanged and still not mine.