Enable commodity and OSS evaluation lanes #191

Closed
opened 2026-08-05 03:44:24 +00:00 by coilyco-ops · 1 comment
Member

PR #190 makes commodity and OSS evaluation cases explicit but leaves both lanes disabled until evidence exists. Collect independently reviewed DeepSeek commodity and Ornith or Mistral OSS results, then enable each lane only after its active role matrices are complete. Portfolio Strategist remains frontier-only unless a separate compatibility change is validated.

PR #190 makes commodity and OSS evaluation cases explicit but leaves both lanes disabled until evidence exists. Collect independently reviewed DeepSeek commodity and Ornith or Mistral OSS results, then enable each lane only after its active role matrices are complete. Portfolio Strategist remains frontier-only unless a separate compatibility change is validated.
Author
Member

Superseded: OSS testing is retired, commodity is the only lane

Closing rather than completing. Kai's decision on 2026-08-11 is that the board tests against commodity, DeepSeek, only. Roles are still used on OSS models. They are not tested against them.

This issue asked to collect DeepSeek commodity and Ornith or Mistral OSS evidence and then enable each lane. Half of it no longer has a target, and the other half is no longer a lane to enable because it is the only lane there is.

Landed in 0c94735 on evals/262-board-run:

  • Each scenario becomes one case on the commodity subject tier. The three-lane expansion is gone.
  • Coverage rejects a case off the subject tier and a repeated scenario, rather than demanding three lanes and comparing prompts between them.
  • The scorecard collapses from a frontier-by-commodity-by-OSS cross product to one lane. Its header names the tier the records carry, so an archived frontier record still renders in historical mode. A record mixing two tiers is rejected rather than summed.
  • disabled_model_tiers is removed. It existed to explain why two lanes never ran.

Deployment tier is untouched. docs/model-tiers.md keeps every role's declaration, Content Creator's OSS Discord seat included, and now states the separation directly: the board does not read a role's declared tier, because tier does not change selected context.

The coverage cost is recorded rather than implied. The board produces no evidence about frontier or OSS behaviour, including for the four roles declared frontier-only. A tier comparison is a separate arm from the release gate, and running it would mean running the same board against another subject. Both evaluation/ornith-35b and evaluation/ministral-3-14b are live on Agent Proxy if that arm is ever wanted.

Reopen or file fresh if a tier-comparison arm becomes worth its own board run.

## Superseded: OSS testing is retired, commodity is the only lane Closing rather than completing. Kai's decision on 2026-08-11 is that the board tests against commodity, DeepSeek, only. Roles are still **used** on OSS models. They are not tested against them. This issue asked to collect DeepSeek commodity and Ornith or Mistral OSS evidence and then enable each lane. Half of it no longer has a target, and the other half is no longer a lane to enable because it is the only lane there is. Landed in `0c94735` on `evals/262-board-run`: * Each scenario becomes one case on the commodity subject tier. The three-lane expansion is gone. * Coverage rejects a case off the subject tier and a repeated scenario, rather than demanding three lanes and comparing prompts between them. * The scorecard collapses from a frontier-by-commodity-by-OSS cross product to one lane. Its header names the tier the records carry, so an archived frontier record still renders in historical mode. A record mixing two tiers is rejected rather than summed. * `disabled_model_tiers` is removed. It existed to explain why two lanes never ran. Deployment tier is untouched. `docs/model-tiers.md` keeps every role's declaration, Content Creator's OSS Discord seat included, and now states the separation directly: the board does not read a role's declared tier, because tier does not change selected context. The coverage cost is recorded rather than implied. The board produces no evidence about frontier or OSS behaviour, including for the four roles declared frontier-only. A tier comparison is a separate arm from the release gate, and running it would mean running the same board against another subject. Both `evaluation/ornith-35b` and `evaluation/ministral-3-14b` are live on Agent Proxy if that arm is ever wanted. Reopen or file fresh if a tier-comparison arm becomes worth its own board run.
Sign in to join this conversation.
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set

Reference
coilyco-flight-deck/agent-compose#191
No description provided.