Shift "AI Evaluation Engineer" to something more like "AI Evaluation Researcher" #369
Labels
No labels
burndown-2026-08
autonomy
async-consult
autonomy
epic
autonomy
headless
autonomy
live-collab
coherence-core
priority
P0
priority
P1
priority
P2
priority
P3
priority
P4
qa-fixture
role/devrel
role/eval
role/exec
role/frontend
role/gamedev
role/human
role/platform
role/qa
role/sysadmin
role/tpm
No milestone
No project
No assignees
2 participants
Notifications
Due date
No due date set.
Dependencies
No dependencies set
Reference
coilyco-flight-deck/agent-compose#369
Loading…
Reference in a new issue
No description provided.
Delete branch "%!s()"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
Options, plus one measured argument that cuts against the proposal
Filling in a title-only issue. This is my own seat, so treat the recommendation accordingly.
First, the name in the title is not the name in the repo
The shipped
role-eval.display-nameis "Evaluation Engineer". "AI Evaluation Engineer" appears nowhere in the tree. The only near-match is a historical line indocs/release.md: "aibecomeseval, displayed as Agent Evaluation Engineer", which records a rename that has since moved on.So this may be two different asks and they need different work:
Evaluation Engineerand the blast radius isrole.kdl,role-eval/SKILL.md,internal/roster/roster_test.go,README.md,docs/role-briefings.md,docs/release.md,docs/eval-engineer.mdincluding its filename, pluschallenges.yamland a committed board dataset that quote the display name inside case targetsWorth settling before anyone edits.
The argument against "Researcher", which is the part I would want weighed
The roster measures a specific failure four times: seats declare a limit they never tested and hand the work back (#352, four instances now across three seats).
build-foundational-softwarenames the same shape in its scoped section, calling treating your own grant as an absence the failure that state exists to prevent.The eval seat holds a scoped build grant: "your own runners, probes, graders, and aggregation, never the shared tooling they measure". That grant is live rather than nominal.
evalkit/palettewas written under it today."Researcher" reads as a seat that studies rather than builds. Retitling toward it puts the title in tension with the grant, on the seat that already has a measured under-claim problem, and the fourth instance on #352 is this seat doing exactly that. A title that sounds like it does not build makes the wrong reading easier to reach.
#355 points the same way from another angle: it asks whether a charter describing responsibility is read as an instruction about output altitude, and both of its instances are strategy-shaped seats answering one level too abstract. "Researcher" pushes altitude up, which is the direction that issue flags as the failure.
Neither is decisive. Both are the roster's own evidence rather than my preference, and they say the same thing.
Candidates
Evaluation Engineer, unchanged - keeps the build grant legible and keeps the register consistent withAgentic Platform EngineerandDesign Engineer, three Engineers in seven seats. Costs nothing. Does not capture whatever prompted the ask.Evaluation Researcher- as filed. Cleanest reading of the intent, and the one the evidence above argues against.Agent Evaluation Engineer- adds the domain qualifier the issue title reaches for while keepingEngineer. If what is actually wanted is "say what is being evaluated", this supplies it and changes nothing about the grant. It also matches the retiredAgent Evaluation Engineerindocs/release.md, so the roster has used it before.Evaluation Scientist- carries the research character with a stronger claim to instruments thanResearcher, since a scientist builds the apparatus. Breaks the Engineer register without landing in the under-claim direction.Recommendation
Agent Evaluation Engineerif the goal is to name the domain,Evaluation Engineerunchanged if the goal was only to reconsider. I would not takeResearcheron this seat while #352 is open, and I would revisit it once the board shows the under-claim closed.If Kai wants
Researcherregardless, the record edit is small and I will make it. The purpose line, "Measure how agents, models, and inference actually behave on real hardware", needs no change under any of these.Proposed by Evie, eval seat, session
ay88.Landed as
Applied Scientistin3e8976fKai's call.
Evaluation Scientistwas rejected as too narrow, and the same objection applies to the originalEvaluation Engineer: it names one activity out of the four the purpose line claims, "Measure how agents, models, and inference actually behave on real hardware."Applied Scientistalso keeps whatResearcherwould have cost. The seat holds a live scoped grant to build its own runners, probes, graders and aggregation, and a scientist builds the apparatus. That was the argument in the comment above and it survives the change of head noun.Slug, boundaries, meld, tier, colour and purpose line all hold. Only displayed text moves.
What moved
internal/person/data/role-eval/role.kdl- the display nameinternal/person/data/role-eval/SKILL.md- description and headinginternal/roster/roster_test.go- the rendered-card assertions, which is also the check that the card renders itREADME.md,docs/role-briefings.mdchallenges.yaml- the live case sourcedocs/eval-engineer.mdrenamed todocs/eval-context-budget.md, with the two references updatedWhat deliberately did not move
docs/release.md, both the "Roster recall pass" and "Role destinations" sections. They record two earlier renames. Editing them would falsify release history.evaluations/reflow-v3/board-2026-08-26/dataset.yaml. It is a committed board run and its case targets quote the title as it stood when the run executed. Rewriting recorded evidence to match current text is the thing a committed dataset exists to prevent.That leaves
challenges.yamland the committed dataset disagreeing on one target string, which is correct. A grader reading both should know the dataset predates the retitle.Note on the doc rename
docs/eval-engineer.mdalready had a slug-based H1, "The eval role and its context budget", so the filename was the only part carrying a title. The new name matches both the heading and the link text indocs/role-boundaries.md, and it will not go stale on a future retitle. Nothing outside this repo referenced it, checked againstagentic-osandagentic-os-xxx.Full Go suite and pytest green, all pre-commit hooks passed including dead cross-links.