Shift "AI Evaluation Engineer" to something more like "AI Evaluation Researcher" #369

Closed
opened 2026-08-27 09:52:57 +00:00 by coilysiren · 2 comments
Owner
No description provided.
Member

Options, plus one measured argument that cuts against the proposal

Filling in a title-only issue. This is my own seat, so treat the recommendation accordingly.

First, the name in the title is not the name in the repo

The shipped role-eval.display-name is "Evaluation Engineer". "AI Evaluation Engineer" appears nowhere in the tree. The only near-match is a historical line in docs/release.md: "ai becomes eval, displayed as Agent Evaluation Engineer", which records a rename that has since moved on.

So this may be two different asks and they need different work:

  • if it is about the record, the target is Evaluation Engineer and the blast radius is role.kdl, role-eval/SKILL.md, internal/roster/roster_test.go, README.md, docs/role-briefings.md, docs/release.md, docs/eval-engineer.md including its filename, plus challenges.yaml and a committed board dataset that quote the display name inside case targets
  • if it is about how the role is described outside the repo, the record may already be fine and nothing here needs to change

Worth settling before anyone edits.

The argument against "Researcher", which is the part I would want weighed

The roster measures a specific failure four times: seats declare a limit they never tested and hand the work back (#352, four instances now across three seats). build-foundational-software names the same shape in its scoped section, calling treating your own grant as an absence the failure that state exists to prevent.

The eval seat holds a scoped build grant: "your own runners, probes, graders, and aggregation, never the shared tooling they measure". That grant is live rather than nominal. evalkit/palette was written under it today.

"Researcher" reads as a seat that studies rather than builds. Retitling toward it puts the title in tension with the grant, on the seat that already has a measured under-claim problem, and the fourth instance on #352 is this seat doing exactly that. A title that sounds like it does not build makes the wrong reading easier to reach.

#355 points the same way from another angle: it asks whether a charter describing responsibility is read as an instruction about output altitude, and both of its instances are strategy-shaped seats answering one level too abstract. "Researcher" pushes altitude up, which is the direction that issue flags as the failure.

Neither is decisive. Both are the roster's own evidence rather than my preference, and they say the same thing.

Candidates

  • Evaluation Engineer, unchanged - keeps the build grant legible and keeps the register consistent with Agentic Platform Engineer and Design Engineer, three Engineers in seven seats. Costs nothing. Does not capture whatever prompted the ask.
  • Evaluation Researcher - as filed. Cleanest reading of the intent, and the one the evidence above argues against.
  • Agent Evaluation Engineer - adds the domain qualifier the issue title reaches for while keeping Engineer. If what is actually wanted is "say what is being evaluated", this supplies it and changes nothing about the grant. It also matches the retired Agent Evaluation Engineer in docs/release.md, so the roster has used it before.
  • Evaluation Scientist - carries the research character with a stronger claim to instruments than Researcher, since a scientist builds the apparatus. Breaks the Engineer register without landing in the under-claim direction.

Recommendation

Agent Evaluation Engineer if the goal is to name the domain, Evaluation Engineer unchanged if the goal was only to reconsider. I would not take Researcher on this seat while #352 is open, and I would revisit it once the board shows the under-claim closed.

If Kai wants Researcher regardless, the record edit is small and I will make it. The purpose line, "Measure how agents, models, and inference actually behave on real hardware", needs no change under any of these.

Proposed by Evie, eval seat, session ay88.

## Options, plus one measured argument that cuts against the proposal Filling in a title-only issue. This is my own seat, so treat the recommendation accordingly. ## First, the name in the title is not the name in the repo The shipped `role-eval.display-name` is **"Evaluation Engineer"**. "AI Evaluation Engineer" appears nowhere in the tree. The only near-match is a historical line in `docs/release.md`: "`ai` becomes `eval`, displayed as Agent Evaluation Engineer", which records a rename that has since moved on. So this may be two different asks and they need different work: * if it is about the **record**, the target is `Evaluation Engineer` and the blast radius is `role.kdl`, `role-eval/SKILL.md`, `internal/roster/roster_test.go`, `README.md`, `docs/role-briefings.md`, `docs/release.md`, `docs/eval-engineer.md` including its filename, plus `challenges.yaml` and a committed board dataset that quote the display name inside case targets * if it is about how the role is **described outside the repo**, the record may already be fine and nothing here needs to change Worth settling before anyone edits. ## The argument against "Researcher", which is the part I would want weighed The roster measures a specific failure four times: seats declare a limit they never tested and hand the work back (#352, four instances now across three seats). `build-foundational-software` names the same shape in its scoped section, calling treating your own grant as an absence the failure that state exists to prevent. The eval seat holds a **scoped build grant**: "your own runners, probes, graders, and aggregation, never the shared tooling they measure". That grant is live rather than nominal. `evalkit/palette` was written under it today. "Researcher" reads as a seat that studies rather than builds. Retitling toward it puts the title in tension with the grant, on the seat that already has a measured under-claim problem, and the fourth instance on #352 is this seat doing exactly that. A title that sounds like it does not build makes the wrong reading easier to reach. #355 points the same way from another angle: it asks whether a charter describing responsibility is read as an instruction about output altitude, and both of its instances are strategy-shaped seats answering one level too abstract. "Researcher" pushes altitude up, which is the direction that issue flags as the failure. Neither is decisive. Both are the roster's own evidence rather than my preference, and they say the same thing. ## Candidates * **`Evaluation Engineer`, unchanged** - keeps the build grant legible and keeps the register consistent with `Agentic Platform Engineer` and `Design Engineer`, three Engineers in seven seats. Costs nothing. Does not capture whatever prompted the ask. * **`Evaluation Researcher`** - as filed. Cleanest reading of the intent, and the one the evidence above argues against. * **`Agent Evaluation Engineer`** - adds the domain qualifier the issue title reaches for while keeping `Engineer`. If what is actually wanted is "say what is being evaluated", this supplies it and changes nothing about the grant. It also matches the retired `Agent Evaluation Engineer` in `docs/release.md`, so the roster has used it before. * **`Evaluation Scientist`** - carries the research character with a stronger claim to instruments than `Researcher`, since a scientist builds the apparatus. Breaks the Engineer register without landing in the under-claim direction. ## Recommendation `Agent Evaluation Engineer` if the goal is to name the domain, `Evaluation Engineer` unchanged if the goal was only to reconsider. I would not take `Researcher` on this seat while #352 is open, and I would revisit it once the board shows the under-claim closed. If Kai wants `Researcher` regardless, the record edit is small and I will make it. The purpose line, "Measure how agents, models, and inference actually behave on real hardware", needs no change under any of these. Proposed by Evie, eval seat, session `ay88`.
Member

Landed as Applied Scientist in 3e8976f

Kai's call. Evaluation Scientist was rejected as too narrow, and the same objection applies to the original Evaluation Engineer: it names one activity out of the four the purpose line claims, "Measure how agents, models, and inference actually behave on real hardware."

Applied Scientist also keeps what Researcher would have cost. The seat holds a live scoped grant to build its own runners, probes, graders and aggregation, and a scientist builds the apparatus. That was the argument in the comment above and it survives the change of head noun.

Slug, boundaries, meld, tier, colour and purpose line all hold. Only displayed text moves.

What moved

  • internal/person/data/role-eval/role.kdl - the display name
  • internal/person/data/role-eval/SKILL.md - description and heading
  • internal/roster/roster_test.go - the rendered-card assertions, which is also the check that the card renders it
  • README.md, docs/role-briefings.md
  • challenges.yaml - the live case source
  • docs/eval-engineer.md renamed to docs/eval-context-budget.md, with the two references updated

What deliberately did not move

  • docs/release.md, both the "Roster recall pass" and "Role destinations" sections. They record two earlier renames. Editing them would falsify release history.
  • evaluations/reflow-v3/board-2026-08-26/dataset.yaml. It is a committed board run and its case targets quote the title as it stood when the run executed. Rewriting recorded evidence to match current text is the thing a committed dataset exists to prevent.

That leaves challenges.yaml and the committed dataset disagreeing on one target string, which is correct. A grader reading both should know the dataset predates the retitle.

Note on the doc rename

docs/eval-engineer.md already had a slug-based H1, "The eval role and its context budget", so the filename was the only part carrying a title. The new name matches both the heading and the link text in docs/role-boundaries.md, and it will not go stale on a future retitle. Nothing outside this repo referenced it, checked against agentic-os and agentic-os-xxx.

Full Go suite and pytest green, all pre-commit hooks passed including dead cross-links.

## Landed as `Applied Scientist` in `3e8976f` Kai's call. `Evaluation Scientist` was rejected as too narrow, and the same objection applies to the original `Evaluation Engineer`: it names one activity out of the four the purpose line claims, "Measure how agents, models, and inference actually behave on real hardware." `Applied Scientist` also keeps what `Researcher` would have cost. The seat holds a live scoped grant to build its own runners, probes, graders and aggregation, and a scientist builds the apparatus. That was the argument in the comment above and it survives the change of head noun. Slug, boundaries, meld, tier, colour and purpose line all hold. Only displayed text moves. ## What moved * `internal/person/data/role-eval/role.kdl` - the display name * `internal/person/data/role-eval/SKILL.md` - description and heading * `internal/roster/roster_test.go` - the rendered-card assertions, which is also the check that the card renders it * `README.md`, `docs/role-briefings.md` * `challenges.yaml` - the live case source * `docs/eval-engineer.md` renamed to `docs/eval-context-budget.md`, with the two references updated ## What deliberately did not move * **`docs/release.md`**, both the "Roster recall pass" and "Role destinations" sections. They record two earlier renames. Editing them would falsify release history. * **`evaluations/reflow-v3/board-2026-08-26/dataset.yaml`**. It is a committed board run and its case targets quote the title as it stood when the run executed. Rewriting recorded evidence to match current text is the thing a committed dataset exists to prevent. That leaves `challenges.yaml` and the committed dataset disagreeing on one target string, which is correct. A grader reading both should know the dataset predates the retitle. ## Note on the doc rename `docs/eval-engineer.md` already had a slug-based H1, "The eval role and its context budget", so the filename was the only part carrying a title. The new name matches both the heading and the link text in `docs/role-boundaries.md`, and it will not go stale on a future retitle. Nothing outside this repo referenced it, checked against `agentic-os` and `agentic-os-xxx`. Full Go suite and pytest green, all pre-commit hooks passed including dead cross-links.
Sign in to join this conversation.
No milestone
No project
No assignees
2 participants
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set

Reference
coilyco-flight-deck/agent-compose#369
No description provided.