Evals Engineer

Main mission

Builds evaluation systems that measure the actual quality of models.

5 key responsibilities

  • Design evaluation suites that measure the actual quality of AI systems.
  • Build business-specific benchmarks for each use case.
  • Automate evaluations: LLM-as-judge, regression tests, A/B.
  • Detect regressions before each production deployment.
  • Spread the culture of measurement within AI teams.

Key skills

Evaluation design, statistics, LLM-as-judge, testing pipelines, critical thinking.

What's expected

Replace 'it looks good' with reliable measurements: one of the booming jobs with the industrialization of AI.

Career paths

Lead AI Quality, Evaluation Researcher, AI Product Manager.

Openings right now

No opening for this role at the moment.

Get alerted as soon as a Evals Engineer position is published.

Create a job alert →