Evals Engineer
Main mission
Builds evaluation systems that measure the actual quality of models.
5 key responsibilities
- Design evaluation suites that measure the actual quality of AI systems.
- Build business-specific benchmarks for each use case.
- Automate evaluations: LLM-as-judge, regression tests, A/B.
- Detect regressions before each production deployment.
- Spread the culture of measurement within AI teams.
Key skills
Evaluation design, statistics, LLM-as-judge, testing pipelines, critical thinking.
What's expected
Replace 'it looks good' with reliable measurements: one of the booming jobs with the industrialization of AI.
Career paths
Lead AI Quality, Evaluation Researcher, AI Product Manager.
Openings right now
No opening for this role at the moment.
Get alerted as soon as a Evals Engineer position is published.
Create a job alert →