AI Safety & Alignment Researcher

Main mission

Ensures advanced systems remain reliable, honest, and aligned with human values.

5 key responsibilities

  • Study risky behaviors of advanced models (deception, manipulation).
  • Develop interpretability methods to understand models from the inside.
  • Design red teaming and adversarial testing protocols.
  • Build scalable supervision methods and reward modeling.
  • Contribute to industry and research safety standards.

Key skills

Mechanistic interpretability, red teaming, scalable supervision, reward modeling.

What's expected

Anticipate risks before they materialize. Safety expertise has seen a significant pay premium.

Career paths

Alignment Research Lead, high-level AI governance, policy and regulation.

Openings right now

No opening for this role at the moment.

Get alerted as soon as a AI Safety & Alignment Researcher position is published.

Create a job alert →