Site Reliability Engineer (SRE)

Main mission

Ensures the reliability and availability of services at scale.

5 key responsibilities

  • Define and meet reliability objectives (SLO, error budgets).
  • Design systems for failure: redundancy, degradation, chaos.
  • Manage major incidents and conduct post-mortems without blame.
  • Automate operations to eliminate toil.
  • Balance team velocity and production stability.

Key skills

Distributed systems, Kubernetes, observability, incident management, programming.

What's expected

Quantified and maintained reliability: SRE transforms stability into an engineering discipline.

Career paths

Staff SRE, Platform Architect, Head of Infrastructure.

Openings right now

No opening for this role at the moment.

Get alerted as soon as a Site Reliability Engineer (SRE) position is published.

Create a job alert →