Site Reliability Engineer (SRE)
Main mission
Ensures the reliability and availability of services at scale.
5 key responsibilities
- Define and meet reliability objectives (SLO, error budgets).
- Design systems for failure: redundancy, degradation, chaos.
- Manage major incidents and conduct post-mortems without blame.
- Automate operations to eliminate toil.
- Balance team velocity and production stability.
Key skills
Distributed systems, Kubernetes, observability, incident management, programming.
What's expected
Quantified and maintained reliability: SRE transforms stability into an engineering discipline.
Career paths
Staff SRE, Platform Architect, Head of Infrastructure.
Openings right now
No opening for this role at the moment.
Get alerted as soon as a Site Reliability Engineer (SRE) position is published.
Create a job alert →