AI Infrastructure Engineer
Main mission
Manages GPU clusters, distributed training, and computing costs.
5 key responsibilities
- Design and operate GPU training and inference clusters.
- Optimize computing usage: scheduling, parallelism, queues.
- Manage computing costs and trade-offs between cloud and on-premise.
- Ensure reliability of long jobs: checkpointing, recovery, monitoring.
- Equip ML teams to be autonomous on infrastructure.
Key skills
Kubernetes, Slurm, GPU (CUDA), high-performance networking, cloud, FinOps for computing.
What's expected
Fast, available computing at the right cost: infrastructure is the backbone of model warfare.
Career paths
AI Platform Architect, Head of Infrastructure, HPC Engineer.
Openings right now
No opening for this role at the moment.
Get alerted as soon as a AI Infrastructure Engineer position is published.
Create a job alert →