AI Infrastructure Engineer

Main mission

Manages GPU clusters, distributed training, and computing costs.

5 key responsibilities

  • Design and operate GPU training and inference clusters.
  • Optimize computing usage: scheduling, parallelism, queues.
  • Manage computing costs and trade-offs between cloud and on-premise.
  • Ensure reliability of long jobs: checkpointing, recovery, monitoring.
  • Equip ML teams to be autonomous on infrastructure.

Key skills

Kubernetes, Slurm, GPU (CUDA), high-performance networking, cloud, FinOps for computing.

What's expected

Fast, available computing at the right cost: infrastructure is the backbone of model warfare.

Career paths

AI Platform Architect, Head of Infrastructure, HPC Engineer.

Openings right now

No opening for this role at the moment.

Get alerted as soon as a AI Infrastructure Engineer position is published.

Create a job alert →