Get more replies from employers
Send a job-specific resume in minutes.
Lambda is seeking an experienced Site Reliability Engineer to monitor and optimize AI compute clusters. You will deploy, scale, and automate HPC environments, building reliable operations for GPU-rich workloads across on-site teams.
You will implement robust tooling in Python/Go, use Prometheus and Grafana for observability, and drive incident response with a focus on stability and efficiency. This role requires presence in SF or Bellevue with a hybrid setup.
Lambda is seeking an experienced Site Reliability Engineer to monitor and optimize AI compute clusters. You will deploy, scale, and automate HPC environments, building reliable operations for GPU-rich workloads across on-site teams.
You will implement robust tooling in Python/Go, use Prometheus and Grafana for observability, and drive incident response with a focus on stability and efficiency. This role requires presence in SF or Bellevue with a hybrid setup.