AI Cloud SRE Lead — Scale High-Performance GPU Infra

GMI Cloud

United States

On-site

USD 120,000 - 180,000

Full time

4 days ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

GMI Cloud is seeking a Site Reliability Engineer to join the Global Infrastructure team in the United States. This hands-on role ensures the stability, efficiency, and reliability of large-scale AI/ML clusters in our data centers.

You will design, deploy, and maintain scalable AI infrastructure, monitor GPU cluster health, and automate deployment and lifecycle management of GPU nodes. Experience with automation tools and Kubernetes is valued.

Qualifications

  • Bachelor’s degree in Computer Science or related field.
  • 3+ years of experience in data center operations, infrastructure, or systems engineering.
  • Experience with site reliability engineering and infrastructure automation (Ansible, Terraform).
  • Familiarity with containers orchestration platforms (Kubernetes, Nvidia GPU operator, Nvidia Network operator, CNI, CSI) and job scheduling systems (Slurm).
  • Familiarity with Linux system administration and scripting (Python, Bash).
  • Familiarity with logging and monitoring tools such as Prometheus, Grafana, Loki.
  • Knowledge of GPU architecture, CUDA, NCCL or related AI/ML frameworks is a plus.

Responsibilities

  • Design, implement and maintain scalable AI/ML infrastructure solutions.
  • Proactively monitor GPU cluster health, performance and troubleshoot issues across compute, accelerator, and storage systems.
  • Automate deployment, configuration and management of infrastructure resources.
  • Manage GPU node lifecycle workflows including provisioning, scaling, maintenance, decommissioning and upgrades.
  • Implement CI/CD pipelines for infrastructure deployment and orchestration.
  • Ensure security, compliance and best practices across infrastructure.
  • Manage incident response related to infrastructure resources.
  • Handle customer provisioning requests for GPU resources and troubleshoot to ensure high satisfaction.
  • Stay current with emerging GPU hardware and software technologies and integrate improvements.

Skills

Linux administration
Python scripting
Bash scripting
Troubleshooting
CI/CD
Communication
Teamwork

Education

Bachelor’s degree in Computer Science or related field

Tools

Ansible
Terraform
Kubernetes
NVIDIA GPU operator
NVIDIA Network operator
CNI
CSI
Prometheus
Grafana
Loki

Job description

GMI Cloud is seeking a Site Reliability Engineer to join the Global Infrastructure team in the United States. This hands-on role ensures the stability, efficiency, and reliability of large-scale AI/ML clusters in our data centers.

You will design, deploy, and maintain scalable AI infrastructure, monitor GPU cluster health, and automate deployment and lifecycle management of GPU nodes. Experience with automation tools and Kubernetes is valued.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Kubernetes SRE for AI Infra & GPU Clusters
Kubernetes SRE for AI Infra & GPU Clusters

GMI Cloud • United States

On-site
USD 100,000 - 130,000
Infra Engineer - SRE(Kubernetes)
Infra Engineer - SRE(Kubernetes)

GMI Cloud • United States

On-site
USD 100,000 - 130,000
Site Reliability Lead
Site Reliability Lead

GMI Cloud • United States

On-site
USD 120,000 - 180,000
Site Reliability Engineer
Site Reliability Engineer

GMI Cloud • United States

On-site
USD 110,000 - 170,000
AI Infra Engineer – SRE (Kubernetes)
AI Infra Engineer – SRE (Kubernetes)

Berrybytes • United States

On-site
USD 110,000 - 150,000
Lead Cloud SRE Architect for Private Cloud & AI CI/CD
Lead Cloud SRE Architect for Private Cloud & AI CI/CD

NVIDIA Gruppe • Santa Clara (CA)

On-site
USD 272,000 - 431,000
Equity
Benefits
Senior AI Infra SRE: GPU Clusters & High-Perf Networking
Senior AI Infra SRE: GPU Clusters & High-Perf Networking

Andromeda • San Francisco (CA)

Hybrid
USD 150,000 - 200,000
Significant ownership and autonomy
Inclusive environment
Opportunity to shape AI infrastructure
Senior AI Platform SRE: Scale Cloud Infra & Kubernetes
Senior AI Platform SRE: Scale Cloud Infra & Kubernetes

GCS Recruitment • Mount Laurel Township (NJ)

On-site
USD 110,000 - 170,000
Cloud SRE Architect — AI-Driven CI/CD & Scale
Cloud SRE Architect — AI-Driven CI/CD & Scale

NVIDIA Corporation • Santa Clara (CA)

On-site
USD 272,000 - 431,000
Equity
Benefits
AI Infra DevOps & Backend Engineer
AI Infra DevOps & Backend Engineer

GMI Cloud • Mountain View (CA)

On-site
USD 160,000 - 210,000