Site Reliability Lead

GMI Cloud

United States

On-site

USD 120,000 - 180,000

Full time

2 days ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

GMI Cloud is seeking a Site Reliability Engineer to join the Global Infrastructure team in the United States. This hands-on role ensures the stability, efficiency, and reliability of large-scale AI/ML clusters in our data centers.

You will design, deploy, and maintain scalable AI infrastructure, monitor GPU cluster health, and automate deployment and lifecycle management of GPU nodes. Experience with automation tools and Kubernetes is valued.

Qualifications

  • Bachelor’s degree in Computer Science or related field.
  • 3+ years of experience in data center operations, infrastructure, or systems engineering.
  • Experience with site reliability engineering and infrastructure automation (Ansible, Terraform).
  • Familiarity with containers orchestration platforms (Kubernetes, Nvidia GPU operator, Nvidia Network operator, CNI, CSI) and job scheduling systems (Slurm).
  • Familiarity with Linux system administration and scripting (Python, Bash).
  • Familiarity with logging and monitoring tools such as Prometheus, Grafana, Loki.
  • Knowledge of GPU architecture, CUDA, NCCL or related AI/ML frameworks is a plus.

Responsibilities

  • Design, implement and maintain scalable AI/ML infrastructure solutions.
  • Proactively monitor GPU cluster health, performance and troubleshoot issues across compute, accelerator, and storage systems.
  • Automate deployment, configuration and management of infrastructure resources.
  • Manage GPU node lifecycle workflows including provisioning, scaling, maintenance, decommissioning and upgrades.
  • Implement CI/CD pipelines for infrastructure deployment and orchestration.
  • Ensure security, compliance and best practices across infrastructure.
  • Manage incident response related to infrastructure resources.
  • Handle customer provisioning requests for GPU resources and troubleshoot to ensure high satisfaction.
  • Stay current with emerging GPU hardware and software technologies and integrate improvements.

Skills

Linux administration
Python scripting
Bash scripting
Troubleshooting
CI/CD
Communication
Teamwork

Education

Bachelor’s degree in Computer Science or related field

Tools

Ansible
Terraform
Kubernetes
NVIDIA GPU operator
NVIDIA Network operator
CNI
CSI
Prometheus
Grafana
Loki

Job description

GMI Cloud is a fast-growing, AI-native infrastructure company delivering high-performance GPU compute, inference services, and infrastructure for AI agents.

Following 8x ARR growth, GMI Cloud continues to scale rapidly across the U.S. and APAC. As a Reference Platform NVIDIA Cloud Partner (NCP) and a validated leading NCP across both markets, we power production AI for leading AI-native companies including Fireworks AI, Cartesia, Reflection, and OpenRouter.

From large-scale compute to optimized inference and agentic workloads, GMI Cloud gives AI teams the infrastructure they need to build, deploy, and scale on one unified cloud.

One cloud for compute, inference, and agents.

Role Overview

We are seeking a skilled Site Reliability Engineer to join the GMI Global Infrastructure team. This role is hands-on and critical to ensuring the stability, efficiency, and reliability of the large-scale high performance AI/ML clusters in our data center. The ideal candidate will bring expertise in system-level troubleshooting, AI cluster maintenance, and operational excellence to ensure maximum performance for our infrastructure. Experience with large-scale infrastructure automation is considered a strong plus.

Responsibilities
  • Design, implement and maintain scalable AI/ML infrastructure solutions.
  • Proactively monitor GPU cluster health, performance and troubleshoot issues across compute, accelerator, and storage systems.
  • Automate deployment, configuration and management of infrastructure resources.
  • Manage GPU node lifecycle workflows, including provisioning, scaling, maintenance, decommissioning and upgrades of GPU nodes.
  • Implement CI/CD pipelines for infrastructure deployment and orchestration.
  • Ensure security, compliance and best practices across infrastructure.
  • Manage incident response related to Infrastructure resources (GPU, CPU, Storage, Network).
  • Handle customer provisioning requests for GPU resources, including onboarding, configuration and troubleshooting; resolve customer service requests related to GPU infrastructure, ensuring high customer satisfaction.
  • Stay current with emerging GPU hardware and software technologies, integrating improvements as appropriate.
  • Regional/international travel to GMI data center locations.
Qualifications
  • Bachelor’s degree in Computer Science or related field.
  • Over 3+ years of experience in data center operations, infrastructure, or systems engineering.
  • Proven experience in site reliability engineering and infrastructure automation (e.g. Ansible, Terraform)
  • Familiarity with containers orchestration platform (e.g. Kubernetes, Nvidia GPU operator, Nvidia Network operator, CNI, CSI) and job scheduling systems (e.g. Slurm).
  • Familiarity with Linux system administration and scripting (Python, Bash).
  • Familiarity with logging and monitoring tools such as Prometheus, Grafana, Loki.
  • Good knowledge of GPU architecture, Nvidia CUDA, NCCL, or related AI/ML frameworks - added advantage.
  • Strong troubleshooting skills and ability to analyze system logs and performance metrics.
  • Excellent communication and teamwork abilities.

Meeting every qualification is not required—if you’re excited about this role, we’d love to hear from you. We believe diverse perspectives and experiences strengthen our team.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer
Site Reliability Engineer

GMI Cloud • United States

On-site
USD 110,000 - 170,000
Infra Engineer - SRE(Kubernetes)
Infra Engineer - SRE(Kubernetes)

GMI Cloud • United States

On-site
USD 100,000 - 130,000
Infra DevOps and Backend Engineer
Infra DevOps and Backend Engineer

GMI Cloud • Mountain View (CA)

On-site
USD 160,000 - 210,000
Senior Solution Architect – AI / GPU Cloud
Senior Solution Architect – AI / GPU Cloud

GMI Cloud • Mountain View (CA)

On-site
USD 190,000 - 260,000
Influence product roadmap
Career growth opportunities
Work with advanced AI organizations
Technical Program Manager – AI Infrastructure / GPU Clusters
Technical Program Manager – AI Infrastructure / GPU Clusters

GMI Cloud • United States

On-site
USD 140,000 - 210,000
AI Infra Engineer – SRE (Kubernetes)
AI Infra Engineer – SRE (Kubernetes)

Berrybytes • United States

On-site
USD 110,000 - 150,000
Senior Site Reliability Engineer, DGX Cloud
Senior Site Reliability Engineer, DGX Cloud

NVIDIA • California (MO)

On-site
USD 168,000 - 334,000
Equity
Benefits
Senior Site Reliability Engineer, DGX Cloud
Senior Site Reliability Engineer, DGX Cloud

NVIDIA AI • Santa Clara (CA)

On-site
USD 168,000 - 334,000
Senior Site Reliability Engineer, DGX Cloud
Senior Site Reliability Engineer, DGX Cloud

NVIDIA • Town of Texas (WI)

On-site
USD 168,000 - 334,000
Senior Site Reliability Engineer, DGX Cloud
Senior Site Reliability Engineer, DGX Cloud

NVIDIA • Santa Clara (CA)

On-site
USD 190,000 - 320,000