Kubernetes SRE for AI Infra & GPU Clusters

GMI Cloud

United States

On-site

USD 100,000 - 130,000

Full time

14 days+
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

GMI Cloud, a fast-growing AI infrastructure startup in Silicon Valley, seeks a dynamic Site Reliability Engineer with expertise in Kubernetes. This hands-on role is essential for managing the stability and efficiency of high-performance AI/ML clusters.

The ideal candidate will have a Bachelor's degree in Computer Science and over 3 years of experience in data center operations. Responsibilities include designing scalable solutions, monitoring system performance, and automating infrastructure resource management.

Qualifications

  • Over 3 years of experience in data center operations or systems engineering.
  • Proven experience in infrastructure automation.
  • Familiarity with logging and monitoring tools.

Responsibilities

  • Design and maintain AI/ML infrastructure solutions.
  • Monitor GPU cluster performance and troubleshoot issues.
  • Automate infrastructure resource management.

Skills

Site reliability engineering
Infrastructure automation
Kubernetes
Troubleshooting skills
Linux system administration
Python scripting
Ansible
Terraform
Monitoring tools (Prometheus, Grafana)

Education

Bachelor’s degree in Computer Science or related field

Tools

Ansible
Terraform
Kubernetes
Prometheus
Grafana

Job description

GMI Cloud, a fast-growing AI infrastructure startup in Silicon Valley, seeks a dynamic Site Reliability Engineer with expertise in Kubernetes. This hands-on role is essential for managing the stability and efficiency of high-performance AI/ML clusters.

The ideal candidate will have a Bachelor's degree in Computer Science and over 3 years of experience in data center operations. Responsibilities include designing scalable solutions, monitoring system performance, and automating infrastructure resource management.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Infra Engineer - SRE(Kubernetes)
Infra Engineer - SRE(Kubernetes)

GMI Cloud • United States

On-site
USD 100,000 - 130,000
AI Infra Engineer – SRE (Kubernetes)
AI Infra Engineer – SRE (Kubernetes)

Berrybytes • United States

On-site
USD 110,000 - 150,000
AI Infra SRE - Kubernetes & GPU Cluster Reliability
AI Infra SRE - Kubernetes & GPU Cluster Reliability

Berrybytes • United States

On-site
USD 110,000 - 150,000
Global Remote SRE for AI Infrastructure & Kubernetes
Global Remote SRE for AI Infrastructure & Kubernetes

Andromeda Cluster • San Francisco (CA)

Hybrid
USD 120,000 - 160,000
Senior SRE: Kubernetes, GPU Infra & ML Ops Leader
Senior SRE: Kubernetes, GPU Infra & ML Ops Leader

Gruve • Redwood City (CA)

On-site
USD 120,000 - 150,000
Senior Kubernetes SRE for AI Infrastructure (Bare-Metal)
Senior Kubernetes SRE for AI Infrastructure (Bare-Metal)

Moonlite • Chicago (IL)

Hybrid
USD 165,000 - 225,000
Competitive total compensation
401(k) match
Fully covered health insurance premiums
Senior AI GPU Infra SRE - Scale, Automation & Equity
Senior AI GPU Infra SRE - Scale, Automation & Equity

Hamilton Barnes Associates Limited • San Francisco (CA)

On-site
USD 270,000 - 330,000
Equity
SRE: Scalable ML Infra & CI/CD Architect
SRE: Scalable ML Infra & CI/CD Architect

Baseten • San Francisco (CA)

On-site
USD 165,000 - 330,000
Site Reliability Engineer
Site Reliability Engineer

Amiri Recruiting • Mountain View (CA)

On-site
USD 130,000 - 160,000
Senior Kubernetes SRE — Scale AI Clusters & Automation
Senior Kubernetes SRE — Scale AI Clusters & Automation

Neura Market • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 240,000
Equity compensation
Health, dental and vision coverage
Wellness and commuter stipends
+2