Infra Engineer - SRE(Kubernetes)

GMI Cloud

United States

On-site

USD 100,000 - 130,000

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

GMI Cloud, a fast-growing AI infrastructure startup in Silicon Valley, seeks a dynamic Site Reliability Engineer with expertise in Kubernetes. This hands-on role is essential for managing the stability and efficiency of high-performance AI/ML clusters.

The ideal candidate will have a Bachelor's degree in Computer Science and over 3 years of experience in data center operations. Responsibilities include designing scalable solutions, monitoring system performance, and automating infrastructure resource management.

Qualifications

  • Over 3 years of experience in data center operations or systems engineering.
  • Proven experience in infrastructure automation.
  • Familiarity with logging and monitoring tools.

Responsibilities

  • Design and maintain AI/ML infrastructure solutions.
  • Monitor GPU cluster performance and troubleshoot issues.
  • Automate infrastructure resource management.

Skills

Site reliability engineering
Infrastructure automation
Kubernetes
Troubleshooting skills
Linux system administration
Python scripting
Ansible
Terraform
Monitoring tools (Prometheus, Grafana)

Education

Bachelor’s degree in Computer Science or related field

Tools

Ansible
Terraform
Kubernetes
Prometheus
Grafana

Job description

We are a fast-growing AI infrastructure startup based in Silicon Valley, working on cutting-edge technologies that power the future of artificial intelligence.We power developers, startups, and enterprises with scalable GPU cloud and inference solutions, helping AI builders turn ideas into reality. As we expand globally, we are looking for a dynamic and hands-on Site Reliability Engineer

Role Overview

We are seeking a skilled Site Reliability Engineer in the area of Kubernetes to join the GMI Global Infrastructure team. This role is hands-on and critical to ensuring the stability, efficiency, and reliability of the large-scale high performance AI/ML clusters in our data center. The ideal candidate will bring expertise in system-level troubleshooting, AI cluster maintenance, and operational excellence to ensure maximum performance for our infrastructure. Experience with large-scale infrastructure automation is considered a strong plus.

Responsibilities
  • Design, implement and maintain scalable AI/ML infrastructure solutions.
  • Proactively monitor GPU cluster health, performance and troubleshoot issues across compute, accelerator, and storage systems.
  • Automate deployment, configuration and management of infrastructure resources.
  • Manage GPU node lifecycle workflows, including provisioning, scaling, maintenance, decommissioning and upgrades of GPU nodes.
  • Implement CI/CD pipelines for infrastructure deployment and orchestration.
  • Ensure security, compliance and best practices across infrastructure.
  • Manage incident response related to Infrastructure resources (GPU, CPU, Storage, Network).
  • Handle customer provisioning requests for GPU resources, including onboarding, configuration and troubleshooting; resolve customer service requests related to GPU infrastructure, ensuring high customer satisfaction.
  • Stay current with emerging GPU hardware and software technologies, integrating improvements as appropriate.
  • Regional/international travel to GMI data center locations.
Qualifications
  • Bachelor’s degree in Computer Science or related field.
  • Over 3+ years of experience in data center operations, infrastructure, or systems engineering.
  • Proven experience in site reliability engineering and infrastructure automation (e.g. Ansible, Terraform)
  • Familiarity with containers orchestration platform (e.g. Kubernetes, Nvidia GPU operator, Nvidia Network operator, CNI, CSI) and job scheduling systems (e.g. Slurm).
  • Familiarity with Linux system administration and scripting (Python, Bash).
  • Familiarity with logging and monitoring tools such as Prometheus, Grafana, Loki.
  • Good knowledge of GPU architecture, Nvidia CUDA, NCCL, or related AI/ML frameworks - added advantage.
  • Strong troubleshooting skills and ability to analyze system logs and performance metrics.
  • Excellent communication and teamwork abilities.

Meeting every qualification is not required—if you’re excited about this role, we’d love to hear from you. We believe diverse perspectives and experiences strengthen our team.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI Infra Engineer – SRE (Kubernetes)
AI Infra Engineer – SRE (Kubernetes)

Berrybytes • United States

On-site
USD 110,000 - 150,000
Senior Site Reliability Engineer (SRE) - AI Inftastructure
Senior Site Reliability Engineer (SRE) - AI Inftastructure

Hamilton Barnes Associates Limited • San Francisco (CA)

On-site
USD 270,000 - 330,000
Equity
Staff Site Reliability Engineer - AI Infrastructure
Staff Site Reliability Engineer - AI Infrastructure

Hamilton Barnes Associates Limited • San Francisco (CA)

On-site
USD 297,500 - 402,500
Huge stock options
Company bonus
Unlimited PTO
+1
Kubernetes SRE for AI Infra & GPU Clusters
Kubernetes SRE for AI Infra & GPU Clusters

GMI Cloud • United States

On-site
USD 100,000 - 130,000
Site Reliability Engineer
Site Reliability Engineer

Amiri Recruiting • Mountain View (CA)

On-site
USD 130,000 - 160,000
Senior SRE - AI Infrastructure
Senior SRE - AI Infrastructure

Hamilton Barnes Associates Limited • San Francisco (CA)

On-site
USD 225,000 - 275,000
IPO Equity
10% comapny bonus
401K 4% match
Infrastructure Engineer
Infrastructure Engineer

ITCAPS LLC • St. Louis (MO)

On-site
USD 140,000 - 200,000
Senior Software Engineer, Infrastructure Software for AI (Centralized AI Data Centers & Distrib[...]
Senior Software Engineer, Infrastructure Software for AI (Centralized AI Data Centers & Distrib[...]

Intelliswift - An LTTS Company • Sunnyvale (CA)

On-site
USD 120,000 - 150,000
Competitive salary
Health insurance
Flexible work hours
AI/ML Infra Engineer - Hosting
AI/ML Infra Engineer - Hosting

Hamilton Barnes Associates Limited • San Francisco (CA)

On-site
USD 225,000 - 275,000
Stock options
AI Kernel / Cluster Engineer
AI Kernel / Cluster Engineer

Blue Signal Search • Santa Clara (CA)

On-site
USD 150,000 - 210,000