AI Infra SRE - Kubernetes & GPU Cluster Reliability

Berrybytes

United States

On-site

USD 110,000 - 150,000

Full time

14 days+
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Berrybytes is seeking an experienced AI Infra Engineer - SRE (Kubernetes) to join our Global Infrastructure team. You will be responsible for designing, operating, and optimizing Kubernetes-based infrastructure for AI workloads, ensuring maximum uptime and efficiency.

The ideal candidate has a Bachelor's degree in Computer Science, 3+ years of experience in data center operations, strong skills in Kubernetes, and familiarity with automation tools such as Terraform and Ansible. Join us to empower AI developers through cutting-edge technology.

Qualifications

  • 3+ years of experience in data center operations or site reliability engineering.
  • Strong background in automation tools like Terraform and Ansible.
  • Experience with observability stacks like Prometheus, Grafana, and Loki.

Responsibilities

  • Design and maintain scalable AI/ML infrastructure using Kubernetes.
  • Monitor GPU cluster performance and perform root-cause analysis.
  • Implement automation for infrastructure provisioning and management.
  • Lead incident response for issues related to GPUs and high-speed networks.

Skills

Infrastructure automation (Terraform, Ansible)
Kubernetes
Linux system administration
Observability stacks (Prometheus, Grafana, Loki)
GPU architecture knowledge

Education

Bachelor’s degree in Computer Science or related field

Tools

Kubernetes
NVIDIA GPU Operator
NVIDIA Network Operator
Slurm

Job description

Berrybytes is seeking an experienced AI Infra Engineer - SRE (Kubernetes) to join our Global Infrastructure team. You will be responsible for designing, operating, and optimizing Kubernetes-based infrastructure for AI workloads, ensuring maximum uptime and efficiency.

The ideal candidate has a Bachelor's degree in Computer Science, 3+ years of experience in data center operations, strong skills in Kubernetes, and familiarity with automation tools such as Terraform and Ansible. Join us to empower AI developers through cutting-edge technology.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI Infra Engineer – SRE (Kubernetes)
AI Infra Engineer – SRE (Kubernetes)

Berrybytes • United States

On-site
USD 110,000 - 150,000
Kubernetes SRE for AI Infra & GPU Clusters
Kubernetes SRE for AI Infra & GPU Clusters

GMI Cloud • United States

On-site
USD 100,000 - 130,000
AI Infra Engineer: Kubernetes on Bare Metal GPUs (Remote)
AI Infra Engineer: Kubernetes on Bare Metal GPUs (Remote)

vCluster • New York (NY)

Hybrid
USD 150,000 - 200,000
Competitive Salary
Platinum-Level Insurance
Flexible Working Schedule
+1
Senior SRE – AI Infra: Scale GPU Clusters & HPC
Senior SRE – AI Infra: Scale GPU Clusters & HPC

Hamilton Barnes Associates Limited • San Francisco (CA)

On-site
USD 225,000 - 275,000
IPO Equity
10% comapny bonus
401K 4% match
AI Infrastructure Engineer - GPU & Kubernetes
AI Infrastructure Engineer - GPU & Kubernetes

HCLTech • California (MO)

On-site
USD 150,000 - 210,000
Medical Insurance
Dental Insurance
Vision Insurance
+2
Senior Kubernetes SRE for AI Infrastructure (Bare-Metal)
Senior Kubernetes SRE for AI Infrastructure (Bare-Metal)

Moonlite • Chicago (IL)

Hybrid
USD 165,000 - 225,000
Competitive total compensation
401(k) match
Fully covered health insurance premiums
Senior AI Infrastructure Engineer — Kubernetes & GPU
Senior AI Infrastructure Engineer — Kubernetes & GPU

Seekr • Austin (TX)

Hybrid
USD 180,000 - 240,000
Equity ownership
Unlimited PTO
14 holidays
+4
Infra Engineer - SRE(Kubernetes)
Infra Engineer - SRE(Kubernetes)

GMI Cloud • United States

On-site
USD 100,000 - 130,000
AI Infrastructure Engineer — GPU Kubernetes for Production
AI Infrastructure Engineer — GPU Kubernetes for Production

vCluster • Germany (OH)

On-site
USD 150,000 - 200,000
Competitive Salary
Platinum-Level Insurance
Flexible Working Schedule
+1
AI Infra & Cluster Engineer — Scale GPU/CPU Orchestration
AI Infra & Cluster Engineer — Scale GPU/CPU Orchestration

Linuxcareers • San Francisco (CA)

On-site
USD 120,000 - 160,000