AI Infra Engineer – SRE (Kubernetes)

Berrybytes

United States

On-site

USD 110,000 - 150,000

Full time

14 days+
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Berrybytes is seeking an experienced AI Infra Engineer - SRE (Kubernetes) to join our Global Infrastructure team. You will be responsible for designing, operating, and optimizing Kubernetes-based infrastructure for AI workloads, ensuring maximum uptime and efficiency.

The ideal candidate has a Bachelor's degree in Computer Science, 3+ years of experience in data center operations, strong skills in Kubernetes, and familiarity with automation tools such as Terraform and Ansible. Join us to empower AI developers through cutting-edge technology.

Qualifications

  • 3+ years of experience in data center operations or site reliability engineering.
  • Strong background in automation tools like Terraform and Ansible.
  • Experience with observability stacks like Prometheus, Grafana, and Loki.

Responsibilities

  • Design and maintain scalable AI/ML infrastructure using Kubernetes.
  • Monitor GPU cluster performance and perform root-cause analysis.
  • Implement automation for infrastructure provisioning and management.
  • Lead incident response for issues related to GPUs and high-speed networks.

Skills

Infrastructure automation (Terraform, Ansible)
Kubernetes
Linux system administration
Observability stacks (Prometheus, Grafana, Loki)
GPU architecture knowledge

Education

Bachelor’s degree in Computer Science or related field

Tools

Kubernetes
NVIDIA GPU Operator
NVIDIA Network Operator
Slurm

Job description

About the Role

We are a fast-growing AI infrastructure company building cutting-edge GPU cloud platforms and high-performance inference solutions that empower AI developers, startups, and enterprises worldwide. As we scale our global operations, we are looking for a skilled and hands-on AI Infra Engineer - SRE (Kubernetes) to join our Global Infrastructure team.

Role Overview

This is a critical hands-on position focused on the reliability, performance, and operational excellence of large-scale, high-performance AI/ML GPU clusters in our data centers. As an AI Infra Engineer - SRE (Kubernetes), you will design, operate, and optimize Kubernetes-based infrastructure to ensure maximum uptime, efficiency, and scalability for demanding AI workloads. You will bring deep expertise in system-level troubleshooting, GPU cluster management, and automation to keep our platforms running at peak performance.

Key Responsibilities
  • Design, build, and maintain scalable, production-grade AI/ML infrastructure using Kubernetes.
  • Proactively monitor GPU cluster health, performance, and utilization across compute, accelerators, storage, and networking layers, performing root-cause analysis and resolution.
  • Develop and implement automation for infrastructure provisioning, configuration, and ongoing management.
  • Own the complete GPU node lifecycle — including provisioning, dynamic scaling, maintenance, decommissioning, and zero-downtime upgrades of GPU-enabled nodes in Kubernetes environments.
  • Build and improve CI/CD pipelines for reliable infrastructure deployment and orchestration.
  • Enforce security best practices, compliance standards, and operational excellence across the infrastructure stack.
  • Lead incident response and post-incident improvements for issues related to GPUs, CPUs, high-speed storage, and networks.
  • Manage end-to-end customer GPU resource provisioning — from request intake and configuration to onboarding, troubleshooting, and support — ensuring high levels of customer satisfaction.
  • Stay up to date with the latest GPU hardware, software, and orchestration technologies, integrating relevant advancements into our platforms.
  • Be available for occasional regional or international travel to data center locations as required.
Requirements
  • Bachelor’s degree in Computer Science, Engineering, or a related technical field.
  • 3+ years of practical experience in data center operations, infrastructure engineering, or site reliability engineering.
  • Strong background in infrastructure automation using tools such as Terraform and Ansible.
  • Deep hands-on experience with Kubernetes in large-scale environments, including:
    • NVIDIA GPU Operator for GPU driver management, device plugins, container toolkit, and monitoring (DCGM).
    • NVIDIA Network Operator for high-performance networking, RDMA, and GPUDirect support.
    • CNI (Container Network Interface) and CSI (Container Storage Interface) plugins tailored for AI/ML workloads.
    • Integration with job schedulers such as Slurm in Kubernetes clusters.
  • Proficiency in Linux system administration and scripting (Python, Bash).
  • Experience with observability stacks including Prometheus, Grafana, and Loki.
  • Solid understanding of GPU architecture, NVIDIA CUDA, NCCL, and AI/ML frameworks is a strong plus.
  • Excellent troubleshooting skills with the ability to analyze complex system logs and performance metrics.
  • Strong communication and collaboration skills to work effectively with engineering and operations teams.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Infra Engineer - SRE(Kubernetes)
Infra Engineer - SRE(Kubernetes)

GMI Cloud • United States

On-site
USD 100,000 - 130,000
Cluster Engineer
Cluster Engineer

STN Inc • San Francisco (CA)

On-site
USD 180,000 - 240,000
AI Kernel / Cluster Engineer
AI Kernel / Cluster Engineer

Blue Signal Search • Santa Clara (CA)

On-site
USD 150,000 - 210,000
Senior Site Reliability Engineer (SRE) - AI Inftastructure
Senior Site Reliability Engineer (SRE) - AI Inftastructure

Hamilton Barnes Associates Limited • San Francisco (CA)

On-site
USD 270,000 - 330,000
Equity
Senior SRE - AI Infrastructure
Senior SRE - AI Infrastructure

Hamilton Barnes Associates Limited • San Francisco (CA)

On-site
USD 225,000 - 275,000
IPO Equity
10% comapny bonus
401K 4% match
Staff Site Reliability Engineer - AI Infrastructure
Staff Site Reliability Engineer - AI Infrastructure

Hamilton Barnes Associates Limited • San Francisco (CA)

On-site
USD 297,500 - 402,500
Huge stock options
Company bonus
Unlimited PTO
+1
Head of AI Data Center Infrastructure Platforms and Software
Head of AI Data Center Infrastructure Platforms and Software

Summit Group Solutions, LLC • United States

On-site
USD 150,000 - 350,000
Platform Engineer (GPU)
Platform Engineer (GPU)

Vero • United States

On-site
USD 136,000 - 160,000
Medical, dental, and vision insurance
Equity Scheme
401(k) with employer match
+3
AI Infrastructure Engineer — GPU Kubernetes for Production
AI Infrastructure Engineer — GPU Kubernetes for Production

vCluster • Germany (OH)

On-site
USD 150,000 - 200,000
Competitive Salary
Platinum-Level Insurance
Flexible Working Schedule
+1
Senior/Staff Software Engineer, Kubernetes Infrastructure
Senior/Staff Software Engineer, Kubernetes Infrastructure

Kindredventures • United States

On-site
USD 140,000 - 190,000