Lead Cloud Infrastructure Engineer - GPU AI Kubernetes

FriendliAI

San Francisco (CA)

On-site

USD 150,000 - 190,000

Full time

14 days+
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Benefits offered by this job

Flexible working hours
Lunch and dinner provided
Health check-up support with top-tier硬
Top-tier equipment/hardware support

Job summary

FriendliAI is building the fastest inference cloud for agents, delivering low-latency, scalable GPU-accelerated AI workloads. We seek a Cloud Infrastructure Engineer to own cluster architecture, extend Kubernetes, and manage the network path for inference traffic.

The role demands hands-on experience with large multi-cluster deployments, designing topology, upgrades, and capacity across tenants. You will shape autoscaling and service mesh, collaborating across SRE, security, and engine teams to

Qualifications

  • Extensive experience designing and operating large-scale Kubernetes infra.
  • Strong programming skills in Go or Python for tooling.
  • Proven track record debugging distributed systems and networks.

Responsibilities

  • Own architecture of multi-cluster, multi-tenant Kubernetes fleet.
  • Extend Kubernetes with custom controllers, operators, and CRDs.
  • Design GPU scheduling, topology-aware placement, and quotas across tenants.
  • Build autoscaling for inference traffic, including queue-driven pod scaling and scale-to-zero.
  • Own the Kubernetes network data plane: CNI, IPAM, DNS, ingress, and L4/L7 load balancing.
  • Design cross-AZ/region connectivity and operate service mesh (mTLS, routing).
  • Define SLOs and lead post-incident hardening; deliver IaC with Terraform, Helm, GitOps.
  • Collaborate with inference, platform, SRE, and security teams to turn requirements into platform capabilities.

Skills

Kubernetes
Go or Python
Distributed systems debugging
Communication

Education

Bachelors or Masters degree in CS/CE/EE

Tools

AWS
Terraform
Helm
Ansible

Job description

FriendliAI is building the fastest inference cloud for agents, delivering low-latency, scalable GPU-accelerated AI workloads. We seek a Cloud Infrastructure Engineer to own cluster architecture, extend Kubernetes, and manage the network path for inference traffic.

The role demands hands-on experience with large multi-cluster deployments, designing topology, upgrades, and capacity across tenants. You will shape autoscaling and service mesh, collaborating across SRE, security, and engine teams to

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Software Engineer – Cloud Infrastructure
Software Engineer – Cloud Infrastructure

FriendliAI • San Francisco (CA)

On-site
USD 150,000 - 190,000
Flexible working hours
Lunch and dinner provided
Health check-up support with top-tier硬
+1
Platform Engineer - GPU Infra & Kubernetes
Platform Engineer - GPU Infra & Kubernetes

Together AI • San Francisco (CA)

On-site
USD 160,000 - 280,000
Equity
Health insurance
Competitive benefits
AI Infra Engineer: GPU Cloud & Kubernetes POC Leader
AI Infra Engineer: GPU Cloud & Kubernetes POC Leader

Vcluster • Northern (KY)

Hybrid
USD 140,000 - 165,000
Competitive Salary
Equity participation
Health, dental, vision, life Insurance
+2
AI Infra Engineer — GPU Cloud, Kubernetes/Slurm
AI Infra Engineer — GPU Cloud, Kubernetes/Slurm

Blue Signal Search • San Francisco (CA)

On-site
USD 180,000 - 240,000
Annual bonus
Equity participation
Comprehensive benefits
+1
Senior AI Infrastructure Engineer - GPU & Kubernetes
Senior AI Infrastructure Engineer - GPU & Kubernetes

HCL Technologies Limited • California (MO)

On-site
USD 120,000 - 180,000
401(k) retirement plan
Paid time off (PTO)
Paid holidays
+1
Kubernetes Platform Engineer for GPU Inference
Kubernetes Platform Engineer for GPU Inference

Togetherai • San Francisco (CA)

On-site
USD 160,000 - 280,000
Health insurance
AI Infrastructure Engineer: Kubernetes & GPU Clusters
AI Infrastructure Engineer: Kubernetes & GPU Clusters

NVIDIA • United States

Remote
USD 184,000 - 288,000
Senior AI Infrastructure Engineer — Global GPU Cloud
Senior AI Infrastructure Engineer — Global GPU Cloud

Together Computer Inc • United States

Remote
USD 160,000 - 230,000
Startup equity
Health insurance
Remote work flexibility
Senior Cloud AI Infra Engineer: Go/C, Kubernetes, GPUs
Senior Cloud AI Infra Engineer: Go/C, Kubernetes, GPUs

Nvidia • Santa Clara (CA)

On-site
USD 240,000 - 340,000
Competitive salaries
Comprehensive benefits package
Equity
AI Infra Engineer: Kubernetes on Bare Metal GPUs (Remote)
AI Infra Engineer: Kubernetes on Bare Metal GPUs (Remote)

vCluster • New York (NY)

Hybrid
USD 150,000 - 200,000
Competitive Salary
Platinum-Level Insurance
Flexible Working Schedule
+1