Lead Cloud Infrastructure Engineer - GPU AI Kubernetes

FriendliAI

San Francisco (CA)

On-site

USD 150,000 - 190,000

Full time

11 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Flexible working hours
Lunch and dinner provided
Health check-up support with top-tier硬
Top-tier equipment/hardware support

Job summary

FriendliAI is building the fastest inference cloud for agents, delivering low-latency, scalable GPU-accelerated AI workloads. We seek a Cloud Infrastructure Engineer to own cluster architecture, extend Kubernetes, and manage the network path for inference traffic.

The role demands hands-on experience with large multi-cluster deployments, designing topology, upgrades, and capacity across tenants. You will shape autoscaling and service mesh, collaborating across SRE, security, and engine teams to

Qualifications

  • Extensive experience designing and operating large-scale Kubernetes infra.
  • Strong programming skills in Go or Python for tooling.
  • Proven track record debugging distributed systems and networks.

Responsibilities

  • Own architecture of multi-cluster, multi-tenant Kubernetes fleet.
  • Extend Kubernetes with custom controllers, operators, and CRDs.
  • Design GPU scheduling, topology-aware placement, and quotas across tenants.
  • Build autoscaling for inference traffic, including queue-driven pod scaling and scale-to-zero.
  • Own the Kubernetes network data plane: CNI, IPAM, DNS, ingress, and L4/L7 load balancing.
  • Design cross-AZ/region connectivity and operate service mesh (mTLS, routing).
  • Define SLOs and lead post-incident hardening; deliver IaC with Terraform, Helm, GitOps.
  • Collaborate with inference, platform, SRE, and security teams to turn requirements into platform capabilities.

Skills

Kubernetes
Go or Python
Distributed systems debugging
Communication

Education

Bachelors or Masters degree in CS/CE/EE

Tools

AWS
Terraform
Helm
Ansible

Job description

FriendliAI is building the fastest inference cloud for agents, delivering low-latency, scalable GPU-accelerated AI workloads. We seek a Cloud Infrastructure Engineer to own cluster architecture, extend Kubernetes, and manage the network path for inference traffic.

The role demands hands-on experience with large multi-cluster deployments, designing topology, upgrades, and capacity across tenants. You will shape autoscaling and service mesh, collaborating across SRE, security, and engine teams to

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Software Engineer – Cloud Infrastructure
Software Engineer – Cloud Infrastructure

FriendliAI • San Francisco (CA)

On-site
USD 150,000 - 190,000
Flexible working hours
Lunch and dinner provided
Health check-up support with top-tier硬
+1
Staff Engineer, AI Cloud Infra (Kubernetes + GPUs)
Staff Engineer, AI Cloud Infra (Kubernetes + GPUs)

Lambda • San Francisco (CA)

Hybrid
USD 314,000 - 465,000
Health, dental, and vision coverage
401k with 2% company match
Wellness stipend
+1
AI Infra Engineer — GPU Cloud, Kubernetes/Slurm
AI Infra Engineer — GPU Cloud, Kubernetes/Slurm

Blue Signal Search • San Francisco (CA)

On-site
USD 180,000 - 240,000
Annual bonus
Equity participation
Comprehensive benefits
+1
AI Infrastructure Engineer — GPU Kubernetes for Production
AI Infrastructure Engineer — GPU Kubernetes for Production

vCluster • Germany (OH)

On-site
USD 150,000 - 200,000
Competitive Salary
Platinum-Level Insurance
Flexible Working Schedule
+1
AI Infra Engineer: Kubernetes on Bare Metal GPUs (Remote)
AI Infra Engineer: Kubernetes on Bare Metal GPUs (Remote)

vCluster • New York (NY)

Hybrid
USD 150,000 - 200,000
Competitive Salary
Platinum-Level Insurance
Flexible Working Schedule
+1
AI Infra & Cluster Engineer — Scale GPU/CPU Orchestration
AI Infra & Cluster Engineer — Scale GPU/CPU Orchestration

Linuxcareers • San Francisco (CA)

On-site
USD 120,000 - 160,000
Staff Engineer, AI Cloud Orchestration
Staff Engineer, AI Cloud Orchestration

Lambda • United States

Remote
USD 180,000 - 240,000
Lead AI Infrastructure Engineer: Kubernetes & GPU
Lead AI Infrastructure Engineer: Kubernetes & GPU

Seekr • Washington

Hybrid
USD 180,000 - 230,000
Equity ownership
Unlimited PTO
Hybrid work (Reston, VA & Austin, TX)
Lead AI Infrastructure Engineer: GPU Clusters & Reliability
Lead AI Infrastructure Engineer: GPU Clusters & Reliability

Luma AI • San Francisco (CA)

On-site
USD 300,000 - 420,000
Senior Cloud Kubernetes Engineer - GPU AI Infra
Senior Cloud Kubernetes Engineer - GPU AI Infra

NVIDIA AI • Seattle (WA)

On-site
USD 184,000 - 357,000
Equity
Benefits
Competitive salary