Software Engineer – Cloud Infrastructure

FriendliAI

San Francisco (CA)

On-site

USD 150,000 - 190,000

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

Flexible working hours
Lunch and dinner provided
Health check-up support with top-tier硬
Top-tier equipment/hardware support

Job summary

FriendliAI is building the fastest inference cloud for agents, delivering low-latency, scalable GPU-accelerated AI workloads. We seek a Cloud Infrastructure Engineer to own cluster architecture, extend Kubernetes, and manage the network path for inference traffic.

The role demands hands-on experience with large multi-cluster deployments, designing topology, upgrades, and capacity across tenants. You will shape autoscaling and service mesh, collaborating across SRE, security, and engine teams to

Qualifications

  • Extensive experience designing and operating large-scale Kubernetes infra.
  • Strong programming skills in Go or Python for tooling.
  • Proven track record debugging distributed systems and networks.

Responsibilities

  • Own architecture of multi-cluster, multi-tenant Kubernetes fleet.
  • Extend Kubernetes with custom controllers, operators, and CRDs.
  • Design GPU scheduling, topology-aware placement, and quotas across tenants.
  • Build autoscaling for inference traffic, including queue-driven pod scaling and scale-to-zero.
  • Own the Kubernetes network data plane: CNI, IPAM, DNS, ingress, and L4/L7 load balancing.
  • Design cross-AZ/region connectivity and operate service mesh (mTLS, routing).
  • Define SLOs and lead post-incident hardening; deliver IaC with Terraform, Helm, GitOps.
  • Collaborate with inference, platform, SRE, and security teams to turn requirements into platform capabilities.

Skills

Kubernetes
Go or Python
Distributed systems debugging
Communication

Education

Bachelors or Masters degree in CS/CE/EE

Tools

AWS
Terraform
Helm
Ansible

Job description

About the job

FriendliAI is looking for a Cloud Infrastructure Engineer to own the architecture and evolution of the cluster platform behind our GPU-accelerated AI inference cloud. As a Software Engineer, Cloud Infrastructure, you will design how our clusters are built and connected, extend Kubernetes where its defaults fall short, and own the network path that inference traffic depends on.

Inference is an unforgiving workload for Kubernetes. Traffic is bursty and latency-sensitive, GPU capacity is scarce and inelastic, tenants must stay isolated, and multi-node serving depends on the network holding up under sustained load. This is a hands-on architecture role for an engineer who has already run large clusters in production and wants to push them further.

Key Responsibilities

Cluster Architecture

  • Own the architecture of our multi-cluster, multi-tenant Kubernetes fleet across both managed and self-managed clusters: cluster topology, control plane and etcd lifecycle, and zero-downtime upgrades.
  • Extend Kubernetes with custom controllers, operators, and CRDs so platform behavior is encoded in software rather than runbooks.
  • Design GPU scheduling and capacity strategy, including topology-aware placement, node pools, priority and preemption, and quota across tenants.
  • Build autoscaling that matches inference traffic: queue-driven pod scaling, node autoscaling, scale-to-zero, and cold-start reduction.

Networking

  • Own the Kubernetes network data plane: CNI, IPAM, DNS, ingress, and L4/L7 load balancing.
  • Design cross-AZ, cross-region, and cross-cluster connectivity, and operate the service mesh for routing, mTLS, and traffic policy.
  • Debug production network issues (packet loss, conntrack exhaustion, MTU mismatches, DNS latency, load balancer behavior) and drive permanent fixes.

Reliability & Collaboration

  • Define SLOs for platform-critical systems and lead post-incident hardening.
  • Deliver infrastructure as code with Terraform, Helm, and GitOps.
  • Partner with the inference engine, platform, SRE, and security teams to turn serving requirements into platform capabilities.
Qualifications
  • 5+ years designing, building, and operating large-scale Kubernetes infrastructure in production.
  • Bachelor’s or Master’s degree in Computer Science, Computer Engineering, Electrical Engineering, or equivalent.
  • Proven experience operating large-scale, high-traffic network services in production.
  • Deep understanding of Kubernetes internals: API server, scheduler, controller loops, kubelet, and etcd.
  • Strong command of Kubernetes and cloud networking: CNI, kube-proxy/eBPF datapaths, DNS, load balancing, service mesh, and VPC routing.
  • Proficiency with AWS, Terraform, Helm, and Ansible.
  • Programming skills in Go or Python, with the ability to build infrastructure tooling and automation.
  • Strong debugging skills across distributed systems, containers, and the Linux networking stack.
  • Clear written and verbal communication, including the ability to document architectural decisions for other engineers.
Preferred Experience
  • Large-scale Kubernetes operations in a high-traffic domain such as gaming, e-commerce, or public cloud.
  • Cilium and eBPF, including kube-proxy replacement or upstream contributions.
  • Cluster provisioning and lifecycle management with Kubespray or similar Ansible-based tooling.
  • GPU orchestration: NVIDIA GPU Operator, device plugins, or Dynamic Resource Allocation (DRA).
  • High-performance networking for distributed workloads: RDMA/RoCE, InfiniBand, EFA, SR-IOV, or NCCL tuning.
  • Multi-cloud, hybrid-cloud, or bare-metal Kubernetes operations.
  • Contributions to Kubernetes, Cilium, Istio, or other CNCF projects.
Benefits
  • Flexible working hours
  • Daily lunch and dinner provided; unlimited snacks and beverages
  • Supportive and highly collaborative work environment
  • Health check-up support and top-tier equipment/hardware support
  • A front-row seat to the generative AI infrastructure revolution
  • Competitive compensation, startup equity, health insurance, and other benefits.
About FriendliAI

FriendliAI is the fastest inference cloud for agents, built to run frontier open-weight models in production at scale. It delivers up to 7x faster output token speed, up to 90% lower inference costs, and 99.99% uptime across the most demanding agent workloads — long-context inference, real-time streaming, and accurate tool calling.

We are a small, fast-moving team doing work that matters at one of the most exciting moments in the history of technology. With our world-class inference stack, we are building the platform teams can actually rely on.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Solutions Architect - AI Inference Specialist
Solutions Architect - AI Inference Specialist

FriendliAI • San Francisco (CA)

On-site
USD 120,000 - 160,000
Competitive compensation and benefits package
Daily lunch and dinner
Unlimited snacks and beverages
+2
Director of Product
Director of Product

FriendliAI • San Francisco (CA)

On-site
USD 210,000 - 280,000
Flexible working hours
Daily meals provided
Unlimited snacks and beverages
+3
Lead Cloud Infrastructure Engineer - GPU AI Kubernetes
Lead Cloud Infrastructure Engineer - GPU AI Kubernetes

FriendliAI • San Francisco (CA)

On-site
USD 150,000 - 190,000
Flexible working hours
Lunch and dinner provided
Health check-up support with top-tier硬
+1
Director of Product Management
Director of Product Management

FriendliAI Inc. • San Francisco (CA)

On-site
USD 180,000 - 280,000
Flexible working hours
Lunch and dinner provided
Collaborative work environment
+3
Software Engineer - Full Stack
Software Engineer - Full Stack

FriendliAI • San Francisco (CA)

On-site
USD 120,000 - 160,000
Flexible working hours
Daily lunch and dinner provided
Health check-up support
+2
Software Engineer - AI Inference Engine
Software Engineer - AI Inference Engine

FriendliAI • San Francisco (CA)

On-site
USD 120,000 - 160,000
Flexible working hours
Daily lunch and dinner provided; unlimited snacks and beverages
Health check-up support and top-tier equipment/hardware support
+2
Founding Engineer, AI Infrastructure
Founding Engineer, AI Infrastructure

Piris Labs • San Francisco (CA)

On-site
USD 100,000 - 200,000
Equity
401(k)
Health insurance
+1
[Junior / Senior / Staff] Software Engineer, Inference / Compute Infrastructure Engineering
[Junior / Senior / Staff] Software Engineer, Inference / Compute Infrastructure Engineering

Together AI • San Francisco (CA)

On-site
USD 160,000 - 280,000
Equity
Health insurance
Competitive benefits
Staff Inference Engineer
Staff Inference Engineer

Designworks Talent LLC • Bellevue (KY)

Hybrid
USD 180,000 - 240,000
Health insurance
401(k) plan with company match
Paid holidays
Solutions Architect - AI Model Specialist
Solutions Architect - AI Model Specialist

FriendliAI • San Francisco (CA)

On-site
USD 110,000 - 150,000
Daily lunch and dinner provided
Unlimited snacks and beverages
Health check-up
+2