Software Engineer - Managed Kubernetes

Lambda Labs

United States

Hybrid

USD 180,000 - 230,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Cash & equity compensation
Health, dental, and vision coverage
Wellness stipend
401k with company match
Flexible PTO

Job summary

Lambda Labs, The Superintelligence Cloud, is a leader in AI cloud infrastructure, seeking a senior software engineer to design, build, and maintain scalable Kubernetes control plane services and automation for cluster lifecycle management. You'll create internal tools and CLIs to deploy inference services and define SLOs/SLIs for the platform.

The role requires strong distributed systems and full-stack skills, with cloud-native experience and a history of contributing to open-source projects.

Qualifications

  • Bachelor's degree in Computer Engineering, Computer Science, Electrical Engineering, or related field.
  • 5 years of progressive experience as Software Engineer or related occupation.
  • At least 1 year in distributed systems software development for networking infra.
  • Full-stack development across Python, Go, Java, or C/C++.
  • Cloud-native tech experience with Docker, Kubernetes, Helm, Etcd, and gRPC.
  • Development and CI/CD with Gerrit, Git, Jenkins.
  • Experience with open-source projects and Kubernetes operators, CRDs, CSI, CNI.

Responsibilities

  • Design, build, and maintain scalable control plane services, operators, and custom controllers for Kubernetes.
  • Develop automation for cluster lifecycle management (provisioning, upgrades, patching, deletion).
  • Develop internal tools, APIs, and CLIs to deploy and monitor inference services.
  • Write resilient systems for large-scale distributed environments.
  • Define and implement SLOs and SLIs for Kubernetes services, workloads, platform.
  • Investigate cluster problems and document findings.

Skills

Distributed systems
Full-stack development
Languages: Python/Go/Java/C/C++
Cloud-native tooling
Open-source contributions
Kubernetes tuning

Education

Bachelor's degree in Computer Engineering/CS/EE

Tools

Docker
Kubernetes
Helm
Etcd
gRPC
Gerrit
Git
Jenkins

Job description

Lambda, The Superintelligence Cloud, is a leader in AI cloud infrastructure serving tens of thousands of customers. Our customers range from AI researchers to enterprises and hyperscalers. Lambda's mission is to make compute as ubiquitous as electricity and give everyone the power of superintelligence. One person, one GPU.

If you'd like to build the world's best AI cloud, join us.

Note: This position requires presence in our San Francisco office location 4 days per week; Lambda’s designated work from home day is currently Tuesday.

What You’ll Do
  • Design, build, and maintain scalable control plane services, operators, and custom controllers for Kubernetes.
  • Develop automation for cluster lifecycle management (provisioning, upgrades, patching, and deletion).
  • Develop internal tools, APIs, and command-line interfaces (CLIs) that enable customers and ML and AI teams to deploy and monitor inference services effectively.
  • Write resilient systems that gracefully handle failure across large-scale distributed environments.
  • Define and implement Service-Level Objectives (SLOs) and Service-Level Indicators (SLIs) for Kubernetes services, workloads, and the platform.
  • Drive into systems at a low level to solve unique cluster problems and write up the findings.
  • Assist customers with high-level Kubernetes questions and integrations with applications, storage, and authentication.
  • Assist with initial cluster buildouts and validation to help identify failed hardware before customer delivery.
  • Work closely with our HPC Ops and Datacenter Ops teams on issues that require lower-level expertise or cross-functional solutions.
  • Participate in a well-managed, sustainable on-call rotation.

Some telecommuting permitted (hybrid).

Qualifications
  • Have a Bachelor’s degree or foreign equivalent in Computer Engineering, Computer Science, Electrical Engineering, or related field
  • Have 5 years of progressive experience as Software Engineer or related occupation
  • Must have at least 1 year of prior work experience in each of the following:
    • Distributed systems software development as applied to computer networking infrastructure.
    • Full-stack development in multiple software languages including Python, Go, Java, or C/C++.
    • Software development using cloud-native technologies including Docker, Kubernetes, Helm, Etcd, and gRPC.
    • Development and CI/CD workflow using tools including Gerrit, Git, and Jenkins.
    • Working with and developing open-source projects.
    • Tuning Kubernetes configuration and working with Operators, CRDs, CSI, and CNI.
Benefits
  • We offer generous cash & equity compensation
  • Health, dental, and vision coverage for you and your dependents
  • Wellness and commuter stipends for select roles
  • 401k Plan with 2% company match (USA employees)
  • Flexible paid time off plan that we all actually use
Equal Opportunity Employer

Lambda is an Equal Opportunity employer. Applicants are considered without regard to race, color, religion, creed, national origin, age, sex, gender, marital status, sexual orientation and identity, genetic information, veteran status, citizenship, or any other factors prohibited by local, state, or federal law.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Software Engineer - Managed Kubernetes
Software Engineer - Managed Kubernetes

Neura Market • San Francisco (CA)

Hybrid
USD 180,000 - 240,000
Cash & equity compensation
Health, dental, and vision coverage
Wellness and commuter stipends
+2
Software Engineer - Managed Kubernetes
Software Engineer - Managed Kubernetes

Socket.dev • San Francisco (CA)

On-site
USD 180,000 - 240,000
Cash & equity compensation
Health, dental, and vision coverage
Wellness and commuter stipends
+2
Senior Site Reliability Engineer - Managed Kubernetes
Senior Site Reliability Engineer - Managed Kubernetes

AI Chopping Block • San Francisco (CA), Northern (KY)

Hybrid
USD 190,000 - 270,000
Health, dental, vision
401k with company match
Wellness stipends
+1
Senior Site Reliability Engineer - Managed Kubernetes
Senior Site Reliability Engineer - Managed Kubernetes

Lambda • San Francisco (CA)

Hybrid
USD 170,000 - 260,000
401k Plan with company match (USA)
Health, dental, and vision coverage
Wellness and commuter stipends
+1
Senior Platform Engineer - Core Infrastructure
Senior Platform Engineer - Core Infrastructure

Lambda • San Francisco (CA)

Hybrid
USD 180,000 - 240,000
Health, dental, vision coverage
401k with company match
Wellness and commuter stipends
+1
Senior Site Reliability Engineer - Managed Kubernetes
Senior Site Reliability Engineer - Managed Kubernetes

Neura Market • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 240,000
Equity compensation
Health, dental and vision coverage
Wellness and commuter stipends
+2
Senior Site Reliability Engineer - Managed Kubernetes
Senior Site Reliability Engineer - Managed Kubernetes

Lambda Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 240,000
Health insurance
Dental insurance
Vision insurance
+3
Senior Site Reliability Engineer - Managed Kubernetes
Senior Site Reliability Engineer - Managed Kubernetes

Socket.dev • San Francisco (CA)

Hybrid
USD 150,000 - 230,000
Health, dental, vision
4-day in-office work week
Wellness stipend
+1
Senior Platform Engineer – Core Infrastructure
Senior Platform Engineer – Core Infrastructure

Lambda • San Francisco (CA)

On-site
USD 180,000 - 240,000
Health, dental, and vision coverage
Wellness and commuter stipends
401k Plan with 2% company match
+1
Senior Platform Engineer - Core Infrastructure
Senior Platform Engineer - Core Infrastructure

Lambda Labs • United States

On-site
USD 180,000 - 260,000
Health insurance
401k plan
Flexible PTO
+2