Senior Backend Engineer — Kubernetes & Slurm Infra (Go/Python)

Lightning AI

New York (NY)

Hybride

USD 180 000 - 250 000

Plein temps

14 jours+
Générateur de candidature

Une candidature conçue pour ce poste — un CV et une lettre de motivation personnalisés qui correspondent à l’offre.

Passez les filtres ATS

Avantages offerts par ce poste

Health Coverage
Equity
401(k) matching
Unlimited PTO
Winter Break
Parental Leave
Learning allowance
Wellness stipend
Sabbatical
Flexible Work
In-Office Meals

Résumé du poste

Lightning AI, the company behind PyTorch Lightning, seeks a Senior Backend Engineer for its Managed Services team. You will design and build backend services, automation, and control planes to provision and operate Kubernetes and Slurm clusters across our GPU fleet.

This role emphasizes distributed systems, cloud-native infra, and large-scale operations, solving challenging problems while improving reliability, security, and observability. Hybrid US-based with preferred hubs.

Qualifications

  • Significant professional experience designing, building, and operating production backend systems using Go or Python.
  • Deep hands-on experience with Kubernetes or Slurm, including operating large-scale production environments.
  • Strong understanding of distributed systems, cloud-native architectures, and production infrastructure.
  • Experience designing and building scalable backend services, APIs, and automation for infrastructure or platform operations.
  • Strong understanding of networking, storage, and cloud infrastructure fundamentals.
  • Familiarity with observability, CI/CD, testing, production operations, and incident response.
  • Ability to own complex technical projects while collaborating effectively across engineering teams.

Responsabilités

  • Design, build, and operate backend services in Go or Python that power Lightning AI's managed infrastructure platform.
  • Develop control plane services that provision, orchestrate, and manage Kubernetes and Slurm clusters across large-scale GPU infrastructure.
  • Build distributed systems that automate cluster lifecycle management, workload scheduling, infrastructure provisioning, and platform operations.
  • Develop platform capabilities using Kubernetes APIs, controllers, operators, and other cloud-native technologies.
  • Improve the reliability, scalability, security, and observability of our managed platform through automation and operational excellence.
  • Diagnose and resolve complex production issues across Kubernetes, distributed systems, networking, and cloud infrastructure.
  • Collaborate with infrastructure, AI, and platform engineering teams to shape the future of our cloud platform.
  • Contribute to technical design, architecture, mentoring, engineering best practices, and on-call operations.

Connaissances

Go
Python
Kubernetes
Slurm
Distributed systems
Cloud-native
Networking

Outils

Terraform
Crossplane
Helm
Argo CD
Flux
Kustomize
CI/CD tooling

Description du poste

Lightning AI, the company behind PyTorch Lightning, seeks a Senior Backend Engineer for its Managed Services team. You will design and build backend services, automation, and control planes to provision and operate Kubernetes and Slurm clusters across our GPU fleet.

This role emphasizes distributed systems, cloud-native infra, and large-scale operations, solving challenging problems while improving reliability, security, and observability. Hybrid US-based with preferred hubs.

Obtenez votre examen gratuit et confidentiel de votre CV.

ou faites glisser et déposez votre fichier ici.

Similar jobs

Postes similaires à comparer

Senior Backend Engineer - GPU Cloud Platform
Senior Backend Engineer - GPU Cloud Platform

Lightning AI • San Francisco (CA), Seattle (WA), New York (NY)

Hybride
USD 180 000 - 250 000
Comprehensive Health Coverage
Meaningful Equity
401(k) matching
+3
Senior Backend Engineer, Managed Services
Senior Backend Engineer, Managed Services

Lightning AI • New York (NY)

Sur place
USD 180 000 - 250 000
Health Coverage
Equity
401(k) matching
+8
Senior GPU Infra Lead: Slurm, Kubernetes & Platform
Senior GPU Infra Lead: Slurm, Kubernetes & Platform

Jobgether SRL • États-Unis

À distance
USD 170 000 - 250 000
Platform Engineer: DevOps for Scalable Cloud & Kubernetes
Platform Engineer: DevOps for Scalable Cloud & Kubernetes

Lightning Labs • Palo Alto (CA)

Sur place
USD 180 000 - 240 000
Senior Infrastructure Software Engineer — Remote/Hybrid
Senior Infrastructure Software Engineer — Remote/Hybrid

Lightning AI • New York (NY)

Sur place
USD 180 000 - 220 000
Health coverage
Equity (RSUs)
401(k) matching
+7
Senior Kubernetes & GPU Infra Engineer for AI-scale Compute
Senior Kubernetes & GPU Infra Engineer for AI-scale Compute

Kindredventures • États-Unis

Sur place
USD 140 000 - 190 000
Member of Technical Staff (AI Infrastructure Engineer)
Member of Technical Staff (AI Infrastructure Engineer)

Perplexity • Palo Alto (CA)

Sur place
USD 220 000 - 405 000
Member of Technical Staff (AI Infrastructure Engineer)
Member of Technical Staff (AI Infrastructure Engineer)

Perplexity • San Francisco (CA)

Sur place
USD 243 000 - 405 000
Senior Backend Engineer, Core Platform & Scale
Senior Backend Engineer, Core Platform & Scale

Lightningai • New York (NY), San Francisco (CA)

Hybride
USD 180 000 - 250 000
Comprehensive Health Coverage
Equity (RSUs)
401(k) matching
+9
Senior Cloud-Native AI Data Center Engineer Kubernetes/Slurm
Senior Cloud-Native AI Data Center Engineer Kubernetes/Slurm

NVIDIA • California (MO)

Sur place
USD 184 000 - 357 000
Equity
Benefits