Senior Backend Engineer, GPU Cloud Infra & Kubernetes

Socket.dev

New York (NY)

Hybrid

USD 180,000 - 250,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Health insurance
Equity
401(k) matching
Unlimited PTO
Winter break
Parental leave
Learning allowance
Wellness stipend
Sabbatical

Job summary

Lightning AI in the United States is seeking a Senior Backend Engineer for its Managed Services team. You will design, build, and operate backend services powering Kubernetes and Slurm‑based infrastructure to support GPU clusters.

The role emphasizes distributed systems, cloud‑native platforms, and on‑call operations, with offices in SF, NYC, or Seattle and a hybrid schedule requiring at least two in‑office days per week.

Qualifications

  • Significant professional experience designing, building, and operating production backend systems using Go or Python.
  • Deep hands‑on experience with Kubernetes or Slurm, including operating large-scale production environments.
  • Strong understanding of distributed systems, cloud‑native architectures, and production infrastructure.
  • Experience designing and building scalable backend services, APIs, and automation for infrastructure or platform operations.
  • Strong understanding of networking, storage, and cloud infrastructure fundamentals.
  • Familiarity with observability, CI/CD, testing, production operations, and incident response.
  • Ability to own complex technical projects while collaborating effectively across engineering teams.

Responsibilities

  • Design, build, and operate backend services in Go or Python that power Lightning AI's managed infrastructure platform.
  • Develop control plane services that provision, orchestrate, and manage Kubernetes and Slurm clusters across large-scale GPU infrastructure.
  • Build distributed systems that automate cluster lifecycle management, workload scheduling, infrastructure provisioning, and platform operations.
  • Develop platform capabilities using Kubernetes APIs, controllers, operators, and other cloud-native technologies.
  • Improve the reliability, scalability, security, and observability of our managed platform through automation and operational excellence.
  • Diagnose and resolve complex production issues across Kubernetes, distributed systems, networking, and cloud infrastructure.
  • Collaborate with infrastructure, AI, and platform engineering teams to shape the future of our cloud platform.
  • Contribute to technical design, architecture, mentoring, engineering best practices, and on‑call operations.

Skills

Go/Python backend
Kubernetes
Distributed systems
APIs & automation
Networking basics
Observability/CI‑CD
Leadership/collaboration

Tools

Kubernetes
Slurm
Terraform
Crossplane
Argo CD
Helm
Kustomize

Job description

Lightning AI in the United States is seeking a Senior Backend Engineer for its Managed Services team. You will design, build, and operate backend services powering Kubernetes and Slurm‑based infrastructure to support GPU clusters.

The role emphasizes distributed systems, cloud‑native platforms, and on‑call operations, with offices in SF, NYC, or Seattle and a hybrid schedule requiring at least two in‑office days per week.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Backend Engineer: Kubernetes & Slurm Platform
Senior Backend Engineer: Kubernetes & Slurm Platform

Lightning-Ai • New York (NY)

Hybrid
USD 180,000 - 250,000
Health coverage
Equity / RSUs
401(k) matching
+8
Senior Backend Engineer — Kubernetes & Slurm GPU Infra
Senior Backend Engineer — Kubernetes & Slurm GPU Infra

Neura Market • New York (NY)

Hybrid
USD 180,000 - 250,000
Health coverage
Equity
401(k) matching
+2
Senior Backend Engineer — Kubernetes & Slurm Infra (Go/Python)
Senior Backend Engineer — Kubernetes & Slurm Infra (Go/Python)

Lightning AI • New York (NY)

Hybrid
USD 180,000 - 250,000
Health Coverage
Equity
401(k) matching
+8
Senior Software Engineer — Managed GPU Kubernetes
Senior Software Engineer — Managed GPU Kubernetes

Front Door Defense • San Jose (CA), Northern (KY)

Hybrid
USD 266,000 - 395,000
Remote Senior Fullstack Engineer - AI Cloud & GPU Infra
Remote Senior Fullstack Engineer - AI Cloud & GPU Infra

Framework Ventures • United States

Remote
USD 140,000 - 180,000
Senior Cloud Platform Engineer - GPU Infrastructure
Senior Cloud Platform Engineer - GPU Infrastructure

Socket.dev • San Francisco (CA)

On-site
USD 180,000 - 240,000
Health, dental, and vision coverage
Wellness stipend
Commuter stipend
+1
Senior Fullstack Engineer - AI Cloud & Distributed Systems
Senior Fullstack Engineer - AI Cloud & Distributed Systems

Framework Ventures • United States

Remote
USD 150,000 - 210,000
GPU & Compute Infra Engineer - Bare-Metal & AI
GPU & Compute Infra Engineer - Bare-Metal & AI

Lightning-Ai • New York (NY)

Hybrid
USD 180,000 - 200,000
Medical coverage (US)
Dental coverage (US)
Vision coverage (US)
+6
Senior Cloud-Native Engineer — Kubernetes & Slurm for Multi-Tenant GPUs
Senior Cloud-Native Engineer — Kubernetes & Slurm for Multi-Tenant GPUs

NVIDIA Gruppe • Santa Clara (CA)

On-site
USD 184,000 - 288,000
Equity options
Comprehensive benefits
Senior Cloud Kubernetes Engineer - GPU AI Infra
Senior Cloud Kubernetes Engineer - GPU AI Infra

NVIDIA AI • Seattle (WA)

On-site
USD 184,000 - 357,000
Equity
Benefits
Competitive salary