On-Site Infra Engineer: Kubernetes, GPUs & AI

Nebula

El Segundo (CA)

On-site

USD 100,000 - 250,000

Full time

5 days ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

Equity participation
Health, dental, vision
401(k) with matching
Nearby housing bonus
On-site work

Job summary

Nebula is seeking an Infrastructure Engineer to run and harden the company’s self-hosted Linux, Kubernetes, networking, storage, and security stack on hardware we own. You will help make the estate reproducible via version control, plan capacity, and manage GPU AI compute workloads on-site in El Segundo.

Experience with Linux, Kubernetes in production, backups, and security discipline is required. GPU infra, Ansible or Terraform, and self-hosted CI tooling are a plus.

Qualifications

  • Degree in computer engineering or computer science, or equivalent experience.
  • Hands-on Linux administration on owned machines: systemd, networking, storage and filesystems.
  • Kubernetes in production: workloads, storage, ingress, upgrades.
  • Fluent with routing, firewalls, DNS, VPNs, TLS and certificates.
  • Backups proven by restores; security discipline expected.
  • Ready to put a hand-built estate into code and diagnose outages.

Responsibilities

  • Bare-metal Linux hosts provisioning, systemd, disks/RAID, encryption at rest, out-of-band management.
  • Kubernetes cluster: scheduling, ingress and TLS, PVs, network policies, upgrades/rollbacks.
  • Network, access and identity: routing, firewalling, VPN, DNS, certs, mTLS, SSO.
  • Storage and backups: capacity planning, retention, offsite copies, restore drills.
  • Monitoring and incident response: metrics, logs, alerts, on-call, postmortems.
  • Security: hardening, patching, secrets, key rotation, least privilege.
  • Infrastructure as code: version control, self-hosted SCM, registry and CI runners.
  • Capacity/hardware planning: hosts, storage, network, GPU AI workloads.

Skills

Linux admin
Kubernetes prod
Networking
Storage mgmt
Security discipline
IaC tooling
GPU infra planning

Education

Bachelors in CS/EE or equivalent

Tools

Ansible
Terraform
Git
CI runners

Job description

Nebula is seeking an Infrastructure Engineer to run and harden the company’s self-hosted Linux, Kubernetes, networking, storage, and security stack on hardware we own. You will help make the estate reproducible via version control, plan capacity, and manage GPU AI compute workloads on-site in El Segundo.

Experience with Linux, Kubernetes in production, backups, and security discipline is required. GPU infra, Ansible or Terraform, and self-hosted CI tooling are a plus.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Kubernetes & GPU Infra Engineer for AI-scale Compute
Senior Kubernetes & GPU Infra Engineer for AI-scale Compute

Kindredventures • United States

On-site
USD 140,000 - 190,000
AI Infra Engineer – SRE (Kubernetes)
AI Infra Engineer – SRE (Kubernetes)

Berrybytes • United States

On-site
USD 110,000 - 150,000
Staff Engineer, AI Cloud Infra (Kubernetes + GPUs)
Staff Engineer, AI Cloud Infra (Kubernetes + GPUs)

Lambda • San Francisco (CA)

Hybrid
USD 314,000 - 465,000
Health, dental, and vision coverage
401k with 2% company match
Wellness stipend
+1
Infra Engineer - SRE(Kubernetes)
Infra Engineer - SRE(Kubernetes)

GMI Cloud • United States

On-site
USD 100,000 - 130,000
AI Infra Engineer: Kubernetes on Bare Metal GPUs (Remote)
AI Infra Engineer: Kubernetes on Bare Metal GPUs (Remote)

vCluster • New York (NY)

Hybrid
USD 150,000 - 200,000
Competitive Salary
Platinum-Level Insurance
Flexible Working Schedule
+1
Kubernetes & AI Infra Engineer — On-Prem & Cloud
Kubernetes & AI Infra Engineer — On-Prem & Cloud

Discernis • New York (NY)

On-site
USD 150,000 - 190,000
Kubernetes Platform Engineer - GPU & AI Infra
Kubernetes Platform Engineer - GPU & AI Infra

GTN Technical Staffing • Dallas (TX)

Hybrid
USD 165,000 - 210,000
Relocation available
Hybrid work model
On-site Linux Systems Ops Engineer – GPU & AI Infra
On-site Linux Systems Ops Engineer – GPU & AI Infra

Vast.ai • Los Angeles (CA)

On-site
USD 90,000 - 160,000
Health insurance
401(k) with company match
Meaningful equity
+2
Senior Cloud AI Infra Engineer: Go/C, Kubernetes, GPUs
Senior Cloud AI Infra Engineer: Go/C, Kubernetes, GPUs

Nvidia • Santa Clara (CA)

On-site
USD 240,000 - 340,000
Competitive salaries
Comprehensive benefits package
Equity
Senior Systems Software Engineer: AI Infra & Kubernetes
Senior Systems Software Engineer: AI Infra & Kubernetes

NVIDIA • Seattle (WA)

Hybrid
USD 184,000 - 357,000
Equity
Health benefits
Flexible work arrangement
+1