Senior Infrastructure Engineer - GPU Compute

Boundless

Northern (KY)

Hybrid

USD 175,000 - 250,000

Full time

14 days+
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Benefits offered by this job

Equity allocation
Remote-first with off-sites
Health, dental, vision

Job summary

Boundless is coordinating GPU compute at scale as it becomes a leader in AI. As a Senior Infrastructure Engineer (GPU Compute), you'll build and operate the compute fabric powering our AI inference workloads across a large, heterogeneous GPU fleet including consumer RTX 5090 and datacenter hardware.

You will optimize scheduling, ensure always-on availability, drive down $/GPU-hour, and work with a remote-first, globally distributed team with a bias for action.

Qualifications

  • 5+ years of infrastructure/DevOps experience operating large-scale production systems
  • Deep expertise in Kubernetes, Docker, and container orchestration at scale
  • Strong Linux systems administration skills
  • Proficiency in infrastructure-as-code tools (Terraform, Ansible, Pulumi)
  • Track record of managing mission-critical, high-throughput systems
  • Strong IaC background in heterogeneous environments
  • Proficiency in at least one scripting/programming language (Python, Bash, TypeScript, Go)
  • Comfort navigating ambiguity with a bias for action

Responsibilities

  • GPU Fleet Orchestration: operate a heterogeneous, multi-region GPU fleet using SkyPilot, Kubernetes/k3s, and cloud+on-prem providers
  • Compute Scheduling & Utilization: maximize GPU utilization across inference workloads and manage placement across spot/on-prem/cloud
  • Bare-Metal & GPU Optimization: optimize PCIe, ReBAR, NUMA, CUDA/driver tuning, memory config, and network topology

Job description

Boundless is coordinating GPU compute at scale as it becomes a leader in AI. As a Senior Infrastructure Engineer (GPU Compute), you'll build and operate the compute fabric that powers our AI inference workloads — a large, heterogeneous, globally distributed GPU fleet spanning consumer cards (including RTX 5090) and datacenter hardware. Your job is to keep that fleet full, fast, cheap, and always on: orchestrating workloads across regions and providers, squeezing every bit of performance out of the hardware, and driving down cost per GPU-hour. This role rewards engineers who want to go deep on bare-metal and GPU optimization.

You should be comfortable operating with a high degree of autonomy, navigating ambiguity, and defaulting to a strong bias for action.

What You'll Do

GPU Fleet Orchestration: Operate a heterogeneous, multi-region GPU fleet (consumer + datacenter, including RTX 5090) using tools like SkyPilot, Kubernetes/k3s, and cloud + on-prem providers. Build the patterns that let us schedule inference workloads across the entire fleet reliably.

Compute Scheduling & Utilization: Maximize GPU utilization across inference workloads. Own workload placement across spot, on-prem, and cloud capacity, keeping the "always-on inference substrate" saturated and economical.

Bare-Metal & GPU Optimization: Go deep on GPU performance — PCIe P2P, ReBAR, NUMA topology (e.g. EPYC SP5), CUDA/driver tuning, memory configuration, and network topology — to push throughput per node.

Reliability, Access & Observability: Build secure fleet access (Tailscale, Teleport), robust observability and alerting, and zero-downtime rollouts across a distributed node fleet.

Cost Optimization: Drive down $/GPU-hr through spot instance management, intelligent workload placement between on-prem and cloud, and resource scheduling — without sacrificing reliability.

  • 5+ years of infrastructure/DevOps experience operating large-scale production systems
  • Deep expertise in Kubernetes, Docker, and container orchestration at scale
  • Strong Linux systems administration skills
  • Proficiency in infrastructure-as-code tools (Terraform, Ansible, Pulumi)
  • Track record of managing mission-critical, high-throughput systems
  • Strong infrastructure-as-code background in heterogeneous environments
  • Proficiency in at least one common scripting or programming language (Python, Bash, TypeScript, Go, etc.)
  • Comfort navigating ambiguity with a strong bias for action
Nice to Have
  • Experience with GPU computing infrastructure (CUDA, bare-metal optimization, kernel tuning)
  • Experience operating ML training or other large-scale distributed compute infrastructure
  • Experience with GPU fleet orchestration (SkyPilot, Ray, Slurm)
  • Familiarity with fleet access and networking tooling (Tailscale, Teleport)
  • Knowledge of network optimization and topology design
  • Experience with multi-region, globally distributed systems
  • Proficiency in Rust or low-level systems programming
  • Experience with on-premises data center operations
Additional Requirements
  • Candidates must include a public GitHub profile in their application.
  • The GitHub profile should demonstrate a minimum of 1 year of activity/history.
  • Applications that do not include a GitHub profile, or show insufficient activity, will not be considered.

At Boundless, we take care of our people, because building the future of AI compute starts with an empowered team. Here's what you can expect when you join us:

  • Competitive salary (proposed band b/t US$175k and $250k annually) + equity allocation
  • Health, dental, vision (for U.S. employees; region-adjusted globally)
  • Flexible PTO
  • Professional development and conference travel budget
  • Remote-first with regular off-sites and a high-trust, high-velocity team environment

We are a global team, and applicants from around the world are welcome to apply.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Applied AI/ML Engineer
Applied AI/ML Engineer

Boundless • Northern (KY)

Hybrid
USD 175,000 - 250,000
Competitive salary + equity
Health/dental/vision
Flexible PTO
+2
Senior GPU Compute Infrastructure Engineer (Remote)
Senior GPU Compute Infrastructure Engineer (Remote)

Boundless • Northern (KY)

Hybrid
USD 175,000 - 250,000
Equity allocation
Remote-first with off-sites
Health, dental, vision
Technical Business Developer
Technical Business Developer

Boundless • San Francisco (CA), Northern (KY)

Hybrid
USD 100,000 - 150,000
Equity
Health insurance
Remote-friendly
+1
Senior GPU Compute Fleet Engineer | Remote-First
Senior GPU Compute Fleet Engineer | Remote-First

Boundless Networks • United States

Remote
USD 175,000 - 250,000
Equity
Health, dental, vision
Flexible PTO
+2
Founding Product Manager
Founding Product Manager

Boundless • Northern (KY)

Hybrid
USD 100,000 - 150,000
Equity
Health, dental, vision
Flexible PTO
+2
Principal Software Engineer - Compute Infrastructure
Principal Software Engineer - Compute Infrastructure

Nvidia Corporation in • Santa Clara (CA)

Hybrid
USD 248,000 - 391,000
Equity
Benefits
Member of Technical Staff - Bare Metal & Fleet Provisioning at Prime Intellect
Member of Technical Staff - Bare Metal & Fleet Provisioning at Prime Intellect

Matcha • Northern (KY)

Hybrid
USD 150,000 - 300,000
Head of AI Data Center Infrastructure Platforms and Software
Head of AI Data Center Infrastructure Platforms and Software

Summit Group Solutions, LLC • United States

On-site
USD 150,000 - 350,000
Member of Technical Staff - Bare Metal & Fleet Provisioning
Member of Technical Staff - Bare Metal & Fleet Provisioning

Prime-Intellect • San Francisco (CA)

On-site
USD 150,000 - 300,000
Member of Technical Staff - Bare Metal & Fleet Provisioning
Member of Technical Staff - Bare Metal & Fleet Provisioning

AI Chopping Block • San Francisco (CA), Northern (KY)

Hybrid
USD 150,000 - 300,000