Lead GPU Fabric Observability Engineer

Baseten

New York (NY)

On-site

USD 200,000 - 380,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Competitive compensation and equity
Medical, dental, and vision coverage
Flexible PTO
Parental leave
401(k)
Learning opportunities

Job summary

Baseten is building its own GPU infrastructure for large-scale inference in the United States. As Lead Software Engineer, you will own the observability and root-cause analysis architecture for the GPU fabric, collecting signals from switches, hosts, and inference services to diagnose incidents in real time.

You will bridge networking and inference software, crafting topology models, service ownership maps, and proactive triage workflows to guide operators on remediation steps.

Qualifications

  • Staff-level or senior staff-level experience building production infrastructure software.
  • Strong distributed systems background, especially streaming systems, telemetry pipelines, diagnostics, or control-plane software.
  • Experience building systems that process high-volume, high-cardinality, noisy operational data.
  • Understanding of networking fundamentals and high-performance networks.
  • Ability to work with low-level infrastructure signals and build practical correlation, anomaly detection, or root-cause analysis systems.

Responsibilities

  • Own Baseten’s GPU fabric observability and root-cause analysis architecture.
  • Build telemetry pipelines across switches, NICs, hosts, GPUs, Kubernetes, and inference services.
  • Model topology, flow paths, service ownership, and failure domains.
  • Separate true fabric faults from host, NIC, GPU, kernel, driver, RDMA, scheduler, and workload failures.

Skills

Distributed systems
Telemetry pipelines
Networking fundamentals
Observability
Root-cause analysis

Tools

Kubernetes

Job description

Baseten is building its own GPU infrastructure for large-scale inference in the United States. As Lead Software Engineer, you will own the observability and root-cause analysis architecture for the GPU fabric, collecting signals from switches, hosts, and inference services to diagnose incidents in real time.

You will bridge networking and inference software, crafting topology models, service ownership maps, and proactive triage workflows to guide operators on remediation steps.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Lead GPU Fabric Observability Engineer
Lead GPU Fabric Observability Engineer

Baseten • San Francisco (CA)

On-site
USD 200,000 - 380,000
Competitive compensation
Meaningful equity
Medical, dental, and vision insurance
+4
Lead Software Engineer - GPU Fabric Telemetry & Diagnostics
Lead Software Engineer - GPU Fabric Telemetry & Diagnostics

BaseTen • New York (NY), San Francisco (CA)

On-site
USD 160,000 - 260,000
Competitive compensation
100% health insurance coverage
Flexible PTO including Winter Break
+4
Software Engineer - GPU Fabric Observability
Software Engineer - GPU Fabric Observability

BaseTen • New York (NY), San Francisco (CA)

On-site
USD 160,000 - 260,000
Competitive compensation
100% health insurance coverage
Flexible PTO including Winter Break
+4
GPU Networking Engineer — RDMA & Distributed Inference
GPU Networking Engineer — RDMA & Distributed Inference

Baseten • United States

Remote
USD 180,000 - 240,000
Software Engineer - GPU Fabric Observability
Software Engineer - GPU Fabric Observability

Baseten • San Francisco (CA)

On-site
USD 200,000 - 380,000
Competitive compensation
Meaningful equity
Medical, dental, and vision insurance
+4
Software Engineer - GPU Fabric Observability
Software Engineer - GPU Fabric Observability

Baseten • New York (NY)

On-site
USD 200,000 - 380,000
Competitive compensation and equity
Medical, dental, and vision coverage
Flexible PTO
+3
Senior GPU Fabric Architect for AI Cloud Infra
Senior GPU Fabric Architect for AI Cloud Infra

Bitdeer (NASDAQ: BTDR) • Austin (TX)

On-site
USD 180,000 - 240,000
Senior GPU Fabric Engineer for AI Cloud Infrastructure
Senior GPU Fabric Engineer for AI Cloud Infrastructure

Bitdeer • San Jose (CA)

On-site
USD 140,000 - 210,000
GPU Networking Engineer for Large-Scale AI Fabric
GPU Networking Engineer for Large-Scale AI Fabric

Thinking Machines Lab Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 350,000 - 475,000
Health, dental, and vision benefits
Unlimited PTO
Parental leave
+1
Senior GPU Compute Fleet Engineer | Remote-First
Senior GPU Compute Fleet Engineer | Remote-First

Boundless Networks • United States

Remote
USD 175,000 - 250,000
Equity
Health, dental, vision
Flexible PTO
+2