AI Training Infrastructure Engineer

United States Digital Space LLC

United States

Remote

USD 150,000 - 190,000

Full time

12 days ago
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

United States Digital Space LLC is seeking an AI Systems Engineer to scale infrastructure behind training and evaluation workflows. You will own projects from bottleneck identification to deployment, collaborating with researchers to improve reliability, throughput, and resource efficiency.

The role combines distributed systems engineering, performance optimization, and building self-service tools to empower researchers to launch experiments with minimal manual intervention.

Qualifications

  • Experience building or operating large-scale distributed systems.
  • Experience in ML infrastructure, inference systems, GPU performance, or infrastructure tooling.

Responsibilities

  • Build and operate infrastructure for large-scale training and evaluation, improving reliability, throughput, and resource efficiency.
  • Develop shared inference and grading platforms with automated capacity management, health monitoring, and visibility into performance.
  • Improve compute scheduling and resource allocation to reduce idle GPU time and help workloads recover quickly from failures.
  • Diagnose bottlenecks across training, inference, and orchestration, and work across teams to improve end-to-end performance.
  • Build self-service tools, automated validation, and observability that help researchers launch experiments, diagnose issues, and compare results with less manual intervention.

Skills

Distributed systems
Software engineering
ML infrastructure
GPU performance

Job description

United States Digital Space LLC is seeking an AI Systems Engineer to scale infrastructure behind training and evaluation workflows. You will own projects from bottleneck identification to deployment, collaborating with researchers to improve reliability, throughput, and resource efficiency.

The role combines distributed systems engineering, performance optimization, and building self-service tools to empower researchers to launch experiments with minimal manual intervention.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

AI Infrastructure Engineer: Scale Training & Systems
AI Infrastructure Engineer: Scale Training & Systems

Precision Labs • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 240,000
AI Systems Engineer - Scalable Training Infra
AI Systems Engineer - Scalable Training Infra

OpenAI • San Francisco (CA)

On-site
USD 160,000 - 210,000
AI Training Optimization Scientist: Scale Efficiently
AI Training Optimization Scientist: Scale Efficiently

United States Digital Space LLC • United States

Remote
USD 140,000 - 230,000
Senior AI Automation Engineer
Senior AI Automation Engineer

United States Digital Space LLC • United States

Hybrid
USD 120,000 - 170,000
AI Infrastructure Engineer — Scale ML Training & Inference
AI Infrastructure Engineer — Scale ML Training & Inference

Triwill Group • San Francisco (CA), Northern (KY)

Hybrid
USD 140,000 - 190,000
Staff AI Platform Engineer – Model Infrastructure
Staff AI Platform Engineer – Model Infrastructure

United States Digital Space LLC • United States

Remote
USD 193,000 - 290,000
Senior AI Infrastructure Engineer - Edge & API Systems
Senior AI Infrastructure Engineer - Edge & API Systems

United States Digital Space LLC • United States

Remote
USD 194,000 - 266,000
AI Infra Engineer – Internal Platform & Telemetry (Remote)
AI Infra Engineer – Internal Platform & Telemetry (Remote)

United States Digital Space LLC • United States

Remote
USD 153,000 - 296,000
Health insurance
Dental coverage
Vision coverage
+5
Senior AI Infrastructure Engineer — HPC & MLOps
Senior AI Infrastructure Engineer — HPC & MLOps

United States Digital Space LLC • Town of Charlotte (NY)

On-site
USD 128,000 - 182,000
AI Training Infrastructure Engineer - Scale & Performance
AI Training Infrastructure Engineer - Scale & Performance

SupportFinity™ • San Francisco (CA)

On-site
USD 175,000 - 220,000
Meaningful equity
Competitive salary
Comprehensive benefits package