LLM Inference Runtime Architect for AI Accelerator

United States Digital Space LLC

United States

Remote

USD 180,000 - 250,000

Full time

6 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

United States Digital Space LLC is hiring a production-grade LLM inference runtime engineer to join the Hardware AI team. You will design and optimize the end-to-end runtime that maps frontier models onto our AI accelerator, balancing latency, throughput, and hardware utilization.

Work spans model architecture, systems software, and silicon interfaces, with emphasis on reliability, observability, and scalable performance in production environments.

Qualifications

  • Strong systems programming experience in C++, Rust, Python, or comparable performance-oriented environments.
  • Built or optimized runtimes, distributed systems, compilers, kernels, model-serving infrastructure.
  • Understand modern LLM inference including prefill and decode behavior, batching, KV-cache tradeoffs, and model parallelism.

Responsibilities

  • Design and implement the LLM inference runtime for frontier models running on custom silicon.
  • Build scheduling, continuous batching, memory management, KV-cache management, and execution orchestration for high-performance inference.
  • Develop distributed execution strategies across chips, hosts, and racks, including model partitioning, communication, and synchronization.
  • Optimize end-to-end latency, throughput, memory efficiency, and hardware utilization across diverse model architectures and serving workloads.
  • Partner with kernel, compiler, architecture, and silicon teams to co-design interfaces and remove performance bottlenecks across the stack.
  • Enable new model features, execution patterns, numerical formats, and hardware capabilities in a reliable production runtime.
  • Create profiling, observability, benchmarking, and performance-modeling tools that make runtime behavior measurable and actionable.
  • Debug complex correctness, performance, and reliability issues spanning model code, runtime software, communication layers, and hardware.
  • Turn workload insights into clear requirements for future generations of silicon and system architecture.

Skills

C++
Rust
Python
Distributed systems
Model-serving

Job description

United States Digital Space LLC is hiring a production-grade LLM inference runtime engineer to join the Hardware AI team. You will design and optimize the end-to-end runtime that maps frontier models onto our AI accelerator, balancing latency, throughput, and hardware utilization.

Work spans model architecture, systems software, and silicon interfaces, with emphasis on reliability, observability, and scalable performance in production environments.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Production-Grade LLM Inference Runtime Engineer
Production-Grade LLM Inference Runtime Engineer

AI Chopping Block • San Francisco (CA), Northern (KY)

Hybrid
USD 250,000 - 360,000
Senior LLM Inference Architect — Heterogeneous Hardware
Senior LLM Inference Architect — Heterogeneous Hardware

d-Matrix inc. • Santa Clara (CA)

On-site
USD 130,000 - 170,000
Competitive compensation
Equity
Inclusive work environment
LLM Inference Systems Engineer — Custom Accelerator
LLM Inference Systems Engineer — Custom Accelerator

Amazon Web Services (AWS) • Cupertino (CA)

On-site
USD 165,000 - 224,000
Health insurance
RSUs
401(k) matching
+2
Senior LLM Inference Engineer: Performance & Optimization
Senior LLM Inference Engineer: Performance & Optimization

Confidential • United States

On-site
USD 180,000 - 240,000
Senior LLM Inference Engineer - End-to-End, Equity
Senior LLM Inference Engineer - End-to-End, Equity

Entrada Ventures • Santa Clara (CA)

On-site
USD 180,000 - 280,000
Competitive compensation
Equity opportunities
Benefits in Santa Clara, CA
Inference Runtime Engineer for LLMs & Diffusion
Inference Runtime Engineer for LLMs & Diffusion

Inferact • San Francisco (CA)

Hybrid
USD 200,000 - 400,000
Health, dental, and vision benefits
401(k) company match
Visa sponsorship on case-by-case basis
ML Engineer—LLM Inference & GPU Optimization (Equity)
ML Engineer—LLM Inference & GPU Optimization (Equity)

IC Resources • San Francisco (CA)

On-site
USD 200,000 - 290,000
401(k)
Unlimited PTO
Modern engineering workspace
+1
LLM Inference Systems Performance Engineer
LLM Inference Systems Performance Engineer

3M HEALTHCARE • Austin (TX)

On-site
USD 150,000 - 210,000
Medical, dental, and vision coverage
Income protection benefits
Paid family leave
+1
Frontier LLM Inference Runtime Engineer
Frontier LLM Inference Runtime Engineer

Triwill Group • United States

On-site
USD 180,000 - 240,000
Edge-to-Cloud LLM Inference Engineer
Edge-to-Cloud LLM Inference Engineer

Insider, Inc. • United States

On-site
USD 100,000 - 140,000