Rust ML Systems Engineer: Real-Time GPU Inference

Archetype AI Inc.

San Mateo (CA)

On-site

USD 180,000 - 280,000

Full time

3 days ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Archetype AI Inc. is building the world’s first physical AI platform to bring AI into the real world, with Newton as the core model.

You will own the serving path for Newton and related multimodal models, driving GPU utilization and low-latency inference from kernel to service. You will work on Rust-native inference nodes, CUDA kernels, and an end-to-end serving stack, collaborating with researchers to translate architecture changes into production serving work.

Qualifications

  • 6+ years of software engineering with ML systems, inference, or high-performance GPU computing.
  • Must have owned a production serving path end-to-end, not just benchmarking.
  • Expert in PyTorch with production-shipped models; anticipate slowdowns before profiling.
  • Strong CUDA or equivalent GPU depth: memory hierarchy, occupancy, Nsight profiling.
  • Rust or C++ alongside Python, Linux performance, production ops; daily Rust work expected.
  • Ability to translate architecture changes from researchers into serving work.

Responsibilities

  • Build and own model nodes in a Rust inference runtime: loading, warmup, batching, streaming, GPU memory pools.
  • Optimize kernels and the GPU path: custom CUDA kernels, mixed precision, quantization, parity against research.
  • Own inference routing and serving: streaming path API to GPU node, request batching, SLO latency & cost.
  • Productionize research checkpoints: export, compilation, quantization, parity evals, rollout.
  • Build observability for inference: latency histograms, GPU metrics, OOM signals, replayable traces.

Skills

6+ yrs software eng
End-to-end serving
PyTorch in production
CUDA depth
Rust or C++
Researcher collaboration

Tools

Python
Linux performance

Job description

Archetype AI Inc. is building the world’s first physical AI platform to bring AI into the real world, with Newton as the core model.

You will own the serving path for Newton and related multimodal models, driving GPU utilization and low-latency inference from kernel to service. You will work on Rust-native inference nodes, CUDA kernels, and an end-to-end serving stack, collaborating with researchers to translate architecture changes into production serving work.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Machine Learning Systems Engineer
Machine Learning Systems Engineer

Archetype AI Inc. • San Mateo (CA)

On-site
USD 180,000 - 280,000
AI Inference Engineer (GPU/Rust/CUDA)
AI Inference Engineer (GPU/Rust/CUDA)

Perplexity • New York (NY)

On-site
USD 220,000 - 485,000
AI Inference Engineer: GPU Performance & Rust Stack
AI Inference Engineer: GPU Performance & Rust Stack

Perplexity • San Francisco (CA)

On-site
USD 140,000 - 210,000
Staff GPU Inference Engineer — Real-Time AI Systems
Staff GPU Inference Engineer — Real-Time AI Systems

Cerebras • United States

Remote
USD 150,000 - 230,000
Performance Engineer, Inference Engine - High-Performance AI
Performance Engineer, Inference Engine - High-Performance AI

EngineersOfAI • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 240,000
Realtime ML Systems Engineer, Networking & AIOps
Realtime ML Systems Engineer, Networking & AIOps

NVIDIA Corporation • Santa Clara (CA), Northern (KY)

Hybrid
USD 152,000 - 288,000
ML Systems Engineer: Inference & GPU-Driven Distributed Workloads
ML Systems Engineer: Inference & GPU-Driven Distributed Workloads

Bake AI • San Mateo (CA), Northern (KY)

Hybrid
USD 180,000 - 240,000
Senior GPU ML Inference Engineer — Edge AI Platforms
Senior GPU ML Inference Engineer — Edge AI Platforms

NVIDIA • Westford (MA)

On-site
USD 224,000 - 432,000
Senior Rust Systems Engineer - Edge AI Inference Platform
Senior Rust Systems Engineer - Edge AI Inference Platform

Webhosting • Austin (TX)

Hybrid
USD 180,000 - 240,000
AI Inference Engineer — High-Performance GPU Systems
AI Inference Engineer — High-Performance GPU Systems

Perplexity • Palo Alto (CA)

On-site
USD 220,000 - 485,000