Machine Learning Systems Engineer

Archetype AI Inc.

San Mateo (CA)

On-site

USD 180,000 - 280,000

Full time

3 days ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Archetype AI Inc. is building the world’s first physical AI platform to bring AI into the real world, with Newton as the core model.

You will own the serving path for Newton and related multimodal models, driving GPU utilization and low-latency inference from kernel to service. You will work on Rust-native inference nodes, CUDA kernels, and an end-to-end serving stack, collaborating with researchers to translate architecture changes into production serving work.

Qualifications

  • 6+ years of software engineering with ML systems, inference, or high-performance GPU computing.
  • Must have owned a production serving path end-to-end, not just benchmarking.
  • Expert in PyTorch with production-shipped models; anticipate slowdowns before profiling.
  • Strong CUDA or equivalent GPU depth: memory hierarchy, occupancy, Nsight profiling.
  • Rust or C++ alongside Python, Linux performance, production ops; daily Rust work expected.
  • Ability to translate architecture changes from researchers into serving work.

Responsibilities

  • Build and own model nodes in a Rust inference runtime: loading, warmup, batching, streaming, GPU memory pools.
  • Optimize kernels and the GPU path: custom CUDA kernels, mixed precision, quantization, parity against research.
  • Own inference routing and serving: streaming path API to GPU node, request batching, SLO latency & cost.
  • Productionize research checkpoints: export, compilation, quantization, parity evals, rollout.
  • Build observability for inference: latency histograms, GPU metrics, OOM signals, replayable traces.

Skills

6+ yrs software eng
End-to-end serving
PyTorch in production
CUDA depth
Rust or C++
Researcher collaboration

Tools

Python
Linux performance

Job description

About Job

At Archetype AI, we’re building the world’s first physical AI platform to bring artificial intelligence into the real world. Our foundation model, Newton, understands the physical world through objective sensor data and generates real-time insights into complex physical behaviors, from industrial machinery and systems to wearable devices and smart environments.

Formed by a high-caliber team from Google and backed by one of Silicon Valley’s most renowned venture funds, Archetype AI is in a Series A phase and rapidly advancing its technology for the next big leap. This is a unique opportunity to join an exciting, fast-growing AI team based in the heart of Silicon Valley.

About the Role

You will own the serving path for Newton and related multimodal models. Much of our inference stack is Rust-native: model nodes in our agent runtime, built on Rust ML stacks (candle, Burn) with custom GPU kernels, plus the routing layer that streams real-time inference to GPU nodes. You will drive GPU utilization, numerical precision, and low-latency serving from the kernel up.

What You'll Own
  • Build and own model nodes in our Rust inference runtime: loading, warmup, batching, streaming, GPU memory pools.

  • Optimize kernels and the GPU path: custom CUDA kernels, mixed precision, quantization, parity against research.

  • Own inference routing and serving: streaming path API to GPU node, request batching, SLO-backed latency and cost.

  • Productionize research checkpoints: export, compilation, quantization, parity evals, and rollout.

  • Build the observability inference needs: latency histograms, GPU metrics, OOM signatures, replayable traces.

Key Qualifications
  • 6+ years software engineering, several of them in ML systems, inference, or high-performance GPU computing.

  • Has owned a production serving path end to end, not only benchmarked models.

  • Expert in PyTorch with models shipped to production; knows what will be slow before the profiler runs.

  • Strong CUDA or equivalent GPU depth: memory hierarchy, occupancy, Nsight or equivalent profiling.

  • Rust or C++ alongside Python, Linux performance, production ops; ready to work in Rust daily.

  • Works with researchers: can translate an architecture change into serving work.

Nice to Have
  • Rust ML stacks: candle, Burn, or comparable GPU compute in Rust.

  • Custom kernels and compiler stacks: Triton, CUTLASS, TorchInductor, TensorRT.

  • Quantization and mixed precision in production with a numerical-correctness suite.

  • Multimodal, video, embedding or time-series serving, not only decoder-only chat LLMs.

  • High-performance serving stacks (vLLM, SGLang, TensorRT-LLM): continuous batching, paged KV cache.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Product Engineer, Platform
Product Engineer, Platform

Archetype AI Inc. • San Mateo (CA)

On-site
USD 150,000 - 210,000
Rust ML Systems Engineer: Real-Time GPU Inference
Rust ML Systems Engineer: Real-Time GPU Inference

Archetype AI Inc. • San Mateo (CA)

On-site
USD 180,000 - 280,000
ML Architect
ML Architect

Blue Signal Search • United States

On-site
USD 180,000 - 280,000
Health insurance
Dental insurance
Life insurance
+1
AI Researcher
AI Researcher

Archetype AI • San Mateo (CA)

On-site
USD 140,000 - 210,000
Staff Machine Learning Engineer
Staff Machine Learning Engineer

Unity • Mountain View (CA)

On-site
USD 180,000 - 280,000
Staff Machine Learning Engineer
Staff Machine Learning Engineer

Unity • San Francisco (CA)

On-site
USD 180,000 - 240,000
Staff Machine Learning Engineer
Staff Machine Learning Engineer

Unity • California (MO)

On-site
USD 180,000 - 240,000
Member of Technical Staff - ML Performance
Member of Technical Staff - ML Performance

Veeda Innovation • Northern (KY)

Hybrid
USD 150,000 - 230,000
Senior AI Research Scientist — Sensor Data / Robotics
Senior AI Research Scientist — Sensor Data / Robotics

STRATOS Search • Palo Alto (CA)

On-site
USD 180,000 - 280,000
Machine Learning Systems Engineer
Machine Learning Systems Engineer

Strativ Group • Palo Alto (CA)

On-site
USD 450,000 - 550,000
Founding equity
Direct exposure to founders
Competitive compensation