Research Engineer - Infrastructure, Inference

MBN Solutions

San Francisco (CA)

On-site

USD 150,000 - 210,000

Full time

29 hours ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Relocation support
Visa transfer assistance
On-site in San Francisco

Job summary

MBN Solutions is seeking a Research Engineer to build and optimize high-throughput, low-latency inference systems for evaluating large-scale physical-world models. You will push the limits of hardware-aware acceleration, profiling bottlenecks, and implementing fast, scalable solutions in a predominantly on-site SF setting.

You will work closely with researchers to develop novel architectures and distributed evaluation pipelines, ensuring reproducible results while scaling to petabyte-scale data

Qualifications

  • Experience building or optimizing inference and serving systems for throughput and latency.
  • Fluent in distributed compute and GPU parallelism, optimize for hardware rather than around it.
  • Deep knowledge of PyTorch or JAX to reason about what they're doing underneath, not just call them.
  • Ability to profile an unfamiliar system, identify bottlenecks, and ship fixes that survive integration.
  • Comfort working across research and engineering with minimal handoff.

Responsibilities

  • Throughput at evaluation scale for large-scale evaluation, backtesting, and scoring against historical observations.
  • Low-latency real-time inference design and implementation for practical use.
  • Real utilisation of FLOPs, bandwidth, and memory beyond reported figures.
  • Distributed orchestration with Kubernetes, Ray, or Slurm for large-batch evaluation sweeps.
  • Establish trustworthy evaluation with reliability, observability, and reproducibility.
  • Collaborate with researchers to develop novel architectures and speed up inference often before reference implementations exist.

Skills

Inference systems
Distributed compute
GPU parallelism
Profiling and debugging

Tools

TensorRT
vLLM
SGLang
PyTorch
JAX
Kubernetes
Ray
Slurm

Job description

Research Engineer - Inference | Foundation Models for the Physical World
San Francisco | On-site, five days a week | Relocation support available

Make inference so fast and cheap that evaluation never gates research.

A team of fewer than 15 people is building a new class of foundation model - one that learns cause and effect from physical systems rather than text. Models are trained from scratch, on novel architectures, across hundreds of GPUs, against petabyte-scale multimodal data drawn from one of the largest collections of real-world observational data available.

Every research decision they make depends on how quickly and cheaply those models can be evaluated. That's the problem you own.

Why this is a different problem

Most inference work today follows a well-mapped road: known architectures, known kernels, a decade of published tricks.

This isn't that. The architectures are new, the modalities are physical rather than textual, and the workload - scoring against historical observations, backtesting, large-batch evaluation sweeps - looks nothing like serving a chat endpoint. Much of what you build won't have a published answer to copy from.

If that reads as an opportunity rather than a risk, keep going.

What you'll own
  • Throughput at evaluation scale. High-throughput inference systems for large-scale evaluation, backtesting, and scoring against historical physical observations.
  • Latency for real-time inference. Designing and implementing the techniques that make it fast enough to be useful.
  • The hardware. Real utilisation of FLOPs, bandwidth, and memory - not just the numbers the framework reports.
  • Distributed orchestration. Extending frameworks like Kubernetes, Ray, or Slurm for distributed inference and large-batch evaluation sweeps.
  • Trustworthy evaluation. Standards for reliability, observability, and reproducibility, so a number from a sweep means the same thing twice.
  • Novel architectures. Working directly with researchers to make new architectures fast, often before there's a reference implementation to work from.
What we're looking for
  • You've built or optimised inference and serving systems for throughput and latency — TensorRT, vLLM, SGLang, or something you wrote yourself.
  • You're fluent in distributed compute and GPU parallelism, and you optimise against the hardware rather than around it.
  • You know PyTorch or JAX deeply enough to reason about what they're doing underneath, not just call them.
  • You can profile an unfamiliar system, find the real bottleneck, and ship a fix that survives contact with other people's code.
  • You're comfortable spanning research and engineering, and you don't need a clean handoff to make progress.

Useful, not essential: open-source contributions to inference or systems infrastructure (vLLM, SGLang, Triton). Distributed training with PyTorch or FSDP. Background in computer vision, robotics, sensor fusion, physics-informed ML, scientific AI, or multimodal learning.

We don't expect all of it. We do expect you to learn the rest quickly.

The team

This is not a large corporate research lab. It's a highly funded early-stage company, fewer than 15 people today, scaling rapidly over the next year.

Everyone writes code. Everyone contributes to research. There's no separate infrastructure org to hand things to and no one to translate the research for you - you'll be in the room where it's decided.

Five days a week in San Francisco. That's deliberate, and it isn't negotiable - the work is too tightly coupled for anything else at this stage. Relocation and visa transfer support is available.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Research Engineer, Infrastructure, Inference
Research Engineer, Infrastructure, Inference

Thinking Machines Lab Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 350,000 - 475,000
Health, dental, vision benefits
Unlimited PTO
Parental leave
+1
Member of Technical Staff - Inference Research
Member of Technical Staff - Inference Research

United States Digital Space LLC • New York (NY)

On-site
USD 150,000 - 210,000
Member of Technical Staff — Inference Infrastructure
Member of Technical Staff — Inference Infrastructure

Causal • San Francisco (CA)

On-site
USD 180,000 - 240,000
Member of Technical Staff — Inference Infrastructure
Member of Technical Staff — Inference Infrastructure

Kindredventures • San Francisco (CA)

On-site
USD 180,000 - 240,000
Member of Technical Staff — Inference Infrastructure
Member of Technical Staff — Inference Infrastructure

Causal Labs • San Francisco (CA)

On-site
USD 150,000 - 210,000
Senior Software Engineer - Model Performance
Senior Software Engineer - Model Performance

inference.net • San Francisco (CA)

Hybrid
USD 220,000 - 320,000
Equity in a high-growth startup
Comprehensive benefits
Inference Engineer
Inference Engineer

techire.® • San Francisco (CA)

On-site
USD 140,000 - 210,000
Medical insurance (including dental &视
Dental insurance
Vision insurance
+4
Founding Engineer, ML Inference
Founding Engineer, ML Inference

Reactor • San Francisco (CA)

On-site
USD 180,000 - 280,000
Competitive salary
Early equity
Health, dental, and vision coverage
+1
Member of Technical Staff - Inference Research
Member of Technical Staff - Inference Research

Mixpeek • New York (NY)

On-site
USD 180,000 - 260,000
ML Engineer, Inference Optimization
ML Engineer, Inference Optimization

Build AI • San Francisco (CA)

On-site
USD 180,000 - 240,000
Competitive pay
Medical coverage
Dental coverage
+8