Get more replies from employers
Send a job-specific resume in minutes.
MBN Solutions is seeking a Research Engineer to build and optimize high-throughput, low-latency inference systems for evaluating large-scale physical-world models. You will push the limits of hardware-aware acceleration, profiling bottlenecks, and implementing fast, scalable solutions in a predominantly on-site SF setting.
You will work closely with researchers to develop novel architectures and distributed evaluation pipelines, ensuring reproducible results while scaling to petabyte-scale data
Make inference so fast and cheap that evaluation never gates research.
A team of fewer than 15 people is building a new class of foundation model - one that learns cause and effect from physical systems rather than text. Models are trained from scratch, on novel architectures, across hundreds of GPUs, against petabyte-scale multimodal data drawn from one of the largest collections of real-world observational data available.
Every research decision they make depends on how quickly and cheaply those models can be evaluated. That's the problem you own.
Most inference work today follows a well-mapped road: known architectures, known kernels, a decade of published tricks.
This isn't that. The architectures are new, the modalities are physical rather than textual, and the workload - scoring against historical observations, backtesting, large-batch evaluation sweeps - looks nothing like serving a chat endpoint. Much of what you build won't have a published answer to copy from.
If that reads as an opportunity rather than a risk, keep going.
Useful, not essential: open-source contributions to inference or systems infrastructure (vLLM, SGLang, Triton). Distributed training with PyTorch or FSDP. Background in computer vision, robotics, sensor fusion, physics-informed ML, scientific AI, or multimodal learning.
We don't expect all of it. We do expect you to learn the rest quickly.
This is not a large corporate research lab. It's a highly funded early-stage company, fewer than 15 people today, scaling rapidly over the next year.
Everyone writes code. Everyone contributes to research. There's no separate infrastructure org to hand things to and no one to translate the research for you - you'll be in the room where it's decided.
Five days a week in San Francisco. That's deliberate, and it isn't negotiable - the work is too tightly coupled for anything else at this stage. Relocation and visa transfer support is available.