Stand out for this role — generate a tailored resume and cover letter in about a minute.
Acceler8 Talent is seeking a Senior Pretraining Engineer to own large-scale video and multimodal pre-training. You will train models across distributed GPU infrastructure and push research into working systems for autonomous robotics applications.
The role focuses on video, vision, diffusion, flow matching and other generative-model approaches, requiring strong software engineering and training-systems fundamentals. Onsite in Bay Area.
Senior Pretraining Engineer – Video & Multimodal AI
Location:Bay Area, CA | Onsite
Focus:Large-scale pre-training, video, multimodal models, diffusion, flow matching
I'm working with an early-stage robotics company building a deeply integrated AI and robotics stack from first principles.
Their long-term goal is ambitious: create highly autonomous factories capable of manufacturing physical goods with dramatically less human manual labour. That means solving problems across robot learning, perception, simulation, rendering, GPU performance and large-scale multimodal model training.
They’re now hiring aSenior Pretraining Engineerto take ownership of large-scale video and multimodal pre-training.
This is not a role for someone who has only fine-tuned existing models or operated clean, established training pipelines.
You’ll be expected to understand what happens whenbig models, big datasets and big compute collide, including the failure modes that only become visible once training runs become genuinely expensive.
You’ll work across model architecture, training infrastructure, data and experimentation to build large-scale vision and multimodal systems that can ultimately contribute to robotic intelligence in complex physical environments.
Strong candidates will have experience with:
Just as importantly, you should be able to talk openly about training runs thatdidn’t work: what failed, how you diagnosed it, what it cost, and what you changed afterward.
The wider engineering environment spans:
That creates an unusually broad technical surface area. Your models won’t exist purely to improve benchmark scores. The longer-term objective is intelligence that can operate through physical systems and contribute to genuinely autonomous manufacturing.
The company is onsite in the Bay Area, with its R&D operation to be based in San Jose.
If you’ve personally taken large multimodal or video models through expensive pre‑training runs, including the painful ones, this is one of the more unusual opportunities to apply that experience to physical‑world AI.