Stand out for this role — generate a tailored resume and cover letter in about a minute.
Acceler8 Talent in the Bay Area is hiring a Senior Pretraining Engineer to lead large-scale video and multimodal pre-training efforts. You will shape architecture, training infrastructure, data and experiments to build systems for robotic intelligence in real-world environments.
This onsite role sits in Bay Area with our R&D team in San Jose. You should have experience with large-scale video or multimodal pre-training, distributed training across substantial GPU compute, and familiarity with
Senior Pretraining Engineer – Video & Multimodal AI
Location: Bay Area, CA | Onsite
Focus: Large-scale pre-training, video, multimodal models, diffusion, flow matching
I'm working with an early-stage robotics company building a deeply integrated AI and robotics stack from first principles.
Their long-term goal is ambitious: create highly autonomous factories capable of manufacturing physical goods with dramatically less human manual labour. That means solving problems across robot learning, perception, simulation, rendering, GPU performance and large-scale multimodal model training.
They’re now hiring a Senior Pretraining Engineer to take ownership of large-scale video and multimodal pre-training.
This is not a role for someone who has only fine-tuned existing models or operated clean, established training pipelines.
You’ll be expected to understand what happens when big models, big datasets and big compute collide, including the failure modes that only become visible once training runs become genuinely expensive.
You’ll work across model architecture, training infrastructure, data and experimentation to build large-scale vision and multimodal systems that can ultimately contribute to robotic intelligence in complex physical environments.
Strong candidates will have experience with:
Just as importantly, you should be able to talk openly about training runs that didn’t work: what failed, how you diagnosed it, what it cost, and what you changed afterward.
The wider engineering environment spans:
That creates an unusually broad technical surface area. Your models won’t exist purely to improve benchmark scores. The longer-term objective is intelligence that can operate through physical systems and contribute to genuinely autonomous manufacturing.
The company is onsite in the Bay Area, with its R&D operation to be based in San Jose.
If you’ve personally taken large multimodal or video models through expensive pre-training runs, including the painful ones, this is one of the more unusual opportunities to apply that experience to physical-world AI.