A complete application in a minute — tailored resume and cover letter, ready to send.
Aionia Group in Menlo Park, CA is hiring a Member of Technical Staff, ML Systems to accelerate model training and inference across image, video, and world-model workloads. You will work with a founding team on kernels, runtimes, and distributed engines that power production-scale ML stacks.
You’ll optimize GPU performance, profile bottlenecks with Nsight, and implement low-level CUDA and Triton improvements.
GPU & ML Systems · Below the Application Layer
About the Company
A deep-tech AI infrastructure startup rebuilding the training and inference stack for world models. Today's ML infrastructure was built for language models — this team is rebuilding it for video, image, and world-model workloads, co-designing across three layers at once: low-level GPU kernel optimization, distributed systems, and the algorithms and models themselves.
The company came out of stealth with public benchmarks already in hand: a leading open video-generation model running roughly 10x faster at half the cost on its stack, a 2K image-generation model running in about four seconds for three cents, and a real-time video model running faster than real time. The founding team — with prior experience across leading AI labs, hyperscalers, and infrastructure companies — works out of Menlo Park in person and has an API already in production. The company raised a $10M seed round and is approaching a Series A.
"You report to the CEO. He runs every screen himself and makes the hiring decision — there is no layer between the work and the person who decides. The work sits below the application layer: kernels, runtimes, and distributed engines for video and world models. Nothing here is agents or RAG."
The Opportunity
You’ll own speed and efficiency across the full ML systems stack — low-level kernels, distributed inference engines, and multi-node training and serving systems for image, video, and world-model workloads. You’ll work directly alongside a founding team that between them covers distributed systems, kernel optimization, cloud infrastructure, and research.
You’ll feel at home here if you’d rather make a video model ten times faster than train one.
What You’ll Do
Requirements
Baseline
CUDA Triton PyTorch Nsight Systems / Compute NCCL RDMA (InfiniBand / RoCE)
Interview Process
1
30-minute conversation with the CEO on background, motivation, and a first read on GPU/distributed systems depth.
2
60-minute technical round with a member of the founding team on kernels, inference, or distributed execution.
3
60-minute systems design session with a member of the founding team.
4
If not already done in person, a chance to meet the full team on-site in Menlo Park.