Get more replies from employers
Send a job-specific resume in minutes.
Inception is seeking engineers and scientists to design, optimize, and maintain compute foundations for large‑scale language model training and inference. You will develop high‑performance ML kernels, enable efficient low‑precision arithmetic, and improve the distributed compute stack powering training and serving of large models.
The role emphasizes CUDA/CuTe/Triton kernel design, memory bandwidth optimization, and scalable infrastructure, with collaboration across ML systems and tooling.
Inception is seeking engineers and scientists to design, optimize, and maintain compute foundations for large‑scale language model training and inference. You will develop high‑performance ML kernels, enable efficient low‑precision arithmetic, and improve the distributed compute stack powering training and serving of large models.
The role emphasizes CUDA/CuTe/Triton kernel design, memory bandwidth optimization, and scalable infrastructure, with collaboration across ML systems and tooling.