A complete application in a minute — tailored resume and cover letter, ready to send.
Luma seeks a seasoned performance engineer to accelerate multimodal models across GPU, CPU, and accelerators. You will write high‑performing PyTorch, Triton, and CUDA kernels and push hardware to the limit while preserving model quality.
You’ll own profiling, optimization, and deployment at scale, develop fused kernels, tensor-core aware code, and build monitoring tools for distributed training and inference.
Luma seeks a seasoned performance engineer to accelerate multimodal models across GPU, CPU, and accelerators. You will write high‑performing PyTorch, Triton, and CUDA kernels and push hardware to the limit while preserving model quality.
You’ll own profiling, optimization, and deployment at scale, develop fused kernels, tensor-core aware code, and build monitoring tools for distributed training and inference.