An application made for this job — a tailored resume and cover letter that speak straight to the posting.
Brahma is conducting this search for a client and is building a team focused on accelerating generative video and image workloads. This early engineering hire reports to the CEO and joins a founding group with deep expertise in distributed systems and ML research.
You’ll optimize GPU performance, implement CUDA/Triton kernels, and build scalable inference and training engines across multiple GPUs. Ideal candidates bring CUDA, Triton, and GPU profiling experience and a passion for deep technical
Brahma is conducting this search on behalf of a client.
We're a small, well-funded founding team rebuilding the training and inference stack for generative video and image models. Today's stack was built for language models. We're co-designing across GPU kernels, distributed systems, and the models themselves to make these workloads dramatically faster and cheaper. Our inference engine is live with customers in generative media and robotics.
This is an early engineering hire. You'll report directly to the CEO and work alongside a founding team with deep expertise in distributed systems, kernel optimization, cloud infrastructure, and ML research.