A cutting-edge AI technology firm in California is seeking an entry-level programmer to design and implement high-performance compute kernels for AI primitives. Responsibilities include optimizing for throughput and memory hierarchy, collaborating with teams, and writing reusable code in C++ and CUDA. Ideal candidates have a Bachelor's or Master's in Computer Science and a strong background in parallel programming and optimization techniques.
Qualifications
Entry-level position in AI compute kernel design and implementation.
Strong background in optimization and parallel programming required.
Hands-on experience profiling and optimizing AI workloads is a plus.
Responsibilities
Design and implement high-performance compute kernels for AI primitives.
Optimize for throughput, latency, and memory hierarchy.
Collaborate with compiler and runtime teams to integrate kernels.
Skills
Parallel programming (CUDA, Triton, SYCL, OpenCL)
C++11 or higher
Performance analysis and parallel debugging
Memory layout and vectorization
Optimization of irregular algorithms
Education
Bachelor's or Master's in Computer Science or related field
Tools
Perfetto
VTune
Tracy
Valgrind
GNU Debugger
Job description
A cutting-edge AI technology firm in California is seeking an entry-level programmer to design and implement high-performance compute kernels for AI primitives. Responsibilities include optimizing for throughput and memory hierarchy, collaborating with teams, and writing reusable code in C++ and CUDA. Ideal candidates have a Bachelor's or Master's in Computer Science and a strong background in parallel programming and optimization techniques.