Netpreme is seeking a skilled professional to build highly optimized ML kernels for GPUs, FPGAs, and custom silicon. Responsibilities include designing and implementing compute and data movement kernels, targeting multiple hardware platforms, and collaborating with hardware architects and ML researchers. The ideal candidate should have solid experience in programming accelerators and knowledge of performance optimizations. The role offers equity awards, comprehensive insurance, 401(k) matching, and flexible PTO.
Qualifications
Solid experience in programming accelerators (GPUs, FPGAs).
Deep knowledge of at least one accelerator programming model (e.g. CUDA).
Interest in performance optimizations and enthusiasm about extracting maximum efficiency.
Knowledge of high-level synthesis (HLS) is a strong plus.
Responsibilities
Design and implement highly-optimized compute and data movement kernels for ML workloads.