Daily lunch and dinner provided; unlimited snacks and beverages
Health check-up support and top-tier equipment/hardware support
Supportive and highly collaborative work environment
Startup equity and health insurance
Job summary
FriendliAI in San Francisco is looking for an experienced Inference Engine Engineer to optimize GPU kernel performance for AI workloads. The position requires a strong background in GPU programming and collaboration with cloud infrastructure teams. The ideal candidate should have 5+ years of experience and expertise in Python, C++. This role offers flexible hours and various benefits, including competitive compensation and health insurance.
Qualifications
5+ years of experience in production or high-impact research environments.
Production-level expertise in Python and C++.
Experience developing machine learning frameworks or performance-critical runtime systems.
Hands-on experience writing and optimizing GPU kernels.
Experience working with generative AI models such as transformer and diffusion models.
Responsibilities
Design and optimize custom GPU kernels for AI workloads.
Contribute to the development of the kernel compiler and other core components.
Collaborate with cloud and infrastructure engineers for end-to-end inference performance.
Analyze performance bottlenecks and implement optimizations.
Bachelor's or Master's degree in Computer Science, Computer Engineering, Electrical Engineering, or equivalent
Tools
Machine learning compilers
Profiling tools
Job description
FriendliAI in San Francisco is looking for an experienced Inference Engine Engineer to optimize GPU kernel performance for AI workloads. The position requires a strong background in GPU programming and collaboration with cloud infrastructure teams. The ideal candidate should have 5+ years of experience and expertise in Python, C++. This role offers flexible hours and various benefits, including competitive compensation and health insurance.