A leading AI research firm in San Francisco is seeking a Member of Technical Staff specialized in Model Efficiency. In this role, you will enhance LLM inference systems by tackling performance issues and collaborating with cross-functional teams. Ideal candidates have over 5 years of coding experience in C++ or Python and a solid understanding of the LLM inference environment. This position offers a remote-friendly work model, a competitive salary, and extensive benefits including a generous vacation policy.
Qualifications
5+ years of experience writing high-performance, production-quality code.
Strong programming skills in languages such as C++ or Python.
Ability to diagnose and resolve performance bottlenecks.
Responsibilities
Work across the inference stack to improve performance metrics.
Identify bottlenecks and develop optimizations.
Collaborate closely with modeling and systems teams.
Skills
High-performance coding
C++ programming
Python programming
LLM inference knowledge
Diagnosing performance bottlenecks
Fast shipping and iteration
Tools
CUDA programming
Distributed systems
Job description
A leading AI research firm in San Francisco is seeking a Member of Technical Staff specialized in Model Efficiency. In this role, you will enhance LLM inference systems by tackling performance issues and collaborating with cross-functional teams. Ideal candidates have over 5 years of coding experience in C++ or Python and a solid understanding of the LLM inference environment. This position offers a remote-friendly work model, a competitive salary, and extensive benefits including a generous vacation policy.