An application made for this job — a tailored resume and cover letter that speak straight to the posting.
Get past ATS filters
Benefits offered by this job
Inclusive work culture
Weekly lunch stipend
Full health and dental benefits
Mental health budget
Parental leave top-up
Personal enrichment benefits
6 weeks of vacation
Job summary
A leading AI research firm in San Francisco is seeking a Member of Technical Staff specialized in Model Efficiency. In this role, you will enhance LLM inference systems by tackling performance issues and collaborating with cross-functional teams. Ideal candidates have over 5 years of coding experience in C++ or Python and a solid understanding of the LLM inference environment. This position offers a remote-friendly work model, a competitive salary, and extensive benefits including a generous vacation policy.
Qualifications
5+ years of experience writing high-performance, production-quality code.
Strong programming skills in languages such as C++ or Python.
Ability to diagnose and resolve performance bottlenecks.
Responsibilities
Work across the inference stack to improve performance metrics.
Identify bottlenecks and develop optimizations.
Collaborate closely with modeling and systems teams.
Skills
High-performance coding
C++ programming
Python programming
LLM inference knowledge
Diagnosing performance bottlenecks
Fast shipping and iteration
Tools
CUDA programming
Distributed systems
Job description
A leading AI research firm in San Francisco is seeking a Member of Technical Staff specialized in Model Efficiency. In this role, you will enhance LLM inference systems by tackling performance issues and collaborating with cross-functional teams. Ideal candidates have over 5 years of coding experience in C++ or Python and a solid understanding of the LLM inference environment. This position offers a remote-friendly work model, a competitive salary, and extensive benefits including a generous vacation policy.