Staff ML Engineer: Build Ultra-Fast AI at Scale (Relocation)
Inworld AI
Mountain View (CA)
On-site
USD 270,000 - 500,000
Full time
14 days+
Application generator
A complete application in a minute — tailored resume and cover letter, ready to send.
Get past ATS filters
Benefits offered by this job
Relocation assistance
Equity options
Comprehensive benefits package
Job summary
A tech-driven AI firm in Mountain View is seeking a high-performance systems engineer. This role demands expertise in inference optimization, model acceleration, and a deep understanding of performance tuning in C++, CUDA, and more. Candidates should possess a PhD or equivalent experience. Compensation ranges from $270,000 to $500,000 plus bonuses and equity. Relocation assistance is provided for successful candidates looking to innovate in a collaborative atmosphere.
Qualifications
Deep understanding of modern serving frameworks like vLLM or TRT-LLM.
Hands-on experience with quantization and caching strategies.
Proficiency in profiling code and optimizing performance for NVIDIA GPUs.
Responsibilities
Take models from research, containerize and optimize for production.
Communicate and collaborate closely with the team.
Design prototypes to explore unclear problems.
Skills
Inference Optimization
Model Acceleration
High-Performance Systems
Distributed Systems & Scaling
Public work
Full-cycle ownership
PhD in CS, Physics, Math or equivalent
Education
PhD in Computer Science or equivalent
Tools
C++
CUDA
Rust
Python
Kubernetes
Job description
A tech-driven AI firm in Mountain View is seeking a high-performance systems engineer. This role demands expertise in inference optimization, model acceleration, and a deep understanding of performance tuning in C++, CUDA, and more. Candidates should possess a PhD or equivalent experience. Compensation ranges from $270,000 to $500,000 plus bonuses and equity. Relocation assistance is provided for successful candidates looking to innovate in a collaborative atmosphere.