On-Device Transformers Tech Lead for Edge Inference
OpenAI
Los Angeles (CA)
Hybrid
USD 400,500 - 489,500
Full time
14 days+
Get more replies from employers
Send a job-specific resume in minutes.
Start fresh or import an existing resume
Job summary
A leading AI research organization in San Francisco seeks an Inference Technical Lead to evaluate silicon platforms and work on model deployment for edge devices. You will collaborate with top machine learning researchers to push the boundaries of model capabilities. This role requires experience with GPUs and NPUs, understanding transformer models, and leading teams on performance-critical software. Competitive compensation package, including equity, is offered, along with a hybrid work model.
Qualifications
Experience evaluating or deploying workloads on GPUs, NPUs, or specialized accelerators.
Understand performance characteristics of transformer models including attention and memory bandwidth.
Design or optimize high-performance compute systems like inference engines and distributed runtimes.
Responsibilities
Evaluate and select silicon platforms for on-device deployment of models.
Work closely with research teams to co-design model architectures.
Analyze system performance tradeoffs between design and hardware capabilities.
Lead a team responsible for implementing the low-level inference stack.
Skills
Experience with GPUs and NPUs
Understanding of transformer model performance
Designing high-performance compute systems
Leading teams in performance-critical software
Job description
A leading AI research organization in San Francisco seeks an Inference Technical Lead to evaluate silicon platforms and work on model deployment for edge devices. You will collaborate with top machine learning researchers to push the boundaries of model capabilities. This role requires experience with GPUs and NPUs, understanding transformer models, and leading teams on performance-critical software. Competitive compensation package, including equity, is offered, along with a hybrid work model.