On-Device Transformers Tech Lead for Edge Inference

OpenAI

Los Angeles (CA)

Hybrid

USD 400,500 - 489,500

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

A leading AI research organization in San Francisco seeks an Inference Technical Lead to evaluate silicon platforms and work on model deployment for edge devices. You will collaborate with top machine learning researchers to push the boundaries of model capabilities. This role requires experience with GPUs and NPUs, understanding transformer models, and leading teams on performance-critical software. Competitive compensation package, including equity, is offered, along with a hybrid work model.

Qualifications

  • Experience evaluating or deploying workloads on GPUs, NPUs, or specialized accelerators.
  • Understand performance characteristics of transformer models including attention and memory bandwidth.
  • Design or optimize high-performance compute systems like inference engines and distributed runtimes.

Responsibilities

  • Evaluate and select silicon platforms for on-device deployment of models.
  • Work closely with research teams to co-design model architectures.
  • Analyze system performance tradeoffs between design and hardware capabilities.
  • Lead a team responsible for implementing the low-level inference stack.

Skills

Experience with GPUs and NPUs
Understanding of transformer model performance
Designing high-performance compute systems
Leading teams in performance-critical software

Job description

A leading AI research organization in San Francisco seeks an Inference Technical Lead to evaluate silicon platforms and work on model deployment for edge devices. You will collaborate with top machine learning researchers to push the boundaries of model capabilities. This role requires experience with GPUs and NPUs, understanding transformer models, and leading teams on performance-critical software. Competitive compensation package, including equity, is offered, along with a hybrid work model.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Edge Transformer Inference Tech Lead
Edge Transformer Inference Tech Lead

OpenAI • San Francisco (CA)

Hybrid
USD 130,000 - 180,000
Inference Technical Lead, On-Device Transformers
Inference Technical Lead, On-Device Transformers

OpenAI • Los Angeles (CA)

Hybrid
USD 445,000
Inference Technical Lead, On-Device Transformers
Inference Technical Lead, On-Device Transformers

OpenAI • San Francisco (CA)

On-site
USD 130,000 - 180,000
Inference Software Engineer: High-Performance Transformers
Inference Software Engineer: High-Performance Transformers

Etched.ai, Inc. • San Jose (CA)

On-site
USD 120,000 - 180,000
Medical, dental, and vision packages
Housing subsidy of $2k per month
Relocation support
+2
Senior AI Inference Engineer — Edge ML, Porting Models
Senior AI Inference Engineer — Edge ML, Porting Models

quadric.io, Inc • Burlingame (CA)

On-site
USD 120,000 - 150,000
Health Care Plan (Medical, Dental & Vision)
Retirement Plan (401k, IRA)
Life Insurance (Basic, Voluntary & AD&D)
+7
Remote AI Inference Engineer — Edge Model Deployment & Optimization
Remote AI Inference Engineer — Edge Model Deployment & Optimization

Quadric Inc. • Burlingame (CA)

On-site
USD 180,000 - 260,000
Health Care Plan (Medical, Dental & Vision)
Retirement Plan (401k, IRA)
Life Insurance (Basic, Voluntary & AD&D)
+7
Senior AI/ML Architect — Edge & On‑Device Inference
Senior AI/ML Architect — Edge & On‑Device Inference

Via Licensing Corporation • Atlanta (GA)

On-site
USD 152,000 - 209,000
Flexible work approach
Excellent compensation and benefits
Opportunity for recognition
Senior Staff Tech Lead — Inference & ML Performance
Senior Staff Tech Lead — Inference & ML Performance

fal • San Francisco (CA)

On-site
USD 150,000 - 200,000
Senior AI Inference Engineer — GPU & Edge Optimization
Senior AI Inference Engineer — GPU & Edge Optimization

Jobtailor • California (MO)

On-site
USD 140,000 - 210,000
Inference Systems Engineer for Transformers & Low-Latency HPC
Inference Systems Engineer for Transformers & Low-Latency HPC

Etched • San Jose (CA)

On-site
USD 180,000 - 240,000
Medical, dental, and vision packages
Housing subsidy of $2k per month
Relocation support for those moving to San Jose
+1