New Grad Staff Engineer — Real-Time Model Optimization

Nuance Labs

Seattle (WA)

On-site

USD 200,000 - 300,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Health Savings Account with $2,000 annual contributions
15 days of PTO plus public holidays
Lunch, drinks, and snacks provided daily

Job summary

Nuance Labs in Seattle is seeking a Member of Technical Staff focused on model optimization and inference. This role demands expertise in refining AI models for real-time interactions, requiring a strong foundation in ML systems and familiarity with frameworks like vLLM and SGLang.

Ideal candidates will have a BS, MS, or PhD-related field and exhibit a passion for optimizing AI performance. Benefits include a competitive salary and a supportive, innovative work environment.

Qualifications

  • Excitement about real-time model optimization.
  • Strong fundamentals in ML systems.
  • Exposure to inference serving frameworks.

Responsibilities

  • Contribute to end-to-end inference optimization.
  • Implement KV cache strategies for long-context conversations.
  • Profile and benchmark latency and throughput.

Skills

Python
PyTorch
LLM inference
KV caching
CUDA

Education

BS, MS, or PhD in CS, ML, or related field

Tools

vLLM
SGLang
TensorRT-LLM

Job description

Nuance Labs in Seattle is seeking a Member of Technical Staff focused on model optimization and inference. This role demands expertise in refining AI models for real-time interactions, requiring a strong foundation in ML systems and familiarity with frameworks like vLLM and SGLang.

Ideal candidates will have a BS, MS, or PhD-related field and exhibit a passion for optimizing AI performance. Benefits include a competitive salary and a supportive, innovative work environment.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior MTS, RL & Post-Training for Real-Time AI
Senior MTS, RL & Post-Training for Real-Time AI

Nuance Labs • Seattle (WA)

On-site
USD 300,000 - 500,000
Health Savings Account plan
15 days PTO plus holidays
Lunch and snacks provided
Staff ML Engineer, Voice AI - Real-Time Inference Lead
Staff ML Engineer, Voice AI - Real-Time Inference Lead

Together AI • San Francisco (CA)

On-site
USD 220,000 - 280,000
Health insurance
Startup equity
Competitive benefits
Member of Technical Staff — Model Optimization and Inference (New Grad)
Member of Technical Staff — Model Optimization and Inference (New Grad)

Nuance Labs • Seattle (WA)

On-site
USD 200,000 - 300,000
Health Savings Account with $2,000 annual contributions
15 days of PTO plus public holidays
Lunch, drinks, and snacks provided daily
Staff ML Engineer: Real-Time Voice Inference Lead
Staff ML Engineer: Real-Time Voice Inference Lead

Togetherai • San Francisco (CA)

On-site
USD 220,000 - 280,000
Competitive compensation
Startup equity
Health insurance
+1
RL Researcher - Post-Training for Omni-Model AI
RL Researcher - Post-Training for Omni-Model AI

Nuance Labs • Seattle (WA)

On-site
USD 250,000 - 350,000
Health savings account contributions
15 days of PTO
Lunch, drinks, and snacks provided
Senior ML Engineer - Low-Latency Inference & Systems
Senior ML Engineer - Low-Latency Inference & Systems

Inworld • Germany (OH)

Hybrid
USD 120,000 - 160,000
Senior ML Engineer: AI Inference & Performance Optimizer
Senior ML Engineer: AI Inference & Performance Optimizer

Nebius • Palo Alto (CA)

Hybrid
USD 195,000 - 263,000
Health insurance
401(k) plan
Parental leave
+2
Senior ML Engineer - Real-Time Inference & Systems
Senior ML Engineer - Real-Time Inference & Systems

Inworld AI • Germany (OH)

On-site
USD 120,000 - 180,000
Staff Engineer - ML Inference & Model Efficiency
Staff Engineer - ML Inference & Model Efficiency

Cohere • San Francisco (CA)

Remote
USD 180,000 - 240,000
Inclusive work culture
Weekly lunch stipend
Full health and dental benefits
+4
Inference Performance Engineer: Optimize Model Serving
Inference Performance Engineer: Optimize Model Serving

Adaption • San Francisco (CA)

On-site
USD 180,000 - 240,000
Lunch stipend
Travel stipend (Adaption Passport)
Well-being benefits
+1