Staff AI Inference & Systems Engineer

Mixpeek

San Francisco (CA)

On-site

USD 180,000 - 240,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Modal is building an infrastructure layer for AI workloads, covering training, deployment, observation, and inference. You will perform hands-on inference research, selecting high-impact bets with the research lead and owning them end-to-end to reduce cost per token and tail latency on customer workloads.

You will work with customers alongside Forward Deployed Engineers to deploy and tune models, while expanding collaborations with external labs and turning frontier techniques into usable

Qualifications

  • Research-leaning or systems background in LLM inference with demonstrable work.
  • Fluency across LLM serving stack from kernels to autoscaling.
  • Track record of shipping research or systems used by others.
  • Ability to independently take a research idea to results in the open.
  • Willingness to work in NYC or San Francisco offices.

Responsibilities

  • Own end-to-end inference research bets and investigate high-impact ideas.
  • Train speculators against production traffic and learn from results.
  • Collaborate with Forward Deployed Engineers to deploy and tune models.
  • Collaborate with outside labs (e.g., ZLab, SGLang, Flash Attention 4 kernels).
  • Turn frontier serving techniques into products with engineering teams.
  • Help shape the research agenda and guide future work.

Skills

LLM inference
Research mindset
Shipping research

Tools

Quantization FP8/INT4
Autoscaling

Job description

Modal is building an infrastructure layer for AI workloads, covering training, deployment, observation, and inference. You will perform hands-on inference research, selecting high-impact bets with the research lead and owning them end-to-end to reduce cost per token and tail latency on customer workloads.

You will work with customers alongside Forward Deployed Engineers to deploy and tune models, while expanding collaborations with external labs and turning frontier techniques into usable

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Staff AI Inference Researcher
Staff AI Inference Researcher

AI Chopping Block • New York (NY)

On-site
USD 180,000 - 260,000
Staff Research Engineer, LLM Inference & Systems
Staff Research Engineer, LLM Inference & Systems

modal • New York (NY)

On-site
USD 180,000 - 240,000
Member of Technical Staff - Research, Inference
Member of Technical Staff - Research, Inference

Mixpeek • San Francisco (CA)

On-site
USD 180,000 - 240,000
Member of Technical Staff - Research, Inference
Member of Technical Staff - Research, Inference

modal • New York (NY)

On-site
USD 180,000 - 240,000
Member of Technical Staff - Inference Research
Member of Technical Staff - Inference Research

Mixpeek • New York (NY)

On-site
USD 180,000 - 260,000
Senior AI Infra Engineer — Real-Time Multimodal Inference
Senior AI Infra Engineer — Real-Time Multimodal Inference

Ambient AI, Inc. • Redwood City (CA)

Hybrid
USD 190,000 - 270,000
Stock options
Health, dental, vision
401(k)
+1
Inference Infra Architect for Multimodal ML Systems
Inference Infra Architect for Multimodal ML Systems

Elorian AI • San Francisco (CA)

On-site
USD 200,000 - 400,000
Health, dental, and vision benefits
Unlimited PTO
Parental leave
Member of Technical Staff - Inference Research
Member of Technical Staff - Inference Research

AI Chopping Block • New York (NY)

On-site
USD 180,000 - 260,000
Software Engineer, Inference - Multi Modal
Software Engineer, Inference - Multi Modal

OpenAI • San Francisco (CA)

On-site
USD 310,000 - 460,000
Realtime Multimodal Inference Architect
Realtime Multimodal Inference Architect

techire.® • San Francisco (CA)

On-site
USD 140,000 - 210,000
Medical insurance (including dental &视
Dental insurance
Vision insurance
+4