AI Engineer - Model Performance

Dormont Manufacturing Co

San Francisco (CA)

On-site

USD 120,000 - 150,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Competitive compensation and benefits
Supportive environment for innovation and growth

Job summary

Dormont Manufacturing Co is seeking a Model Performance Engineer to enhance the efficiency of our model inference stack and improve our AI team's operational capabilities. You will focus on optimizing system performance for large-scale applications while troubleshooting issues in production environments.

The ideal candidate will possess deep expertise in LLM serving frameworks, strong skills in Python programming, and a solid understanding of GPU performance analysis. Our company offers competitive compensation, a dynamic work environment, and opportunities for professional growth.

Qualifications

  • Deep experience with LLM serving frameworks such as vLLM or SGLang.
  • Hands-on experience with quantization techniques and tuning.
  • Experience in production fine-tuning processes with frameworks like LoRA or QLoRA.
  • Strong proficiency in Python programming for developing infrastructure.
  • Ability to conduct GPU profiling and performance evaluations.

Responsibilities

  • Optimize model inference stack for speed and reliability.
  • Benchmark and evaluate GPU families for optimal performance.
  • Debug and troubleshoot production inference issues.
  • Construct fine-tuning pipelines for AI models.
  • Enhance GPU cost modeling and resource allocation strategies.

Skills

Deep experience with LLM serving frameworks
Hands‑on quantization experience
Production fine-tuning experience
Strong Python
Comfort with GPU profiling and performance analysis

Job description

ROLE OVERVIEW

We’re hiring a Model Performance Engineer to own the speed, cost, and reliability of our model inference stack, and to build the fine-tuning infrastructure that makes the rest of the AI team faster.

This is not a research role. You’ll be optimizing real systems serving millions of meetings — choosing between quantization trade-offs, debugging speculative decoding, or figuring out why one GPU family’s tail latency explodes at high concurrency while another stays stable.

You’ll own two things:

HOW YOU’LL HELP US WIN
  • Benchmark FP8 quantization across GPU families, find that FP8 KV cache causes catastrophic repetition loops, identify static quantization as 6% faster than dynamic on certain hardware, and ship a production config that gets 1.3x speedup with <1% quality degradation.
  • Evaluate serving frameworks (vLLM vs SGLang) with speculative decoding — discover that ngram speculation degrades ASR quality while EAGLE3 draft models don’t, and that torch.compile makes certain GPUs 7% slower.
  • Build a fine-tuning pipeline that takes a JSONL dataset and produces an optimized tune ready for serving, so a teammate can train a small classifier in an afternoon instead of a week.
  • Optimize GPU spend — know which GPU families are best for batch workloads (stable under high concurrency) vs latency-sensitive paths (40% faster, but tail latency blows up under load), and when a 30% cost premium isn’t worth it.
  • Debug production inference issues — trace a quality regression to a serving framework upgrade that changed the default attention backend, or find that audio format handling in the multimodal pipeline silently drops segments.
REQUIREMENTS

Hard Skills:

  • Deep experience with LLM serving frameworks (vLLM, SGLang, TensorRT-LLM, or similar) — not just deploying them, but tuning them: attention backends, scheduling strategies, CUDA graph warmup, prefix caching.
  • Hands‑on quantization experience — you’ve gone beyond “apply FP8 and hope.” You understand weight vs activation quantization, per-channel vs per-tensor scaling, and when dynamic quantization introduces more overhead than it saves.
  • Production fine-tuning experience — LoRA/QLoRA SFT, familiarity with training frameworks (ms‑swift, Axolotl, torchtune, or similar), understanding of data formatting, learning rate schedules, and how to diagnose training failures.
  • Strong Python. You’ll write serving infrastructure, benchmarking harnesses, and training pipelines — not notebooks.
  • Comfort with GPU profiling and performance analysis. You should be able to look at a benchmark result and know whether the bottleneck is compute, memory bandwidth, or scheduling overhead.

Strong signal:

  • Cost modeling for GPU infrastructure — you’ve had to choose between GPU types and justify the tradeoff.
  • Experience with multimodal models (audio/vision encoders + LLM decoders).
  • Experience with Modal, Ray Serve, or similar serverless GPU platforms.
  • Understanding of audio processing (codecs, chunking, sample rates).
  • Experience building internal tooling that other engineers use — this role succeeds when the rest of the team ships faster.

Not required:

  • ML research background or publications.
  • Prompt engineering expertise (we have a team for that).
  • Frontend or full-stack experience.
  • Masters/PhD (though it’s fine if you have one).
WHAT’S IN IT FOR YOU
  • The opportunity to shape the foundational software services of a growing company.
  • A role that balances innovation and incremental improvement.
  • A dynamic and collaborative engineering team.
  • Competitive compensation and benefits.
  • A supportive environment that encourages innovation and personal growth.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI Engineer - Model Performance
AI Engineer - Model Performance

Pantera Capital • San Francisco (CA)

Hybrid
USD 120,000 - 180,000
Competitive compensation and benefits
Supportive environment for personal growth
Dynamic and collaborative engineering team
Senior Software Engineer - Model Performance
Senior Software Engineer - Model Performance

inference.net • San Francisco (CA)

Hybrid
USD 220,000 - 320,000
Equity in a high-growth startup
Comprehensive benefits
Senior Software Engineer - Model Performance
Senior Software Engineer - Model Performance

Inference • San Francisco (CA)

On-site
USD 220,000 - 320,000
Competitive compensation
Equity in a high-growth startup
Comprehensive benefits
Software Engineer- Model Performance Systems
Software Engineer- Model Performance Systems

Baseten • San Francisco (CA)

On-site
USD 160,000 - 200,000
Competitive compensation
Equity
Medical/dental/vision insurance
+4
Member of Technical Staff, Performance Optimization
Member of Technical Staff, Performance Optimization

Fireworks AI • San Mateo (CA)

On-site
USD 175,000 - 220,000
Competitive compensation
Inclusive environment
Ownership & impact
Software Engineer, Model Performance Systems
Software Engineer, Model Performance Systems

Baseten • New York (NY)

On-site
USD 160,000 - 200,000
Competitive compensation with equity
100% medical, dental, and vision insurance coverage
Generous PTO including Winter Break
+3
Member of Technical Staff — Model Optimization and Inference (New Grad)
Member of Technical Staff — Model Optimization and Inference (New Grad)

Nuance Labs • Seattle (WA)

On-site
USD 200,000 - 300,000
Health Savings Account with $2,000 annual contributions
15 days of PTO plus public holidays
Lunch, drinks, and snacks provided daily
Senior LLM Inference Engineer — Performance & GPU Optimization
Senior LLM Inference Engineer — Performance & GPU Optimization

Confidential • United States

On-site
USD 180,000 - 240,000
AI Inference Performance Engineer - New College Grad 2026
AI Inference Performance Engineer - New College Grad 2026

NVIDIA Gruppe • Santa Clara (CA)

On-site
USD 120,000 - 160,000
Staff / Principal Machine Learning Engineer, Serving
Staff / Principal Machine Learning Engineer, Serving

Inworld AI • Mountain View (CA)

On-site
USD 270,000 - 500,000
Relocation assistance
Equity options
Comprehensive benefits package