Get more replies from employers
Send a job-specific resume in minutes.
Dormont Manufacturing Co is seeking a Model Performance Engineer to enhance the efficiency of our model inference stack and improve our AI team's operational capabilities. You will focus on optimizing system performance for large-scale applications while troubleshooting issues in production environments.
The ideal candidate will possess deep expertise in LLM serving frameworks, strong skills in Python programming, and a solid understanding of GPU performance analysis. Our company offers competitive compensation, a dynamic work environment, and opportunities for professional growth.
We’re hiring a Model Performance Engineer to own the speed, cost, and reliability of our model inference stack, and to build the fine-tuning infrastructure that makes the rest of the AI team faster.
This is not a research role. You’ll be optimizing real systems serving millions of meetings — choosing between quantization trade-offs, debugging speculative decoding, or figuring out why one GPU family’s tail latency explodes at high concurrency while another stays stable.
You’ll own two things:
Hard Skills:
Strong signal:
Not required: