We’re looking for an exceptional Senior Principal Software Engineer to lead the optimization of Large Language Model inference pipelines across data centers, edge, and embedded platforms.
Responsibilities
- Optimize and deploy high‑performance LLM inference pipelines.
- Own inference runtimes across data center, edge, and embedded platforms.
- Push model performance through quantization, kernel fusion, and cache optimization.
- Drive latency and throughput improvements that directly impact production products.
- Enable efficient, reliable deployment without external vendor dependency.
- Build deep expertise and ownership of vLLM, TensorRT‑LLM, llama.cpp, QAIRT.
- Extend and tune inference engines using custom CUDA kernels.
- Adapt runtimes for constrained and embedded deployment environments.
- Implement and evaluate quantisation strategies: INT8, INT4, FP4, FP8, mixed precision, AWQ, GPTQ.
- Balance accuracy, latency, memory footprint, and throughput.
- Optimize key‑value cache performance through paging, prefix caching, and cache‑aware memory layout design.
- Reduce memory pressure while sustaining high throughput.
- Design and tune batching strategies, continuous batching, speculative decoding.
- Optimize latency and tokens per second under real production traffic patterns.
- Deploy models efficiently on edge and embedded devices.
- Minimize end‑to‑end latency and ensure it is predictable.
- Reduce inference cost per request.
Qualifications
- Proven experience optimizing ML inference performance in production.
- Deep understanding of GPU architecture and memory hierarchies.
- Hands‑on experience with CUDA and low‑level performance tuning.
- Experience deploying models beyond research environments.
- Proficiency with inference engines: vLLM, TensorRT‑LLM, llama.cpp, QAIRT.
- CUDA kernel development and profiling.
- Quantisation techniques: INT8/INT4/FP4/FP8, AWQ, GPTQ.
- KV cache optimisation and memory layout design.
- Latency optimisation: batching, speculative decoding, continuous batching.
- Ability to solve deployment challenges on edge or embedded targets, achieve competitive tokens per second, and reduce inference latency.
Benefits
- Salary range $185,000 - $280,000 USD (determined by experience).
- Annual bonus opportunity.
- Insurance coverage (medical, dental, vision, life, and disability).
- Paid time off.
- Paid holidays.
- Company contribution to the RRSP.
- Equity awards for certain positions and levels.
- Remote and/or hybrid work available depending on the position.
Equal Opportunity Employer
Cerence is firmly committed to Equal Employment Opportunity (EEO) and to compliance with all federal, state and local laws that prohibit employment discrimination on the basis of age, race, color, gender, gender identity, gender expression, sex, sex stereotyping, pregnancy, national origin, ancestry, religion, physical or mental disability, medical condition, marital status, citizenship status, sexual orientation, protected military or veteran status, genetic information and other protected classifications. Cerence Equal Employment Opportunity Policy Statement.
All prospective and current Employees need to remain vigilant when it comes to executing security policies in the workplace. This includes:
- Following workplace security protocols and training programs to familiarize with the ways to maintain a safe workplace.
- Following security procedures to report any suspicious activity.
- Having respect for corporate security procedures to allow those procedures to be effective.
- Adhering to company's compliance and regulations.
- Encouraging to follow a zero tolerance for workplace violence.
- Basic knowledge of information security and data privacy requirements (e.g., how to protect data & how to be handling this data).
- Demonstrative knowledge of information security through internal training programs.