Senior ML Engineer - Low-Latency Inference & Systems
Inworld
Germany (OH)
Hybrid
USD 120,000 - 160,000
Full time
14 days+
Get more replies from employers
Send a job-specific resume in minutes.
Start fresh or import an existing resume
Job summary
A leading AI technology firm in the United States is seeking an experienced engineer to optimize model performance. The role requires expertise in inference optimization, model acceleration, and proficiency in C++, CUDA, and Python, among other skills. You'll work collaboratively with teams to tackle complex problems, ensuring high-quality outcomes. Professional fluency in English is essential for daily collaboration. The company provides potential relocation support for candidates interested in the San Francisco Bay Area.
Qualifications
Strong understanding of modern serving frameworks and techniques.
Hands-on experience with quantization, distillation, and continuous batching.
Proficiency in C++, CUDA, Rust, or optimised Python.
Responsibilities
Optimize model serving and ensure reliability in production.
Collaborate with US-based teams on unclear problems.
Take ownership from research to production.
Skills
Inference Optimization
Model Acceleration
High-Performance Systems
Distributed Systems & Scaling
Public work
Full-cycle ownership
Background in CS, Physics, or Math
Professional fluency in English
Education
PhD in CS, Physics, Math, or equivalent
Job description
A leading AI technology firm in the United States is seeking an experienced engineer to optimize model performance. The role requires expertise in inference optimization, model acceleration, and proficiency in C++, CUDA, and Python, among other skills. You'll work collaboratively with teams to tackle complex problems, ensuring high-quality outcomes. Professional fluency in English is essential for daily collaboration. The company provides potential relocation support for candidates interested in the San Francisco Bay Area.