Inference Runtime Engineer - On-Device & Cloud AI, Flexible WFH

EngRadar

New York (NY)

On-site

USD 150,000 - 230,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Equity grants
Medical plan
Vision plan
Dental plan
Catered team lunch
Expensed dinners in the office
Flexible PTO
Flexible WFH

Job summary

EngRadar seeks a senior ML infrastructure engineer to develop and optimize the inference stack for both on-device and cloud AI, integrating new engines and multimodal models. You will work on latency and throughput improvements across CPU/GPU runtimes and contribute to open-source projects.

Required is extensive experience building production ML systems with strong Python and C++ skills, plus deep transformer knowledge and CPU/GPU profiling expertise.

Qualifications

  • Proven experience building production ML systems or performance-sensitive infrastructure.
  • Strong proficiency in Python and C++.
  • Deep knowledge of transformer architectures and profiling CPU/GPU workloads.

Responsibilities

  • Develop and optimize the inference stack for on-device and cloud AI.
  • Integrate new engines and multimodal models into the stack.
  • Improve latency and throughput across hardware runtimes and contribute to open-source projects.

Skills

Python
C++
Transformer Architectures
Model Inference
CPU Profiling
GPU Profiling
PyTorch
Llama.cpp
MLX
ExecuTorch
vLLM
SGLang
TensorRT-LLM
CUDA
Metal
Vulkan

Tools

PyTorch
Llama.cpp
CUDA
Vulkan

Job description

EngRadar seeks a senior ML infrastructure engineer to develop and optimize the inference stack for both on-device and cloud AI, integrating new engines and multimodal models. You will work on latency and throughput improvements across CPU/GPU runtimes and contribute to open-source projects.

Required is extensive experience building production ML systems with strong Python and C++ skills, plus deep transformer knowledge and CPU/GPU profiling expertise.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Inference Runtime Engineer — On-Device & Cloud ML
Inference Runtime Engineer — On-Device & Cloud ML

Lm-Studio • New York (NY)

Hybrid
USD 140,000 - 210,000
Competitive salary and equity grants
Excellent medical, vision, dental care
Catered team lunch / expensed dinners
+3
Software Engineer, Inference Runtime
Software Engineer, Inference Runtime

EngRadar • New York (NY)

On-site
USD 150,000 - 230,000
Equity grants
Medical plan
Vision plan
+5
Senior Staff Engineer Inference Runtime — Flexible Hours
Senior Staff Engineer Inference Runtime — Flexible Hours

jobr.pro • San Francisco (CA)

Hybrid
USD 405,000 - 485,000
Competitive compensation
Generous vacation
Flexible working hours
Embedded ML Inference & Optimization Engineer
Embedded ML Inference & Optimization Engineer

Applied Intuition Inc. • Sunnyvale (CA)

Hybrid
USD 159,000 - 200,000
Inference Systems Engineer — High-Performance ML Runtime
Inference Systems Engineer — High-Performance ML Runtime

The Consensus • San Jose (CA)

On-site
USD 180,000 - 240,000
Medical/dental/vision benefits
Housing subsidy
Relocation support
+2
Machine Learning Engineer- Inference Optimization | Experienced Hire
Machine Learning Engineer- Inference Optimization | Experienced Hire

Susquehanna International Group, LLP • Bala Cynwyd (PA)

On-site
USD 110,000 - 150,000
Remote ML Inference Infrastructure Engineer
Remote ML Inference Infrastructure Engineer

United States Digital Space LLC • United States

Remote
USD 150,000 - 230,000
Senior Real-Time Multimodal Inference Engineer
Senior Real-Time Multimodal Inference Engineer

Amazon • Sunnyvale (CA)

On-site
USD 193,000 - 262,000
Health insurance
401(k) matching
Paid time off
+1
Member of Technical Staff - ML Systems & Inference
Member of Technical Staff - ML Systems & Inference

Acceler8 Talent • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 280,000
Senior Machine Learning Engineer
Senior Machine Learning Engineer

Acceler8 Talent • San Francisco (CA), Northern (KY)

Hybrid
USD 170,000 - 240,000