Inference Runtime Engineer - On-Device & Cloud AI, Flexible WFH

EngRadar

New York (NY)

On-site

USD 150,000 - 230,000

Full time

14 days+
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Benefits offered by this job

Equity grants
Medical plan
Vision plan
Dental plan
Catered team lunch
Expensed dinners in the office
Flexible PTO
Flexible WFH

Job summary

EngRadar seeks a senior ML infrastructure engineer to develop and optimize the inference stack for both on-device and cloud AI, integrating new engines and multimodal models. You will work on latency and throughput improvements across CPU/GPU runtimes and contribute to open-source projects.

Required is extensive experience building production ML systems with strong Python and C++ skills, plus deep transformer knowledge and CPU/GPU profiling expertise.

Qualifications

  • Proven experience building production ML systems or performance-sensitive infrastructure.
  • Strong proficiency in Python and C++.
  • Deep knowledge of transformer architectures and profiling CPU/GPU workloads.

Responsibilities

  • Develop and optimize the inference stack for on-device and cloud AI.
  • Integrate new engines and multimodal models into the stack.
  • Improve latency and throughput across hardware runtimes and contribute to open-source projects.

Skills

Python
C++
Transformer Architectures
Model Inference
CPU Profiling
GPU Profiling
PyTorch
Llama.cpp
MLX
ExecuTorch
vLLM
SGLang
TensorRT-LLM
CUDA
Metal
Vulkan

Tools

PyTorch
Llama.cpp
CUDA
Vulkan

Job description

EngRadar seeks a senior ML infrastructure engineer to develop and optimize the inference stack for both on-device and cloud AI, integrating new engines and multimodal models. You will work on latency and throughput improvements across CPU/GPU runtimes and contribute to open-source projects.

Required is extensive experience building production ML systems with strong Python and C++ skills, plus deep transformer knowledge and CPU/GPU profiling expertise.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Inference Runtime Engineer — On-Device & Cloud ML
Inference Runtime Engineer — On-Device & Cloud ML

Lm-Studio • New York (NY)

Hybrid
USD 140,000 - 210,000
Competitive salary and equity grants
Excellent medical, vision, dental care
Catered team lunch / expensed dinners
+3
Software Engineer, Inference Runtime
Software Engineer, Inference Runtime

EngRadar • New York (NY)

On-site
USD 150,000 - 230,000
Equity grants
Medical plan
Vision plan
+5
Senior On-Device ML Engineer – Mobile Inference Expert
Senior On-Device ML Engineer – Mobile Inference Expert

Unity • California (MO)

On-site
USD 180,000 - 240,000
Inference Optimization Engineer: Fast, Cost-Effective ML
Inference Optimization Engineer: Fast, Cost-Effective ML

Build AI • San Francisco (CA)

On-site
USD 150,000 - 210,000
Competitive pay
Medical, dental, and vision packages
Housing subsidy $2k/month near SF offi
+6
Senior Inference Systems Engineer (GPU/On-Device)
Senior Inference Systems Engineer (GPU/On-Device)

Genesis AI • Northern (KY)

Hybrid
USD 150,000 - 210,000
Senior Inference Systems Engineer - Low-Latency & On-Device
Senior Inference Systems Engineer - Low-Latency & On-Device

Genesis AI • United States

On-site
USD 180,000 - 240,000
On-Device AI Inference Engineer — Ultra-Low Latency
On-Device AI Inference Engineer — Ultra-Low Latency

Hark • San Jose (CA), Northern (KY)

Hybrid
USD 200,000 - 450,000
Remote AI Inference Engineer: Edge & Performance
Remote AI Inference Engineer: Edge & Performance

ITRex Group • United States

Remote
USD 120,000 - 170,000
Remote flexibility
Competitive salary + medical benefits
Learning opportunities
Machine Learning Engineer- Inference Optimization | Experienced Hire
Machine Learning Engineer- Inference Optimization | Experienced Hire

Susquehanna International Group, LLP • Bala Cynwyd (PA)

On-site
USD 110,000 - 150,000
Senior On-Device ML Engineer for Real-Time Multimodal AI
Senior On-Device ML Engineer for Real-Time Multimodal AI

LE130 Unity Technologies SF • United States

Remote
USD 218,000 - 284,000
Health and life insurance
Commuter subsidy
Employee stock ownership
+8