Software Engineer, Inference Runtime

EngRadar

New York (NY)

On-site

USD 150,000 - 230,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Equity grants
Medical plan
Vision plan
Dental plan
Catered team lunch
Expensed dinners in the office
Flexible PTO
Flexible WFH

Job summary

EngRadar seeks a senior ML infrastructure engineer to develop and optimize the inference stack for both on-device and cloud AI, integrating new engines and multimodal models. You will work on latency and throughput improvements across CPU/GPU runtimes and contribute to open-source projects.

Required is extensive experience building production ML systems with strong Python and C++ skills, plus deep transformer knowledge and CPU/GPU profiling expertise.

Qualifications

  • Proven experience building production ML systems or performance-sensitive infrastructure.
  • Strong proficiency in Python and C++.
  • Deep knowledge of transformer architectures and profiling CPU/GPU workloads.

Responsibilities

  • Develop and optimize the inference stack for on-device and cloud AI.
  • Integrate new engines and multimodal models into the stack.
  • Improve latency and throughput across hardware runtimes and contribute to open-source projects.

Skills

Python
C++
Transformer Architectures
Model Inference
CPU Profiling
GPU Profiling
PyTorch
Llama.cpp
MLX
ExecuTorch
vLLM
SGLang
TensorRT-LLM
CUDA
Metal
Vulkan

Tools

PyTorch
Llama.cpp
CUDA
Vulkan

Job description

The role involves developing and optimizing the inference stack for on-device and cloud AI, integrating new engines and multimodal models. Responsibilities include improving latency and throughput across various hardware runtimes and contributing to open-source projects.

Requirements: Candidates need significant experience building production ML systems or performance-sensitive infrastructure with strong proficiency in Python and C++. Deep knowledge of transformer architectures and experience profiling CPU/GPU workloads are essential.

Key Skills: Python, C++, Transformer Architectures, Model Inference, CPU Profiling, GPU Profiling, PyTorch, Llama.cpp, MLX, ExecuTorch, vLLM, SGLang, TensorRT-LLM, CUDA, Metal, Vulkan

Benefits: Competitive salary, Equity grants, Medical healthcare plan, Vision healthcare plan, Dental healthcare plan, Catered team lunch, Expensed dinners in the office, Flexible PTO, Flexible WFH

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Software Engineer, Inference Runtime
Software Engineer, Inference Runtime

Lm-Studio • New York (NY)

Hybrid
USD 140,000 - 210,000
Competitive salary and equity grants
Excellent medical, vision, dental care
Catered team lunch / expensed dinners
+3
Inference Runtime Engineer - On-Device & Cloud AI, Flexible WFH
Inference Runtime Engineer - On-Device & Cloud AI, Flexible WFH

EngRadar • New York (NY)

On-site
USD 150,000 - 230,000
Equity grants
Medical plan
Vision plan
+5
Senior Software Engineer - AI Inference
Senior Software Engineer - AI Inference

NVIDIA AI • Town of Santa Clara (NY)

On-site
USD 150,000 - 230,000
Equity
Health Insurance
Member of Technical Staff (AI Inference Engineer)
Member of Technical Staff (AI Inference Engineer)

Kindredventures • Palo Alto (CA)

On-site
USD 190,000 - 250,000
Comprehensive health insurance
Dental insurance
Vision insurance
+1
ML Inference Engineer San Francisco · Engineering · Full Time →
ML Inference Engineer San Francisco · Engineering · Full Time →

Reactor • San Francisco (CA)

On-site
USD 120,000 - 160,000
Visa sponsorship
Relocation support
Generous health, dental, and vision coverage
Software Engineer, Systems ML
Software Engineer, Systems ML

Meta • San Francisco (CA)

On-site
USD 180,000 - 240,000
Bonus
Equity
Machine Learning Engineer- Inference Optimization | Experienced Hire
Machine Learning Engineer- Inference Optimization | Experienced Hire

Susquehanna International Group, LLP • Bala Cynwyd (PA)

On-site
USD 110,000 - 150,000
Software Engineer, Inference
Software Engineer, Inference

Luma AI • San Francisco (CA)

On-site
USD 180,000 - 240,000
Member of Technical Staff, Inference
Member of Technical Staff, Inference

Reactor • San Francisco (CA)

On-site
USD 180,000 - 240,000
Competitive SF salary
Early equity
Visa sponsorship
+2
Inference Performance Engineer
Inference Performance Engineer

adaption • United States

Hybrid
USD 180,000 - 240,000
Flexible in-person collaboration in BA
Adaption Passport travel stipend
Lunch stipend
+1