Low-Latency AI Inference Engineer

OP Recruiting

Chicago (IL)

On-site

USD 150,000 - 210,000

Full time

27 hours ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Benefits offered by this job

Health benefits

Job summary

OP Recruiting in New York, NY is seeking an AI Research Engineer to advance deep learning inference with a focus on ultra-low latency and scalable deployment. You will bridge ML research and hardware acceleration, building low-latency inference systems that handle continuous global data streams to inform core business decisions.

The role emphasizes optimizing kernels, collaborating with researchers on neural architectures, and accelerating inference across diverse hardware platforms for

Qualifications

  • Hands-on experience engineering production-grade deep learning systems in complex domains.
  • Strong lower-level engineering foundation with custom compute kernels.
  • Proficiency with framework compilation internals and low-level runtimes such as PyTorch/JAX/XLA/CUDA Graphs.

Responsibilities

  • Drive performance optimizations for large-scale model execution, including custom kernel creation and data streaming pipelines.
  • Collaborate with researchers to co-design neural architectures for ultra-real-time execution.
  • Develop low-level primitives and custom ops to maximize compute throughput on modern hardware.
  • Evaluate and deploy hardware acceleration technologies, from accelerators to specialized silicon.
  • Identify and address bottlenecks in ultra-low-latency deep learning inference.

Skills

Deep learning systems
CUDA kernels
PyTorch internals
Hardware acceleration

Tools

CUDA
JAX
XLA
CuTe DSLs

Job description

OP Recruiting in New York, NY is seeking an AI Research Engineer to advance deep learning inference with a focus on ultra-low latency and scalable deployment. You will bridge ML research and hardware acceleration, building low-latency inference systems that handle continuous global data streams to inform core business decisions.

The role emphasizes optimizing kernels, collaborating with researchers on neural architectures, and accelerating inference across diverse hardware platforms for

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Low-Latency ML Inference Engineer
Low-Latency ML Inference Engineer

Career Techniques • New York (NY)

Hybrid
USD 200,000 - 300,000
AI Research Engineer, Inference
AI Research Engineer, Inference

OP Recruiting • Chicago (IL)

On-site
USD 150,000 - 210,000
Health benefits
Microsecond ML Inference Architect
Microsecond ML Inference Architect

Long Ridge Partners • New York (NY)

On-site
USD 600,000 - 1,500,000
Generous PTO
Hybrid work options
Free breakfast, lunch, snacks
Infrastructure Engineer: Scale Low-Latency AI Compute (NYC)
Infrastructure Engineer: Scale Low-Latency AI Compute (NYC)

Harnham • New York (NY)

On-site
USD 200,000 - 440,000
Inference Runtime Engineer — On-Device & Cloud ML
Inference Runtime Engineer — On-Device & Cloud ML

Lm-Studio • New York (NY)

Hybrid
USD 140,000 - 210,000
Competitive salary and equity grants
Excellent medical, vision, dental care
Catered team lunch / expensed dinners
+3
Scale Low-Latency ML Inference Engineer (GPU/CUDA)
Scale Low-Latency ML Inference Engineer (GPU/CUDA)

AI Breaking Wire • San Francisco (CA), Northern (KY)

Hybrid
USD 140,000 - 190,000
Equity
Medical benefits (full coverage)
PTO & Hybrid work policy
+1
On-Device AI Inference Engineer — Ultra-Low Latency
On-Device AI Inference Engineer — Ultra-Low Latency

Hark • San Jose (CA), Northern (KY)

Hybrid
USD 200,000 - 450,000
Senior ML Infra Engineer - Low-Latency Scale
Senior ML Infra Engineer - Low-Latency Scale

Patreon • New York (NY), San Francisco (CA)

Hybrid
USD 180,000 - 230,000
Healthcare
401k with matching
Paid time off
Inference Infra Engineer: Scale Low-Latency AI Serving
Inference Infra Engineer: Scale Low-Latency AI Serving

Elorian • Palo Alto (CA)

On-site
USD 200,000 - 400,000
Health, dental, and vision benefits
Unlimited PTO
Parental leave
+1
Inference Optimization Engineer: Fast, Cost-Effective ML
Inference Optimization Engineer: Fast, Cost-Effective ML

Build AI • San Francisco (CA)

On-site
USD 150,000 - 210,000
Competitive pay
Medical, dental, and vision packages
Housing subsidy $2k/month near SF offi
+6