AI Research Engineer, Inference

OP Recruiting

Chicago (IL)

On-site

USD 150,000 - 210,000

Full time

26 hours ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Benefits offered by this job

Health benefits

Job summary

OP Recruiting in New York, NY is seeking an AI Research Engineer to advance deep learning inference with a focus on ultra-low latency and scalable deployment. You will bridge ML research and hardware acceleration, building low-latency inference systems that handle continuous global data streams to inform core business decisions.

The role emphasizes optimizing kernels, collaborating with researchers on neural architectures, and accelerating inference across diverse hardware platforms for

Qualifications

  • Hands-on experience engineering production-grade deep learning systems in complex domains.
  • Strong lower-level engineering foundation with custom compute kernels.
  • Proficiency with framework compilation internals and low-level runtimes such as PyTorch/JAX/XLA/CUDA Graphs.

Responsibilities

  • Drive performance optimizations for large-scale model execution, including custom kernel creation and data streaming pipelines.
  • Collaborate with researchers to co-design neural architectures for ultra-real-time execution.
  • Develop low-level primitives and custom ops to maximize compute throughput on modern hardware.
  • Evaluate and deploy hardware acceleration technologies, from accelerators to specialized silicon.
  • Identify and address bottlenecks in ultra-low-latency deep learning inference.

Skills

Deep learning systems
CUDA kernels
PyTorch internals
Hardware acceleration

Tools

CUDA
JAX
XLA
CuTe DSLs

Job description

Job Title:

AI Research Engineer – Deep Learning Inference

Location:

New York, NY | London, UK

About The Opportunity

Join an industry-leading global quantitative technology organization at the forefront of machine learning innovation. We are seeking a High-Performance AI Research Engineer to drive speed and scalability across our real-time predictive modeling infrastructure. In this role, you will bridge the gap between machine learning research and hardware acceleration, building low-latency inference systems that process continuous global data streams to drive core business decisions.

Responsibilities
  • Drive performance optimizations across all aspects of large-scale model execution, including custom kernel creation, data streaming pipelines, and novel hardware integration.
  • Partner directly with machine learning researchers to co-design neural architectures optimized for extreme real-time execution.
  • Author low-level primitives and custom operations to extract maximum compute throughput from modern hardware architectures.
  • Evaluate, benchmark, and deploy novel hardware acceleration tech, ranging from off-the-shelf accelerators to custom specialized silicon.
  • Formulate and execute engineering initiatives that address complex, non-obvious bottlenecks in ultra-low-latency deep learning inference.
Requirements (Must-Have)
  • At least two years of hands-on experience engineering production-grade deep learning systems within any complex domain (such as robotics, computer vision, audio, NLP, physics, or recommender platforms).
  • Strong lower-level engineering foundation, including experience writing custom compute kernels (e.g., CUDA, Triton, Pallas, or CuTe DSLs).
  • Proficiency with framework compilation internals and low-level runtime environments (PyTorch, JAX, XLA, or CUDA Graphs).
  • Practical exposure to hardware acceleration tech, such as FPGAs, ASICs, or specialized AI processors.
  • Proven ability to adapt algorithms and technical concepts across different domain applications.
Preferred Qualifications
  • Experience optimizing or serving Large Language Models (LLMs) and foundation architectures.
  • Note: Prior background in quantitative finance or trading is explicitly NOT required.
Compensation & Benefits
  • Highly competitive base salary, performance-based bonus incentive, and premium health/wellness benefits package. Equal Opportunity Employer.

Accepting Candidates

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

AI Research Engineer, Pre-Training
AI Research Engineer, Pre-Training

OP Recruiting • Chicago (IL)

On-site
USD 160,000 - 260,000
Research Engineer, Infrastructure, Inference
Research Engineer, Infrastructure, Inference

Thinking Machines Lab Inc. • San Francisco (CA), Northern (KY)

On-site
USD 350,000 - 475,000
Health, dental, vision benefits
Unlimited PTO
Parental leave
+1
Senior Software Engineer - Model Performance
Senior Software Engineer - Model Performance

Inference • San Francisco (CA)

On-site
USD 220,000 - 320,000
Competitive compensation
Equity in a high-growth startup
Comprehensive benefits
Senior/Staff AI Engineer
Senior/Staff AI Engineer

Data Direct Networks • California (MO)

On-site
USD 150,000 - 230,000
Vacation plans
Paid holidays
Bonus programs
+5
Senior Software Engineer - Model Performance
Senior Software Engineer - Model Performance

inference.net • San Francisco (CA)

On-site
USD 220,000 - 320,000
Equity in a high-growth startup
Comprehensive benefits
Member of Technical Staff (Inference) - AI Infrastructure
Member of Technical Staff (Inference) - AI Infrastructure

Hamilton Barnes • United States

On-site
USD 225,000 - 275,000
Full Benefits
AI Inference Engineer
AI Inference Engineer

Premier Global Links • Palo Alto (CA)

On-site
USD 230,000 - 350,000
Machine Learning Performance Engineer (Inference)
Machine Learning Performance Engineer (Inference)

Career Techniques • New York (NY)

On-site
USD 200,000 - 300,000
Low-Latency AI Inference Engineer
Low-Latency AI Inference Engineer

OP Recruiting • Chicago (IL)

On-site
USD 150,000 - 210,000
Health benefits
Machine Learning Engineer - Inference
Machine Learning Engineer - Inference

Fintal Partners • New York (NY)

On-site
USD 140,000 - 220,000