Realtime ML Inference Engineer — Scalable Serving

Yobi AI

New York (NY)

Remote

USD 120,000 - 150,000

Full time

14 days+
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

Competitive base salary
Meaningful equity
Annual bonus
Health, dental, vision plans
Unlimited PTO
401(k) with company match

Job summary

A leading Behavioral AI company is seeking a Machine Learning Engineer focused on inference and serving. In this role, you will design and optimize systems to operationalize AI models. The ideal candidate has deep expertise in model deployment, a strong low-latency mindset, and is proficient in languages like Go, Rust, or Java. You will work in either a fully remote or hybrid environment. Competitive salary and equity options are offered.

Qualifications

  • Deep expertise in model deployment and production ML serving.
  • Strong low-latency mindset and knowledge of inference optimization techniques.
  • Operational maturity with monitoring, drift detection, and observability.

Responsibilities

  • Build and scale production ML serving systems.
  • Ensure low-latency inference by optimizing model graphs.
  • Write robust, high-performance code.

Skills

Model deployment and production ML serving
Inference optimization techniques
High-performance code writing
Monitoring and drift detection
Understanding of custom runtimes
Applied ML reasoning

Tools

Go
Rust
C++
Java
Python

Job description

A leading Behavioral AI company is seeking a Machine Learning Engineer focused on inference and serving. In this role, you will design and optimize systems to operationalize AI models. The ideal candidate has deep expertise in model deployment, a strong low-latency mindset, and is proficient in languages like Go, Rust, or Java. You will work in either a fully remote or hybrid environment. Competitive salary and equity options are offered.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Real-Time ML Inference Engineer for Scalable Serving
Real-Time ML Inference Engineer for Scalable Serving

Yobi • New York (NY)

Hybrid
USD 100,000 - 150,000
Competitive Base Salary
Meaningful equity
Annual performance bonus
+3
Senior AI Model Serving Engineer — Low-Latency Inference
Senior AI Model Serving Engineer — Low-Latency Inference

Menlo Ventures • San Francisco (CA)

On-site
USD 166,000 - 225,000
Annual performance bonus
Equity options
Comprehensive benefits package
Inference Systems Engineer — Scalable, Reliable ML Serving
Inference Systems Engineer — Scalable, Reliable ML Serving

Speedrun Talent Network • San Francisco (CA)

On-site
USD 300,000 - 400,000
Health, dental, and vision benefits
Unlimited PTO
Paid parental leave
+1
Senior ML Systems Engineer - Distributed AI Inference
Senior ML Systems Engineer - Distributed AI Inference

Amazon Web Services (AWS) • Seattle (WA)

On-site
USD 120,000 - 160,000
Senior AI Inference Infrastructure Engineer
Senior AI Inference Infrastructure Engineer

Modular • United States

Hybrid
USD 167,000 - 273,000
Competitive salary
Premier insurance plans
Flexible paid time off
+1
Senior ML Engineer — Real-Time ML at Scale (Remote)
Senior ML Engineer — Real-Time ML at Scale (Remote)

SeatGeek • United States

Hybrid
USD 145,000 - 209,000
Equity stake
Flexible work environment
Unlimited PTO
+2
Senior ML Systems Engineer - Model Inference & Efficiency
Senior ML Systems Engineer - Model Inference & Efficiency

Cohere • New York (NY)

Hybrid
USD 100,000 - 150,000
Inclusive culture and work environment
Weekly lunch stipend, in-office lunches & snacks
Full health and dental benefits
+4
Senior ML Serving Engineer for LLMs & Inference
Senior ML Serving Engineer for LLMs & Inference

Alldus • San Jose (CA)

On-site
USD 180,000 - 220,000
ML Engineer: Distributed Inference & Model Serving
ML Engineer: Distributed Inference & Model Serving

ByteDance • Seattle (WA)

On-site
USD 177,688 - 416,100
Medical, dental, and vision insurance
401(k) savings plan with company match
Paid parental leave
+2
Staff Engineer, Scalable AI Inference Infrastructure
Staff Engineer, Scalable AI Inference Infrastructure

Inferact • San Francisco (CA)

Hybrid
USD 200,000 - 400,000
Generous health, dental, and vision benefits
401(k) company match
Equity options