Realtime ML Inference Engineer — Scalable Serving

Yobi AI

New York (NY)

Remote

USD 120,000 - 150,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Competitive base salary
Meaningful equity
Annual bonus
Health, dental, vision plans
Unlimited PTO
401(k) with company match

Job summary

A leading Behavioral AI company is seeking a Machine Learning Engineer focused on inference and serving. In this role, you will design and optimize systems to operationalize AI models. The ideal candidate has deep expertise in model deployment, a strong low-latency mindset, and is proficient in languages like Go, Rust, or Java. You will work in either a fully remote or hybrid environment. Competitive salary and equity options are offered.

Qualifications

  • Deep expertise in model deployment and production ML serving.
  • Strong low-latency mindset and knowledge of inference optimization techniques.
  • Operational maturity with monitoring, drift detection, and observability.

Responsibilities

  • Build and scale production ML serving systems.
  • Ensure low-latency inference by optimizing model graphs.
  • Write robust, high-performance code.

Skills

Model deployment and production ML serving
Inference optimization techniques
High-performance code writing
Monitoring and drift detection
Understanding of custom runtimes
Applied ML reasoning

Tools

Go
Rust
C++
Java
Python

Job description

A leading Behavioral AI company is seeking a Machine Learning Engineer focused on inference and serving. In this role, you will design and optimize systems to operationalize AI models. The ideal candidate has deep expertise in model deployment, a strong low-latency mindset, and is proficient in languages like Go, Rust, or Java. You will work in either a fully remote or hybrid environment. Competitive salary and equity options are offered.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Real-Time ML Inference Engineer for Scalable Serving
Real-Time ML Inference Engineer for Scalable Serving

Yobi • New York (NY)

Hybrid
USD 100,000 - 150,000
Competitive Base Salary
Meaningful equity
Annual performance bonus
+3
Senior AI Model Serving Engineer — Low-Latency Inference
Senior AI Model Serving Engineer — Low-Latency Inference

Menlo Ventures • San Francisco (CA)

On-site
USD 166,000 - 225,000
Annual performance bonus
Equity options
Comprehensive benefits package
Senior ML Engineer - Real-Time Inference & Systems
Senior ML Engineer - Real-Time Inference & Systems

Inworld AI • Germany (OH)

On-site
USD 120,000 - 180,000
Senior ML Engineer - Real-Time Inference & Scalable Systems
Senior ML Engineer - Real-Time Inference & Scalable Systems

careers.bitkraft.vc - Jobboard • Germany (OH)

On-site
USD 120,000 - 160,000
Senior ML Systems Engineer - Distributed AI Inference
Senior ML Systems Engineer - Distributed AI Inference

Amazon Web Services (AWS) • Seattle (WA)

On-site
USD 120,000 - 160,000
Senior Engineer, Model Serving & Inference
Senior Engineer, Model Serving & Inference

Databricks • San Francisco (CA)

On-site
USD 166,000 - 225,000
AI Inference Engineer: Real-Time ML, Hybrid, Equity
AI Inference Engineer: Real-Time ML, Hybrid, Equity

Pantera Capital • Palo Alto (CA)

Hybrid
USD 190,000 - 250,000
Comprehensive health insurance
Dental insurance
Vision insurance
+1
Senior AI Inference Infrastructure Engineer
Senior AI Inference Infrastructure Engineer

Modular • United States

Hybrid
USD 167,000 - 273,000
Senior ML Engineer — Real-Time ML at Scale (Remote)
Senior ML Engineer — Real-Time ML at Scale (Remote)

SeatGeek • United States

Hybrid
USD 145,000 - 209,000
Engineering Manager, AI Inference & Scale (Hybrid)
Engineering Manager, AI Inference & Scale (Hybrid)

Menlo Ventures • United States

Hybrid
USD 425,000 - 560,000
Competitive compensation
Generous vacation and parental leave
Flexible working hours