Real-Time ML Inference Engineer for Scalable Serving

Yobi

New York (NY)

Hybrid

USD 100,000 - 150,000

Full time

14 days+
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Benefits offered by this job

Competitive Base Salary
Meaningful equity
Annual performance bonus
Comprehensive health benefits
Unlimited PTO
401k with company match

Job summary

A Behavioral AI company is seeking a Machine Learning Engineer to design and optimize systems for bringing their models to life. The role involves ensuring ML models are efficient and reliable, requiring experience in model deployment and robust coding skills. Candidates should be familiar with low-latency techniques and operational maturity in ML systems. This position can be remote or hybrid from several hubs.

Qualifications

  • Deep expertise in model deployment and scaling production ML serving systems.
  • Understand low-latency model inference techniques.
  • Robust coding skills in Go, Rust, C++, or Java.
  • Experience in monitoring and observing models.

Responsibilities

  • Design, optimize, and operate systems for Behavioral AI models.
  • Package, version, and roll out ML models in production environments.
  • Ensure models are fast, accurate, and accountable.

Skills

Model deployment
Low-latency optimization
High-performance coding in Go, Rust, C++, Java
Monitoring model drift
Infrastructure design
Applied ML understanding

Job description

A Behavioral AI company is seeking a Machine Learning Engineer to design and optimize systems for bringing their models to life. The role involves ensuring ML models are efficient and reliable, requiring experience in model deployment and robust coding skills. Candidates should be familiar with low-latency techniques and operational maturity in ML systems. This position can be remote or hybrid from several hubs.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Realtime ML Inference Engineer — Scalable Serving
Realtime ML Inference Engineer — Scalable Serving

Yobi AI • New York (NY)

Remote
USD 120,000 - 150,000
Competitive base salary
Meaningful equity
Annual bonus
+3
Senior ML Engineer — Real-Time ML at Scale (Remote)
Senior ML Engineer — Real-Time ML at Scale (Remote)

SeatGeek • United States

Hybrid
USD 145,000 - 209,000
Equity stake
Flexible work environment
Unlimited PTO
+2
Inference Systems Engineer — Scalable, Reliable ML Serving
Inference Systems Engineer — Scalable, Reliable ML Serving

Speedrun Talent Network • San Francisco (CA)

On-site
USD 300,000 - 400,000
Health, dental, and vision benefits
Unlimited PTO
Paid parental leave
+1
Senior ML Systems Engineer - Model Inference & Efficiency
Senior ML Systems Engineer - Model Inference & Efficiency

Cohere • New York (NY)

Hybrid
USD 100,000 - 150,000
Inclusive culture and work environment
Weekly lunch stipend, in-office lunches & snacks
Full health and dental benefits
+4
Senior AI Model Serving Engineer — Low-Latency Inference
Senior AI Model Serving Engineer — Low-Latency Inference

Menlo Ventures • San Francisco (CA)

On-site
USD 166,000 - 225,000
Annual performance bonus
Equity options
Comprehensive benefits package
Production AI/ML Engineer - Scalable Models & Real-Time AI
Production AI/ML Engineer - Scalable Models & Real-Time AI

Devitechs • Town of Texas (WI)

On-site
USD 80,000 - 120,000
Senior ML Engineer — Scalable AI (Hybrid)
Senior ML Engineer — Scalable AI (Hybrid)

ZipRecruiter • Los Angeles (CA)

Hybrid
USD 140,000 - 225,000
Competitive compensation
Exceptional benefits package
Flexible Vacation & Paid Time Off
+1
AI/ML Engineer - Build & Deploy Scalable Models
AI/ML Engineer - Build & Deploy Scalable Models

10xTalents • Fremont (CA)

On-site
USD 100,000 - 150,000
Senior ML Serving Engineer for LLMs & Inference
Senior ML Serving Engineer for LLMs & Inference

Alldus • San Jose (CA)

On-site
USD 180,000 - 220,000
ML Performance Engineer – Real-Time Inference
ML Performance Engineer – Real-Time Inference

Odyssey • Palo Alto (CA)

On-site
USD 130,000 - 160,000