AI Backend Engineer: Low-Latency Inference Systems

A1

Palo Alto (CA)

On-site

USD 255,000 - 405,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

A leading technology company based in California is seeking candidates to build backend systems for AI-powered features. The role involves optimizing service performance and ensuring reliable production operations. Ideal candidates will have experience with high-throughput services, be familiar with AI inference patterns, and have proficiency in Python and NodeJs among other tools. This position offers a compensation range from $255K to $405K, emphasizing collaboration and direct contributions to the company's mission.

Qualifications

  • Experience running high-throughput, low-latency services.
  • Familiarity with AI inference patterns (LLMs, embeddings, multimodal).
  • Bias toward shipping and learning from production behavior.

Responsibilities

  • Build and operate backend systems that serve AI-powered features in production.
  • Design inference pipelines, orchestration layers, and service boundaries around models.
  • Own production concerns: monitoring, logging, alerting, and incident response.
  • Optimize latency and throughput across inference, caching, batching, and streaming.

Skills

High-throughput service experience
AI inference pattern familiarity
Production behavior learning
Python
NodeJs
Pytorch
OpenAI / Anthropic / open-source LLMs
SQL
noSQL
Docker

Job description

A leading technology company based in California is seeking candidates to build backend systems for AI-powered features. The role involves optimizing service performance and ensuring reliable production operations. Ideal candidates will have experience with high-throughput services, be familiar with AI inference patterns, and have proficiency in Python and NodeJs among other tools. This position offers a compensation range from $255K to $405K, emphasizing collaboration and direct contributions to the company's mission.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Backend Engineer, AI Inference Platform
Senior Backend Engineer, AI Inference Platform

Together AI • San Francisco (CA)

On-site
USD 160,000 - 250,000
Competitive compensation
Startup equity
Health insurance
Backend Engineer, AI: Build Fast, Reliable AI APIs
Backend Engineer, AI: Build Fast, Reliable AI APIs

Bjak • San Francisco (CA)

On-site
USD 120,000 - 160,000
Senior AI Inference Infrastructure Engineer
Senior AI Inference Infrastructure Engineer

Modular • United States

Hybrid
USD 167,000 - 273,000
Senior ML Engineer - Low-Latency Inference & Systems
Senior ML Engineer - Low-Latency Inference & Systems

Inworld • Germany (OH)

Hybrid
USD 120,000 - 160,000
Senior Backend Engineer, AI Inference Platform
Senior Backend Engineer, AI Inference Platform

MongoDB • Seattle (WA)

Hybrid
USD 120,000 - 170,000
Senior AI Model Serving Engineer — Low-Latency Inference
Senior AI Model Serving Engineer — Low-Latency Inference

Menlo Ventures • San Francisco (CA)

On-site
USD 166,000 - 225,000
Annual performance bonus
Equity options
Comprehensive benefits package
Senior Backend Engineer: Real-Time AI Platform
Senior Backend Engineer: Real-Time AI Platform

Apiphany Corporation • San Francisco (CA)

Hybrid
USD 160,000 - 300,000
Equity: 0.1%–1.0%
401(k) plan
Medical, Dental, and Vision insurance
+1
AI Inference Engineer: Real-Time ML, Hybrid, Equity
AI Inference Engineer: Real-Time ML, Hybrid, Equity

Pantera Capital • Palo Alto (CA)

Hybrid
USD 190,000 - 250,000
Comprehensive health insurance
Dental insurance
Vision insurance
+1
Staff Engineer, Scalable AI Inference Infrastructure
Staff Engineer, Scalable AI Inference Infrastructure

Inferact • San Francisco (CA)

Hybrid
USD 200,000 - 400,000
AI Backend Engineer — Inference & Orchestration
AI Backend Engineer — Inference & Orchestration

Bjak • Germany (OH)

On-site
USD 110,000 - 150,000