AI Backend Engineer: Inference & Orchestration

A1

Palo Alto (CA)

On-site

USD 150,000 - 210,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

A1 is hiring a Backend Engineer, AI, in Palo Alto to own the inference and orchestration layer powering AI interactions in our product. You will build and operate production systems that turn model capabilities into fast, stable APIs accessed by mobile and desktop clients.

You will design multi-step inference pipelines, manage tool calls and retries, and optimize routing, caching, batching, streaming, and state management to meet latency and throughput goals.

Qualifications

  • Experience building backend systems for AI-powered features.
  • Design and operate inference pipelines and orchestration layers.
  • Strong emphasis on latency, reliability, and observability.

Responsibilities

  • Build and operate backend systems that serve AI-powered features in production.
  • Design inference pipelines and orchestration layers that handle multi-step workflows, tool calls, and retries.
  • Manage the full lifecycle of AI requests: routing, caching, batching, streaming, and state management.
  • Optimize latency, throughput, and cost across model inference and downstream systems.
  • Design systems that remain reliable despite non-deterministic model behavior and external dependencies.
  • Implement observability for AI systems, including logging, tracing, and debugging of model outputs and failures.
  • Collaborate with ML and product teams to translate model capabilities into stable, production-grade APIs.

Skills

Python
NodeJs
Pytorch
OpenAI / Anthropic / open-source LLMs
SQL & NoSQL
Docker

Job description

A1 is hiring a Backend Engineer, AI, in Palo Alto to own the inference and orchestration layer powering AI interactions in our product. You will build and operate production systems that turn model capabilities into fast, stable APIs accessed by mobile and desktop clients.

You will design multi-step inference pipelines, manage tool calls and retries, and optimize routing, caching, batching, streaming, and state management to meet latency and throughput goals.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI Backend Engineer: Inference & Orchestration at Scale
AI Backend Engineer: Inference & Orchestration at Scale

Bjak • Town of Sweden (NY)

On-site
USD 120,000 - 160,000
AI Backend Engineer: Scalable Inference & Orchestration
AI Backend Engineer: Scalable Inference & Orchestration

Bjak • Town of Poland (NY)

On-site
USD 140,000 - 190,000
AI Backend Engineer — Inference & Orchestration
AI Backend Engineer — Inference & Orchestration

Bjak • Germany (OH)

On-site
USD 110,000 - 150,000
Backend AI Engineer - Inference & Orchestration
Backend AI Engineer - Inference & Orchestration

United States Digital Space LLC • Germany (OH)

On-site
USD 120,000 - 180,000
Backend AI Engineer: Production-Grade Inference
Backend AI Engineer: Production-Grade Inference

European Recruitment BV • United States

On-site
USD 120,000 - 160,000
Backend Engineer, AI (Agent Systems)
Backend Engineer, AI (Agent Systems)

European Recruitment BV • United States

On-site
USD 120,000 - 160,000
AI Backend Engineer: Build Scalable Inference Pipelines
AI Backend Engineer: Build Scalable Inference Pipelines

NTIATIVE IT Recruitment • Town of Poland (NY)

On-site
USD 120,000 - 190,000
Autonomy from day one
Work with AI technologies
High ownership and impact
+1
AI Backend Engineer: Scale Enterprise AI Systems
AI Backend Engineer: Scale Enterprise AI Systems

Within • San Francisco (CA)

On-site
USD 140,000 - 190,000
Equity
Employer-Paid Medical
Parental Leave
+5
Backend Engineer, AI (Agent Systems)
Backend Engineer, AI (Agent Systems)

Bjak • Germany (OH)

On-site
USD 110,000 - 150,000
Backend Engineer, AI (Agent Systems)
Backend Engineer, AI (Agent Systems)

Bjak • Town of Poland (NY)

On-site
USD 140,000 - 190,000