AI Backend Engineer: Low-Latency Inference Systems
A1
Palo Alto (CA)
On-site
USD 255,000 - 405,000
Full time
14 days+
Get more replies from employers
Send a job-specific resume in minutes.
Start fresh or import an existing resume
Job summary
A leading technology company based in California is seeking candidates to build backend systems for AI-powered features. The role involves optimizing service performance and ensuring reliable production operations. Ideal candidates will have experience with high-throughput services, be familiar with AI inference patterns, and have proficiency in Python and NodeJs among other tools. This position offers a compensation range from $255K to $405K, emphasizing collaboration and direct contributions to the company's mission.
Familiarity with AI inference patterns (LLMs, embeddings, multimodal).
Bias toward shipping and learning from production behavior.
Responsibilities
Build and operate backend systems that serve AI-powered features in production.
Design inference pipelines, orchestration layers, and service boundaries around models.
Own production concerns: monitoring, logging, alerting, and incident response.
Optimize latency and throughput across inference, caching, batching, and streaming.
Skills
High-throughput service experience
AI inference pattern familiarity
Production behavior learning
Python
NodeJs
Pytorch
OpenAI / Anthropic / open-source LLMs
SQL
noSQL
Docker
Job description
A leading technology company based in California is seeking candidates to build backend systems for AI-powered features. The role involves optimizing service performance and ensuring reliable production operations. Ideal candidates will have experience with high-throughput services, be familiar with AI inference patterns, and have proficiency in Python and NodeJs among other tools. This position offers a compensation range from $255K to $405K, emphasizing collaboration and direct contributions to the company's mission.