Staff AI Engineer

Eleven Recruiting

United States

Remote

USD 180,000 - 240,000

Full time

11 hours ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Eleven Recruiting represents a growing SaaS client seeking a Staff AI Engineer to design and own high-stakes backend systems for long-running, agentic workflows. You will lead architecture decisions for multi-agent coordination with human-in-the-loop oversight, and build durable state management across distributed services.

The role requires deep experience with TypeScript, Python, and AWS, plus hands-on work with LangGraph, LangChain, and vector databases.

Qualifications

  • 8+ years building and operating backend systems in production environments where uptime and correctness are critical.
  • A track record of leading the design of complex, distributed, or high-scale systems from architecture through deployment and ongoing operations.
  • Hands-on production experience with agent orchestration frameworks (such as LangGraph) and long-running, stateful, multi-step agentic workflows.
  • Demonstrated technical leadership beyond your own commits: setting patterns and standards, driving cross-functional initiatives, mentoring engineers, and influencing decisions.
  • Strong experience with relational databases (PostgreSQL or similar): schema design, query optimization, data modeling, and migrations.
  • Hands-on work with event-driven architectures: message queues, async processing, and distributed job execution.
  • Production experience with AWS (Lambda, SQS, S3, ECS) or an equivalent cloud platform.
  • Comfort reading and writing both TypeScript and Python, or clear evidence of picking up a second language quickly.
  • Experience across the full software delivery lifecycle: design, implementation, testing, deployment, monitoring, and incident response.
  • Sound judgment on build-versus-buy decisions, with the ability to make and defend architectural trade-offs under real time and cost constraints.
  • Familiarity with modern agent/LLM tooling such as LangChain, vector databases, and cloud-based model platforms (e.g., AWS Bedrock).

Responsibilities

  • Design and build the most technically demanding parts of the system yourself, working in an AI native, agentic coding workflow as your daily practice, not just directing others.
  • Set the architecture for multi-agent coordination across complex, long-running workflows with human-in-the-loop oversight, including agent state machines, tool routing, context windowing, and retry logic for processes that can run for minutes or hours.
  • Own the durable state and run-record layer for long-running agentic workflows: state hydration, failure handling, checkpoint/resume, and recovery across distributed services.
  • Build production-grade pipelines that extract, classify, and validate structured documents across a wide range of formats and quality levels, and define the validation layer that turns raw extraction into output the business can trust.
  • Build the measurement backbone for the platform: task completion rates, accuracy attribution, cost tracking, and regression detection, so that "is this good enough to ship" is backed by data, not opinion.
  • Design the supervision layers, validation rules, human review gates, and audit trail needed to make non-deterministic AI output trustworthy for professionals who can't tolerate errors.
  • Set the framework the team uses to reason about cost, accuracy, and latency trade-offs across use cases and peak-volume periods.
  • Mentor engineers, raise the bar in design and code review, and help spread an AI native, agentic engineering culture across the broader team, reducing key-person risk by leaving systems more legible than you found them.

Skills

Backend systems
Distributed systems
Agent orchestration
Technical leadership
PostgreSQL
AWS
TypeScript & Python

Tools

LangGraph
LangChain
vector databases
cloud platforms

Job description

We are a specialized technology staffing agency supporting professional and financial services companies. Why do we stand out in technology staffing? We listen and act as advisors for our candidates on how they can best add value, find interesting projects, and pave a path for career advancement. We advocate for the best pay, diversity in tech, and the best job fit for every candidate we place.

Our client, a growing SaaS company, is seeking a fully remote Staff AI Engineer to join their team. Candidates must be based in states under Central and/or Eastern time zones.

Responsibilities

  • Design and build the most technically demanding parts of the system yourself, working in an AI native, agentic coding workflow as your daily practice, not just directing others.
  • Set the architecture for multi-agent coordination across complex, long-running workflows with human-in-the-loop oversight, including agent state machines, tool routing, context windowing, and retry logic for processes that can run for minutes or hours.
  • Own the durable state and run-record layer for long-running agentic workflows: state hydration, failure handling, checkpoint/resume, and recovery across distributed services.
  • Build production-grade pipelines that extract, classify, and validate structured documents across a wide range of formats and quality levels, and define the validation layer that turns raw extraction into output the business can trust.
  • Build the measurement backbone for the platform: task completion rates, accuracy attribution, cost tracking, and regression detection, so that "is this good enough to ship" is backed by data, not opinion.
  • Design the supervision layers, validation rules, human review gates, and audit trail needed to make non-deterministic AI output trustworthy for professionals who can't tolerate errors.
  • Set the framework the team uses to reason about cost, accuracy, and latency trade-offs across use cases and peak-volume periods.
  • Mentor engineers, raise the bar in design and code review, and help spread an AI native, agentic engineering culture across the broader team, reducing key-person risk by leaving systems more legible than you found them.

Qualifications

  • 8+ years building and operating backend systems in production environments where uptime and correctness are critical.
  • A track record of leading the design of complex, distributed, or high-scale systems from architecture through deployment and ongoing operations.
  • Hands-on production experience with agent orchestration frameworks (such as LangGraph) and long-running, stateful, multi-step agentic workflows. You need to have shipped this class of system, not just studied it.
  • Demonstrated technical leadership beyond your own commits: setting patterns and standards, driving cross-functional initiatives, mentoring engineers, and influencing decisions across an organization.
  • Strong experience with relational databases (PostgreSQL or similar): schema design, query optimization, data modeling, and migrations.
  • Hands-on work with event-driven architectures: message queues, async processing, and distributed job execution.
  • Production experience with AWS (Lambda, SQS, S3, ECS) or an equivalent cloud platform.
  • Comfort reading and writing both TypeScript and Python, or clear evidence of picking up a second language quickly.
  • Experience across the full software delivery lifecycle: design, implementation, testing, deployment, monitoring, and incident response.
  • Sound judgment on build-versus-buy decisions, with the ability to make and defend architectural trade-offs under real time and cost constraints.
  • Familiarity with modern agent/LLM tooling such as LangChain, vector databases, and cloud-based model platforms (e.g., AWS Bedrock).

Pluses

  • Created and scaled products from 0-to-1 that also generated revenue.
  • Depth in multi-agent coordination specifically: designing supervision, routing, and hand-off between cooperating agents.
  • Familiarity with RAG architectures, vector databases, or document processing pipelines.
  • Experience with multi-tenant SaaS architecture: schema isolation, tenant-scoped data, and access control.
  • Background in document intelligence: OCR, structured extraction from PDFs, and form understanding.
  • Experience standing up evaluation, observability, or quality systems for ML/LLM products.
  • Work in a regulated or high-trust domain (finance, legal, healthcare, tax) where output correctness is non-negotiable.
  • Open-source contributions, technical writing, or other public evidence of engineering depth.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Lead AI Agent Engineer
Lead AI Agent Engineer

EPAM Systems Inc • United States

Remote
USD 180,000 - 240,000
Technical Architect - ML
Technical Architect - ML

Quantiphi, Inc. • Princeton (NJ)

On-site
USD 140,000 - 190,000
Sr. Staff Software Engineer
Sr. Staff Software Engineer

UKG • Lowell (MA)

On-site
USD 180,000 - 230,000
Agentic AI Engineer
Agentic AI Engineer

SubcontractorHub • Park City (UT), Northern (KY)

Hybrid
USD 140,000 - 180,000
Senior Agentic AI Engineer
Senior Agentic AI Engineer

Build • Northern (KY)

Hybrid
USD 150,000 - 190,000
Senior Agentic AI Engineer
Senior Agentic AI Engineer

ImagineX LLC • Northern (KY)

Hybrid
USD 140,000 - 190,000
Principal AI Engineer
Principal AI Engineer

NAM Info Inc • New York (NY)

Hybrid
USD 210,000 - 320,000
Hybrid work model
Senior AI Agent Engineer
Senior AI Agent Engineer

EPAM Systems Inc • United States

Remote
USD 150,000 - 210,000
Engineering Manager, AI
Engineering Manager, AI

Jobtailor • New York (NY)

On-site
USD 180,000 - 240,000
Hybrid schedule
In-office 3 days/week (MWF) in NYC
AI Engineer
AI Engineer

Hiringhood • United States

Remote
USD 120,000 - 170,000