Staff Engineer / AI Builder

Robots & Pencils

California

On-site

USD 124,000 - 171,000

Full time

6 hours ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Robots & Pencils is seeking a Staff AI Builder to lead the design and delivery of AI/ML systems, owning end-to-end architecture and driving technical leadership across the team.

You’ll build agentic workflows, develop RAG pipelines, integrate with AWS Bedrock, and craft production prompts with strong observability. This role emphasizes reliability, security, and scalable AI across enterprise applications.

Qualifications

  • 6+ years of professional software engineering experience with GenAI/LLM systems in production.
  • Hands-on depth in agentic AI: reasoning loops, tool calling, multi-agent orchestration.
  • Practical RAG expertise: embeddings, vector DBs, re-ranking.
  • Experience building evaluation and observability for LLM systems.
  • Strong prompt engineering with production-grade prompts and constraints.
  • AWS GenAI stack experience: Bedrock, Lambda, DynamoDB, S3, SQS, Step Functions.
  • Strong Python and Node.js skills across full-stack and APIs.
  • Knowledge of production reliability patterns: retries, circuit breakers, fallback models.
  • Experience with Docker and cloud-native deployment.
  • Comfort owning ambiguous, integration-heavy problems.

Responsibilities

  • Design and build agentic workflows with multi-agent orchestration and clear pattern choices.
  • Build and maintain RAG pipelines with chunking, embeddings, and vector search.
  • Integrate with AWS Bedrock and Agent Core, including MCP-based tool design.
  • Write and iterate production prompts with structured role framing and constraints.
  • Add evaluation/observability: golden datasets, LLM-as-judge, tracing tools.
  • Design for LLM failure with retries, backoff, and fallback strategies.
  • Develop backend services in Python/Node.js, using serverless architectures.
  • Design DynamoDB single-table schemas for conversation state and memory.
  • Support event-driven orchestration (Step Functions, SQS, EventBridge).
  • Collaborate with frontend teams to integrate agentic features for users.

Skills

GenAI/LLM
Python
Node.js
RESTful APIs
Docker
AWS Bedrock
DynamoDB
OpenSearch
LangFuse
LangSmith
Prompt engineering
Observability
System design

Tools

AWS Lambda
API Gateway
S3
SQS
EventBridge

Job description

We’re looking for a Staff AI Builder to lead the design and delivery of AI/ML systems. This role is ideal for an experienced engineer who thrives on architectural decisions, can confidently own systems end‑to‑end, and contributes to technical leadership across the team.

In this role, you will work as a key technical leader on a cross‑functional team, defining AI architecture, leading model development and optimization, and solving challenging integration problems. You’ll be joining real, in‑flight work where reliability, security, and scalability are critical. You’ll establish engineering standards, raise the bar on how we build AI systems, and contribute to the technical decisions that shape how our AI work scales over time.

Why This Role Matters

At Robots & Pencils, we design AI systems for a human world. Our name says it all. Robots and pencils means engineering paired with creativity, because every agent we ship has to work for real people in real workflows. That balance is baked into how we operate.

Every role here contributes directly to that mission. Here, you shape how AI systems integrate into enterprise operations, how teams move at real velocity, and how products create measurable impact for clients and the people they serve. We ship production‑ready AI in 30 to 45 days. That pace demands people who take ownership, lead with craft, and care deeply about what they put their name on.

What You’ll Do
Agentic & GenAI Engineering
  • Design and build agentic workflows — reasoning loops, tool/function calling, and orchestration across single- and multi-agent architectures — with clear judgment on which pattern fits the problem and which doesn’t
  • Build and maintain RAG pipelines: chunking strategy, embeddings, vector search (OpenSearch), re‑ranking, and staleness/refresh handling for dynamic knowledge sources
  • Integrate with AWS Bedrock and Agent Core, including MCP‑based tool design — writing tool descriptions precise enough that an orchestrator routes correctly every time
  • Write and iterate on production system prompts — structured role framing, output constraints, and few‑shot design — not just prompt tinkering
  • Build eval and observability into every agent you ship: golden datasets, RAGAS‑style metrics, LLM‑as‑judge (as both a runtime guardrail and an offline eval), and tracing via LangFuse/LangSmith or equivalent
  • Design for LLM failure, not just LLM success: retries with backoff and jitter, circuit breakers, fallback models, and a clear user‑facing story when something upstream degrades
Full-Stack & Platform Engineering
  • Build backend services in Python and Node.js, including serverless architectures (AWS Lambda, API Gateway) that support agentic workflows
  • Design DynamoDB single‑table schemas — composite keys, transactional writes — for conversation state, agent memory, and session history
  • Support event-driven orchestration (Step Functions, SQS, EventBridge) for asynchronous agent operations, with an eye toward reliability and clean error handling
  • Contribute to frontend integration points so agentic features land cleanly for end users, and write clean, well‑tested code across the stack
  • Support deployment, monitoring, and production troubleshooting in cloud-native environments (AWS, Docker), with guidance from senior engineers where needed
Collaboration & Technical Leadership
  • Bring informed technical judgment to architecture discussions — naming the real trade‑offs of a pattern (e.g., gateway centralization vs. latency, single‑vs multi‑agent design) rather than defaulting to what's familiar
  • Partner with product, design, and delivery leads across global teams to scope and ship features end‑to‑end
  • Mentor other engineers on the pod and help raise the bar on agentic engineering practices, including AI‑assisted development tools like Claude Code
  • Own assigned features and releases end‑to‑end — including the unglamorous parts: debugging, hardening, and keeping production systems trustworthy
What You’ll Bring
  • 6+ years of professional software engineering experience, including meaningful time shipping GenAI/LLM-powered systems in production — not just prototypes
  • Real, hands‑on depth in agentic AI: reasoning loops, tool/function calling, multi‑agent orchestration, and a considered point of view on when single‑agent design beats multi‑agent, and why
  • Practical RAG expertise: chunking strategies, embeddings, vector databases (OpenSearch or similar), cosine similarity search, and re‑ranking — you can explain the mechanism, not just the terminology
  • Experience building evaluation and observability for LLM systems: golden datasets, LLM‑as‑judge, RAGAS or comparable metrics, and tracing tools like LangFuse/LangSmith
  • Strong prompt engineering skills — you can write a production system prompt with real structure and constraints on request, not just describe the concept
  • Hands‑on experience with the AWS GenAI stack: Bedrock, Agent Core, Lambda, DynamoDB (single‑table design), S3, SQS, EventBridge, Step Functions
  • Strong Python and Node.js skills, with experience building full‑stack applications and RESTful APIs
  • Solid understanding of production reliability patterns for LLM-backed systems: retries/backoff, circuit breakers, fallback models — a bundle, not a single tactic
  • Experience with containerization (Docker) and cloud-native deployment
  • Comfort owning ambiguous, integration‑heavy problems, and the ability to go a level deeper when someone pushes on your answer instead of staying abstract
Helpful Extras And Unique Skills
  • Direct experience with Amazon Bedrock Agent Core — agent registration, MCP tool wiring, memory/session management, gateway provisioning
  • Experience in a regulated or high‑stakes domain (education, healthcare, financial services) where evaluation rigor and reliability carry real consequences
  • Familiarity with workflow orchestration tools such as Step Functions, Sequencer, or similar state‑machine-based systems
  • A point of view on LLM observability tooling, gained from actually running it in production rather than reading about it

Our salary range is $176,612 - $243,680 CAD

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Principal AI Engineering Architect
Principal AI Engineering Architect

Robots and Pencils • United States

Remote
USD 180,000 - 231,000
AI Full Stack Architect
AI Full Stack Architect

CirrusLabs, LLC • Atlanta (GA)

On-site
USD 180,000 - 240,000
Competitive compensation
Flexible work arrangements
Staff AI Engineer
Staff AI Engineer

Robots & Pencils • California (MO)

On-site
USD 125,558 - 173,239
Forward Deployed Engineer
Forward Deployed Engineer

Robots & Pencils • California (MO)

On-site
USD 125,989 - 173,834
Forward Deployed Engineer
Forward Deployed Engineer

Robots & Pencils • United States

On-site
USD 177,375 - 209,625
EG - Agentic AI Engineer
EG - Agentic AI Engineer

CodeRoad • United States

Remote
USD 120,000 - 180,000
100% Remote
Holidays Off
Paid Time Off
+3
Senior Agentic AI Engineer
Senior Agentic AI Engineer

Build • Northern (KY)

Hybrid
USD 150,000 - 190,000
Senior Agentic AI Engineer
Senior Agentic AI Engineer

ImagineX LLC • Northern (KY)

Hybrid
USD 140,000 - 190,000
Agentic AI Engineer
Agentic AI Engineer

Tactical Edge • Washington

On-site
USD 150,000 - 210,000
Principal Data Scientist, Agentic AI Technical Lead
Principal Data Scientist, Agentic AI Technical Lead

Bot Jobs • Myrtle Point (OR)

Remote
USD 180,000 - 280,000