Stand out for this role — generate a tailored resume and cover letter in about a minute.
Oscar is seeking a Senior Agentic AI Engineer to lead autonomous agent development and digital twin infrastructure for advanced manufacturing. You will own the core agent harness and evaluation infra to enable trustworthy, stateful AI decisions in complex environments.
You will architect agent orchestration, implement RAG pipelines, and collaborate with security to ensure RBAC and secure deployments in hybrid SF offices.
Title: Senior Agentic AI Engineer (Hybrid - San Francisco)
Position Type: Full-Time
A high-growth enterprise AI platform is seeking a Senior Agentic AI Engineer to drive the development of autonomous AI agents and digital twin infrastructure for advanced manufacturing and complex physical systems. You will own the core agent harness and evaluation infrastructure that powers autonomous decision-making, designing trustworthy, stateful AI applications that translate complex technical data into high-impact operational decisions.
Architect and build production-grade agent orchestration layers and stateful multi-agent workflows leveraging or extending frameworks like LangGraph, Google ADK, or OpenAI Agents SDK.
Develop MCP servers and tool-registration APIs to give agents persistent context and shared state across complex enterprise workflows.
Construct human-in-the-loop systems, including review tooling, checkpoints, and judge-gated flows to ensure system output accuracy and reliability.
Build first-class evaluation infrastructure using LLM-as-judge gating, golden datasets, replay runs, and full-stack tracing.
Implement knowledge-graph-backed agent context (Neo4j) and specialized retrieval-augmented generation (RAG) pipelines for complex domain retrieval.
Partner with security leads to ensure proper RBAC, audit trails, and network-isolated serving for enterprise customer environments.
5+ years of production software engineering experience, including 2+ years shipping and operating LLM or agentic systems in live environments.
Deep architectural knowledge of stateful agent orchestration frameworks, including execution models, state management, tool orchestration, and failure handling.
Demonstrated experience building evaluation infrastructure (judge thresholds, golden-set regressions, replay runs) and end-to-end RAG architectures.
Hands-on proficiency with Docker, Kubernetes, and major cloud infrastructure (Azure, AWS, or GCP).
Track record of taking complex, stateful systems from design into highly observable production deployment.
Experience with programmatic prompt optimization (DSPy or similar) and graph databases (Neo4j) in production.
Background in uncertainty quantification, model calibration, or network-isolated/air-gapped enterprise deployments.
Exposure to complex engineering domains, hardware systems, or advanced manufacturing operations.