Uses AI-assisted coding tools like Claude Code, Codex, and Cursor and focuses on building agent-driven, vibe-style creative experiences.
About the Role
Lead the architecture and implementation of OpenArt's agent layer to power large-scale, creative AI workflows. Own end-to-end agent infrastructure (harness, orchestration, memory, evals, tooling) and ship production frontend and tooling while setting engineering standards and mentoring the team.
Job Description
Role
OpenArt is hiring a Senior/Staff Software Engineer to lead the agentic architecture behind their next-generation creative products. You will own the end-to-end agent harness — planning, tool use, memory/context management, orchestration, evaluation, and the production systems and frontends that expose these capabilities to millions of creators.
Key Responsibilities
- Own the architecture and implementation of the agent harness powering conversational and autonomous creative workflows (tool use, context/memory, multi-step planning, sub-agent orchestration).
- Design long-horizon planning and execution for creative tasks (scene planning, storyboarding, asset generation, revision loops) across multiple model vendors with clear failure and recovery modes.
- Build and maintain MCP servers and CLI-based agent tooling to connect agents to internal services, external model vendors, and creative asset pipelines.
- Author reusable Agent Skills and structured workflows; define patterns and standards other engineers follow.
- Build evaluation and observability: automated evals, regression suites, tracing, and quality gates to detect regressions before users see them.
- Drive reliability, latency, and cost-per-successful-run as first-class metrics.
- Ship polished, production frontend experiences end-to-end (React/Next.js) for agent-driven products.
- Set technical standards via architecture reviews, engineering standards, and mentorship.
- Track the agent ecosystem (MCP, Agent Skills, orchestration frameworks, model capabilities) and decide what belongs in OpenArt’s stack.
- Partner with founders, product, design, and GTM to translate user needs into technical direction and communicate trade-offs.
Requirements
- 7+ years of full-stack engineering experience shipping and owning production systems at scale, including 2+ years building LLM-powered agents in production (not just prototypes).
- Deep, hands-on experience with agent architecture: tool use, context/memory management, multi-step planning, sub-agent orchestration, and long-running workflows.
- Track record designing evaluation frameworks and observability for agentic systems.
- Strong system design and data modeling skills; comfortable with state, schemas, and APIs for real product use cases.
- Experience authoring reusable Agent Skills, structured agent workflows, or equivalent abstractions.
- Demonstrated technical leadership setting direction for systems or teams.
- A strong AI-assisted coding practice (uses tools like Claude Code, Codex, or Cursor and has opinions on them).
- High product sense and strong communication skills to explain agent behavior and trade-offs to technical and non-technical stakeholders.
Nice to Have
- Practical experience building MCP servers and CLI tooling for agents.
- Experience with agent orchestration frameworks (Claude Agent SDK, LangGraph, or similar) and LLM evaluation/observability tooling.
- Hands-on experience with image/video generation models and creative asset pipelines.
- Experience with guardrails, safety, and cost controls for autonomous systems.
- Startup experience or ownership of a product surface from 0 to 1.
Compensation & Work Setup
- Compensation range: $500,000 - $600,000 total compensation (includes base, bonus, equity).
- Equity: meaningful ownership in what you build.
- Work setup: Bay Area preferred, hybrid available. Visa sponsorship available.
What Success Looks Like in 3–6 Months
- A clear, documented architecture for the agent harness that the team builds against.
- A production agent workflow you designed is shipped and running for real users.
- An eval and observability loop that catches regressions before users see them.
- Measurable improvement in reliability, latency, or cost-per-successful-run on an important workflow.
- Tight feedback loop established with design, founders, and users; setting direction for agent engineering.
Skills
System Design Data Modeling Agent Architecture Observability Evaluation Frameworks Reliability Engineering Latency and Cost Optimization Full-stack Development Frontend Development Technical Leadership Mentorship Communication Product Sense AI-assisted Coding CLI Tooling Orchestration Memory/Context Management Workflow and Planning Design Testing and Regression Suites