Member of Technical Staff - Engineering

Engg

San Francisco (CA)

On-site

USD 190,000 - 280,000

Full time

7 days ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Patronus AI is seeking a hands-on Staff Engineer to build scalable simulation infrastructure and ML systems. You’ll own end-to-end delivery across frontend, backend, and ML workloads, collaborating with researchers to translate ambitious AI ideas into robust production platforms.

You should ship code with strong product judgment, operate effectively in ambiguity, and mentor teammates while advancing tooling for agent environments and evaluation.

Qualifications

  • Track record of shipping non-trivial software end-to-end as an individual contributor.
  • Strong fundamentals with depth in backend/infrastructure, frontend/product engineering, or ML systems.
  • Experience building production systems in Python, Go, or TypeScript; quick to adapt to new stacks.
  • Experience with modern LLMs and agents at the application level and tool calling concepts.
  • Good architectural judgment around design, reliability, and operational complexity.

Responsibilities

  • Build agent environments and simulations end-to-end, including frontend interfaces and APIs.
  • Develop infrastructure for agent gym, orchestration, sandboxing, and benchmarking.
  • Create internal platforms and developer tools for researchers and engineers.
  • Operate ML infrastructure, including model deployment and GPU workloads.
  • Own systems from idea to production and iterate based on performance.
  • Collaborate with researchers to productionize experiments and scale tooling.

Skills

Shipping end-to-end software
Backend/infrastructure depth
Frontend/product engineering
ML systems
Python
Go
TypeScript
LLMs and agents
System design judgment
Independent work in ambiguity
AI tooling fluency

Education

BS/MS/PhD in Computer Science, Machine Learning, Software Engineering, or related field

Tools

React
TypeScript
Next.js
Python
Relational databases
APIs / API frameworks
ML tooling

Job description

About Patronus AI

About Patronus AI Patronus AI is a frontier lab developing simulation research and infrastructure to accelerate progress toward human-aligned AGI. We are on a mission to simulate all of the world’s intelligence. We are the team behind some of the earliest and most influential research in AI evaluation like FinanceBench , Lynx , SimpleSafetyTests , CopyrightCatcher , Humanity’s Last Exam , and more. We are formerly AI researchers and engineers from companies like Meta AI, Amazon AGI, and Google. Our customers include foundation model labs and Fortune 500 enterprises like Adobe. We are backed by top-tier investors like Lightspeed Venture Partners, Notable Capital, Stanford University, Noam Brown, Gokul Rajaram, and more.

Responsibilities

Responsibilities As a Member of Technical Staff – Engineering at Patronus AI, you’ll build the systems, infrastructure, and products that power our simulation research and agent training work. This is a broad engineering role for people who like operating across boundaries. Depending on the problem, you might build a realistic RL environment end-to-end, design infrastructure for running thousands of agent trajectories, ship internal platforms used by researchers, deploy and serve models, or build AI-powered developer tools that make the entire team faster. You’ll work across the stack — frontend interfaces, backend services, infrastructure, and ML/agent integrations — and partner closely with researchers and engineers to turn ambiguous problems into robust systems. We’re looking for engineers with high ownership, strong product and technical judgment, and the ability to learn unfamiliar areas quickly. You don't need to be an expert in every part of the stack. You should have meaningful depth somewhere and the curiosity and engineering range to work wherever the problem requires. In this role, you will:

  • Build agent environments and simulations end-to-end, including frontend interfaces, backend services, APIs, data models, tools, and realistic workflows used to train and evaluate AI agents.
  • Build the infrastructure that powers our agent gym, including orchestration, sandboxing, packaging, benchmarking, and systems for running environments across heterogeneous targets.
  • Develop internal platforms and developer tools used by researchers and engineers, from backends and dashboards to CLIs, SDKs, review agents, codegen helpers, and workflow automations.
  • Build and operate ML infrastructure, including model deployment and serving, evaluation systems, GPU workloads, and the services that make compute accessible to the broader team.
  • Own systems from ambiguous idea through production.
  • Define the problem, make architectural decisions, implement the solution, instrument it, and iterate based on how it performs in practice.
  • Think deeply about correctness and failure modes. Design for edge cases, adversarial agent behavior, reproducibility, observability, and the messy realities of production systems.
  • Partner closely with researchers to productionize experiments and build the software and infrastructure needed to turn research ideas into scalable systems.
  • Be a power user of AI coding tools like Claude Code, Codex, Cursor, and similar tools — and build new tooling and automations on top of them when existing workflows aren't good enough.
  • Move quickly without sacrificing judgment.
  • Make pragmatic decisions about what needs to be robust today, what can evolve later, and where technical investment will create leverage for the team.
Qualifications

Qualifications “The number one qualification to succeed in this machine learning course is gumption” - John Lafferty, CS Professor at Yale We're looking for a hands‑on generalist who ships. Above all, we value strong product instincts, the ability to operate independently in ambiguity, and genuine curiosity about agents and the frontier of AI. The list below is broad, and we don't expect every candidate to tick every box. What matters more is that you ship, stay curious, and use AI tools fluently enough to close gaps in days or weeks, not months. The team leans on AI tooling heavily to ramp on unfamiliar areas, and we expect anyone joining to do the same. You are a strong fit if you have:

  • A track record of shipping non‑trivial software end‑to‑end as an individual contributor, ideally at a startup or on a small, high‑velocity team.
  • Strong engineering fundamentals and meaningful depth in at least one of backend/infrastructure, frontend/product engineering, or ML systems, with the ability and desire to work across boundaries.
  • Experience building production systems in languages such as Python, Go, and/or TypeScript, and the ability to become productive quickly in an unfamiliar stack.
  • Experience working with modern LLMs and agents at the application level — including concepts like tool calling, agent loops, context management, harnesses, and evaluation.
  • Strong engineering judgment around system design, correctness, reliability, failure modes, edge cases, and operational complexity.
  • High independence. You can take an ambiguous goal, determine what needs to be built, find the people or information necessary to unblock yourself, and ship without requiring the work to be fully pre‑scoped.
  • Fluency with modern AI coding tools and a strong instinct for where AI can automate or accelerate engineering workflows.
  • A BS, MS, or PhD in Computer Science, Machine Learning, Software Engineering, or a related quantitative field — or equivalent experience.

Depending on your area of depth, you may also have experience with:

  • Building complex full‑stack products using technologies like React, TypeScript, Next.js, Python, relational databases, and modern API frameworks.
  • Building developer platforms, distributed systems, orchestration systems, sandboxes, or internal infrastructure.
  • Reinforcement learning environments, agent evaluation, verifiers, reward models, or benchmarking infrastructure.
  • Deploying and serving ML models using managed inference providers or self‑operated GPUs.
  • GPU infrastru
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Strategic Projects Lead
Strategic Projects Lead

Patronus AI • San Francisco (CA)

On-site
USD 150,000 - 210,000
Member of Technical Staff (Applied AI Engineer, Agent Capabilities)
Member of Technical Staff (Applied AI Engineer, Agent Capabilities)

United States Digital Space LLC • San Francisco (CA)

On-site
USD 190,000 - 230,000
Member of Technical Staff (Applied AI Engineer, Agent Capabilities)
Member of Technical Staff (Applied AI Engineer, Agent Capabilities)

Perplexity AI • New York (NY)

On-site
USD 150,000 - 210,000
Member of Technical Staff (Applied AI Engineer, Agent Capabilities)
Member of Technical Staff (Applied AI Engineer, Agent Capabilities)

Perplexity • New York (NY), Northern (KY)

On-site
USD 150,000 - 190,000
Member of Technical Staff (Applied AI Engineer, Agent Capabilities)
Member of Technical Staff (Applied AI Engineer, Agent Capabilities)

Precision Labs • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 280,000
Member of Technical Staff (Applied AI Engineer, Agent Capabilities)
Member of Technical Staff (Applied AI Engineer, Agent Capabilities)

Apply • New York (NY), Northern (KY)

Hybrid
USD 180,000 - 240,000
Technical Program Manager
Technical Program Manager

Patronus AI • San Francisco (CA)

On-site
USD 125,000 - 250,000
Competitive salary and equity packages
15 days of paid vacation per annum
Parental & sick leave
+9
AI Engineer
AI Engineer

Valsoft Corporation • United States

On-site
USD 140,000 - 230,000
AI Engineer
AI Engineer

Valsoft Corporation • Northern (KY)

Hybrid
USD 120,000 - 180,000
Staff AI Infrastructure Engineer
Staff AI Infrastructure Engineer

Cassi Home • United States

On-site
USD 180,000 - 260,000