Member of Technical Staff - AI Platform Engineer

Patronus AI, Inc.

San Francisco, Northern (CA, KY)

On-site

USD 125,000 - 200,000

Full time

10 days ago
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

Competitive salary
Equity
15 days paid vacation per annum
Parental & sick leave
Health, dental, and vision insurance
401(k) plan + matching
Whoop band
Monthly meal stipend
Monthly health and wellness stipend
Global offsites

Job summary

Patronus AI, Inc. sits at the frontier of simulation research and infrastructure in San Francisco. The AI Platform Engineer role connects research and engineering to build and operate internal platforms (backends, dashboards, CLI, SDKs) used daily by colleagues.

The position emphasizes deploying and serving models on managed and self-operated GPUs, plus robust CI/CD, observability, and scalable tooling. In-office attendance is required five days a week at our SF headquarters.

Qualifications

  • Several years of hands-on experience building and operating production applications.
  • Strong backend engineering in production-grade Python — API design, async services, relational databases and object storage.
  • Hands-on experience deploying and serving models in production, on managed inference providers and on self-operated GPUs, with at least one modern LLM serving stack.
  • Experience running workloads on shared GPU clusters through a scheduler such as Slurm, and working knowledge of cloud GPU infrastructure.
  • Experience with CI/CD, containerized workflows and Kubernetes.
  • Experience with MLOps and LLMOps tooling — model registries and hubs such as HuggingFace, experiment tracking and monitoring such as Weights & Biases, and deployment telemetry.
  • Experience with production observability and on-call operation — logging, metrics and alerting.
  • Experience with modern agentic frameworks.
  • Clear written and verbal communication, and experience collaborating cross-functionally with research, product and platform teams.

Responsibilities

  • Building and operating the internal platforms end to end — backends, storage, dashboards, and the CLI and SDK surfaces they are driven through — including multi-tenancy, sign-in and access control.
  • Owning workload orchestration on the GPU clusters and the services around them — submission and scheduling, provisioning, quotas and placement.
  • Deploying models and keeping them served, on managed inference providers and on internally operated GPUs — sizing each deployment for its hardware, and writing serving wrappers where no off-the-shelf engine fits.
  • Building the evaluation surface, so that a result stays comparable across months, colleagues and models.
  • Establishing and hardening CI/CD, containerized workflows and release safety.
  • Instrumenting the platform with logging, metrics and alerting, so that divergence between what a service promises and what it serves is caught by a test rather than by a customer.
  • Adding agent surfaces to the platform, with whatever an agent resolves written back as auditable configuration.
  • Partnering with research engineers to productionize their experiments, and carrying production issues through to resolution — including on-call for systems built in this role.

Skills

Backend engineering
Python
ML model deployment
GPU cluster management
CI/CD
Kubernetes
LLMOps tooling
Observability & SRE
Communication

Education

BS, MS, or PhD in Computer Science, Machine Learning, Software Engineering, or a related quantitative field

Tools

Slurm
Weights & Biases
HuggingFace
Docker
Linux

Job description

Member of Technical Staff - AI Platform Engineer

Patronus AI is a frontier lab developing simulation research and infrastructure to accelerate progress toward human-aligned AGI. We are on a mission to simulate all of the world’s intelligence.

We are the team behind some of the earliest and most influential research in AI evaluation like FinanceBench , Lynx , SimpleSafetyTests , CopyrightCatcher , Humanity’s Last Exam , and more. We are formerly AI researchers and engineers from companies like Meta AI, Amazon AGI, and Google. Our customers include foundation model labs and Fortune 500 enterprises like Adobe. We are backed by top-tier investors like Lightspeed Venture Partners, Notable Capital, Stanford University, Noam Brown, Gokul Rajaram, and more.

The AI Platform Engineer sits between research and engineering, turning research workflows into services that other teams can self-serve. The work centres on building and operating the several internal platforms that the research and engineering teams rely on daily. These are production products with real users — backends, dashboards, CLIs and SDKs — built for colleagues inside the company rather than for an abstract audience.

The second half of the role is the foundation those platforms stand on: deploying models and keeping them served, both on managed inference providers and on internally operated GPUs, together with the cluster services that make GPU capacity self-serve rather than a matter of negotiation. This is a hands‑on, delivery‑oriented platform position rather than a research one: the emphasis falls on the systems and tooling that make other teams' work possible. Model training is part of the surrounding environment and remains accessible, but it is not the focus of the position.

In this role, you will:

  • Building and operating the internal platforms end to end — backends, storage, dashboards, and the CLI and SDK surfaces they are driven through — including multi‑tenancy, sign‑in and access control.
  • Owning workload orchestration on the GPU clusters and the services around them — submission and scheduling, provisioning, quotas and placement.
  • Deploying models and keeping them served, on managed inference providers and on internally operated GPUs — sizing each deployment for its hardware, and writing serving wrappers where no off‑the‑shelf engine fits.
  • Building the evaluation surface, so that a result stays comparable across months, colleagues and models.
  • Establishing and hardening CI/CD, containerized workflows and release safety.
  • Instrumenting the platform with logging, metrics and alerting, so that divergence between what a service promises and what it serves is caught by a test rather than by a customer.
  • Adding agent surfaces to the platform, with whatever an agent resolves written back as auditable configuration.
  • Partnering with research engineers to productionize their experiments, and carrying production issues through to resolution — including on‑call for systems built in this role.
Qualifications

"The number one qualification to succeed in this machine learning course is gumption” - John Lafferty, CS Professor at Yale

Above all, we look for a proactive mindset, willingness to learn, unlimited energy, and relentless optimism. You are a great fit if you have a background in the following:

  • Several years of hands‑on experience building and operating production applications.
  • Strong backend engineering in production‑grade Python — API design, async services, relational databases and object storage.
  • Hands‑on experience deploying and serving models in production, on managed inference providers and on self‑operated GPUs, with at least one modern LLM serving stack.
  • Experience running workloads on shared GPU clusters through a scheduler such as Slurm, and working knowledge of cloud GPU infrastructure.
  • Experience with CI/CD, containerized workflows and Kubernetes.
  • Experience with MLOps and LLMOps tooling — model registries and hubs such as HuggingFace, experiment tracking and monitoring such as Weights & Biases, and deployment telemetry.
  • Experience with production observability and on‑call operation — logging, metrics and alerting.
  • Experience with modern agentic frameworks.
  • A BS, MS, or PhD in Computer Science, Machine Learning, Software Engineering, or a related quantitative field.
  • Clear written and verbal communication, and experience collaborating cross‑functionally with research, product and platform teams.
  • Strong integrity, sound judgment, and respect for others.
To support close collaboration, this role is based in our San Francisco headquarters and requires in‑office attendance 5 days a week.

The expected base salary range for this role is$125,000 - $200,000 USD. In addition to base salary, we offer equity and benefits. Actual compensation will be determined based on experience, qualifications, skills, and location.

  • Competitive salary and equity packages
  • 15 days of paid vacation per annum
  • Parental & sick leave
  • Health, dental, and vision insurance plans
  • 401(k) plan + matching
  • Whoop band
  • Monthly meal stipend
  • Monthly health and wellness stipend
  • Fun global offsites!

Patronus AI is an equal opportunity employer. We celebrate diversity in our workplace, and all qualified applicants will receive consideration for employment without regard to age, ancestry, color, family or medical care leave, gender identity or expression, genetic information, marital status, medical condition, national origin, physical or mental disability, political affiliation, protected veteran status, race, religion, sex (including pregnancy), sexual orientation, or other legally protected characteristics.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Technical Program Manager
Technical Program Manager

Patronus AI • San Francisco (CA)

On-site
USD 125,000 - 250,000
Competitive salary and equity packages
15 days of paid vacation per annum
Parental & sick leave
+9
Recruiting Coordinator
Recruiting Coordinator

Patronus AI • San Francisco (CA)

On-site
USD 80,000 - 125,000
Competitive salary and equity packages
15 days of paid vacation per annum
Parental & sick leave
+9
Senior AI Platform Engineer (#5640)
Senior AI Platform Engineer (#5640)

Talanto • Union (NJ), Northern (KY)

Hybrid
USD 140,000 - 210,000
Flexible working format
Competitive salary
Professional development
Member of Technical Staff – Model Training
Member of Technical Staff – Model Training

Inflection AI • Palo Alto (CA)

Hybrid
USD 175,000 - 350,000
Competitive stock options
Diverse medical, dental, and vision options
401k matching program
+3
Software Engineer - Hosted Model Infrastructure
Software Engineer - Hosted Model Infrastructure

Palantir Technologies • Washington

On-site
USD 145,000 - 200,000
Medical, dental, vision insurance
Paid time off
401(k) plan enrollment
+1
AIML Engineer
AIML Engineer

Qubeaxis • San Francisco (CA)

On-site
USD 180,000 - 260,000
Performance bonus (up to 20% of base)
Equity participation
Health, dental, and vision insurance
+3
AI Engineer
AI Engineer

Patlytics • New York (NY)

On-site
USD 120,000 - 160,000
Comprehensive health coverage
Unlimited PTO
Immediate 401(k) enrollment
+2
Pioneering Intelligence | Cambridge, MA Principal Software Engineer - Agentic Platform
Pioneering Intelligence | Cambridge, MA Principal Software Engineer - Agentic Platform

Flagship Pioneering • Cambridge (MA)

On-site
USD 188,000 - 258,000
Healthcare coverage
Annual incentive program
Retirement benefits
Principal AI Platform Engineer
Principal AI Platform Engineer

Booster • United States

Hybrid
USD 117,000 - 207,000
Health insurance
Retirement savings plan
Paid time off
Member of Technical Staff - Training Platform
Member of Technical Staff - Training Platform

Kubelt • San Francisco (CA)

On-site
USD 150,000 - 300,000
Cash compensation $150K–$300K
Flexible work arrangement
Full visa sponsorship
+1