Member of Technical Staff, AI Evaluation

P-1 Ai

United States

Hybrid

USD 110,000 - 180,000

Full time

2 days ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Benefits offered by this job

Competitive salary
Equity ownership
Healthcare
Unlimited PTO

Job summary

P-1 AI is pioneering an AI engineer agent for the physical world named Archie. You will help evaluate Archie against real engineering tasks, work with customers and partners, and shape evaluation methods that guide product development.

You will ensure the underlying engineering reasoning remains sound and actionable. The role sits at the forefront of AI for physical engineering, collaborating with cross-functional teams to expose failure modes and improve Archie’s capabilities toward engineering

Qualifications

  • Technical background in a physical engineering domain.
  • Strong understanding of agentic AI systems to design and reason about evaluations.
  • Experience with AI evaluations, including creating evaluation tasks and judging methods.
  • Comfortable working with customers, engineers, SMEs, partners, and internal teams to extract relevance from workflows.
  • Strong product judgment, turning observations into concrete evaluation work.

Responsibilities

  • Study how engineers and customers use Archie, identify successes and failures from an engineering perspective, and translate observations into actionable evaluations.
  • Recreate real-world engineering tasks and failure modes for repeatable development use.
  • Build and refine evaluation methodologies, including automated judges and human evaluation processes.
  • Assess alignment between automated judges and expert human judgment and improve them when needed.
  • Collaborate across engineering, product, customers, SMEs, and external partners to reflect real workflows in evaluations.

Skills

Physical engineering background
Agentic AI understanding
AI evaluation design
Customer collaboration
Product judgment

Job description

TL;DR:

If you:

  • have mastered extreme systems in a physical engineering domain to know what peak engineering actually looks like;
  • have built or worked with AI evaluations, including judges and human evaluation;
  • can work comfortably with customers, subject matter experts, partners, and technical teams;
  • are motivated by the goal of building superintelligence for engineering…
About P-1 AI:

At P-1 AI, we are building an AI engineer agent for the physical world named Archie. We maximize Archie’s anthropomorphism so that he fits seamlessly into existing engineering teams and workflows in the form factor of a human engineer. Archie today is at the level of a junior mechanical and electrical engineer, with a quantitative intuition over the product design space and the ability to use complex engineering tools—the same tools his human teammates use. Archie's tech stack includes a custom agentic harness, structured design representation, continual skills learning, and small custom post-trained models (SFT and RLVR) using proprietary semi-synthetic training data sets and environments which create a deep competitive moat. Our ultimate aim is to build engineering ASI. We recently announced a $50 million Series A financing led by NEA, which added former General Electric CEO Jeff Immelt to our board. The round also included the addition of several AI luminaries from Anthropic and Nominal to our existing angel investors from Google and OpenAI.

About the opportunity:

In this role, you’ll sit at the forefront of AI for physical engineering, working with subject matter experts, customers, partners, and our internal teams to challenge and evaluate Archie on tough engineering work, understand where it fails, and capture those failures in evaluations that guide product development. You need to understand not only whether outlook is plausible, but whether the underlying engineering reasoning and decisions make sense.

About the role:
  • Study how engineers and customers use Archie in the field, identify where Archie succeeds or fails from an engineering perspective, and turn those observations into actionable evaluations.
  • Recreate real-world engineering tasks and failure modes in environments that can be repeatedly used by our development teams.
  • Build and refine evaluation methodologies for Archie, including automated judges and human evaluation processes.
  • Assess whether automated judges actually align with expert human judgment and improve them when they do not.
  • Work across engineering, product, customers, SMEs, and external partners to make sure our evaluations reflect real engineering workflows rather than abstract benchmarks.
  • Use evaluation results to help the team understand where Archie needs to improve and which problems matter most to users.
About you:
  • Technical background in a physical engineering domain such as power systems, power electronics, mechanical engineering, aerospace, robotics, or a closely related field.
  • Strong enough understanding of agentic AI systems to design and reason about evaluations.
  • Experience with AI evaluations, including creating evaluation tasks, designing or crafting judges, and comparing automated evaluation against human judgment.
  • Comfortable working directly with customers, engineers, subject matter experts, partners, and internal technical teams to extract what matters from ambiguous workflows.
  • Strong product judgment, with the ability to turn observations about how a system is being used or failing into concrete evaluation work.
  • Motivated by the mission of building superintelligence for engineering and excited by the challenge of making AI systems genuinely capable in the physical world.
Location:

Remote (US/Canada) or San Mateo, CA. Remote employees spend one week out of six working together on-site in our San Mateo office. Relocation support available.

Benefits:

Competitive salary, meaningful equity ownership, healthcare, dental, vision, 401(k) match, and unlimited PTO.

Interview process:
  • Introductory call (30 mins)
  • Biographical/behavioural interview (45 mins)
  • Technical interview (60 mins)
  • CEO interview (30 mins)
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Director of Product
Director of Product

P-1 AI Inc. • Northern (KY)

On-site
USD 180,000 - 250,000
Competitive salary
Bonus
Equity ownership
+5
Director of Product
Director of Product

P-1 AI • United States

On-site
USD 180,000 - 240,000
Competitive salary
Bonus
Meaningful equity ownership
+5
Full Stack Super SWE
Full Stack Super SWE

P-1 AI Inc. • San Mateo (CA)

Hybrid
USD 180,000 - 280,000
Equity
Health care
Dental
+3
Full Stack Super SWE
Full Stack Super SWE

P-1 AI • United States

Hybrid
USD 150,000 - 190,000
Competitive salary
Meaningful equity ownership
Healthcare
+4
Full Stack Super SWE
Full Stack Super SWE

Precision Labs • Northern (KY)

Hybrid
USD 140,000 - 200,000
Competitive salary
Equity
Healthcare
+2
AI Research Scientist
AI Research Scientist

Namely • San Francisco (CA)

On-site
USD 150,000 - 210,000
Competitive salary
Meaningful equity ownership
Healthcare
+4
Customer Success Manager
Customer Success Manager

AI Chopping Block • San Mateo (CA)

On-site
USD 120,000 - 180,000
Healthcare
Dental
Vision
+3
Customer Success Manager
Customer Success Manager

P-1 Ai • United States

On-site
USD 90,000 - 130,000
Competitive base salary
Meaningful equity ownership
Healthcare
+2
AI Research Scientist - Adaptive Intelligence
AI Research Scientist - Adaptive Intelligence

P-1 AI • San Francisco (CA)

On-site
USD 120,000 - 160,000
Competitive salary
Meaningful equity ownership
Healthcare benefits
+1
AI Research Scientist/Manager - Adaptive Intelligence
AI Research Scientist/Manager - Adaptive Intelligence

P-1 AI Inc. • San Mateo (CA)

On-site
USD 120,000 - 160,000
Competitive salary
Equity ownership
Healthcare
+3