Founding Machine Learning Engineer

Established Search

San Francisco (CA)

On-site

USD 120,000 - 150,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Established Search is seeking a candidate to evaluate safety-critical AI systems in medical imaging. This role involves designing and executing evaluations, investigating performance across diverse settings, and developing validation methodologies.

As an early technical hire, you'll shape the product's development, automate evidence generation, and engage with stakeholders, establishing best practices for AI system evaluation.

Qualifications

  • Strong investigative skills to evaluate AI behavior in medical imaging.
  • Proficiency in Python for developing evaluation workflows.
  • Knowledge of AI safety and regulatory compliance.

Responsibilities

  • Design and execute evaluations for medical imaging AI systems.
  • Investigate failure modes and robustness limitations.
  • Develop frameworks for structuring claims and evidence.

Skills

Investigative skills
Python programming
Understanding of AI safety
Experience with medical imaging

Education

Degree in a relevant field

Tools

AI development tools

Job description

Medical Imaging AI Evaluation, Reliability & Evidence Infrastructure

About the Opportunity

My client is building the infrastructure layer for evaluating and validating safety-critical AI systems. As AI becomes increasingly embedded in clinical workflows, benchmark performance alone is no longer enough. Healthcare providers, regulators, insurers, and patients need evidence that AI systems behave reliably across real-world environments, populations, scanners, and workflow

s.This company is working with leading medical imaging AI organizations and healthcare institutions to redefine how AI validation is performed, moving beyond static testing towards continuous evidence generation and monitoring.

Their goal is to build the systems, methodologies, and tooling that allow organizations to understand how models behave in practice, identify risk, and generate defensible evidence for deployment and regulatory decisions.

The role:

This is not a traditional machine learning engineering role.

You will not spend your time simply training models or chasing benchmark improvements.

Instead, you will investigate how AI systems behave in real-world environments, determine where validation approaches break down, identify sources of risk, and help define what evidence is required to support safe deployment.

The work sits at the intersection of:

  • Medical Imaging AI
  • Evaluation Model
  • Robustness & Reliability
  • AI Safety & Validation
  • Regulatory Evidence
  • Generation Software Engineering

As one of the earliest technical hires, you will play a key role in shaping both the product and the methodology used to evaluate safety-critical AI systems.

  • Design and execute evaluations for medical imaging
  • AI systemsAnalyze performance across populations, institutions, scanners, imaging protocols, and clinical workflows
  • Investigate failure modes, robustness limitations, and generalization gaps
  • Evaluate distribution shift, demographic bias, subgroup performance, and deployment risks
  • Produce evidence that supports, challenges, or refines claims about model performance and safety

Develop AI Validation Methodology

  • Define frameworks for structuring claims, arguments and evidence
  • Determine what evidence is sufficient for deployment, regulatory, and clinical
  • decision-making
  • Challenge assumptions and identify weaknesses in existing validation approaches
  • Transform recurring investigations into repeatable workflows and reusable methodologies
  • Help establish best practices for evaluating safety-critical AI systems

Build Product & Evaluation Infrastructure

  • Write production-quality Python code supporting evaluation workflows
  • Develop reusable investigation pipelines and benchmarking frameworks
  • Build agentic workflows that automate evidence generation and analysis
  • Prototype customer-facing functionality using modern AI development tools
  • Collaborate with customers, researchers, clinicians, and regulatory stakeholders
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Founding ML Engineer for Medical Imaging Validation
Founding ML Engineer for Medical Imaging Validation

Established Search • San Francisco (CA)

On-site
USD 120,000 - 150,000
Founding Forward Deployed Machine Learning Engineer [33151]
Founding Forward Deployed Machine Learning Engineer [33151]

Stealth Startup • Sunnyvale (CA)

On-site
USD 160,000 - 220,000
0.5-2.0% Equity
Insurance
Founding Member of Technical Staff
Founding Member of Technical Staff

name • United States

On-site
USD 150,000 - 200,000
Founding team equity
Early-stage upside
AI Engineer
AI Engineer

Harnham • San Francisco (CA)

On-site
USD 100,000 - 150,000
Computer Vision Engineer
Computer Vision Engineer

DeepRec.ai • San Francisco (CA)

On-site
USD 170,000 - 250,000
Founding Forward-Deployed ML Engineer
Founding Forward-Deployed ML Engineer

United States Digital Space LLC • United States

Hybrid
USD 150,000 - 230,000
Equity stake
Founding ML Engineer – Safety-Critical AI
Founding ML Engineer – Safety-Critical AI

name • United States

On-site
USD 150,000 - 200,000
Founding team equity
Early-stage upside
Senior Consultant, AI/ML Engineer
Senior Consultant, AI/ML Engineer

Hollstadt Consulting • Minnesota

On-site
USD 140,000 - 200,000
Evals Infrastructure Tech Lead / Manager
Evals Infrastructure Tech Lead / Manager

United States Digital Space LLC • San Francisco (CA)

Hybrid
USD 500,000 - 850,000
Founding Forward-Deployed ML Engineer
Founding Forward-Deployed ML Engineer

Meyandy LLC • Sunnyvale (CA), Northern (KY)

Hybrid
USD 150,000 - 230,000
Equity opportunity
Founding team member