Senior AI Reliability Engineer (Platform)

Flatiron Health

Berlin (NH)

On-site

USD 140,000 - 190,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Flatiron Health seeks a Senior AI Reliability Engineer (Platform) to scale AI-enabled workflows across engineering, product, and business teams. You will help build evaluation infrastructure, observability, and guardrails for reliable AI production systems.

You will shape SLOs/SLIs, monitor risk, and collaborate with data science, security, and cloud/platform teams to drive pragmatic AI adoption decisions. Fluency in English and strong communication are essential.

Qualifications

  • 5+ years of experience in platform engineering, SRE, machine learning, MLOps, or related field with Python production experience.
  • Experience designing experiments, evaluation frameworks, statistical analyses, and quality metrics for ML/AI systems.
  • Strong understanding of AI reliability, observability, including logging, tracing, monitoring, drift detection, and production health.

Responsibilities

  • Design and continuously improve evaluation frameworks, benchmarks, and automated testing pipelines for AI, LLM-powered, and agentic workflows.
  • Define and monitor quality, reliability, safety, performance, and cost metrics for AI systems, including observability and drift detection.
  • Develop reliability engineering practices for AI-enabled systems, including SLOs, SLIs, monitoring, incident response, runbooks, and root-cause analysis of AI failure modes.
  • Design orchestration, governance, and guardrails for multi-agent AI systems, including agent coordination, permissions, auditability, human oversight, and secure deployment patterns.
  • Partner with platform, product, security, engineering, and data science teams to evaluate AI solutions, establish reusable standards, and guide build-vs-buy decisions.
  • Support experimentation with emerging AI technologies while helping the organisation make pragmatic, scalable decisions across global teams.

Skills

Fluent in English
Strong communication

Tools

Python
AWS
Kubernetes
CI/CD
Datadog
OpenTelemetry
Databricks
Airflow
SageMaker

Job description

We’re looking for a Senior AI Reliability Engineer (Platform) to help us accomplish our mission to improve and extend lives by learning from the experience of every person with cancer. Are you ready to be the next changemaker in cancer care?

Flatiron Health is a healthtech company using data for good to power smarter care for every person with cancer, around the world. Flatiron partners with cancer centers in the US, Europe and Asia to transform patients’ real-life experiences into real-world evidence and create a more modern, connected oncology ecosystem. Our multidisciplinary teams include oncologists, data scientists, software engineers, epidemiologists, product experts and more. Flatiron Health is an independent affiliate of the Roche Group.

What You’ll Do

We’re seeking a Senior AI Reliability Engineer (Platform) to help Flatiron safely and effectively scale AI-enabled workflows across our engineering, product, and business teams. This role sits at the intersection of data science, platform engineering, AI evaluation, and production reliability.

As Flatiron’s use of AI grows, we need to move beyond experimentation and build the systems, standards, and feedback loops that allow AI workflows to be evaluated, monitored, trusted, and improved over time. Platform’s strategy is to enable AI adoption without becoming a gatekeeper: building reusable patterns, evaluation infrastructure, observability, and guardrails that help teams move quickly while managing reliability, safety, and cost.

As a Senior AI Reliability Engineer (Platform), you will:

  • Design, build, and continuously improve evaluation frameworks, benchmarks, and automated testing pipelines for AI, LLM-powered, and agentic workflows.
  • Define and monitor quality, reliability, safety, performance, and cost metrics for AI systems, including observability, drift detection, hallucination risk, retrieval quality, and end-to-end workflow behaviour.
  • Develop reliability engineering practices for AI-enabled systems, including SLOs, SLIs, monitoring, alerting, incident response, runbooks, and root-cause analysis of AI failure modes.
  • Design orchestration, governance, and guardrails for multi-agent AI systems, including agent coordination, permissions, auditability, human oversight, and secure deployment patterns.
  • Partner with platform, product, security, engineering, and data science teams to evaluate AI solutions, establish reusable standards, and guide build‑vs‑buy, model selection, and AI adoption decisions.
  • Support experimentation with emerging AI technologies while helping the organisation make pragmatic, scalable decisions in a rapidly evolving landscape, collaborating across global teams and participating in on‑call rotations.
Who You Are

You’re a senior technical practitioner with experience working across data science, machine learning, software engineering, platform engineering, or reliability engineering. You are comfortable operating in ambiguous spaces where the right answer is not always obvious, and you are motivated by turning emerging AI capabilities into production‑ready systems that teams can actually trust.

You understand that AI systems fail differently from traditional software. A model may not crash, but it may silently degrade, become less accurate, respond inconsistently, produce poor outputs, or create business risk in ways that are hard to detect without the right evaluation and observability patterns. This role is focused on that production behaviour and system health, not on pure model research or training.

You likely have:

  • 5+ years of experience in platform engineering, SRE, machine learning, MLOps or a related technical field, with strong Python skills and experience building production‑quality systems.
  • Experience designing experiments, evaluation frameworks, statistical analyses, and quality metrics for ML or AI systems, with familiarity in LLMs, RAG, AI agents, prompt evaluation, and model behaviour.
  • Strong understanding of AI reliability and observability, including logging, tracing, monitoring, drift detection, statistical analysis, uncertainty, alerting, and production system health.
  • Experience with modern cloud and ML infrastructure, including AWS, containers, Kubernetes, CI/CD, data pipelines, workflow orchestration, versioning, and distributed compute platforms.
  • Knowledge of agentic and multi‑agent systems, including orchestration, state management, tool execution, governance, reliability, human‑in‑the‑loop controls, and selecting the appropriate level of AI autonomy for a given problem.
  • Strong communication and collaboration skills, with the ability to explain AI behaviour and tradeoffs to technical and non‑technical stakeholders and thrive in a fast‑moving, ambiguous environment with a pragmatic, enablement‑focused mindset.
  • Fluent in English.
Optional
  • Experience with LLM evaluation, red‑teaming, adversarial testing, AI safety, RAG evaluation, retrieval quality measurement, embedding drift, or AI observability and model monitoring.
  • Hands‑on experience with observability and data/ML platforms such as Datadog, Splunk, OpenTelemetry, Databricks, Spark, Airflow, dbt, Ray, SageMaker, GitLab CI/CD, or similar technologies.
  • Experience working in healthcare, life sciences, or other regulated, privacy‑sensitive environments.
Who We Are

Our people are at the center of everything we do. We strive to foster a culture where our teammates feel equipped and empowered to make meaningful contributions with confidence, compassion, and clarity.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior AI Reliability Engineer — Platform & Observability
Senior AI Reliability Engineer — Platform & Observability

Flatiron Health • Berlin (NH)

On-site
USD 140,000 - 190,000
AI Tech Lead Manager (Platform)
AI Tech Lead Manager (Platform)

Flatiron Health • Berlin (NH)

On-site
USD 159,853 - 216,943
Senior Platform Engineer / Software Development / AI
Senior Platform Engineer / Software Development / AI

HAAR Recruitment • Nashville (TN)

On-site
USD 90,000 - 120,000
Senior Member of Technical Staff
Senior Member of Technical Staff

DeepRec.ai • Palo Alto (CA)

Hybrid
Competitive salary + meaningful equity
Flexible hybrid working model
High-performing, collaborative team environment
Senior AI Platform Engineer
Senior AI Platform Engineer

FactualIQ • Indiana (PA)

On-site
USD 180,000 - 240,000
VP of Engineering
VP of Engineering

Obin AI • New York (NY)

On-site
USD 140,000 - 180,000
Senior Engineering Manager / Platform Tech Lead
Senior Engineering Manager / Platform Tech Lead

Obin AI • New York (NY)

On-site
USD 150,000 - 200,000
Applied AI Data Scientist - Product AI Team
Applied AI Data Scientist - Product AI Team

Flatiron Health • New York (NY)

On-site
USD 110,000 - 150,000
Flexible hours
Comprehensive compensation
401(k)
+2
Director, AI Platform Engineering
Director, AI Platform Engineering

NYC Health + Hospitals • New York (NY)

On-site
USD 150,000 - 200,000
Senior Consultant, AI/ML Engineer
Senior Consultant, AI/ML Engineer

Hollstadt Consulting • Minnesota

On-site
USD 140,000 - 200,000