Senior AI Reliability Engineer (Platform)

Flatiron Health

Berlin

Vor Ort

EUR 100.000 - 150.000

Vollzeit

14 Tage+

Erhalte mehr Antworten von Arbeitgebern

Versende in nur wenigen Minuten einen passgenauen Lebenslauf.

Benefits dieser Stelle

Flexible work hours
Paid time off
Mental well-being tools and services
401(k) contribution
Parental benefits and policies

Zusammenfassung

Flatiron Health is seeking a Senior AI Reliability Engineer (Platform) to scale AI-enabled workflows across engineering, product, and business teams. You’ll design evaluation frameworks, establish observability, and guide release patterns for safe, scalable AI systems.

The role combines data science, platform engineering, AI evaluation, and production reliability, with a focus on SLOs/SLIs, monitoring, and incident response in healthcare-focused environments.

Qualifikationen

  • 5+ years of experience in platform engineering, SRE, ML, or MLOps with strong Python skills.
  • Experience designing experiments, evaluation frameworks, LLMs, RAG, AI agents, prompt evaluation, and model behaviour.
  • Fluent in English with strong communication and collaboration skills.

Aufgaben

  • Design evaluation frameworks, benchmarks, and automated testing pipelines for AI and AI-enabled workflows.
  • Define and monitor quality, reliability, safety, performance, and cost metrics for AI systems, including observability and drift detection.
  • Develop reliability practices for AI-enabled systems: SLOs, SLIs, monitoring, incident response, and root-cause analysis of AI failure modes.
  • Design governance and guardrails for multi-agent AI systems, including human oversight and secure deployment patterns.
  • Collaborate with platform, product, security, engineering, and data science teams to guide build-vs-buy decisions and AI adoption.

Kenntnisse

Python
SRE
Machine Learning
MLOps
Observability
Cloud platforms
English fluency
Communication

Tools

Datadog
Splunk
OpenTelemetry
Databricks
Spark
Airflow
dbt
Ray
SageMaker
GitLab CI/CD

Jobbeschreibung

  • We’re seeking a Senior AI Reliability Engineer (Platform) to help Flatiron safely and effectively scale AI-enabled workflows across our engineering, product, and business teams
  • This role sits at the intersection of data science, platform engineering, AI evaluation, and production reliability
  • As Flatiron’s use of AI grows, we need to move beyond experimentation and build the systems, standards, and feedback loops that allow AI workflows to be evaluated, monitored, trusted, and improved over time
  • Platform’s strategy is to enable AI adoption without becoming a gatekeeper: building reusable patterns, evaluation infrastructure, observability, and guardrails that help teams move quickly while managing reliability, safety, and cost
  • Design, build, and continuously improve evaluation frameworks, benchmarks, and automated testing pipelines for AI, LLM-powered, and agentic workflows
  • Define and monitor quality, reliability, safety, performance, and cost metrics for AI systems, including observability, drift detection, hallucination risk, retrieval quality, and end-to-end workflow behaviour
  • Develop reliability engineering practices for AI-enabled systems, including SLOs, SLIs, monitoring, alerting, incident response, runbooks, and root-cause analysis of AI failure modes
  • Design orchestration, governance, and guardrails for multi-agent AI systems, including agent coordination, permissions, auditability, human oversight, and secure deployment patterns
  • Partner with platform, product, security, engineering, and data science teams to evaluate AI solutions, establish reusable standards, and guide build-vs-buy, model selection, and AI adoption decisions
  • Support experimentation with emerging AI technologies while helping the organisation make pragmatic, scalable decisions in a rapidly evolving landscape, collaborating across global teams and participating in on-call rotations
Benefits
  • Work/life autonomy via flexible work hours and flexible paid time off
  • Financial health resources including 1:1 financial advice
  • Mental well-being tools and services
  • 401(k) contribution to help you reach your retirement planning goals
  • Parental benefits and policies including family-building care and generous leave
  • Path to parenthood programs supporting fertility, adoption and surrogacy
  • Travel support for safe healthcare services

We’re looking for a Senior AI Reliability Engineer (Platform) to help us accomplish our mission to improve and extend lives by learning from the experience of every person with cancer. Are you ready to be the next changemaker in cancer care?

A model may not crash, but it may silently degrade, become less accurate, respond inconsistently, produce poor outputs, or create business risk in ways that are hard to detect without the right evaluation and observability patterns.

You are comfortable operating in ambiguous spaces where the right answer is not always obvious, and you are motivated by turning emerging AI capabilities into production-ready systems that teams can actually trust. You understand that AI systems fail differently from traditional software.

You’re a senior technical practitioner with experience working across data science, machine learning, software engineering, platform engineering, or reliability engineering. This role is focused on that production behaviour and system health, not on pure model research or training.

Experience with modern cloud and ML infrastructure, including AWS, containers, Kubernetes, CI/CD, data pipelines, workflow orchestration, versioning, and distributed compute platforms. Strong understanding of AI reliability and observability, including logging, tracing, monitoring, drift detection, statistical analysis, uncertainty, alerting, and production system health.

5+ years of experience in platform engineering, SRE, machine learning, MLOps or a related technical field, with strong Python skills and experience building production-quality systems.

Experience designing experiments, evaluation frameworks, statistical analyses, and quality metrics for ML or AI systems, with familiarity in LLMs, RAG, AI agents, prompt evaluation, and model behaviour. Fluent in English. Strong communication and collaboration skills, with the ability to explain AI behaviour and tradeoffs to technical and non-technical stakeholders and thrive in a fast-moving, ambiguous environment with a pragmatic, enablement-focused mindset.

Knowledge of agentic and multi-agent systems, including orchestration, state management, tool execution, governance, reliability, human-in-the-loop controls, and selecting the appropriate level of AI autonomy for a given problem.

Experience with LLM evaluation, red-teaming, adversarial testing, AI safety, RAG evaluation, retrieval quality measurement, embedding drift, or AI observability and model monitoring.

Hands-on experience with observability and data/ML platforms such as Datadog, Splunk, OpenTelemetry, Databricks, Spark, Airflow, dbt, Ray, SageMaker, GitLab CI/CD, or similar technologies. Experience working in healthcare, life sciences, or other regulated, privacy-sensitive environments.

Hol dir deinen kostenlosen, vertraulichen Lebenslauf-Check.
oder ziehe deine Datei hierhin.
Similar jobs

Ähnliche Jobs, die dir auch gefallen könnten

AI Tech Lead Manager (Platform)
AI Tech Lead Manager (Platform)

Flatiron Health • Berlin

Vor Ort
EUR 140.000 - 190.000
Flexible work hours
Flexible PTO
401(k) contribution
Principal ML Platform Engineer Europe
Principal ML Platform Engineer Europe

SLAMcore • Deutschland

Vor Ort
EUR 70.000 - 100.000
Senior AI Engineer - Agentic AI Evaluation
Senior AI Engineer - Agentic AI Evaluation

Resaro • München

Vor Ort
EUR 90.000 - 140.000
Senior AI Engineer - Agentic AI Evaluation Brain Team · Munich, Singapore ·
Senior AI Engineer - Agentic AI Evaluation Brain Team · Munich, Singapore ·

Resaro • München

Hybrid
EUR 90.000 - 130.000
Director of AI Operations – Governance
Director of AI Operations – Governance

Jobtailor • Deutschland

Hybrid
EUR 120.000 - 180.000
Senior Applied AI Engineer (all genders)
Senior Applied AI Engineer (all genders)

Accenture DACH • Kronberg im Taunus

Vor Ort
EUR 90.000 - 130.000
Senior Software Engineer- AI Platform(US)
Senior Software Engineer- AI Platform(US)

Embedded Shishya • Deutschland

Hybrid
EUR 70.000 - 110.000
Senior Generative AI Operations (GenAI Ops) Engineer
Senior Generative AI Operations (GenAI Ops) Engineer

EPAM Systems • Deutschland

Hybrid
EUR 70.000 - 90.000
VP Engineering
VP Engineering

Jobtailor • Berlin

Vor Ort
EUR 130.000 - 190.000
Senior Applied AI Engineer (all genders)
Senior Applied AI Engineer (all genders)

United States Digital Space LLC • Berlin

Hybrid
EUR 90.000 - 140.000
Hybrid work model
Competitive rewards & bonuses
Employee share purchase opportunities
+1