Director, Platform Engineering

Danaher

Kraków

On-site

PLN 400,000 - 560,000

Full time

10 days ago
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Danaher, hosted by Cytiva in Kraków, Poland, seeks a Senior Director to own end-to-end reliability of the AI platform. You will lead DevOps, CI/CD, and incident response for large-scale ML workloads, partnering with AI teams across the business to ensure availability and performance.

You will shape the reliability strategy for LLM-driven systems, establish SLOs, guardrails, and scale the platform while managing cross-functional teams across time zones in a corporate setting.

Qualifications

  • Bachelor's degree in Computer Science, Engineering, or related field; advanced degree welcome.
  • 10+ years in engineering with production reliability ownership and large-scale platforms.

Responsibilities

  • Own end-to-end platform reliability for AI compute and cloud stack.
  • Lead incident management lifecycle and establish SLOs and on-call practices.
  • Design and operate GPU/accelerator fleets and data pipelines for ML workloads.
  • Own deployment, runtime health, and operations between infra and apps.
  • Drive CI/CD, IaC, and release engineering to ensure rapid, safe changes.
  • Collaborate with AI teams as internal customers and scale capabilities.

Skills

Platform reliability
Incident management
SRE ownership
Observability
Cloud architecture
CI/CD
Automation
Stakeholder communication

Education

Bachelor's degree in CS/Engineering

Tools

Docker
Kubernetes
Terraform
Pulumi

Job description

Bring more to life. At Danaher, our work saves lives. And each of us plays a part. Fueled by our culture of continuous improvement, we turn ideas into impact - innovating at the speed of life.

Our 60,000+ associates work across the globe at more than 15 unique businesses within life sciences, diagnostics, and biotechnology.

Are you ready to accelerate your potential and make a real difference? At Danaher, you can build an incredible career at a leading science and technology company, where we’re committed to hiring and developing from within. You’ll thrive in a culture of belonging where you and your unique viewpoint matter.

Learn about the Danaher Business System which makes everything possible.

The Director, Platform Engineering role is responsible for end-to-end reliability of the platform behind our AI initiatives via accelerated compute, cloud infrastructure, DevOps and release engineering, application support, and incident response. Your customers are the teams building AI into consequential work across the company: molecular design, autonomous labs, supply chain, professional services, and more. You will keep their systems available and performant today, and define what reliability looks like as workloads become increasingly LLM-driven and agentic.

This position reports to the Senior Director, Data and AI Platform, part of the Chief Information Officer (CIO) office and will be located onsite in Krakow, Poland.

This is a Danaher Corporate role, hosted by our Cytiva operating company in Krakow.

In This Role, You Will Have The Opportunity To
  • End-to-end platform reliability. Own availability, performance, and recovery for the compute and cloud platform underpinning the company's AI initiatives - you are accountable for whether the systems our AI teams depend on are working, not just for the components your team builds.
  • Incident management and response. Lead the incident lifecycle across the platform: detection, triage, mitigation, and post-incident review. Establish SLOs, error budgets, and on-call practices that hold across a diverse set of workloads and drive measurable reduction in time-to-detect and time-to-mitigate.
  • Compute and cloud infrastructure for AI workloads. Design and operate the GPU/accelerator fleet, scheduling and capacity management, storage, and networking that support large-scale training, inference, and simulation - balancing utilization, cost, and the reliability expectations of production-facing teams.
  • Application support and operations. Own the operational layer between infrastructure and the applications running on it: deployment, configuration, runtime health, and the support model that gets AI teams unblocked quickly when something breaks.
  • DevOps and delivery pipelines. Own CI/CD, infrastructure as code, environment management, and release engineering so that changes ship rapidly and safely, and so that reliability is enforced in the pipeline rather than discovered in production.
  • Serving internal AI customer teams. Treat corporate-wide AI initiatives - molecular design, autonomous labs, supply chain, professional services, and others - as your customers. Gather their reliability and capacity requirements, translate friction into scoped platform work, and prove impact through before/after measurement.
  • Building toward agentic systems. Set the reliability strategy for the next generation of LLM-based and agentic workloads: observability and evaluation for non-deterministic systems, guardrails and governance for agents acting in production, and the infrastructure patterns that make agentic execution safe, traceable, and dependable at scale.
The Essential Requirements Of This Job Include
  • Education: Bachelor's degree in Computer Science, Engineering, or a related technical field, or equivalent practical experience; advanced degree welcome but not required.
  • Functional experience: 10+ years in engineering, with significant time owning production reliability, SRE, or platform operations for large-scale distributed systems, and demonstrated ownership of reliability outcomes at organizational scale - you can point to concrete before/after results in availability, MTTR, incident volume, or error-budget performance from programs you led.
  • Leadership experience: Proven success leading and scaling multidisciplinary engineering organizations, including platform tech leads and senior engineers across distributed time zones - recruiting, coaching, setting technical direction, and holding a team accountable for operational outcomes.
  • Deep, hands-on experience running incident management for business-critical systems: on-call design, escalation paths, blameless post-incident review, and the follow-through that turns findings into permanent fixes.
  • Production experience with major cloud platforms (AWS, Azure, or GCP), containerization and orchestration (Docker, Kubernetes), and infrastructure as code (Terraform / Pulumi / etc), operating infrastructure for ML/AI workloads at scale - accelerated compute, distributed training or high-throughput inference, job scheduling, and the associated capacity and cost management.
  • Strong observability expertise: metrics, logging, tracing, and SLO instrumentation, plus a track record of driving toil reduction and automation rather than adding headcount to absorb operational load.
  • Excellent written and verbal communication, and the ability to work with demanding internal customers - negotiating trade-offs, setting expectations, and representing platform reliability to senior leadership.
Travel, Motor Vehicle Record & Physical/Environment Requirements
  • Ability to travel – up to 20%
It would be a plus if you also possess previous experience in:
  • Hands-on experience with LLM and agentic systems in production - inference serving, evaluation and guardrails, or observability for non-deterministic workloads.
  • Familiarity with scientific or research computing environments (HPC, lab automation, instrument data pipelines) or other domains where AI workloads sit close to physical systems.
  • Experience operating in compliance-constrained or isolated environments (SOC 2, FedRAMP, GxP, or similar), or across multiple clouds and on-premise infrastructure.

Together, we’ll accelerate the real-life impact of tomorrow’s science and technology. We partner with customers across the globe to help them solve their most complex challenges, architecting solutions that bring the power of science to life.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior DevOps Engineer
Senior DevOps Engineer

Danaher • Kraków

On-site
PLN 180,000 - 260,000
Site Reliability Architect
Site Reliability Architect

Danaher Corporation • Kraków

Remote
PLN 260,000 - 420,000
Remote work eligibility
Competitive benefits
Lead AI SRE and QA Engineer
Lead AI SRE and QA Engineer

Danaher Corporation • Kraków

On-site
PLN 200,000 - 350,000
Staff Site Reliability Engineer
Staff Site Reliability Engineer

Cytiva • Kraków

On-site
PLN 260,000 - 380,000
Remote work arrangement
Lead AI Scientist
Lead AI Scientist

Danaher Corporation • Kraków

On-site
PLN 180,000 - 240,000
Principal AI Engineering Excellence
Principal AI Engineering Excellence

Danaher Corporation • Kraków

On-site
PLN 300,000 - 420,000
Director of AI Platform Reliability & Cloud Ops
Director of AI Platform Reliability & Cloud Ops

Danaher • Kraków

On-site
PLN 400,000 - 560,000
Principal AI Full Stack Engineer
Principal AI Full Stack Engineer

Danaher Corporation • Kraków

On-site
PLN 450,000 - 650,000
Director AI Architecture, Standardization & Engineering Excellence
Director AI Architecture, Standardization & Engineering Excellence

Danaher Corporation • Kraków

On-site
PLN 260,000 - 380,000
Senior AI Application Support Engineer
Senior AI Application Support Engineer

Danaher Corporation • Kraków

On-site
PLN 216,000 - 264,000