SRE Leader

Kontakt Micro-Location Sp. Z.o.o.

New York (NY)

Hybrid

USD 180,000 - 260,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Equity in a high-growth company
Health, dental, and vision coverage
401k
Paid time off
Parental leave
Hybrid schedule in NYC

Job summary

Kontakt.io in New York City is seeking an experienced SRE Leader to own the reliability, performance, and automation of a cloud‑based real‑time platform. You will lead resilience initiatives, scale the SRE team, and ensure 99.99% uptime for healthcare customers.

The role focuses on observability, incident response, self‑healing automation, and IaC adoption with Terraform, Kubernetes, and AWS. You’ll collaborate with product and security to meet HIPAA/SOC 2 requirements and drive the technical

Qualifications

  • 10+ years of experience in Site Reliability Engineering or Cloud Infrastructure.
  • Proven success scaling high-traffic, mission-critical platforms in SaaS/IoT/healthcare.
  • Deep expertise in cloud platforms (AWS) and distributed systems.
  • Strong background in monitoring, logging, and observability with Prometheus/OpenTelemetry/Datadog.
  • Hands-on incident management, postmortems, and building resilient systems.
  • Deep knowledge of CI/CD automation, GitOps, and infrastructure as code (Terraform).
  • A mature leadership approach with the ability to guide a high-performance SRE team.
  • Strong understanding of network security, access management, and compliance (HIPAA, SOC 2).

Responsibilities

  • Ensure 99.99% uptime across the cloud platform.
  • Design self-healing, fault-tolerant systems to prevent failures.
  • Define SLIs, SLOs, and SLAs and monitor performance.
  • Architect and manage scalable AWS cloud infrastructure for real-time data.
  • Optimize containerized environments (Kubernetes, Docker) for multi-region deployments.
  • Lead the adoption of Terraform and IaC to automate infrastructure.
  • Build and refine monitoring/logging with Prometheus, Grafana, OpenTelemetry, Datadog.
  • Lead incident response and on-call operations to reduce MTTR.
  • Conduct blameless postmortems and improve system resilience.
  • Reduce manual interventions via automated deployment and failover mechanisms.
  • Collaborate with Security & Compliance to meet HIPAA/SOC 2 standards.
  • Lead disaster recovery planning and business continuity measures.
  • Drive technical strategy and roadmap for scalability and reliability.
  • Collaborate with Product, Engineering, and Infrastructure teams to align SRE initiatives with business priorities.

Skills

Site Reliability Engineering
Cloud Infrastructure
AWS
Kubernetes
Observability
CI/CD
Terraform
GitOps
Incident management
Leadership
HIPAA/SOC 2
OpenTelemetry
Prometheus
Datadog

Tools

Docker
Terraform

Job description

We are makers and builders.
We challenge the status quo and have fun doing it.
We bring passion and purpose to making a difference in healthcare, elevating how care is delivered and experienced.

Every day, intelligent software orchestrates the physical world around us – from matching drivers with riders to optimizing global supply chains. Yet inside hospitals, where every second and every decision can affect a patient's life, operations are still spread across dozens of disconnected systems.

At Kontakt.io, we're changing that. By combining proprietary hardware, AI-powered intelligence, and deep integrations with the systems hospitals already rely on, we're creating a real-time understanding of hospital operations that software alone can't deliver. That intelligence powers the execution layer hospitals have been missing – helping care teams make smarter decisions and deliver better patient care.

Backed by Goldman Sachs and trusted by leading health systems including HCA Healthcare, Sutter Health, AdventHealth, Trinity Health, Northwell Health, Cleveland Clinic, and the U.S. Department of Veterans Affairs, we're pioneering the next generation of healthcare operations. We've more than doubled our revenue over the past year and are on track to surpass $70M in annual recurring revenue – not because we're following a market, but because we're defining one.

If you're excited to solve hard problems, work with a team of builders, and help hospitals deliver better care every day, we'd love to meet you!

We’re looking for an SRE Leader to own the reliability, performance, and automation of our cloud‑based, real‑time platform. This role will focus on keeping our platform running smoothly 24/7, minimizing downtime, improving observability, incident response, and self‑healing automation. You will lead and scale the SRE team to ensure our infrastructure stays ahead of demand, operates efficiently, and meets the needs of our growing healthcare customers.

Responsibilities
  • Ensure 99.99% uptime across our cloud platform, meeting strict SLAs for healthcare customers.
  • Design and implement self‑healing, fault‑tolerant systems to prevent failures before they happen.
  • Define SLIs, SLOs, and SLAs, ensuring proactive performance monitoring and incident resolution.
  • Architect and manage scalable cloud infrastructure (AWS) for massive real‑time data processing.
  • Optimize containerized environments (Kubernetes, Docker) to support multi‑region deployments.
  • Lead the adoption of infrastructure as code (Terraform) to fully automate infrastructure management.
  • Build and refine a world‑class monitoring, alerting, and logging system using Prometheus, Grafana, OpenTelemetry, and Datadog.
  • Lead incident response and on‑call operations, reducing mean time to detection (MTTD) and mean time to resolution (MTTR).
  • Conduct blameless postmortems and continuously improve system resilience.
  • Reduce manual intervention through automated deployment, scaling, and failover mechanisms.
  • Partner with Security & Compliance teams to ensure infrastructure meets HIPAA and SOC 2 standards.
  • Lead disaster recovery and business continuity planning to ensure critical healthcare services are always available.
  • Drive technical strategy and roadmap for scalability, monitoring, and reliability engineering.
  • Collaborate with Product, Engineering, and Infrastructure teams to align SRE initiatives with business priorities.
What You Bring
  • 10+ years of experience in Site Reliability Engineering or Cloud Infrastructure.
  • Proven success scaling high‑traffic, mission‑critical platforms in SaaS, IoT, or healthcare.
  • Deep expertise in cloud platforms (AWS), Kubernetes, and distributed systems.
  • Strong background in monitoring, logging, and observability with Prometheus, OpenTelemetry, or similar tools.
  • Hands‑on experience with incident management, postmortems, and building resilient systems.
  • Deep knowledge of CI/CD automation, GitOps, and infrastructure as code (Terraform, etc.).
  • A mature leadership approach, with the ability to drive technical strategy while growing and mentoring a high‑performance SRE team.
  • Strong understanding of network security, access management, and compliance frameworks (HIPAA, SOC 2).
Bonus Points If You Have:
  • Experience with healthcare IT, including EHR data, FHIR, and HL7 interoperability.
  • Prior experience leading on‑call rotations and major incident management processes.
Why You'll Love It Here
  • Own Mission‑Critical Reliability – Ensure hospitals and care facilities always stay online with a 99.99% uptime healthcare platform.
  • Scale AI‑Powered Infrastructure – Work on real‑time automation and self‑healing cloud systems that orchestrate care delivery.
  • Drive Big Impact in Healthcare – Help reduce waste, optimize resources, and improve patient care with technology that delivers 10X ROI.
  • Automation‑First Culture – Minimize manual ops with cutting‑edge automation, observability, and incident response strategies.
  • Join a High‑Performing Team – Work with top engineers, AI experts, and healthcare innovators solving real‑world challenges.

Built for collaboration - our team a hybrid schedule of 3 days/week minimum from our New York City office

Equity in a high‑growth company scaling toward $400M+ ARR and backed by leading investors

Full health, dental, and vision coverage, a 401k, paid time off, paid parental leave and all the tools you need to do your best work

Autonomy to solve meaningful problems with work that ships quickly and makes a difference


you are not like everybody else.
We are not like any other company.
We’d love to be part of your growth.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Lead Site Reliability Engineer
Lead Site Reliability Engineer

Kontakt.io • New York (NY)

Hybrid
USD 200,000 - 250,000
Hybrid work 3 days/week in NYC office.
Equity in a high-growth company
Health, dental, vision insurance
+1
Senior Site Reliability Engineer (SRE)
Senior Site Reliability Engineer (SRE)

Practice by Numbers • United States

On-site
USD 120,000 - 160,000
High ownership and autonomy
Strong engineering culture
Impactful work on healthcare infrastructure
Site Reliability Engineer
Site Reliability Engineer

Latent • San Francisco (CA)

On-site
USD 140,000 - 200,000
Sr. Site Reliability Engineer
Sr. Site Reliability Engineer

Practice By Numbers, Inc. • Redmond (WA)

On-site
USD 130,000 - 160,000
High impact work on healthcare infrastructure
Strong engineering culture focused on automation
Small team with high ownership and autonomy
Principal Site Reliability Engineer (SRE)
Principal Site Reliability Engineer (SRE)

Symmetrio • United States

Hybrid
USD 120,000 - 160,000
Health Care Plan (Medical, Dental & Vision)
Retirement Plan (401k, IRA)
Paid Time Off (Vacation, Sick & Public Holidays)
Site Reliability Engineer - 7 Month Contract
Site Reliability Engineer - 7 Month Contract

Orion Health group • Fort Worth (TX), Town of Texas (WI)

On-site
USD 110,000 - 170,000
Site Reliability Engineering Manager
Site Reliability Engineering Manager

NationsBenefits, LLC • Plantation (FL)

Remote
USD 140,000 - 180,000
Fully remote
Unlimited PTO
Competitive compensation
+1
Sr. Site Reliability Engineer
Sr. Site Reliability Engineer

Practice by Numbers • Bellevue (WA)

On-site
USD 120,000 - 150,000
Senior SRE Lead — Real-Time Healthcare Platform
Senior SRE Lead — Real-Time Healthcare Platform

Kontakt.io • New York (NY)

Hybrid
USD 200,000 - 250,000
Hybrid work 3 days/week in NYC office.
Equity in a high-growth company
Health, dental, vision insurance
+1
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

OutSolve • Mission (KS)

Remote
USD 90,000 - 130,000
100% remote work environment
Competitive compensation
Professional development opportunities
+1