Sr. Data Platform Engineer

tripstack

Toronto

On-site

CAD 140,000 - 170,000

Full time

14 days+
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Benefits offered by this job

AI in hiring process

Job summary

Tripstack is seeking a senior infrastructure engineer to migrate its data stack from bare-metal VMs to Kubernetes on OpenStack. You will own the day‑to‑day health of the data platform, managing Druid, Spark, Redpanda, Airflow, PostgreSQL, Elasticsearch end‑to‑end.

You will design stateful patterns, codify migrations with IaC, and partner with SRE on hardware and networking. Strong English and time-zone collaboration are essential.

Qualifications

  • Strong Kubernetes experience with stateful workloads (StatefulSets, PVCs, rolling upgrades) and real migrations.
  • Infrastructure as Code at a senior level (Terraform, Helm/Kustomize) and GitOps tooling.
  • Observability and RCA discipline (Prometheus, Grafana, distributed tracing, SLOs, error budgets).
  • Production operations experience with Druid, Kafka/Redpanda, Spark or Elasticsearch, on-prem or self-hosted.
  • 7+ years building and operating production data/platform systems; on-call experience.
  • Excellent written and verbal English; comfortable collaborating across multiple time zones (Toronto, Kraków, Pune, Stockholm).

Responsibilities

  • Lead the data-stack migration from bare-metal VMs to Kubernetes on OpenStack with zero downtime for stateful workloads.
  • Design StatefulSet, PVC, PodDisruptionBudget, and rolling upgrade patterns for production data systems.
  • Codify migration with IaC (Terraform, Helm/Kustomize; GitOps via ArgoCD/Flux) for reproducibility.
  • Own operational health of Druid, Spark, Redpanda, Airflow, PostgreSQL, and Elasticsearch; tune JVMs, ingestion, replication, and partitioning.
  • Build KPIs, alerting, dashboards, and runbooks to detect cluster stress before incidents.
  • Lead query, report, segment, and tiering optimisations for cost-effective analytics.
  • Build Prometheus, Grafana, and tracing coverage; enforce SLOs and post-incident discipline.
  • Collaborate with SRE on hardware, networking, and Kubernetes fundamentals; own data applications end-to-end.

Skills

Kubernetes
Stateful workloads
Terraform
Helm
Kustomize
GitOps
ArgoCD
Flux
Prometheus
Grafana
Distributed tracing
SLOs
Error budgets

Tools

Terraform
Helm
Kustomize
OpenStack
ArgoCD
Flux

Job description

About Tripstack

Founded in Toronto, Canada in 2016, Tripstack has been part of Etraveli Group since 2019. It is a B2B Flights as a Service provider and a world leader in virtual interlining. Operating from offices in Canada, India, and Poland, Tripstack is the gateway into Etraveli Group’s world leading tech platform – giving partners access to global flight content, virtual interlining, and a full suite of services including payments, fraud prevention, pricing, and customer support.

Its technology ingest over 30B price points and handles over 240 million searches daily.

Through partnerships with airlines, OTAs, and other distribution channels across the globe, Tripstack expands networks, drives new revenue streams, and offers more choice at competitive prices, all backed by robust technology and traveler protection.

The Role

Tripstack is moving its entire data stack from bare‑metal VMs to Kubernetes on OpenStack in a new data centre. We are looking for a senior infrastructure engineer who has done stateful migrations before, who can plan, codify, and execute this one safely, and who will own the day‑to‑day operational health of the data platform once we are there. This is a hands‑on platform and SRE role with a clear, time‑bounded mission. You will partner closely with our SRE team on networking, hardware, and Kubernetes fundamentals, and own the data applications — Druid, Spark, Redpanda, Airflow, PostgreSQL, Elasticsearch — end‑to‑end. It is not an ML role. We have a separate plan for evolving our MLFlow platform, and the right hire here may grow into more of that work over time, but day‑one impact is the migration and the operational health of the platform.

Responsibilities
Lead the Data‑Stack Migration
  • Plan and execute the migration of Druid, Spark, Redpanda, and our orchestration layer from bare‑metal VMs to Kubernetes on OpenStack, with no downtime on stateful workloads.
  • Design StatefulSet, PVC, pod‑disruption‑budget, and rolling‑upgrade patterns that are safe for production data systems.
  • Codify the migration with Infrastructure as Code — Terraform for OpenStack, Helm or Kustomize for Kubernetes, GitOps via ArgoCD or Flux — so the result is reproducible and supportable by the whole team.
Operate the Data Platform
  • Own the operational health of Druid, Spark, Redpanda, Airflow, PostgreSQL, and Elasticsearch as production systems — segment lifecycle, JVM tuning, ingestion specs, broker/coordinator/overlord internals, partition design, consumer lag, replication tuning.
  • Build the KPIs, alerting, dashboards, and runbooks that let us see cluster exhaustion before it becomes an incident, and diagnose it quickly when it does.
  • Own the query, report, segment, and tiering optimisations that keep our analytics cost‑effective and responsive under load.
Raise the Bar on Observability and Reliability
  • Build the Prometheus, Grafana, and distributed‑tracing coverage our data systems need. Treat SLOs, error budgets, and post‑incident discipline as table stakes.
  • Partner with SRE on hardware, networking, and Kubernetes fundamentals — but own the data applications themselves end‑to‑end.
Requirements
  • Strong Kubernetes experience with stateful workloads — StatefulSets, PVCs, pod disruption budgets, and rolling upgrades for data systems. You have done a real stateful migration before and can talk through what went wrong.
  • Infrastructure as Code at a senior level — Terraform, Helm or Kustomize, GitOps with ArgoCD or Flux. You have shipped production infrastructure this way, not just experimented with it.
  • Observability and RCA discipline — Prometheus, Grafana, distributed tracing, SLOs, error budgets, and the habit of writing the runbook that stops the next incident.
  • Production operations experience with at least one of Apache Druid, Apache Kafka or Redpanda, Apache Spark, or Elasticsearch — deep enough to be credible on internals and willing to learn the others.
  • 7+ years building and operating production data or platform systems, at least 2 of them on self‑hosted or bare‑metal infrastructure. You have been on‑call for what you built.
  • Clear written and verbal English; comfortable working across Kraków, Toronto, Pune, and Stockholm time zones.
Additional Experience That Would Be Considered an Asset
  • Deep Apache Druid production operations — segment lifecycle, JVM tuning, ingestion spec authoring, and RCA on broker/coordinator‑class failures.
  • Apache Spark at scale — DAG execution, shuffle optimisation, memory tuning, Spark‑on‑Kubernetes.
  • OpenStack familiarity — Neutron networking, Cinder/Ceph storage.
  • Fluency with agentic coding tools (Claude Code, Gemini, or equivalent) as a real part of your workflow — including the judgement to know when not to trust an AI‑generated config for a production data system.
  • Python that is production‑grade, plus working knowledge of at least one JVM or Go‑family language so you can integrate with our services directly.
  • Security and secrets management in Kubernetes — Vault, network policies, encryption at rest.
  • Change Data Capture patterns, especially PostgreSQL‑to‑Druid streaming.
  • Delta Lake or Apache Iceberg experience and architectural judgement about when to introduce a table format.
  • Exposure to travel, flights, or large‑scale search and cache systems.
  • Interest in growing into ownership of our MLFlow‑based ML platform over time. We have a separate plan for that work, but a curious operator is welcome.
Compensation

Canada – Toronto Office: 140,000 – 170,000 CAD / Annual.

The pay range shown is based on our compensation structure in place at the time of posting and may be updated periodically based on business needs. Individual pay is based on additional factors including job‑related skills, experience, and relevant education and/or training.

The targeted pay range listed reflects the base pay only and does not include bonus, or other benefits.

We use AI in our hiring process.

Benefits

We offer an opportunity to work with a young, dynamic, and a growing team composed of high‑caliber professionals. We value professionalism and promote a culture where individuals are encouraged to do more and be more. If you feel you share our passion for excellence, and growth, then look no further. We have an ambitious mission, and we need a world‑class team to make it a reality.

At Tripstack, we proudly believe in embracing diversity. This is true for our team, clients, communities and stakeholders. We are an equal opportunity employer and committed to creating a safe, healthy and accessible environment. We encourage applications regardless of race, colour, ancestry, religion, sex, national origin, sexual orientation, age, citizenship, marital status, disability, gender identity or any other grounds protected by law. Please let us know if you need any accommodations during any part of the recruitment process.

Tripstack thanks all applicants for their interest, however only those selected to continue in the process will be contacted.

Learn more about us at www.tripstack.com

#tripstack

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

HR Operations Specialist
HR Operations Specialist

tripstack • Toronto

Hybrid
CAD 55,000 - 70,000
Senior Data Engineer
Senior Data Engineer

Jobgether • Canada

On-site
CAD 120,000 - 160,000
100% remote within the Americas
One-time USD 500 home-office setup
USD 150 monthly stipend
Sr Software Developer (Full-Stack)
Sr Software Developer (Full-Stack)

TripArc • Toronto

On-site
CAD 115,000 - 125,000
Senior People Scientist, Talent Management
Senior People Scientist, Talent Management

StackAdapt • Ottawa

On-site
CAD 100,000 - 138,000
Competitive salary
401K / Pension savings globally
Paid time off
+4
Senior People Scientist, Talent Management
Senior People Scientist, Talent Management

StackAdapt • Vancouver

On-site
CAD 100,000 - 138,000
Healthcare benefits
401K / Pension
Paid time off
+3
Senior Staff People Technologist New
Senior Staff People Technologist New

StackAdapt Inc. • Canada

Hybrid
CAD 118,000 - 163,000
Retirement plan
PTO incl. birthday off
Mental health care program
+3
Cloud Infrastructure & DevSecOps Engineer
Cloud Infrastructure & DevSecOps Engineer

InfoTrack Canada • Toronto

Hybrid
CAD 130,000 - 150,000
Hybrid work model
Senior/Staff Machine Learning Engineer
Senior/Staff Machine Learning Engineer

StackAdapt • Canada

On-site
CAD 100,000 - 150,000
Highly competitive salary
Retirement/401K/Pension Savings
Paid time off including birthday days
+10
Senior Software Engineer, Data Pipelines
Senior Software Engineer, Data Pipelines

United States Digital Space LLC • Toronto

Hybrid
CAD 136,000 - 170,000
Extended health & dental
Mental health benefits
Family building benefits
+6
Product Data Analyst, Mobile App User Acquisition
Product Data Analyst, Mobile App User Acquisition

StackAdapt Inc. • Canada

Hybrid
CAD 76,000 - 105,000
Health benefits
Remote work options
WeWork access
+2