Senior Live Systems Reliability Engineer

Whatnot

Leinster

Hybrid

EUR 90,000 - 130,000

Full time

4 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Flexible time off
Health insurance options (medical, etc
Work-from-home support
Home office setup allowance
Cellphone/internet allowance
Wellness allowance
Childcare assistance
Family planning support
Retirement/pension plans
Dogfood budget
Parental leave

Job summary

Whatnot is seeking a Production Engineer to embed with product, platform, and infrastructure teams, owning reliability and performance at scale. You will hunt anomalies across traffic and latency, link signals to business impact, and drive fixes that span multiple services while preparing for peak events.

You’ll work across critical areas like payments, live video, search, and the bidding path, with strong emphasis on observability, incident response, and cross-team collaboration.

Qualifications

  • Strong systems and distributed systems fundamentals: failure modes, saturation, queueing, cascading failure.
  • Fluency with cloud-native environments (AWS/GCP) and Kubernetes, infrastructure as code.
  • Proficiency in Python/Elixir and Go for performance-sensitive infra.

Responsibilities

  • Go deep with the embed team to improve reliability, performance, and scalability.
  • Hunt anomalies across traffic, latency, error, and cost signals then trace to root cause.
  • Connect operational signals to business impact across cohorts.
  • Work on high-risk areas: payments, live video, search, bidding, security.
  • Eliminate scale bottlenecks in production and generalize fixes.
  • Prepare for peak events with capacity modeling, load testing, and drills.
  • Improve observability to catch next anomaly quickly.
  • Apply AI to operations: anomaly detection, on-call assistance, remediation.
  • Share on-call with the embedded team and act as escalation point for incidents.
  • Lead incident response for complex cross-team failures and drive systemic fixes.

Skills

Distributed systems
Cloud-native
Programming languages (Python/Elixir/"

Education

Bachelor's degree in CS or related field

Tools

SRE tooling

Job description

Whatnot is seeking a Production Engineer to embed with product, platform, and infrastructure teams, owning reliability and performance at scale. You will hunt anomalies across traffic and latency, link signals to business impact, and drive fixes that span multiple services while preparing for peak events.

You’ll work across critical areas like payments, live video, search, and the bidding path, with strong emphasis on observability, incident response, and cross-team collaboration.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Production Engineer — Real-Time Scale & Reliability
Senior Production Engineer — Real-Time Scale & Reliability

Whatnot • Dublin

Hybrid
EUR 90,000 - 150,000
Flexible Time Off
Health Insurance - options
Work From Home Support
+4
Senior Production Engineer
Senior Production Engineer

Whatnot • Leinster

Hybrid
EUR 90,000 - 130,000
Flexible time off
Health insurance options (medical, etc
Work-from-home support
+8
Senior Production Engineer
Senior Production Engineer

Whatnot • Dublin

Hybrid
EUR 90,000 - 150,000
Flexible Time Off
Health Insurance - options
Work From Home Support
+4
Staff Site Reliability Engineer: Scale & Reliability Lead
Staff Site Reliability Engineer: Scale & Reliability Lead

United States Digital Space LLC • Dublin

On-site
EUR 120,000 - 160,000
Global Benefit programs
Family Planning Support
Mental Health & Coaching Benefits
+1
Senior Performance Test Engineer
Senior Performance Test Engineer

Harvey Nash • Dublin

On-site
EUR 90,000 - 120,000
Senior Site Reliability Engineer – Automation & Resilience
Senior Site Reliability Engineer – Automation & Resilience

Mastercard • Dublin

On-site
EUR 120,000 - 180,000
Senior Performance Engineer: CI/CD & Scalability
Senior Performance Engineer: CI/CD & Scalability

Harvey Nash • Dublin

On-site
EUR 90,000 - 120,000
Site Reliability Engineer III - Eng
Site Reliability Engineer III - Eng

UKG • Leinster

On-site
EUR 60,000 - 80,000
Senior SRE: Lead Incidents, Reliability & Observability
Senior SRE: Lead Incidents, Reliability & Observability

Jobtailor • Ireland

On-site
EUR 70,000 - 120,000
Automation‑Focused Site Reliability Engineer
Automation‑Focused Site Reliability Engineer

ServiceNow, Inc. • Dublin

Hybrid
EUR 90,000 - 120,000
Holidays
Well-being days
Parental leave
+2