Senior SRE: Shape Reliability for Trading Platform

tastytrade, Inc.

Chicago (IL)

Hybrid

USD 180,000 - 200,000

Full time

12 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Performance Bonuses
Stock Purchase Options
401k Plan
20 Paid Vacation Days (plus an extra)
10 Paid Sick Days
Pet Insurance
Wellness & Mental Health Programs
Charitable Donation Matching
Daily catered lunch in office
Full kitchen with snacks
Shuttle to Metra

Job summary

tastytrade, Inc. in Chicago, IL is seeking a Senior Site Reliability Engineer to define reliability standards and build a scalable observability culture across Ruby, Java, and Elixir services.

You will embed with infrastructure and application teams, shaping a production readiness practice on a HashiCorp Nomad-based fabric. You will implement SLOs, error budgets, and instrumentation using OpenTelemetry, Prometheus, and Grafana, while mentoring engineers and driving fault-injection exercises.

Qualifications

  • Production-quality coding in Ruby and/or Java, with Python for automation.
  • A track record embedding SRE practices within engineering teams, including SLOs and burn-rate alerting.
  • Hands-on experience with OpenTelemetry, Prometheus, and Grafana for instrumentation.
  • Strong Linux internals and networking fundamentals (TCP/IP, UDP/multicast, packet capture).
  • On-call experience on production systems and ability to lead blameless post-incident reviews.
  • Experience with HashiCorp Nomad, Consul, or Vault is a plus.

Responsibilities

  • Define customer-meaningful SLOs and set error budgets with multi-window burn-rate alerting for critical brokerage flows.
  • Author tastytrade's first reliability standards including SLO methodology and observability guide.
  • Contribute reliability patterns (circuit breakers, retries with backoff, bulkheads, load-shedding) into services.
  • Extend observability stack and guide teams across Nomad-based service fabric.
  • Design and run tabletop exercises and fault-injection testing to stress-test the platform.
  • Mentor engineers to build a culture of site reliability champions.

Skills

Ruby
Java
Python
SRE practices
OpenTelemetry
Prometheus
Grafana
Linux internals
Networking
TCP/IP
On-call experience
Nomad
Consul
Vault

Tools

Prometheus
Grafana
OpenTelemetry
Nomad
Consul
Vault

Job description

tastytrade, Inc. in Chicago, IL is seeking a Senior Site Reliability Engineer to define reliability standards and build a scalable observability culture across Ruby, Java, and Elixir services.

You will embed with infrastructure and application teams, shaping a production readiness practice on a HashiCorp Nomad-based fabric. You will implement SLOs, error budgets, and instrumentation using OpenTelemetry, Prometheus, and Grafana, while mentoring engineers and driving fault-injection exercises.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior SRE: Shape Reliability for Real-Time Trading Platform
Senior SRE: Shape Reliability for Real-Time Trading Platform

Tastyworks • Chicago (IL), Northern (KY)

Hybrid
USD 180,000 - 200,000
Performance Bonuses
Stock Purchase Options
401k Plan
+6
Senior SRE: Build Reliability & Observability (Hybrid)
Senior SRE: Build Reliability & Observability (Hybrid)

tastytrade • Chicago (IL)

Hybrid
USD 180,000 - 200,000
Performance Bonuses
Stock Purchase Options
401k Plan
+7
Senior SRE — Build a Reliable Trading Platform
Senior SRE — Build a Reliable Trading Platform

tastylive • Chicago (IL)

Hybrid
USD 180,000 - 200,000
Performance Bonuses
Stock Purchase Options
401k Plan
+8
Senior SRE: Trading-Scale Reliability & Automation
Senior SRE: Trading-Scale Reliability & Automation

Acquire Me • New York (NY)

On-site
USD 180,000 - 240,000
Market-leading compensation
Growth-focused engineering culture
Direct impact on trading infra
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Acquire Me • New York (NY)

On-site
USD 180,000 - 240,000
Market-leading compensation
Growth-focused engineering culture
Direct impact on trading infra
Senior SRE & Platform Engineer — Observability & Automation
Senior SRE & Platform Engineer — Observability & Automation

Techunting • United States

On-site
USD 120,000 - 150,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Mission Staffing • New York (NY)

Hybrid
USD 140,000 - 200,000
Senior SRE Lead: Reliability, Observability & Automation
Senior SRE Lead: Reliability, Observability & Automation

Jobtailor • Arizona

On-site
USD 140,000 - 230,000
Senior Application SRE: Reliability & Automation Leader
Senior Application SRE: Reliability & Automation Leader

Integral Development Corp. • Palo Alto (CA)

On-site
USD 120,000 - 160,000
Competitive compensation
Benefits package
Career growth opportunities
Senior SRE: Platform Resilience & Observability (Java)
Senior SRE: Platform Resilience & Observability (Java)

TechDigital Group • San Leandro (CA)

On-site
USD 100,000 - 150,000