Senior Site Reliability Engineer - Linux Systems & Application Observability

tastytrade

Chicago (IL)

Hybrid

USD 180,000 - 200,000

Full time

10 hours ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

Performance Bonuses
Stock Purchase Options
Medical/Vision/Dental Benefits
401k Plan
20 Paid Vacation Days
10 Paid Sick Days
Gym Membership Reimbursement
Commuter Benefits
Pet Insurance
Wellness & Mental Health Programs
Charitable Donation Matching
Two Paid Volunteer Days Off
Daily catered lunch when in the office
Full kitchen with snacks and beverages
In-building gym
Shuttle to/from Metra

Job summary

tastytrade, part of IG Group, is hiring a Senior Site Reliability Engineer to harden our Linux-based infrastructure and observability. You’ll own scalability for Nomad service fabric, extend OpenTelemetry, Prometheus, and Grafana instrumentation, and set SLOs with burn-rate alerts.

You’ll work alongside infrastructure and application teams to keep trading platforms reliable and fast. You’ll mentor engineers, help design fault-tolerant systems, and contribute during on-call rotations with a

Qualifications

  • Hands-on experience designing fault-tolerant, self-healing distributed systems with real deployments.
  • Deep understanding of distributed systems, Linux, cloud-native architectures, and containers.
  • Experience running gap analyses on observability/telemetry systems to close blind spots.
  • Track record scaling systems under production load, incl. capacity planning and bottleneck identification.
  • Hands-on with OpenTelemetry, Prometheus, and Grafana, including instrumentation.
  • Strong Linux internals, networking (TCP/IP, UDP, multicast), and packet capture.
  • On-call experience and blameless post-incident reviews.
  • Familiarity with SLOs and error budgets; Nomad/Consul/Vault is a plus.
  • Strong programming skills in Python, Ruby, Java or similar.

Responsibilities

  • Build self-healing, fault-tolerant infrastructure and internal tooling to reduce toil for Platform and Application teams.
  • Perform gap analyses across observability stack to surface hidden issues before customers are affected.
  • Own scalability work across HashiCorp Nomad service fabric: capacity planning, load testing, bottleneck identification.
  • Extend observability stack (Prometheus, Honeycomb, OpenTelemetry) with instrumentation for failure modes.
  • Set SLOs and burn-rate alerting for critical brokerage flows after establishing fault-tolerance foundations.
  • Mentor engineers to grow a culture of site reliability champions across teams.

Skills

Distributed systems
Linux systems
Cloud-native architectures
Containerization
Observability/telemetry
OpenTelemetry
Prometheus
Grafana
SLOs and error budgets
HashiCorp Nomad
Consul/Vault (HashiCorp)
Programming: Python
Programming: Ruby/Java

Tools

Nomad
Consul
Vault
OpenTelemetry
Prometheus
Grafana

Job description

Company Name: tastytrade

Role: Senior Site Reliability Engineer - Linux Systems & Application Observability

Location: Chicago, IL (Hybrid, 3 days/week in office)

Role Summary

Come join tastytrade, part of IG Group, as we build the reliability practice behind the brokerage platform that active options, futures, and equities traders rely on every market day. As our first Senior Site Reliability Engineer, you'll harden the systems behind order execution and market data delivery — designing for fault tolerance, closing gaps in our telemetry, and making sure our HashiCorp Nomad-based service fabric scales cleanly as trading volume grows. You'll work embedded alongside our infrastructure and application engineering teams, contributing directly to our Ruby, Java, and Elixir services. This is a rare opportunity to shape a practice and a culture from day one, on a platform where every order, quote, and position has to be right, because real client capital is on the line.

What You'll Do (Job Responsibilities)
  • Build self-healing, fault-tolerant infrastructure and internal tooling that automates repetitive operational work and reduces toil for Platform and Application teams.
  • Run gap analysis across our observability stack to find blind spots in telemetry, logging, and alerting coverage, then close them so failures surface before customers feel them.
  • Own scalability work across our HashiCorp Nomad service fabric: capacity planning, load testing, and identifying architectural bottlenecks before they become incidents.
  • Extend our observability stack (Prometheus, Honeycomb, OpenTelemetry) with the instrumentation needed to actually see the failure modes above.
  • Set SLOs and error budgets with multi-window burn-rate alerting for critical brokerage flows, once the fault-tolerance and telemetry foundation is in place.
  • Mentor engineers across teams to build a culture of site reliability champions so the practice outlives any one person.
Who You Are (Skills Needed)
  • Hands‑on experience designing fault‑tolerant, self‑healing distributed systems — not just describing the patterns, but having shipped them.
  • Deep understanding of one or more: distributed systems, Linux systems, cloud-native architectures, containerization.
  • Experience running gap analyses on observability/telemetry systems: identifying what's not instrumented, not alerted on, or not visible until it's too late.
  • A track record scaling systems under real production load, including capacity planning and architectural bottleneck identification.
  • Hands‑on experience with OpenTelemetry, Prometheus, and Grafana, with the ability to instrument services directly.
  • Strong Linux internals and networking fundamentals, including TCP/IP, UDP/multicast, packet capture, and flow analysis.
  • On‑call experience on production systems and comfort building a blameless post‑incident review process.
  • Working knowledge of SLOs and error budgets as a tool, not the job description; HashiCorp Nomad, Consul, or Vault experience is a strong plus.
  • Strong programming skills in a language such as Python, Ruby, Java, or similar.
Company Perks + Benefits
  • Performance Bonuses
  • Stock Purchase Options
  • Medical/Vision/Dental Benefits
  • 401k Plan
  • 20 Paid Vacation Days (plus an additional paid vacation day the month of your birthday!)
  • 10 Paid Sick Days
  • Gym Membership Reimbursement
  • Commuter Benefits
  • Pet Insurance
  • Wellness & Mental Health Programs
  • Charitable Donation Matching
  • Two Paid Volunteer Days Off
  • Daily catered lunch when in the office
  • Full kitchen with snacks and beverages
  • In-building gym
  • Shuttle to/from Metra

Base Salary Range: $180,000-$200,000

The actual salary offered will be based on the candidate's level of experience and qualifications.

Discretionary Performance Bonus: 15-20% of base salary based on individual and company performance.

About IGNA + tasty

IG North America is home to tastytrade, tastylive & tastyfx—a family of brands built to democratize trading and empower individual investors. Founded in Chicago by the creators of thinkorswim, acquired by London-based IG Group in 2021, we combine startup innovation with the backing of a FTSE 100 fintech operating across five continents serving over 1.3m customers and handling billions of dollars in transactions – built on scale, trust, and proof. From our headquarters in Chicago's Fulton Market, our team builds award‑winning trading platforms, produces live financial education content daily, and creates technology that makes complex markets accessible. tastytrade is our retail brokerage for self‑directed investing and trading, with tastylive providing content to educate our customers. tastyfx is the fastest‑growing forex broker in the US. tastycrypto provides self‑custody digital wallets for decentralized finance. We’re a lean, collaborative team that values autonomy, pragmatism, and impact. Whether you are building technology, creating content, serving customers, or supporting operations, you’ll work alongside people who are passionate about disrupting traditional finance and genuinely care about helping traders succeed. Our culture rewards initiative, embraces experimentation, and measures success by the value we create for our users. The bar is high – bring a curious and forward‑thinking mindset and we’ll give you the platform to define what comes next. Join us at IG|tasty – the future gets built here.

Location: Our office is in the West Loop - Chicago's growing center of tech, great cuisine, and high-end bars.

tastytrade | tastylive | tastyfx 1330 W Fulton Market, Chicago, IL 60607

  • Don't meet every single requirement? Studies have shown that women and people of color are less likely to apply to jobs unless they have every single qualification. Our team is dedicated to building a diverse, inclusive, and authentic workplace, so if you’re excited about this role, but your experience doesn’t align perfectly, we encourage you to apply anyway. You may be just the right candidate for this or other roles!
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Site Reliability Engineer - Linux Systems & Application Observability
Senior Site Reliability Engineer - Linux Systems & Application Observability

Tastyworks • Chicago (IL)

Hybrid
USD 180,000 - 200,000
Performance Bonuses
Stock Purchase Options
Medical/Vision/Dental Benefits
+2
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Tastyworks • Chicago (IL), Northern (KY)

Hybrid
USD 180,000 - 200,000
Performance Bonuses
Stock Purchase Options
401k Plan
+6
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Tastylive • United States

Hybrid
USD 180,000 - 200,000
Performance Bonuses
Stock Purchase Options
Medical/Vision/Dental Benefits
+2
Senior Site Reliability Engineer
Senior Site Reliability Engineer

tastytrade, Inc. • Chicago (IL)

Hybrid
USD 180,000 - 200,000
Performance Bonuses
Stock Purchase Options
401k Plan
+8
Senior Linux Infrastructure Engineer
Senior Linux Infrastructure Engineer

tastytrade, Inc. • Chicago (IL), Northern (KY)

Hybrid
USD 140,000 - 180,000
Performance Bonuses
Stock Purchase Options
401k Plan
+5
Senior Linux Infrastructure Engineer
Senior Linux Infrastructure Engineer

Tastyworks • Chicago (IL), Northern (KY)

Hybrid
USD 140,000 - 180,000
Performance Bonuses
Stock Purchase Options
401k Plan
+8
Staff Software Engineer, Order Handling and Ledger Team
Staff Software Engineer, Order Handling and Ledger Team

tastylive • Chicago (IL)

Hybrid
USD 180,000 - 230,000
Staff Software Engineer, Order Handling and Ledger Team
Staff Software Engineer, Order Handling and Ledger Team

Tastyworks • Chicago (IL), Northern (KY)

Hybrid
USD 180,000 - 230,000
Staff Software Engineer, Order Handling and Ledger Team New Chicago, Illinois
Staff Software Engineer, Order Handling and Ledger Team New Chicago, Illinois

Tastytrade • Chicago (IL), Northern (KY)

Hybrid
USD 180,000 - 230,000
Hybrid work in Chicago
Bonus program
Staff Software Engineer, Order Handling and Ledger Team
Staff Software Engineer, Order Handling and Ledger Team

tastytrade, Inc. • Chicago (IL), Northern (KY)

Hybrid
USD 180,000 - 230,000