Sr Site Reliability Engineer (New Relic/Octopus Deploy/Terraform)

The Judge Group

Illinois

Hybrid

USD 140,000 - 170,000

Full time

4 days ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

Competitive Salary
Equity
Comprehensive Benefits

Job summary

The Judge Group is seeking a Senior Site Reliability Engineer / Observability Specialist to lead reliability architecture in a hybrid US role. Based in the Greater Chicago Area with 2-3 days on-site in the Loop, you will own observability, IaC, and incident management across Azure and AWS.

You will partner with platform and product teams, set SLOs/SLIs, optimize log costs, and drive CI/CD automation using Terraform, New Relic, and related tools. Extensive cloud experience required.

Qualifications

  • 7+ years in SRE/DevOps in high-availability environments.
  • Expert-level New Relic skills (NRQL, APM, Logs) and defining actionable SLOs/SLIs.
  • Advanced Terraform and PowerShell scripting in Windows and Linux.
  • Azure cloud services, Azure DevOps or Octopus Deploy pipelines.
  • Experience designing on-call rotations, runbooks, and incident response tooling.

Responsibilities

  • Build and optimize observability instrumentation (New Relic) across multi-cloud environments.
  • Author and maintain Terraform modules for monitoring configurations and alert pipelines.
  • Establish incident architecture with automation to reduce MTTD/MTTR.
  • Define actionable SLOs/SLIs and manage error budgets across teams.
  • Coach platform teams through PIRs and CI/CD automation.

Skills

SRE/DevOps experience
Observability mastery
IaC & automation
Cloud & CI/CD
Incident management

Tools

New Relic
Terraform
PowerShell
Azure DevOps
Octopus Deploy
Incident.IO
PagerDuty
OpsGenie

Job description

To maintain strong team cohesion and operational continuity, candidates must currently be based in the Greater Chicago Area and commit to a hybrid schedule in the downtown/loop (2-3 days/week on-site). All applicants must possess unrestricted work authorization in the United States. Please ONLY apply if you meet these criteria.

Ready to Lead Observability & Reliability Architecture?

At-a-Glance Snapshot

  • Role: Senior Site Reliability Engineer / Observability Specialist
  • Location / Work Model: Hybrid / Remote (US)
  • Compensation / Perks: Competitive Salary + Equity + Comprehensive Benefits

Why Join Our Client?

  • Massive Financial Scale: Build and optimize high-throughput liquidity and corporate financial platforms handling real-time, global payment flows.
  • Greenfield Ownership: Establish core incident management and reliability frameworks from scratch—you'll set the playbook, not just follow one.
  • Engineering-First Consulting: Enjoy the balance of hands-on technical execution (80%) and high-impact cross-functional team coaching (20%).

What You'll Do

  • Engineered Observability: Build and optimize New Relic instrumentation (NRQL, APM, Logs) across Azure and AWS to streamline metrics collection using RED/USE frameworks.
  • Infrastructure as Code (IaC): Author and maintain Terraform modules to manage monitoring configurations, alert pipelines, and cloud resources across multi-cloud environments.
  • Incident Architecture: Establish early-stage incident response foundations, deploying Incident.IO, Slack/OpsGenie integrations, and automated escalation workflows to drive down MTTD/MTTR.
  • Reliability Strategy: Define and govern actionable SLOs/SLIs and error budgets, optimizing log ingestion costs while eliminating alert fatigue for stream-aligned engineering teams.
  • Technical Enablement: Partner directly with platform and product teams through hands-on coaching, post-incident reviews (PIRs), and Azure DevOps CI/CD pipeline automation.

What You Bring

  • 7+ Years in SRE/DevOps: Deep experience scaling enterprise cloud infrastructure and production reliability in high-availability environments.
  • Observability Mastery: Expert-level hands-on skills with New Relic (NRQL, Synthetics, APM) and defining actionable SLOs/SLIs. *MUST HAVE*
  • Strong IaC & Automation: Advanced proficiency in Terraform and PowerShell scripting within enterprise Windows (80%) and Linux (20%) environments. *MUST HAVE*
  • Cloud & CI/CD Expertise: Proven track record in Azure (App Services, Virtual Machines, Azure SQL) combined with Azure DevOps or Octopus Deploy pipelines.
  • Incident Management Focus: Experience designing on-call rotations, runbooks, and incident response operations via Incident.IO, PagerDuty, or OpsGenie.

Ready to reshape global financial infrastructure and take complete ownership of production reliability?

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Site Reliability Engineer, Observability
Senior Site Reliability Engineer, Observability

blockchaincapital.com • New York (NY)

On-site
USD 130,000 - 180,000
Lead Site Reliability Engineer
Lead Site Reliability Engineer

Intone Inc • Northern (KY)

Hybrid
USD 140,000 - 180,000
Senior Site Reliability Engineer, Observability New York, NY, United States
Senior Site Reliability Engineer, Observability New York, NY, United States

Ripple • New York (NY)

On-site
USD 160,000 - 200,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Mission Staffing • New York (NY)

Hybrid
USD 140,000 - 200,000
Senior SRE & Observability Lead — New Relic, Terraform
Senior SRE & Observability Lead — New Relic, Terraform

The Judge Group • Illinois

Hybrid
USD 140,000 - 170,000
Competitive Salary
Equity
Comprehensive Benefits
Senior Site Reliability Engineer
Senior Site Reliability Engineer

United States Digital Space LLC • Charlotte (TX)

On-site
USD 153,000 - 192,000
Discretionary incentive eligible
Benefits package
Senior Site Reliability Engineer
Senior Site Reliability Engineer

SDI International • Chicago (IL)

Hybrid
USD 130,000 - 180,000
Sr. Site Reliability Engineer
Sr. Site Reliability Engineer

MeridianLink, Inc. • Northern (KY)

Hybrid
USD 140,000 - 210,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Optomi • Dallas (TX)

Hybrid
USD 120,000 - 150,000
Sr. Site Reliability Engineer
Sr. Site Reliability Engineer

MeridianLink • United States

On-site
USD 140,000 - 190,000