Sr Site Reliability Engineer (New Relic/Octopus Deploy/Terraform)

The Judge Group

Chicago (IL)

Hybrid

USD 140,000 - 210,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Competitive Salary
Equity
Comprehensive Benefits

Job summary

The Judge Group seeks a Senior Site Reliability Engineer / Observability Specialist to lead the client's reliability architecture and observability strategy in a hybrid US-based role. This position blends hands-on implementation with coaching across Azure and AWS, building New Relic instrumentation and Terraform modules while driving incident management disciplines and SLO/SLI governance.

If you thrive in engineering-first consulting and want to shape production reliability at scale, this role

Qualifications

  • 7+ years in SRE/DevOps with enterprise cloud scale and production reliability.
  • Observability mastery with New Relic (NRQL, Synthetics, APM) and defining actionable SLOs/SLIs.
  • Strong IaC & automation with Terraform and PowerShell in Windows and Linux environments.
  • Cloud & CI/CD expertise in Azure DevOps or Octopus Deploy pipelines.
  • Experience designing on-call rotations, runbooks, and incident response operations.

Responsibilities

  • Build and optimize New Relic instrumentation (NRQL, APM, Logs) across multi-cloud to streamline metrics collection.
  • Author and maintain Terraform modules to manage monitoring configurations and alert pipelines.
  • Establish incident response foundations with Incident.IO, Slack/OpsGenie integrations, and automated escalation.
  • Define and govern actionable SLOs/SLIs and optimize log ingestion costs to reduce alert fatigue.
  • Partner with platform and product teams through hands-on coaching and post-incident reviews.

Skills

New Relic NRQL
Terraform
PowerShell
Azure DevOps
Octopus Deploy
Incident Management
SLOs / SLIs

Tools

Incident.IO
PagerDuty
OpsGenie
Azure
AWS

Job description

Ready to Lead Observability & Reliability Architecture?
At-a-Glance Snapshot
  • Role: Senior Site Reliability Engineer / Observability Specialist
  • Location / Work Model: Hybrid / Remote (US)
  • Compensation / Perks: Competitive Salary + Equity + Comprehensive Benefits
Why Join Our Client?
  • Massive Financial Scale: Build and optimize high-throughput liquidity and corporate financial platforms handling real-time, global payment flows.
  • Greenfield Ownership: Establish core incident management and reliability frameworks from scratch—you'll set the playbook, not just follow one.
  • Engineering-First Consulting: Enjoy the balance of hands-on technical execution (80%) and high-impact cross-functional team coaching (20%).
What You'll Do
  • Engineered Observability: Build and optimize New Relic instrumentation (NRQL, APM, Logs) across Azure and AWS to streamline metrics collection using RED/USE frameworks.
  • Infrastructure as Code (IaC): Author and maintain Terraform modules to manage monitoring configurations, alert pipelines, and cloud resources across multi-cloud environments.
  • Incident Architecture: Establish early-stage incident response foundations, deploying Incident.IO, Slack/OpsGenie integrations, and automated escalation workflows to drive down MTTD/MTTR.
  • Reliability Strategy: Define and govern actionable SLOs/SLIs and error budgets, optimizing log ingestion costs while eliminating alert fatigue for stream-aligned engineering teams.
  • Technical Enablement: Partner directly with platform and product teams through hands-on coaching, post-incident reviews (PIRs), and Azure DevOps CI/CD pipeline automation.
What You Bring
  • 7+ Years in SRE/DevOps: Deep experience scaling enterprise cloud infrastructure and production reliability in high-availability environments.
  • Observability Mastery: Expert-level hands-on skills with New Relic (NRQL, Synthetics, APM) and defining actionable SLOs/SLIs. *MUST HAVE*
  • Strong IaC & Automation: Advanced proficiency in Terraform and PowerShell scripting within enterprise Windows (80%) and Linux (20%) environments. *MUST HAVE*
  • Cloud & CI/CD Expertise: Proven track record in Azure (App Services, Virtual Machines, Azure SQL) combined with Azure DevOps or Octopus Deploy pipelines.
  • Incident Management Focus: Experience designing on-call rotations, runbooks, and incident response operations via Incident.IO, PagerDuty, or OpsGenie.

Ready to reshape global financial infrastructure and take complete ownership of production reliability?

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Site Reliability Engineer, Observability
Senior Site Reliability Engineer, Observability

blockchaincapital.com • New York (NY)

On-site
USD 130,000 - 180,000
Senior Site Reliability Engineer, Observability New York, NY, United States
Senior Site Reliability Engineer, Observability New York, NY, United States

Ripple • New York (NY)

On-site
USD 160,000 - 200,000
Lead Site Reliability Engineer
Lead Site Reliability Engineer

Luxoft • Wilmington (DE)

On-site
USD 140,000 - 190,000
Lead Site Reliability Engineer
Lead Site Reliability Engineer

Luxoft • Buffalo (NY)

On-site
USD 140,000 - 190,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Veloc Inc • Coppell (TX)

On-site
USD 140,000 - 190,000
Site Reliability Engineer
Site Reliability Engineer

TalentDome Staffing • United States

On-site
USD 140,000 - 210,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

The ReWork Group • New York (NY)

On-site
USD 120,000 - 160,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Mission Staffing • New York (NY)

Hybrid
USD 140,000 - 200,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Hobbsnews • Plano (TX)

On-site
USD 152,000 - 192,000
Discretionary incentive eligible
Benefits eligible
Director, Site Reliability Engineering
Director, Site Reliability Engineering

Jobtailor • California (MO)

On-site
USD 180,000 - 260,000