Lead Site Reliability Engineer

Intone Inc

Northern (KY)

Hybrid

USD 140,000 - 180,000

Full time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Intone Inc. seeks a Lead Site Reliability Engineer for a 12-month contract, remote within the United States. You will design, implement, and operate reliable, scalable systems with a strong emphasis on observability, automation, and SRE best practices.

You will drive reliability initiatives, develop IaC with Terraform, build CI/CD pipelines, establish SLOs/SLIs, lead incident management, and pursue self-healing and proactive capacity planning across cloud-native and hybrid environments.

Qualifications

  • Strong hands-on experience with observability and monitoring tools, including Dynatrace and OpenTelemetry (OTel).
  • Expertise in distributed tracing, metrics collection and analysis, centralized logging and log aggregation.
  • Proficiency in alerting and dashboard development.
  • Strong proficiency in Infrastructure as Code using Terraform.
  • Proven experience designing and executing automated regression testing frameworks.
  • Expert knowledge of production systems monitoring, incident management, and operational troubleshooting.

Responsibilities

  • Design and implement observability and monitoring across production systems.
  • Develop and maintain Infrastructure as Code using Terraform for reproducible deployments.
  • Build and manage CI/CD pipelines, deployment automation, and operational tooling for reliable releases.
  • Design and execute automated regression testing frameworks to validate stability after deployments.
  • Implement and operate SRE practices including SLOs, SLIs, and error budgets.
  • Lead incident management, root cause analysis, and problem management initiatives.

Skills

SRE fundamentals
Observability
Distributed tracing
Alerting
Logging
Cloud-native architectures

Tools

Dynatrace
OpenTelemetry
Terraform
Azure Monitor
Application Insights
Log Analytics
CI/CD pipelines

Job description

We are seeking a Lead Site Reliability Engineer for a 12-month contract position, remote within the US. This role focuses on designing, implementing, and operating reliable, scalable systems with a strong emphasis on observability, automation, and SRE best practices.

Responsibilities
  • Design and implement comprehensive observability and monitoring solutions across production systems
  • Develop and maintain Infrastructure as Code using Terraform for reproducible and scalable deployments
  • Build and manage CI/CD pipelines, deployment automation, and operational tooling to enable reliable releases
  • Design and execute automated regression testing frameworks and test suites to validate application and platform stability following deployments
  • Implement and operationalize SRE practices including Service Level Objectives (SLOs), Service Level Indicators (SLIs), and error budgets
  • Lead incident management, root cause analysis, and problem management initiatives to improve system reliability
  • Develop automated recovery mechanisms and self-healing solutions to reduce manual operational overhead
  • Drive reliability engineering initiatives through performance tuning, capacity planning, and proactive issue detection
  • Support hybrid and cloud-native infrastructure environments with a focus on operational excellence
Qualifications
  • Required
    • Strong hands‑on experience with observability and monitoring tools, including Dynatrace and OpenTelemetry (OTel)
    • Expertise in distributed tracing, metrics collection and analysis, centralized logging and log aggregation
    • Proficiency in alerting and dashboard development
    • Strong proficiency in Infrastructure as Code using Terraform
    • Proven experience designing and executing automated regression testing frameworks
    • Expert knowledge of production systems monitoring, incident management, and operational troubleshooting
    • Strong understanding of application performance management, distributed systems, and modern cloud‑native architectures
    • Strong experience with Microsoft Azure, including App Services, Resource Groups, networking concepts, scaling, and performance optimization
    • Experience with Azure‑native operational tooling such as Azure Monitor, Application Insights, Log Analytics, and dashboards with alerting
    • Demonstrated experience implementing and operating SRE practices including SLOs, SLIs, error budgets, incident management, problem management, and root cause analysis
    • Experience with CI/CD pipelines, deployment automation, and operational tooling
  • Preferred
    • Knowledge of resiliency engineering patterns, disaster recovery planning, and high‑availability architectures
    • Experience supporting cloud‑native and hybrid infrastructure environments
    • Experience developing reliability automation and self‑healing solutions
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Site Reliability Engineer
Senior Site Reliability Engineer

Axiom Pursuits • San Francisco (CA)

On-site
USD 150,000 - 180,000
Remote Lead SRE: Observability & Reliability
Remote Lead SRE: Observability & Reliability

Intone Inc • Northern (KY)

Hybrid
USD 140,000 - 180,000
Sr. Site Reliability Engineer
Sr. Site Reliability Engineer

MeridianLink • United States

On-site
USD 140,000 - 190,000
Site Reliability Engineer
Site Reliability Engineer

TalentDome Staffing • United States

On-site
USD 140,000 - 210,000
Site Reliability Engineer
Site Reliability Engineer

Intone Inc • Plano (TX)

On-site
USD 120,000 - 150,000
Sr. Site Reliability Engineer
Sr. Site Reliability Engineer

MeridianLink, Inc. • Northern (KY)

Hybrid
USD 140,000 - 210,000
Sr Site Reliability Engineer (New Relic/Octopus Deploy/Terraform)
Sr Site Reliability Engineer (New Relic/Octopus Deploy/Terraform)

The Judge Group • Illinois

Hybrid
USD 140,000 - 170,000
Competitive Salary
Equity
Comprehensive Benefits
Site Reliability Engineer
Site Reliability Engineer

Ethos Group • Irving (TX)

On-site
USD 110,000 - 160,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Optomi • Dallas (TX)

Hybrid
USD 120,000 - 150,000
Site Reliability Engineering Manager
Site Reliability Engineering Manager

O.C. Tanner • Salt Lake City (UT)

On-site
USD 180,000 - 240,000