Senior Site Reliability Automation Engineer

Jobtailor

Colorado

On-site

USD 140,000 - 200,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Jobtailor is seeking an experienced engineer to design and implement automated reliability solutions for production systems in a distributed SaaS environment. The role emphasizes reducing manual toil through automation and advanced observability.

You will work across AWS/Azure stacks, CI/CD pipelines, and deployment workflows to improve system health, performance, and operability, while employing AI-driven approaches to optimize knowledge management and troubleshooting.

Qualifications

  • 7+ years of experience in Software Engineering, Test Automation, Site Reliability Engineering, Production Engineering, Platform Engineering, DevOps, Performance Engineering, or related technical discipline.
  • Proven software development experience in Python, PowerShell, TypeScript, JavaScript, Java, or C#.
  • Experience troubleshooting complex production issues in large-scale, distributed, or SaaS environments.
  • Experience conducting root cause analysis and implementing long-term remediation solutions.
  • Experience building automation tools, operational workflows, reliability tooling, or engineering platforms.
  • Experience with CI/CD pipelines, deployment automation, release processes, and AWS and/or Azure environments.

Responsibilities

  • Build automation solutions that reduce manual operational work and improve scalability, reliability, and consistency of production environments.
  • Develop intelligent diagnostics, remediation workflows, and self-service tools for faster issue resolution by Engineering and Support teams.
  • Investigate complex production issues to identify root causes and drive long-term reliability improvements.
  • Enhance monitoring, alerting, dashboards, and telemetry to improve visibility into system health and risk.
  • Partner with Engineering teams to improve observability, operational readiness, platform stability, and production support practices.
  • Support performance and scalability initiatives by identifying degradation trends and diagnosing bottlenecks.
  • Build and integrate reliability automation across AWS, Azure, CI/CD pipelines, deployment workflows, and cloud-native environments.
  • Leverage AI-enabled tools and emerging technologies to improve troubleshooting and engineering efficiency.

Skills

Python
PowerShell
TypeScript
JavaScript
Java
C#
SRE
DevOps
CI/CD
Cloud (AWS/Azure)

Tools

CI/CD tools
AWS
Azure

Job description

Responsibilities
  • Build automation solutions that reduce manual operational work and improve the scalability, reliability, and consistency of production environments.
  • Develop intelligent diagnostics, remediation workflows, and self-service tools that enable Engineering and Support teams to resolve issues faster.
  • Strengthen incident response by investigating complex production issues, identifying root causes, and driving long-term reliability improvements.
  • Enhance monitoring, alerting, dashboards, and telemetry to improve visibility into system health, performance, and operational risk.
  • Partner with Engineering teams to improve observability standards, operational readiness, platform stability, and production support practices.
  • Support performance and scalability initiatives by identifying degradation trends, diagnosing bottlenecks, and recommending optimization strategies.
  • Build and integrate reliability automation across AWS, Azure, CI/CD pipelines, deployment workflows, and cloud-native platform environments.
  • Leverage AI-enabled tools and emerging technologies to improve troubleshooting, knowledge management, operational workflows, and engineering efficiency.
Requirements
  • 7+ years of experience in Software Engineering, Test Automation, Site Reliability Engineering, Production Engineering, Platform Engineering, DevOps, Performance Engineering, or a related technical discipline.
  • Strong software development experience using Python, PowerShell, TypeScript, JavaScript, Java, C#, or similar programming languages.
  • Experience troubleshooting complex production issues in large-scale, distributed, or SaaS environments.
  • Experience conducting root cause analysis and implementing long-term remediation solutions.
  • Experience building automation tools, operational workflows, reliability tooling, or engineering platforms.
  • Experience with CI/CD pipelines, deployment automation, release processes, and AWS and/or Microsoft Azure environments.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Site Reliability Engineer
Senior Site Reliability Engineer

Veloc Inc • Coppell (TX)

On-site
USD 140,000 - 190,000
Site Reliability Engineer
Site Reliability Engineer

Jobtailor • New Hampshire

On-site
USD 110,000 - 160,000
Site Reliability Engineer
Site Reliability Engineer

TechDigital Group • Houston (TX), Juno Beach (FL)

On-site
USD 120,000 - 180,000
Site Reliability Engineer
Site Reliability Engineer

Shrive Technologies LLC • Schaumburg (IL)

On-site
USD 110,000 - 170,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Bank of America • Chandler (MN)

On-site
USD 100,000 - 130,000
Site Reliability Engineer
Site Reliability Engineer

TalentDome Staffing • United States

On-site
USD 140,000 - 210,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Madison-Davis, LLC • United States

On-site
USD 120,000 - 150,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

InterEx Group • New York (NY)

On-site
USD 100,000 - 130,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Jobtailor • Arlington (VA)

On-site
USD 140,000 - 200,000
Site Reliability Engineer
Site Reliability Engineer

Jobtailor • California (MO)

Hybrid
USD 120,000 - 160,000