Senior Site Reliability Engineer – Digital Assets

Jobtailor

Arizona

On-site

USD 120,000 - 170,000

Full time

5 days ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Jobtailor is seeking an experienced software reliability engineer to apply engineering, automation, and DevOps principles across services. You will define SLIs/SLOs, improve observability, and drive CI/CD and IaC practices while mentoring engineers and leading incident response.

Candidates should have 5–8 years in software/SRE roles and be eligible to work in the United States. The role emphasizes collaboration with software teams to ensure reliability, scalability, and operational readiness

Qualifications

  • Bachelor's degree in Computer Science, Software Engineering, or related field or equivalent practical experience.
  • Eligibility to work in the United States for any employer; no visa sponsorship.

Responsibilities

  • Apply software engineering, automation, and DevOps principles to improve how services are built, tested, deployed, observed, operated, and recovered.
  • Define, implement, or improve SLIs, SLOs, error budgets, and other service-health measures.
  • Improve observability through metrics, logging, tracing, monitoring, alerting, dashboards, and service-health instrumentation.
  • Drive continuous improvement across CI/CD, observability, deployment practices, Infrastructure as Code, automation, testing, incident response, capacity management, resilience, and operational readiness.
  • Identify recurring production issues and translate operational experience into improvements in code, architecture, automation, tooling, and engineering practices.
  • Partner with Software Engineering teams to incorporate reliability, resiliency, scalability, performance, observability, recoverability, and operational readiness throughout the development lifecycle.
  • Participate in or lead incident response, troubleshooting, service restoration, and blameless post-incident learning.
  • Reduce operational toil through software, automation, reusable patterns, and improved engineering practices.
  • Mentor engineers, share knowledge, and develop reusable solutions and practices.
  • Own reliability outcomes across multiple services and guide less-experienced engineers.

Skills

Programming languages
Analytical thinking
Problem solving
Communication
Collaboration
Mentoring

Education

Bachelor's degree in CS/engineering

Tools

AWS
Azure
GCP
OCI
Linux/Unix
Containers
Orchestration
Monitoring tools
Logging tools
Dashboards

Job description

  • Apply software engineering, automation, and DevOps principles to improve how services are built, tested, deployed, observed, operated, and recovered
  • Use data, evidence, experimentation, and rigorous engineering analysis to identify reliability risks, test assumptions, and guide technical decisions
  • Define, implement, or improve SLIs, SLOs, error budgets, and other service-health measures
  • Improve observability through metrics, logging, tracing, monitoring, alerting, dashboards, and service-health instrumentation
  • Drive continuous improvement across CI/CD, observability, deployment practices, Infrastructure as Code, automation, testing, incident response, capacity management, resilience, and operational readiness
  • Identify recurring or systemic production issues and translate operational experience into improvements in code, architecture, automation, tooling, and engineering practices
  • Partner with Software Engineering teams to incorporate reliability, resiliency, scalability, performance, observability, recoverability, and operational readiness throughout the development lifecycle
  • Participate in or lead incident response, troubleshooting, service restoration, and blameless post-incident learning
  • Reduce operational toil and manual intervention through software, automation, reusable patterns, and improved engineering practices
  • Mentor engineers, share knowledge, and develop reusable solutions and practices
  • Own reliability outcomes across multiple services and guide less-experienced engineers
Requirements
  • Typically 5-8 years of relevant professional experience in Software Engineering, Site Reliability Engineering, Systems Engineering, Cloud/Platform Engineering, DevOps, Infrastructure Engineering, Architecture where applicable, or a comparable technical discipline
  • Experience with software development or scripting using one or more modern programming languages
  • Experience with software engineering principles, distributed systems, production troubleshooting, automation, and observability appropriate to the level
  • Experience with public cloud technologies and architectures, preferably AWS, along with infrastructure, networking, Linux/Unix, and modern application architectures appropriate to the level
  • Demonstrated analytical, problem-solving, communication, and collaboration skills appropriate to the scope of the role
  • Preferred: hands‑on experience with AWS or comparable experience with Microsoft Azure, Google Cloud Platform (GCP), or Oracle Cloud Infrastructure (OCI)
  • Experience developing, deploying, operating, or improving highly available production software or distributed systems
  • Experience with CI/CD, Infrastructure as Code, containers or orchestration, observability, monitoring, alerting, and software-delivery automation
  • Experience with SLIs, SLOs, error budgets, incident management, performance analysis, capacity management, resilience testing, disaster recovery, or operational readiness
  • Experience creating reusable automation, tooling, platforms, patterns, or practices that improve engineering effectiveness
  • Bachelor's degree in Computer Science, Software Engineering, Computer Engineering, Information Systems, or a related technical field, or equivalent practical experience
  • Eligibility to work in the United States for any employer at the date of hire; position is ineligible for employment Visa sponsorship
  • Ability to perform essential functions and physical requirements with or without reasonable accommodation
  • Ability to work in a normal office environment, primarily sedentary, with extensive computer use and occasional standing, walking, kneeling, and reaching; ability to lift 10 pounds occasionally and/or negligible force frequently; visual acuity, dexterity, and communication ability
Core Competencies

Demonstrates expertise in Software Engineering, Site Reliability Engineering and DevOps principles with a strong focus on automation, observability, and continuous improvement. Proficient in leveraging public cloud technologies and implementing reliability measures such as SLIs and SLOs to enhance service performance and operational readiness.

Highest-signal resume keywords
  • Software Engineering
  • Site Reliability Engineering
  • AWS Cloud Technologies
  • CI/CD Automation
  • Observability and Monitoring
Hard Skills
  • Software Development
  • Scripting
  • Distributed Systems
  • Infrastructure as Code
  • Incident Management
  • Performance Analysis
  • Capacity Management
  • Resilience Testing
  • Disaster Recovery
  • Automation
Soft Skills
  • Analytical Skills
  • Problem-Solving
  • Communication
  • Collaboration
  • Mentoring
Industry Keywords
  • DevOps
  • Infrastructure Engineering
  • Cloud/Platform Engineering
  • Operational Readiness
  • Service Health
Tools & Technologies
  • AWS
  • Microsoft Azure
  • Google Cloud Platform
  • Oracle Cloud Infrastructure
  • Linux/Unix
  • Containers
  • Orchestration
  • Monitoring Tools
  • Logging Tools
  • Dashboards
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Principal Site Reliability Engineer
Principal Site Reliability Engineer

Jobtailor • Arizona

On-site
USD 180,000 - 240,000
Senior Application Development Advisor
Senior Application Development Advisor

Jobtailor • Colorado

On-site
USD 130,000 - 170,000
Staff Site Reliability Engineer, SRE
Staff Site Reliability Engineer, SRE

Jobtailor • California (MO)

On-site
USD 120,000 - 210,000
Site Reliability Engineer – Lead
Site Reliability Engineer – Lead

Jobtailor • Arizona

On-site
USD 140,000 - 230,000
Site Reliability Engineer
Site Reliability Engineer

Jobtailor • United States

On-site
USD 120,000 - 180,000
Staff Site Reliability Engineer
Staff Site Reliability Engineer

Jobtailor • Town of Florida (NY)

On-site
USD 180,000 - 240,000
AVP SRE, Cloud Solutions
AVP SRE, Cloud Solutions

Jobtailor • Arlington (TX)

On-site
USD 180,000 - 240,000
Site Reliability Engineer II
Site Reliability Engineer II

Jobtailor • California (MO)

On-site
USD 85,000 - 120,000
Tech Services Delivery Manager – Platform Services
Tech Services Delivery Manager – Platform Services

Jobtailor • North Carolina

On-site
USD 120,000 - 150,000
Reliability Observability Engineer, Level 2
Reliability Observability Engineer, Level 2

Jobtailor • Colorado

On-site
USD 120,000 - 180,000