Site Reliability Engineer

Optomi

United States

On-site

USD 120,000 - 180,000

Full time

10 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Optomi, in partnership with a large national banking institution, is seeking an Application Site Reliability Engineer (SRE) to join their team. You will help deliver reliable, scalable, and high-performing production systems within a highly available enterprise environment.

The role emphasizes automation, observability, incident response, and DevOps best practices while collaborating with application teams to improve stability and accelerate software delivery.

Qualifications

  • 5+ years of SRE/DevOps experience in enterprise production.
  • On-call experience and high availability mindset.
  • Hands-on with GitLab, Git, and Terraform.
  • Strong observability using Dynatrace and Splunk.
  • Experience developing and maintaining SLOs and error budgets.
  • Automation of operational processes.
  • Incident response leadership and blameless postmortems.
  • Experience with AWS and API Gateway.

Responsibilities

  • Support and enhance reliability, scalability, and performance of production systems.
  • Lead incident response efforts and promote blameless postmortems.
  • Collaborate with cross-functional teams to improve platform health.
  • Drive reliability initiatives through automation and observability.

Skills

SRE/DevOps
GitLab & Git
Terraform IaC
Observability
Incident response
SLOs & error budgets
Automation
AWS cloud
Logging & monitoring

Tools

GitLab
Git
Terraform
Dynatrace
Splunk
API Gateway
AWS

Job description

Optomi, in partnership with a large national banking institution, is seeking a Application Site Reliability Engineer (SRE) to join their team. This is an exciting opportunity to support and enhance the reliability, scalability, and performance of critical production systems within a highly available enterprise environment. The ideal candidate will have experience with cloud infrastructure, observability, automation, incident response, and DevOps best practices, while partnering closely with application teams to improve platform stability and accelerate software delivery.

What The Right Professional Will Enjoy

  • Opportunity to support mission-critical production systems for a leading national financial institution.
  • Driving reliability initiatives through automation, observability, and Site Reliability Engineering best practices.
  • Collaborating with cross-functional application, infrastructure, cloud, and engineering teams to improve system health and operational excellence.
  • Leading incident response efforts while promoting blameless postmortems and continuous service improvements.
  • Leveraging modern AI-assisted tooling to accelerate incident triage and reduce mean-time-to-detect (MTTD).

Apply Today If Your Background Includes

  • 5+ years of experience in Site Reliability Engineering, DevOps, Systems Engineering, or a related discipline supporting enterprise production environments.
  • Experience supporting production applications, participating in on-call rotations, and maintaining high system availability and reliability.
  • Hands-on experience with DevOps technologies such as GitLab, Git, and Infrastructure as Code tools including Terraform.
  • Strong troubleshooting skills with experience building and enhancing observability using monitoring and logging platforms such as Dynatrace and Splunk.
  • Experience developing and maintaining Service Level Objectives (SLOs), error budget tracking, production readiness standards, and operational health scoring.
  • Proven ability to automate manual operational processes and identify opportunities to improve platform efficiency.
  • Experience leading incident response efforts, conducting blameless postmortems, and partnering with development teams to improve logging, runbooks, and technical debt.
  • Experience working with public cloud platforms, preferably AWS, including API Gateway technologies.
  • Experience developing APIs, microservices, or frontend applications is a plus.
  • Strong understanding of operational excellence, observability best practices, and continuous improvement within enterprise environments.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Site Reliability Engineer
Senior Site Reliability Engineer

Optomi • Dallas (TX)

Hybrid
USD 120,000 - 150,000
Senior Site Reliability Engineer – AI-Driven Reliability
Senior Site Reliability Engineer – AI-Driven Reliability

Optomi • United States

On-site
USD 120,000 - 180,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Optomi • Detroit (MI)

Hybrid
USD 100,000 - 130,000
Sr. Director, Site Reliability and Platform Engineering
Sr. Director, Site Reliability and Platform Engineering

Optomi • Tacoma (WA)

On-site
USD 150,000 - 200,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Jobgether • United States

Hybrid
USD 120,000 - 160,000
Competitive compensation package
Flexible work arrangements
Professional development opportunities
+2
Sr. Site Reliability Engineer
Sr. Site Reliability Engineer

MeridianLink • United States

Remote
USD 140,000 - 190,000
Lead Site Reliability Engineer
Lead Site Reliability Engineer

Luxoft • Buffalo (NY)

On-site
USD 140,000 - 190,000
Lead Site Reliability Engineer
Lead Site Reliability Engineer

Luxoft • Wilmington (DE)

On-site
USD 140,000 - 190,000
Site Reliability Engineering Manager
Site Reliability Engineering Manager

O.C. Tanner • Salt Lake City (UT)

On-site
USD 180,000 - 260,000
Sr. Site Reliability Engineer
Sr. Site Reliability Engineer

MeridianLink, Inc. • Northern (KY)

Hybrid
USD 140,000 - 210,000