Senior SRE

banyansoftware

United States

Remote

USD 130,000 - 165,000

Full time

10 days ago
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

Remote role (US/Canada)

Job summary

BanyanSoftware is seeking a Senior Site Reliability Engineer who will own operational excellence for a portfolio of modernized, distributed SaaS apps across multiple clouds. It is a hands-on, code-intensive role focused on IaC and AI-driven automation to reduce toil at scale.

Responsibilities include on-call coverage for containerized production apps, expanding observability tooling, Disaster Recovery planning, incident response, and building CI/CD pipelines with Terraform and GitHub Actions

Qualifications

  • 5-7 years of progressive experience in software engineering or SRE focusing on operating distributed systems.
  • Proficient coding in Python, JavaScript, or Go with APIs, authentication, parallelization, and data transformation.
  • Hands-on experience with Docker and Kubernetes for scalable distributed systems.
  • Production Terraform experience with modules at scale.
  • Operational experience with AWS and/or Azure services.
  • Deep CI/CD background with GitHub Actions or GitLab CI and DevSecOps in workflows.
  • Experience operating highly available production systems with logging, troubleshooting, and tracing.

Responsibilities

  • Participate in rotating on-call coverage for containerized production applications across multiple business units.
  • Maintain and extend observability tooling (monitoring, logging, tracing).
  • Develop, test, and execute disaster recovery and business continuity procedures.
  • Respond to security incidents by following runbooks and coordinating remediation.
  • Build and operate IaC (Terraform) and CI/CD pipelines (GitHub Actions, GitLab CI).
  • Design and operate AI agents that automate SRE tasks within a DevSecOps practice.
  • Act as a technical escalation point for infrastructure, network, and automation issues.

Skills

Python
JavaScript
Go
DevSecOps
Distributed systems

Tools

Docker
Kubernetes
Terraform
GitHub Actions
GitLab CI

Job description

Role overview

The Senior SRE owns operational excellence for a portfolio of modernized, distributed SaaS applications running across multiple cloud environments. Working as part of a team that provides round-the-clock coverage, the role blends Tier 1 reliability engineering, observability, disaster recovery, and security incident response on AWS and Azure. It is a hands-on, code-intensive position that leans on Infrastructure-as-Code and AI-driven automation to reduce toil at scale.

Responsibilities
  • Participate in rotating on-call coverage for containerized production applications across multiple business units, ensuring continuous availability and rapid response
  • Maintain and extend observability tooling (monitoring, logging, tracing) to detect degradation early and shorten mean-time-to-detect and mean-time-to-resolve
  • Develop, test, and execute disaster recovery and business continuity procedures for cyber, infrastructure, or geographic disruptions
  • Respond to security incidents against SaaS platforms by following established runbooks and coordinating remediation
  • Build and operate Infrastructure-as-Code (Terraform) and CI/CD pipelines (GitHub Actions, GitLab CI) to automate deployments and reduce manual work
  • Design and operate AI agents that automate SRE tasks such as triage, detection, and remediation within a DevSecOps practice
  • Act as a technical escalation point, applying strong analytical skills to resolve infrastructure, network, and automation issues across distributed multi-tenant SaaS environments
Requirements
  • 5-7 years of progressive experience in software engineering and/or site reliability engineering, focused on operating distributed systems
  • Deep coding expertise in Python, JavaScript, or Go, including APIs, authentication, parallelization, triggering, and data transformation
  • Strong hands-on experience with container technologies (Docker, Kubernetes) supporting highly scalable and resilient distributed systems
  • Production Terraform experience working with modules at scale
  • Operational experience with AWS (for example EC2, Lambda, EKS, S3, RDS) and/or Azure (such as Container Apps, AKS, Container Storage)
  • Deep CI/CD background with GitHub Actions or GitLab CI and a track record of embedding DevSecOps practices into operational workflows
  • Experience operating highly available production systems with application-level logging, troubleshooting, and tracing tools
Nice to have
  • Hands-on experience building or operating AI agents that automate SRE tasks and incident response
  • Familiarity with security incident response procedures and runbook-driven operations
Benefits and work setup
  • Remote role open to candidates based in the United States or Canada
  • US salary band of approximately USD 130,000-165,000 and Canada band of CAD 120,000-145,000, excluding annual bonus and equity where applicable
  • Compensation varies based on market conditions, location, job-related skills, and experience
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Principal Site Reliability Engineer
Principal Site Reliability Engineer

Gen Digital Inc. • United States

Remote
USD 180,000 - 240,000
Site Reliability Engineer
Site Reliability Engineer

Stelvio Inc. • Town of Texas (WI)

On-site
USD 125,000 - 145,000
Senior DevOps Engineer/Site Reliability Engineer-East Coast
Senior DevOps Engineer/Site Reliability Engineer-East Coast

Stellar Cyber • North Carolina

On-site
USD 165,000 - 215,000
Pre‑IPO Stock Options
Medical, Dental & Vision care
401(k)
+2
Senior Site Reliability Engineer
Senior Site Reliability Engineer

GovCIO • Arlington (VA)

On-site
USD 210,000 - 230,000
Site Reliability Engineer
Site Reliability Engineer

Motion Recruitment Partners LLC • Chicago (IL), Northern (KY)

On-site
USD 140,000 - 170,000
SRE (Site Realiability Engineer)
SRE (Site Realiability Engineer)

STRATIS Cloud Tech Solutions INC • Arkansas

On-site
USD 110,000 - 150,000
Competitive salary
Growth and learning opportunities
Friendly, collaborative team
Senior Lead Site Reliability Engineer
Senior Lead Site Reliability Engineer

JPMorgan Chase & Co. • Jersey City (NJ)

On-site
USD 150,000 - 210,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Kovoro • Denver (CO), Northern (KY)

On-site
USD 150,000 - 190,000
Sr. Site Reliability Engineer
Sr. Site Reliability Engineer

Staffing Science • Arizona

On-site
USD 180,000 - 240,000
Senior SRE Engineer
Senior SRE Engineer

Compunnel, Inc. • Alpharetta (GA)

On-site
USD 140,000 - 190,000