Senior SRE

Banyan Software

United States

Remote

USD 145,000 - 170,000

Full time

10 days ago
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Banyan Software seeks an experienced Senior SRE to own the reliability of modernized SaaS applications across AWS and Azure. You will operate in a 24x7 on-call setting, implement robust observability, automate deployments with Terraform and CI/CD, and drive security incident response.

The role emphasizes hands-on engineering, AI-assisted tooling, and disaster recovery planning.

Qualifications

  • 5–7 years of progressive experience in Software Engineering or SRE for distributed systems.
  • Deep expertise in Python, Javascript, or Go with automation and tool integrations.
  • Container technologies (Docker/Kubernetes) for scalable distributed systems.
  • Terraform IaC at scale across multiple modules.
  • Hands-on experience with AWS and/or Azure production workloads.
  • CI/CD platforms (GitHub Actions, GitLab CI) with DevSecOps practices.
  • Observability tooling (logging, tracing, monitoring) for production systems.
  • AI-assisted engineering tools experience (e.g., Claude Code).
  • APM tooling experience (Datadog/New Relic/Dynatrace) to optimize performance.
  • Incident response, on-call rotations, and disaster recovery procedures.

Responsibilities

  • Operate 24x7 with on‑call rotations for OpCo containerized apps.
  • Serve as Tier 1 SRE managing cloud integrations across AWS and Azure.
  • Implement observability tooling to monitor performance and availability.
  • Develop and test disaster recovery and business continuity procedures.
  • Respond to security incidents following established runbooks.
  • Use Terraform IaC and CI/CD pipelines to deploy and automate environments.
  • Build AI agents to automate SRE tasks and incident response.
  • Provide hands-on problem solving across distributed, multi‑tenant SaaS environments.

Skills

Automation coding
Containerization
CI/CD
Incident response
AI‑fluent engineering
Communication
APM

Education

Bachelor's degree in Computer Science

Tools

Terraform
Docker
Kubernetes
GitHub Actions
GitLab CI
Datadog
New Relic
Dynatrace
Claude Code

Job description

Senior SRE (Site Reliability Engineer) – Modernized Application Operations

Banyan Software is the best permanent home for software businesses that serve specialized industries, their employees, and their customers. With a buy-grow-and-hold-for-life approach and a permanent capital base, Banyan acquires and grows companies worldwide, honoring founder legacies and helping portfolio companies modernize through shared AI expertise and operational discipline. Founded in 2016, Banyan operates more than 120 portfolio companies across North and South America, Europe, and APAC, and has appeared on the Inc. 5000 list for six consecutive years. The Banyan Software Foundation, endowed with $100 million in Banyan stock, leverages technology to build a greener and more equitable world.

Remote: US/Canada
Overview

We are seeking a highly experienced and hands-on SRE to own the operational excellence of the modernized SaaS applications produced by the Banyan AI Factory. This is not a role focused on building the factory itself; instead, you will run the reliability of the modernized applications the factory delivers to our Operating Companies (OpCos).

You will join a team that provides 24x7 coverage with rotating on-call responsibilities, serving as Tier 1 Site Reliability Engineering (SRE) for our OpCos’ distributed applications. Day to day this will include: automated deployments, cloud service integration, application performance and availability monitoring/observability, and security incident response across our two target clouds — Amazon Web Services (AWS) and Microsoft Azure. The ideal candidate has a track record of keeping secure, highly available production systems running at scale.

Key Responsibilities
  • 24x7 Operations & On-Call: Operate as part of a team providing round-the-clock coverage of OpCo containerized applications, participating in a rotating on-call schedule to ensure continuous availability and rapid response.
  • Tier 1 SRE & Operations: Serve as Tier 1 SRE for the modernized applications, managing day-to-day cloud integrations across our two target clouds — AWS and Azure — to keep production systems healthy, performant, and secure.
  • Performance & Availability Monitoring/Observability: Implement and maintain robust application observability tooling (monitoring, logging, tracing) to track performance and availability, proactively detect degradation, and drive down mean-time-to-detect and mean-time-to-resolve.
  • Disaster Recovery and Service Restoration: Develop, maintain, test, and execute disaster recovery and business continuity procedures. Ensure the timely recovery and restoration of services following geographic disruptions, cyber incidents, infrastructure failures, or other disaster events.
  • Security Incident Response: Respond to security incidents and operational events affecting OpCo SaaS platforms, executing established runbooks, coordinating remediation
  • Automation & Infrastructure-as-Code : Use Infrastructure-as-Code (Terraform) and CI/CD pipelines (e.g., GitHub Actions, GitLab CI) to manage, deploy, and automate the operational environments of modernized applications, reducing toil and improving consistency.
  • AI Agents & DevSecOps Scale: Build scale in our DevSecOps practice by designing, building, and operating AI agents that automate SRE tasks and incident response, reducing toil and accelerating detection, triage, and remediation.
  • Hands-on Problem Solving: Serve as a technical escalation point for operational challenges, applying strong analytical skills to resolve infrastructure, network, and automation issues across distributed, multi-tenant SaaS environments while navigating technical ambiguity.
Required Qualifications & Experience
  • Experience: 5–7 years of progressive experience in Software Engineering, and/or Site Reliability Engineering, with a focus on operating distributed systems.
  • Automation Coding Experience: Deep expertise in Python, Javascript, or Go. Building automation and integrations between tools. This may be with AI assistance, but you must have a deep understanding of the code and scripting principals such as: authentication, parallelization, triggering, APIs, data transformation, etc.
  • Containerization: Deep expertise in container technologies (Docker/Kubernetes) supporting highly scalable and resilient distributed systems.
  • Infrastructure-as-Code with Terraform: Have experience working with modules at scale. This is a requirement for the role.
  • Cloud Native Services: hands‑on experience operating production workloads on Amazon Web Services (AWS) (e.g., EC2, Lambda, EKS, S3, RDS) and / or Microsoft Azure (e.g., Container Apps, AKS, Container Storage).
  • CI/CD & Automation: Deep history of hands‑on work with CI/CD platforms (GitHub Actions, GitLab CI) and embedding DevSecOps practices directly into operational workflows.
  • Operations, Monitoring & Observability: Experience with application level logging, troubleshooting, and tracing tools, with a proven track record operating highly available production systems.
  • AI‑Fluent Engineering: Experience with AI‑assisted engineering tools such as Claude Code or similar
  • Application Performance Management (APM): Familiarity with APM tooling and practices (e.g., Datadog, New Relic, Dynatrace, or similar) to instrument, profile, and optimize application performance in production.
  • Incident & Security Response: Demonstrated experience participating in on-call rotations, responding to production and security incidents, and executing disaster recovery procedures.
  • Communication & Collaboration: Exceptional communication, presentation, and collaboration skills, with a proven ability to coordinate across teams.
  • Education: Bachelor’s degree in Computer Science or a related technical field.
Preferred Skills (A Plus)

Familiarity with advanced cloud security tools like Wiz, Prisma Cloud, and Checkov.

The expected base salary for this position is approximately USD $145,000 - $170,000 for US-based candidates and CAD $120,000 - $145,000 for Canada-based candidates, excluding annual bonus and equity (when applicable). Salary is based on a number of factors, including market conditions, location, job‑related skills and experience, and may vary accordingly.

Diversity, Equity, Inclusion & Equal Employment Opportunity at Banyan

Banyan affirms that inequality is detrimental to our Global Teams, associates, our Operating Companies, and the communities we serve. As a collective, our goal is to impact lasting change through our actions. Together, we unite for equality and equity. Banyan is committed to equal employment opportunities regardless of any protected characteristic, including race, color, genetic information, creed, national origin, religion, sex, affectional or sexual orientation, gender identity or expression, lawful alien status, ancestry, age, marital status, or protected veteran status and will not discriminate against anyone on the basis of a disability. We support an inclusive workplace where associates excel based on personal merit, qualifications, experience, ability, and job performance.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Lead AI Architect - Data Platform
Lead AI Architect - Data Platform

Banyan Software • Seattle (WA)

On-site
USD 127,967 - 149,295
Human Resources Group Leader
Human Resources Group Leader

Banyan Software • United States

Remote
USD 160,000 - 180,000
VP of Engineering — AI-Driven Platform & Scale
VP of Engineering — AI-Driven Platform & Scale

Banyan Infrastructure • San Francisco (CA)

On-site
Confidential
VP, Engineering
VP, Engineering

Banyan Infrastructure • San Francisco (CA)

On-site
USD 200,000 - 230,000
Top tier health plans
Professional development
Flexible time off policy
+1
Lead AI Architect - Data Platform
Lead AI Architect - Data Platform

Banyan Software • United States

On-site
USD 127,967 - 170,623
Senior SRE: AI-Driven Reliability (Remote)
Senior SRE: AI-Driven Reliability (Remote)

banyansoftware • United States

Remote
USD 130,000 - 165,000
Remote role (US/Canada)
Senior AI Engineer - FinTech for Green Infrastructure
Senior AI Engineer - FinTech for Green Infrastructure

Banyan Infrastructure • San Francisco (CA)

On-site
USD 180,000 - 215,000
Entrepreneur In Residence
Entrepreneur In Residence

Banyan Software • Austin (TX)

On-site
USD 160,000 - 320,000
Equity in venture
Ownership of venture
AI-first environment
Chief Executive Officer | Vertical SaaS - Ongoing & Future Opportunities
Chief Executive Officer | Vertical SaaS - Ongoing & Future Opportunities

Banyan Software • United States

Remote
USD 250,000 - 450,000
AI Engineer (Mid-Senior Level)
AI Engineer (Mid-Senior Level)

Banyan Infrastructure • San Francisco (CA)

On-site
USD 180,000 - 215,000