Senior SRE - Government Cloud Operations

Cato Networks

United States

On-site

USD 140,000 - 200,000

Full time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Benefits offered by this job

Health insurance
401(k)
Stock options
Health Savings Account
Flexible time-off
Parental leave
Disability benefits

Job summary

Cato Networks is seeking a Senior Site Reliability Engineer to build and sustain regulated cloud platforms with FedRAMP High/IL4 deployments. You will oversee production operations, design scalable infrastructure, and lead incident response while ensuring compliance and security across Kubernetes, Linux, and cloud environments.

You will collaborate with Security, Compliance, and Engineering to mature observability, automate controls, and implement secure CI/CD pipelines.

Qualifications

  • 7+ years of experience in SRE/Production Engineering/Cloud Operations.
  • Experience with regulated environments such as FedRAMP Moderate/High, DoD IL4/IL5, or GovCloud.
  • Strong knowledge of NIST 800-53, vulnerability remediation SLAs, and audit evidence generation.
  • Deep experience with Infrastructure as Code (Terraform), GitOps workflows, and secure CI/CD pipelines.

Responsibilities

  • Own production operations for mission-critical services across distributed systems.
  • Design, build, and operate highly available cloud infrastructure for regulated environments.
  • Lead major incident response, root cause analysis, and postmortem remediation.
  • Operationalize compliance requirements across Kubernetes, Linux, container runtimes, and cloud infra.
  • Support audit readiness, vulnerability management, and continuous monitoring.
  • Develop automation to maintain platform compliance and observability.
  • Collaborate with Security, Compliance, and Engineering to improve reliability and deployment safety.

Skills

SRE experience
Incident response
Kubernetes
Prometheus Grafana
Terraform
GitOps
Python/Go/Bash
Security/compliance
Linux

Tools

Terraform
GitOps
Kubernetes
CI/CD pipelines
Prometheus
Grafana

Job description

Welcome to the future of cloud networking and security! Cato Networks is the first company to converge enterprise networking and security into one centralized and global service that is delivered by cloud. It is led by networking and security pioneer Shlomo Kramer (Check Point, Imperva) and early investor (Palo Alto Networks, Exabeam, Trusteer and more). Cato’s unique technology inspired a brand-new product category, later named “SASE” by Gartner and a market expected to reach $28.5 billion by 2028. This is your opportunity to get on the rocket ship and join a company that is building a cutting-edge enterprise network and secure cloud platform, and is on a fast track to becoming the worldwide market leader - don’t miss it.

Description

We’re seeking a Senior Site Reliability Engineer with hands-on experience building and sustaining regulated cloud platforms through FedRAMP High / IL4 operational lifecycles, including continuous monitoring and post-ATO operational management. In this critical role, you will support our growing operations, network, and systems environments. You will play a pivotal role in administering internal platforms while participating in key architectural and operational decisions. This position offers the opportunity to innovate, establish best-practice processes, and continuously improve the reliability, security, and compliance posture of our regulated cloud environments.

Responsibilities
  • Own production operations for mission-critical services, including availability, latency, scalability, and operational health across complex distributed systems.
  • Design, build, and operate highly available cloud infrastructure supporting regulated environments, including FedRAMP High / IL4+ deployments.
  • Lead major incident response, root cause analysis, and postmortem remediation; drive operational maturity through change governance, disaster recovery testing, and service resiliency programs.
  • Operationalize compliance requirements, including NIST 800-53 controls and STIG baselines, across Kubernetes platforms, Linux systems, container runtimes, and cloud infrastructure.
  • Support regulated environment readiness through audit preparation, evidence collection, vulnerability management, configuration management, and continuous monitoring activities.
  • Develop automation and tooling to continuously assess and maintain platform compliance posture; contribute to immutable, reproducible infrastructure patterns that simplify regulatory sustainment.
  • Implement and maintain secure CI/CD pipelines and infrastructure-as-code practices aligned with security and compliance requirements.
  • Improve observability across infrastructure and applications through metrics, logging, tracing, and alerting; integrate compliance telemetry and configuration auditing into operational workflows.
  • Partner with Security, Compliance, and Engineering teams to improve service reliability, deployment safety, and operational maturity throughout the software lifecycle.
Requirements
  • 7+ years of experience in Site Reliability Engineering, Production Engineering, Cloud Operations, or Infrastructure Engineering.
  • Hands-on experience operating cloud infrastructure in regulated environments such as FedRAMP Moderate/High, DoD IL4/IL5, or equivalent, including AWS GovCloud or other isolated government cloud environments.
  • Experience supporting cloud authorization efforts (ATO) and sustaining environments post-authorization through continuous monitoring, including FedRAMP monthly reporting, vulnerability tracking, and control assessment activities.
  • Strong knowledge of NIST 800-53 controls, vulnerability remediation SLAs, secure configuration management, and audit evidence generation.
  • Deep experience with Infrastructure as Code (Terraform preferred), GitOps workflows, and secure CI/CD pipelines, including container hardening and image security practices.
  • Proficiency in Python, Go, or Bash for operational automation and tooling development.
  • Proficiency with cloud-native technologies including Kubernetes, Prometheus, and Grafana, along with a solid understanding of Linux/Unix operating systems.
  • Experience supporting production operations for SaaS, cloud service provider, or multi-tenant platforms at scale.
  • Ability to communicate operational risk and compliance posture clearly to both technical and non-technical stakeholders.
Preferred Qualifications
  • Experience working directly with 3PAOs, auditors, or compliance assessors during authorization and continuous monitoring cycles.
  • Familiarity with STIG implementation across Kubernetes, Linux systems, and container runtimes.
  • Understanding of Zero Trust architectures and secure access platforms.
  • Experience with operational resilience exercises and disaster recovery validation.
Benefits
  • Health/vision/dental insurance
  • 401(k)
  • Stock options
  • Health Savings/Flexible Spending Accounts
  • Flexible time-off
  • Paid parental leave
  • Disability benefits

As an EEO/Affirmative Action Employer all qualified applicants will receive consideration for employment without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, disability, veteran status.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Gov Cloud SRE - FedRAMP, K8s & Secure CI/CD
Senior Gov Cloud SRE - FedRAMP, K8s & Secure CI/CD

Cato Networks • United States

On-site
USD 140,000 - 200,000
Health insurance
401(k)
Stock options
+4
Product Support T3 Engineer
Product Support T3 Engineer

Cato Networks • Northern (KY)

Hybrid
USD 80,000 - 110,000
Health insurance
Vision insurance
Dental insurance
+6
Senior DevOps Engineer/Site Reliability Engineer-East Coast
Senior DevOps Engineer/Site Reliability Engineer-East Coast

Stellar Cyber • North Carolina

On-site
USD 165,000 - 215,000
Pre‑IPO Stock Options
Medical, Dental & Vision care
401(k)
+2
Professional Services Customer Success Engineering Manager
Professional Services Customer Success Engineering Manager

Clutch Canada • United States

Hybrid
USD 110,000 - 150,000
Health/vision/dental insurance
401(k)
Stock options
+3
Manager, Strategic Customer Success - AI Security
Manager, Strategic Customer Success - AI Security

Cato Networks • Atlanta (GA)

On-site
USD 140,000 - 220,000
Health insurance
Vision insurance
Dental insurance
+4
Facilities Operations Specialist
Facilities Operations Specialist

Cato Networks • United States

On-site
USD 80,000 - 100,000
AI Security - Solutions Architect
AI Security - Solutions Architect

Clutch Canada • Atlanta (GA)

On-site
USD 100,000 - 130,000
Customer Success Manager - Chicago
Customer Success Manager - Chicago

Clutch Canada • Chicago (IL)

On-site
USD 160,000 - 200,000
Health insurance
401(k)
Stock options
+4
Principal SRE Engineer (US Citizen)
Principal SRE Engineer (US Citizen)

Palo Alto Networks, Inc. • Santa Clara (CA)

On-site
USD 147,000 - 238,000
Site Reliability Engineer - Top Secret (req-236)
Site Reliability Engineer - Top Secret (req-236)

Cathexisfederal • Tysons (VA)

On-site
USD 100,000 - 160,000
Performance Bonuses
Medical Insurance
Dental Insurance
+4