Senior Cloud Site Reliability Engineer (SRE)

Peraton

United States

Remote

USD 104,000 - 166,000

Full time

4 days ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Peraton is seeking a Senior Cloud Site Reliability Engineer to design and develop Python-based AWS cloud solutions and reliability tools for the Cloud Foundation Services platform. This remote role focuses on applying IaC with Terraform and building scalable, reusable reliability utilities for the Federal Reserve System.

You will implement observability with Grafana and CloudWatch, develop CI/CD pipelines, and define SRE standards while collaborating in Agile teams to deliver secure, reliable

Qualifications

  • U.S. citizen and ability to obtain Public Trust clearance
  • Bachelor's degree with 8+ years of experience or equivalent education/experience
  • 5+ years of advanced Python development for AWS
  • 7+ years of software development focused on reliability
  • 3+ years of hands-on AWS environment experience (EC2, VPC, S3, Lambda, IAM, CloudFormation, EventBridge, Step Functions)
  • 3+ years applying SRE principles including observability, toil automation, SLIs/SLOs
  • Expert-level Terraform IaC with module development and state management
  • Strong CI/CD and automated testing experience
  • Experience with Grafana and AWS CloudWatch
  • Experience defining and managing SLOs/SLIs; RCA/postmortem practices
  • Familiarity with ITSM processes, resilience testing and chaos engineering
  • Experience in Agile/Scaled Agile environments

Responsibilities

  • Design, develop, and maintain reliability solutions and SRE utilities using Python in AWS to improve platform reliability
  • Build automation scripts, APIs, and utilities in Python to reduce toil
  • Implement observability and monitoring solutions (Grafana, CloudWatch) with Python-based metrics and dashboards
  • Develop Terraform IaC to manage AWS resources and optimize costs
  • Create and optimize CI/CD pipelines and automated tests for rapid delivery
  • Define SRE standards, metrics (SLI/SLOs) and guidelines for adoption
  • Apply version control, code reviews, TDD, and documentation across development
  • Participate in incident management; on-call rotation and root-cause postmortems
  • Stay current with AWS services and drive adoption of new cloud-native solutions
  • Collaborate in Agile/Scaled Agile environments to deliver integrated cloud automation
  • Produce blameless postmortems with actionable follow-ups

Skills

Python
Terraform
AWS
SRE principles
Observability
CI/CD
DevOps practices
Agile
On-call support
Cost optimization

Education

Bachelor's Degree
High School diploma or equivalent with 12 years of experience

Tools

Terraform
Grafana
AWS CloudWatch
EC2
S3
Lambda
EventBridge
CloudFormation

Job description

Responsibilities

Peraton is looking for a Senior Cloud Site Reliability Engineer (SRE) who will be responsible for designing and developing advanced Python-based AWS cloud solutions and engineering reliability tools for the Cloud Foundation Services (CFS) platform in the Infrastructure, Platforms & Operations organization. This person will apply software engineering practices including Infrastructure-as-Code (IaC) with Terraform to build scalable, reusable solutions and utilities that enhance platform reliability across the Federal Reserve System.

Work Location: This is a remote position

What You Will Do

  • Design, develop, and maintain reliability solutions and SRE utilities using Python in AWS environments to reduce toil, improve cloud platform reliability, and industrialize SRE practices across the system
  • Build automation scripts, APIs, and utilities in Python to reduce toil and improve platform reliability.
  • Implement observability and monitoring solutions (Grafana, AWS CloudWatch) leveraging Python for custom metrics and dashboards
  • Build and optimize Infrastructure as Code (IaC) using Terraform to manage AWS resources related to SRE solutions, incorporating cost-efficient design principles
  • Optimize Infrastructure as Code (IaC) with Terraform for AWS resources, integrating Python-based workflows.
  • Develop CI/CD pipelines and automated testing to ensure code quality, reliability, and rapid delivery of the solutions
  • Define SRE standards, best practices, and guidelines for adoption across teams; establish SRE metrics like SLI, SLOs, etc.
  • Apply software engineering best practices including version control, code reviews, test-driven development, and documentation to all development
  • Participate in incident management and on-call rotation, providing technical support for SRE tools, troubleshooting production issues, and collaborating with teams to reduce incident recurrence through proactive detection and pattern analysis
  • Stay current with emerging AWS services, SRE methodologies, and cloud-native development technologies, and drive adoption of innovative solutions
  • Collaborate within Agile and Scaled Agile frameworks with cross-functional teams to deliver integrated cloud automation solutions
  • Produce clear, blameless postmortems with actionable items and documented failure scenarios
Qualifications

Basic Qualifications

  • Must be a U.S. Citizen with the ability to obtain and maintain the required Public Trust level Clearance
  • Bachelors Degree and 8 years of experience, or a High School diploma or equivalent and 12 years of experience
  • Must have 5 years of advanced Python development experience, building enterprise-grade, highly available tools, APIs, and utilities for AWS
  • 7 years of extensive experience in software development with focus on reliability and platform engineering
  • 3 years of hands-on experience developing solutions in AWS environments with deep understanding of core services (EC2, VPC, S3, Lambda, IAM, CloudFormation, EventBridge, Step Functions etc.) and resource cost optimization
  • 3 years of experience applying SRE principles including observability, toil automation, SLIs/SLOs and reliability engineering
  • Expert-level proficiency with Infrastructure as Code (IaC) using Terraform, including module development and state management
  • Strong experience with CI/CD pipelines, automated testing frameworks, and DevOps practices
  • Experience with observability tools and practices including Grafana, AWS CloudWatch, AWS Canary
  • Experience defining, implementing, and managing SLOs/SLIs and error budgets; familiarity with conducting RCAs and producing postmortem documentation
  • Working experience in Agile and Scaled Agile environments and familiarity with ITSM processes (incident, change, and problem management), resilience testing and chaos engineering practices

Preferred Qualifications

  • Experience with GoLang or additional programming languages is a plus
  • Bachelors Degree in Computer Science, Information Systems, or similar
Peraton Overview

Peraton is a next-generation national security company that drives missions of consequence spanning the globe and extending to the farthest reaches of the galaxy. As the world's leading mission capability integrator and transformative enterprise IT provider, we deliver trusted, highly differentiated solutions and technologies to protect our nation and allies. Peraton operates at the critical nexus between traditional and nontraditional threats across all domains: land, sea, space, air, and cyberspace. The company serves as a valued partner to essential government agencies and supports every branch of the U.S. armed forces. Each day, our employees do the can't be done by solving the most daunting challenges facing our customers. Visit peraton.com to learn how we're keeping people around the world safe and secure.

Target Salary Range

$104,000 - $166,000. This represents the typical salary range for this position. Salary is determined by various factors, including but not limited to, the scope and responsibilities of the position, the individual's experience, education, knowledge, skills, and competencies, as well as geographic location and business and contract considerations. Depending on the position, employees may be eligible for overtime, shift differential, and a discretionary bonus in addition to base pay.

EEO

EEO: Equal opportunity employer, including disability and protected veterans, or other characteristics protected by law.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Cloud Site Reliability Engineer (SRE)
Senior Cloud Site Reliability Engineer (SRE)

Peraton • Reston (VA)

Remote
USD 104,000 - 166,000
Senior Cloud Site Reliability Engineer (SRE)
Senior Cloud Site Reliability Engineer (SRE)

Peraton • Herndon (VA)

Remote
USD 104,000 - 166,000
Lead Site Reliability Engineer
Lead Site Reliability Engineer

Peraton • Herndon (VA)

On-site
USD 112,000 - 179,000
Lead Site Reliability Engineer
Lead Site Reliability Engineer

Peraton • Reston (VA)

On-site
USD 112,000 - 179,000
Senior Cloud Site Reliability Engineer (SRE)
Senior Cloud Site Reliability Engineer (SRE)

Peraton • Washington

Remote
USD 130,000 - 170,000
Remote work
External Job Posting Title Site Reliability Engineer
External Job Posting Title Site Reliability Engineer

Peraton • Washington

On-site
USD 112,000 - 179,000
External Job Posting Title Lead Site Reliability Engineer
External Job Posting Title Lead Site Reliability Engineer

Peraton • Northern (KY)

Hybrid
USD 112,000 - 179,000
Cloud Platform Technical Lead
Cloud Platform Technical Lead

Peraton • Reston (VA)

Remote
USD 112,000 - 179,000
Software Engineer - Cloud, Lead Associate - TS/SCI w/poly
Software Engineer - Cloud, Lead Associate - TS/SCI w/poly

Peraton • Laurel (MD)

On-site
USD 176,000 - 282,000
Heavily subsidized medical, dental, &/
Vision coverage for employees & depend
25 days PTO annually
+2
External Job Posting Title Cloud Platform Technical Lead
External Job Posting Title Cloud Platform Technical Lead

Peraton • Northern (KY)

On-site
USD 112,000 - 179,000