Senior Site Reliability Engineer (SRE)

Kids for the Future

United States

Hybrid

USD 130,000 - 195,000

Full time

13 days ago
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

Healthcare benefits
401(k) matching
Paid time off

Job summary

CenCore Group seeks a Senior Site Reliability Engineer to implement, secure, and operate the cloud infrastructure behind our enterprise SaaS platform. You will own platform reliability, manage AWS and Kubernetes environments, and drive resilience across production systems.

The role requires Top Secret/SCI-eligible clearance, strong incident response, and collaboration with software engineering to improve performance, stability, and deployment reliability. Remote/hybrid work options apply.

Qualifications

  • Active Top Secret clearance with SCI eligibility.
  • Experience with AWS cloud services and production cloud operations.
  • Experience administering or operating Kubernetes environments.
  • Working knowledge of infrastructure reliability and incident response practices.
  • Experience implementing monitoring, logging, alerting, or observability tools.

Responsibilities

  • Manage, maintain, and improve AWS-based cloud infrastructure.
  • Operate and support Kubernetes environments (EKS).
  • Own platform reliability, scalability, availability, and disaster recovery readiness.
  • Design and support cloud networking, load balancing, and routing components.
  • Implement monitoring, alerting, logging, and observability solutions.
  • Document operational standards, SLOs, incident response processes, and best practices.
  • Collaborate with engineering and product teams to improve performance and deployment reliability.
  • Apply security practices across IAM, encryption, secrets management, and vulnerability remediation.
  • Support production operations and participate in incident resolution as needed.

Skills

Active Top Secret clearance with SCI
AWS cloud experience
Kubernetes administration
Security best practices
Documentation of processes
Incident response coordination

Tools

AWS
Kubernetes (EKS)
Terraform
Datadog/CloudWatch

Job description

  • Job Category Information Technology, Platform Engineering, Site Reliability Engineering
  • Industry Computer Software , SaaS, National Security
  • Employee Type FT Exempt
  • Manage Others No
Description

The Senior Site Reliability Engineer (SRE) will implement, secure, and operate the cloud infrastructure that supports CenCore Group’s proprietary enterprise SaaS platform. This role is responsible for maintaining a scalable, highly available, secure, and reliable cloud environment as the platform grows and supports enterprise customers.

Key Responsibilities
  • Manage, maintain, and improve AWS-based cloud infrastructure supporting enterprise SaaS operations.
  • Operate and support Kubernetes environments, including Amazon EKS.
  • Own platform reliability, scalability, availability, disaster recovery readiness, and operational resilience.
  • Design and support cloud networking, load balancing, routing, traffic management, and related infrastructure components.
  • Implement and maintain monitoring, alerting, logging, and observability solutions to support proactive issue detection and response.
  • Establish and document operational standards, Service Level Objectives (SLOs), incident response processes, and reliability best practices.
  • Partner with software engineering and product teams to improve application performance, platform stability, and deployment reliability.
  • Apply security best practices across IAM, secrets management, encryption, vulnerability remediation, access controls, and production operations.
  • Support production operations, troubleshoot critical issues, and participate in incident resolution as needed.
Required Qualifications
  • Active Top Secret clearance withSCI eligibility
  • Professional experience supporting cloud infrastructure, site reliability, DevOps, platform engineering, or systems engineering functions.
  • Hands-on experience with AWS cloud services and production cloud operations.
  • Experience administering or operating Kubernetes environments.
  • Working knowledge of infrastructure reliability, availability, scalability, incident response, and operational support practices.
  • Experience implementing monitoring, logging, alerting, or observability tools.
  • Ability to troubleshoot complex production issues and coordinate resolution across technical teams.
  • Strong understanding of cloud security fundamentals, including identity and access management, encryption, secrets management, and vulnerability remediation.
  • Ability to document technical processes, standards, and operational procedures.
Preferred Qualifications
  • Experience with AWS services such as EKS, ALB, VPC, CloudFront, Route 53, RDS/Aurora, S3, and IAM.
  • Experience with Terraform or other Infrastructure as Code tools.
  • Experience with monitoring platforms such as Datadog, CloudWatch, Grafana, Prometheus, or similar tools.
  • PostgreSQL administration, performance tuning, or database operations experience.
  • Experience supporting enterprise SaaS, cloud-native applications, or customer-facing production platforms.
  • Experience developing disaster recovery, operational readiness, or production support documentation.
  • Cloud infrastructure operations and automation
  • Platform reliability, scalability, and performance optimization
  • Kubernetes administration and containerized application support
  • Monitoring, observability, and incident response
  • Cloud security and operational risk awareness
  • Technical troubleshooting and root cause analysis
  • Cross-functional collaboration with engineering, product, and operations teams
  • Clear technical documentation and process improvement
Work Environment and Physical Requirements

This role is primarily performed in a professional office or remote technology environment, depending on business needs and position requirements. Work involves regular use of a computer, collaboration tools, and cloud-based systems. The position may require participation in production support, incident response, or after-hours troubleshooting as needed. Physical requirements are generally sedentary and include prolonged periods of sitting, computer use, and communicating with internal teams.

Summary
Company Overview

At CenCore Group, we deliver security solutions that support critical national security missions. We specialize in designing, building, securing, and maintaining advanced technology environments at the intersection of emerging technology and national security. We are seeking a dependable, cleared professional to join our team in support of secure physical security operations.

Compensation Overview

Eligible employees may enroll in company-sponsored medical, dental, and vision benefits in accordance with plan terms and enrollment requirements. The company contributes toward employee healthcare premiums, subject to plan provisions.

Employees are also eligible to participate in the company's 401(k) retirement savings plan, including employer matching contributions in accordance with plan guidelines and eligibility requirements.

Paid time off, sick leave, and holiday benefits are provided in accordance with company policy, applicable contract requirements, and federal, state, and local laws. Benefit eligibility and accrual rates may vary based on position, location, and length of service.

The company is committed to equitable pay practices and does not seek or rely on an applicant's wage history when making compensation decisions.

Benefits may be modified from time to time in accordance with company policy and applicable law.

Equal Opportunity Employer

CenCore Group is an equal opportunity employer. We hire based on merit and qualifications and do not discriminate on the basis of race, color, religion, sex, sexual orientation, gender identity, national origin, age, disability, veteran status, genetic information, pregnancy, childbirth or related medical conditions, or any other status protected by applicable federal, state, or local law.

Additional Information
  • Required Security Clearance Top Secret/SCI
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior SRE - Cloud Infra, AWS & Kubernetes
Senior SRE - Cloud Infra, AWS & Kubernetes

Kids for the Future • United States

Hybrid
USD 130,000 - 195,000
Healthcare benefits
401(k) matching
Paid time off
Site Security Manager (SSM) - QTS / CO / MW RZ
Site Security Manager (SSM) - QTS / CO / MW RZ

Kids for the Future • Denver (CO)

On-site
USD 125,000 - 130,000
Site Security Manager (SSM) - QTS
Site Security Manager (SSM) - QTS

Kids for the Future • Aurora (CO)

On-site
USD 125,000 - 130,000
Site Security Manager
Site Security Manager

Kids for the Future • Cheyenne (WY)

On-site
USD 70,000 - 100,000
Site Reliability Engineer
Site Reliability Engineer

VantageScore® • San Francisco (CA)

On-site
USD 150,000
Medical insurance
Dental insurance
401(k) plan
+1
Site Reliability Engineer - Top Secret (req-236)
Site Reliability Engineer - Top Secret (req-236)

Cathexisfederal • Tysons (VA)

On-site
USD 100,000 - 160,000
Performance Bonuses
Medical Insurance
Dental Insurance
+4
Lead Site Reliability Engineer (SRE)
Lead Site Reliability Engineer (SRE)

IPSecure, Inc • San Antonio (TX)

Hybrid
USD 140,000 - 190,000
Medical
Dental
Vision
+10
Site Lead
Site Lead

CENCORE ASSOCIATES • Bluffdale (UT)

On-site
USD 95,000 - 135,000
Competitive salary
Growth opportunities
OT/shift premiums
K8 platform engineer
K8 platform engineer

Seneca Resources Company, LLC • Washington

On-site
USD 185,000 - 230,000
Performance bonuses
Company-paid training/certifications
Referral bonuses
+1
Site Reliability Engineer
Site Reliability Engineer

Skyhigh Security • Frisco (TX)

Hybrid
USD 110,000 - 140,000
Retirement Plans
Medical, Dental and Vision Coverage
Paid Time Off
+2