Reliability Engineer

SES Corporation

Boston (MA)

Hybrid

USD 120,000 - 150,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Medical, Dental, and Vision
Company paid Life Insurance
401k with employer contribution
Paid Time Off
Pet Insurance

Job summary

SES Corporation is seeking a Reliability Engineer to support the U.S. Air Force Cloud One Architecture contract. This position is hybrid remote and candidates should preferably be located near Boston, MA.

The role focuses on ensuring the availability and scalability of critical systems, involving automation, incident response, and continuous improvement in reliability. The ideal candidate has significant experience in cloud platforms, incident management, and a Bachelor's or Master's degree with security clearance.

Qualifications

  • Bachelors and eight (8) years or more of experience; Masters and six (6) years or more of experience.
  • Active Secret clearance is required.
  • US citizenship is mandatory.
  • Experience with cloud platforms and managed services.
  • Hands-on experience with monitoring/observability tools is essential.

Responsibilities

  • Design, implement, and maintain highly available, fault-tolerant systems.
  • Participate in on-call rotations and lead incident response activities.
  • Automate operational tasks to reduce manual intervention and risk.
  • Ensure systems scale to meet demand through capacity planning.

Skills

Experience with cloud platforms (AWS, Azure, OCI, or GCP)
Containerized environments (Docker, Kubernetes)
Experience supporting incident management and on-call operations
Hands-on experience with monitoring/observability tools
Strong understanding of distributed systems
Capacity modeling and performance testing
Scripting/programming languages (Python, Bash, Go, PowerShell)
CI/CD pipelines and deployment automation

Education

Bachelor's degree (8 years experience) or Master's degree (6 years experience)

Tools

Monitoring/Observability tools (e.g., Prometheus, Grafana)
Infrastructure as Code (Terraform, ARM, CloudFormation)

Job description

This role supports the U.S. Air Force Cloud One Architecture and Common Shared Services contract. The Reliability Engineer is responsible for ensuring the availability, performance, scalability, and resiliency of mission-critical systems. This role applies software engineering principles to infrastructure and operations, with a strong emphasis on automation, monitoring, incident response, and continuous reliability improvement.

Location

This position is hybrid remote. Candidates will be required to work onsite as needed and are preferred to be located near Hanscom AFB (Boston, MA).

Responsibilities
System Reliability & Availability
  • Design, implement, and maintain highly available, fault‑tolerant systems in cloud and hybrid environments.
  • Define, measure, and report Service Level Indicators (SLIs), Service Level Objectives (SLOs), and error budgets.
  • Identify reliability risks and implement mitigation strategies across the system lifecycle.
  • Conduct capacity planning and performance modeling to ensure systems scale to meet demand.
Monitoring, Observability & Alerting
  • Implement and manage monitoring, logging, and tracing solutions to provide full system observability.
  • Define actionable alerting thresholds that minimize noise and enable rapid incident detection.
  • Analyze trends and metrics to proactively identify potential reliability issues.
Incident Response & Problem Management
  • Participate in on-call rotations and lead incident response activities for production systems.
  • Coordinate troubleshooting efforts across development, infrastructure, and security teams.
  • Conduct post-incident reviews (PIRs) and develop corrective and preventive action plans.
  • Track recurring issues and ensure root causes are resolved.
Automation & Engineering Excellence
  • Automate operational tasks to reduce manual intervention and operational risk.
  • Develop scripts, tools, and services that improve system reliability and reduce mean time to recovery (MTTR).
  • Promote "automation over toil" and standardize operational workflows.
Reliability-Focused Engineering
  • Participate in architecture and design reviews with an emphasis on reliability, resiliency, and recoverability.
  • Validate disaster recovery (DR) and business continuity plans; test failover mechanisms.
  • Support chaos engineering, fault injection testing, and resilience validation where appropriate.
Collaboration & Governance
  • Partner with DevOps, Platform, and Security teams to ensure reliability aligns with delivery and compliance objectives.
  • Document system reliability standards, runbooks, and operational procedures.
  • Support compliance and audit activities (e.g., FedRAMP, FISMA, internal operational controls).
Required Skills
  • Bachelors and eight (8) years or more of experience; Masters and six (6) years or more of experience. Additional experience may be accepted in lieu of degree.
  • Active Secret clearance at a minimum required to start.
  • US citizenship required.
  • Experience with cloud platforms (AWS, Azure, OCI, or GCP), including managed services.
  • Experience with containerized environments (Docker, Kubernetes).
  • Familiarity with CI/CD pipelines and deployment automation.
  • Experience with SLOs and error budgets.
  • Capacity modeling and performance testing.
  • Strong understanding of distributed systems, high‑availability architectures, Linux/Windows administration, and networking fundamentals (DNS, TCP/IP, load balancing).
  • Hands‑on experience with monitoring/observability tools (e.g., Prometheus, Grafana, ELK/Elastic, Datadog, Azure Monitor), Infrastructure as Code (Terraform, ARM, CloudFormation), and scripting/programming languages (Python, Bash, Go, PowerShell).
  • Experience supporting incident management and on‑call operations.
Preferred Skills
  • Experience with USAF Cloud One or Platform 1.
  • Experience with Zero Trust Architecture.
  • Cloud certifications in AWS, Azure, Google, or Oracle clouds.
Benefits

SES provides a competitive salary and the following benefits:

  • Medical, Dental, and Vision
  • AD&D, STD, and LTD
  • Company paid Life Insurance
  • 401k with employer contribution
  • Paid Time Off
  • Pet Insurance
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Hybrid Cloud Reliability Engineer (SRE)
Hybrid Cloud Reliability Engineer (SRE)

SES Corporation • Boston (MA)

Hybrid
USD 120,000 - 150,000
Medical, Dental, and Vision
Company paid Life Insurance
401k with employer contribution
+2
Cloud Security Architect
Cloud Security Architect

SES Corporation • Boston (MA)

Hybrid
USD 120,000 - 150,000
Medical
Dental
Vision
+3
Cloud Security Architect Jr
Cloud Security Architect Jr

SES Corporation • Boston (MA)

Hybrid
USD 150,000 - 190,000
Medical insurance
Dental insurance
Vision insurance
+2
Cloud Security Architect Jr
Cloud Security Architect Jr

Internetwork Expert • Boston (MA)

Hybrid
USD 130,000 - 190,000
Medical
Dental
Vision
+3
DevOps Engineer
DevOps Engineer

Systems Engineering Solutions Corporation • Montgomery (AL)

Hybrid
USD 110,000 - 160,000
Medical
Dental
Vision
+7
Cloud Security Architect
Cloud Security Architect

Systems Engineering Solutions Corporation • Boston (MA)

Hybrid
USD 120,000 - 150,000
Medical
Dental
Vision
+2
E01-L03 Reliability Engineer IV
E01-L03 Reliability Engineer IV

Talentwerx.Io • Boston (MA)

Hybrid
USD 90,000 - 120,000
Hybrid Cloud Scrum Master: OCI & GCP for DoD
Hybrid Cloud Scrum Master: OCI & GCP for DoD

SES Corporation • Boston (MA)

On-site
USD 60,000 - 80,000
Reliability Engineer
Reliability Engineer

Jobtailor • Reston (VA)

On-site
USD 140,000 - 190,000
OCI/GCP Scrum Master
OCI/GCP Scrum Master

SES Corporation • Boston (MA)

On-site
USD 100,000 - 130,000
Medical
Dental
Vision
+2