Senior Site Reliability Engineer: Build Resilient, Scalable Systems

Capgemini

New York (NY)

On-site

USD 86,000 - 127,000

Full time

43 hours ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Capgemini is seeking a Site Reliability Engineer to design, build, and operate resilient, scalable services across our platform. This role emphasizes DevOps, SRE principles, and software engineering to improve system reliability and automate operations.

The ideal candidate will drive observability, CI/CD automation, and incident response while partnering with engineering teams to embed reliability throughout the software lifecycle. Strong cloud and on-prem experience with IaC tools is required.

Qualifications

  • 9+ years of demonstrated experience developing and designing products around Site Reliability Engineering principles to improve stability and platform availability for containerized workloads and on-premises services
  • Experience managing and interpreting large datasets using query languages and creating dashboards and reports with Power BI and Grafana
  • Strong background in managing cloud and on-premises infrastructure using IaC tools, including Terraform, and CloudFormation
  • Hands-on experience building, operating, monitoring, logging, and alerting distributed systems at scale using Datadog and Splunk
  • Experience supporting DevOps practices for service delivery and operations using Jenkins, Azure DevOps, Team Foundation Version Control, and CI/CD automation
  • Experience developing software and automation solutions to support application delivery, operations, and repeatable business processes using Python
  • Knowledge of scalability and resiliency practices for applications deployed on AWS and Azure, including Lambda and API Gateway
  • Strong development experience in scripting, automation, and integration across Linux and Windows-based environments

Responsibilities

  • Design, build and operate resilient, scalable systems using DevOps, SRE, and software development best practices
  • Deliver high-availability services through automation, infrastructure as code, and proactive reliability engineering
  • Improve monitoring, logging, alerting, and observability for distributed systems
  • Support CI/CD automation, deployment workflows, and production tooling to reduce operational toil
  • Drive incident response, root cause analysis, and recovery improvements to minimize downtime
  • Partner with engineering teams to embed reliability into the software development lifecycle
  • Automate provisioning, configuration, and self-healing across cloud and on-prem environments
  • Validate resiliency and performance through testing, chaos engineering, and capacity planning

Skills

Node.js
Python
DevOps
Jenkins
AWS
Kubernetes
Terraform
CloudFormation
Datadog
Splunk
CI/CD
Automation

Tools

Docker

Job description

Capgemini is seeking a Site Reliability Engineer to design, build, and operate resilient, scalable services across our platform. This role emphasizes DevOps, SRE principles, and software engineering to improve system reliability and automate operations.

The ideal candidate will drive observability, CI/CD automation, and incident response while partnering with engineering teams to embed reliability throughout the software lifecycle. Strong cloud and on-prem experience with IaC tools is required.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Onsite SRE & DevOps Engineer
Onsite SRE & DevOps Engineer

Capgemini • Town of Texas (WI)

Hybrid
USD 86,000 - 127,000
Benefits package
SRE & DevOps Engineer: Scale, Automate, Secure
SRE & DevOps Engineer: Scale, Automate, Secure

Capgemini • Dallas (TX)

On-site
USD 86,000 - 127,000
Paid time off
Health insurance
401(k)
Senior Site Reliability Engineer: Build Resilient, Scalable Systems
Senior Site Reliability Engineer: Build Resilient, Scalable Systems

Inspire-Brands • Atlanta (GA)

On-site
USD 140,000 - 190,000
Senior Site Reliability Engineer: Build Resilient Systems
Senior Site Reliability Engineer: Build Resilient Systems

IRB USA Inspire Resources • Atlanta (GA)

On-site
USD 130,000 - 180,000
Site Reliability Engineer – Cloud & On-Prem Reliability
Site Reliability Engineer – Cloud & On-Prem Reliability

Charles River Associates International • Boston (MA)

Hybrid
USD 130,000 - 150,000
Healthcare benefits
Paid time off
401(k) matching
Senior SRE: Build Resilient, Scalable Systems
Senior SRE: Build Resilient, Scalable Systems

Methodic • San Francisco (CA)

On-site
USD 140,000 - 210,000
Senior SRE II: Automate & Scale Cloud Reliability
Senior SRE II: Automate & Scale Cloud Reliability

Talentify • Birmingham (AL)

On-site
USD 90,000 - 120,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Axiom Pursuits • San Francisco (CA)

On-site
USD 150,000 - 180,000
Senior Site Reliability Engineer — Cloud, Resilience & Automation
Senior Site Reliability Engineer — Cloud, Resilience & Automation

Compunnel Inc. • Denton (TX)

Hybrid
USD 120,000 - 150,000
Junior Site Reliability Engineer - Automation & Reliability
Junior Site Reliability Engineer - Automation & Reliability

IBM • Tucson (AZ)

On-site
USD 110,000 - 160,000