Sr. Site Reliability Engineering (SRE) Lead

Cognizant

Bridgewater (MA)

Hybrid

USD 63,000 - 100,000

Full time

45 hours ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Medical/Dental/Vision/Life Insurance
401(k) plan and contributions
Paid time off

Job summary

Cognizant is seeking a Sr. Site Reliability Engineering (SRE) Lead to drive stability, resiliency, and performance of mission‑critical enterprise applications across hybrid cloud and on‑prem environments.

You will lead complex incident responses, implement preventive measures, enhance observability with Splunk, automate with Python, Shell, and Ansible, and coordinate with cross‑functional teams to ensure service levels. Note: visa transfer or sponsorship is not available now or in the future.

Qualifications

  • 10+ years in SRE, production support, or related fields.
  • Strong expertise with WebSphere and Tomcat, plus middleware.
  • AWS experience (EC2, ALB, networking, security).
  • Networking basics and load balancing concepts.
  • DB2/Oracle connectivity and performance understanding.
  • Splunk observability and Python/Shell with Ansible.

Responsibilities

  • Drive reliability and availability across hybrid infrastructure.
  • Lead incident management to minimize disruption across stacks.
  • Assess deployments for production risk and safe rollouts.
  • Root cause analyses and long-term corrective actions.
  • Build observability using Splunk and related tools.
  • Automate operations with Python, Shell, and Ansible.
  • Plan resiliency and disaster recovery exercises.
  • Collaborate with cross-functional and offshore teams.

Skills

SRE experience
Incident management
JVM performance tuning
AWS cloud
Networking concepts

Tools

WebSphere
Tomcat
IBM DB2
Oracle
JDBC
Splunk
Ansible
Python
Shell scripting

Job description

Sr. Site Reliability Engineering (SRE) Lead
About the role

As a Senior Site Reliability Engineering (SRE) Lead, you will make an impact by ensuring the stability, resiliency, and performance of mission-critical enterprise applications across hybrid cloud and on-premises environments. You will be a valued member of the Site Reliability Engineering team and work collaboratively with application development, infrastructure, security, operations, vendor partners, and global support teams to drive operational excellence and continuous reliability improvements.

Please note that this position is not eligible for visa transfer or sponsorship now or at any time in the future

In this role, you will:
  • Drive application reliability and availability across hybrid infrastructure environments, proactively identifying risks and implementing preventive measures before they impact business operations.
  • Lead complex incident management efforts, rapidly diagnosing and resolving issues across application, middleware, database, network, and cloud technology stacks to minimize disruption and reduce recovery times.
  • Assess production risk associated with infrastructure, operating system, security, middleware, and platform changes, ensuring safe and successful deployments.
  • Conduct comprehensive root cause analyses, identifying underlying issues and driving long-term corrective actions that improve system stability and prevent recurrence.
  • Build and enhance observability capabilities using Splunk and related monitoring tools to detect anomalies, correlate events, and accelerate issue resolution.
  • Automate operational processes using Python, Shell scripting, and Ansible to improve efficiency, consistency, and operational resilience.
  • Plan and execute resiliency and disaster recovery exercises to validate failover capabilities, recovery procedures, and operational readiness.
  • Partner with cross-functional teams, stakeholders, and offshore support organizations to maintain service levels, manage escalations, and drive continuous improvement initiatives.
Work model

We believe hybrid work is the way forward as we strive to provide flexibility wherever possible. Based on this role’s business requirements, this is a hybrid position requiring associates to work from a client or Cognizant office in Bridgewater, NJ as determined by business needs. Regardless of your working arrangement, we are here to support a healthy work-life balance through our various wellbeing programs.

The working arrangements for this role are accurate as of the date of posting. This may change based on the project you’re engaged in, as well as business and client requirements. Rest assured; we will always be clear about role expectations.

What you need to have to be considered
  • 10+ years of experience in Site Reliability Engineering, production support, systems engineering, or application support for large-scale distributed enterprise environments.
  • Strong expertise administering and troubleshooting Java-based application platforms, including WebSphere Application Server (WAS), Apache Tomcat, and related middleware technologies.
  • Hands-on experience diagnosing JVM performance issues using heap dumps, thread dumps, garbage collection analysis, and memory tuning techniques.
  • Proven experience supporting enterprise applications in AWS environments, including EC2, Application Load Balancers, networking, and security constructs.
  • Strong understanding of networking concepts, including TCP/IP, SSL/TLS, load balancing, firewall interactions, and timeout management.
  • Experience troubleshooting database connectivity and performance issues involving IBM DB2, Oracle, JDBC drivers, and connection pooling technologies.
  • Proficiency with observability and monitoring platforms, particularly Splunk, for log analysis, event correlation, and root cause investigation.
  • Strong automation and scripting skills using Python and Shell scripting, with experience implementing infrastructure automation and configuration management solutions such as Ansible.
These will help you stand out
  • Experience supporting IBM HTTP Server (IHS), F5 BIG-IP LTM, AWS ALB, and enterprise traffic management solutions.
  • Knowledge of IBM MQ, RabbitMQ, and other messaging and middleware technologies.
  • Experience managing TLS/SSL certificates, trust stores, certificate renewals, and enterprise certificate authority migrations.
  • Familiarity with Linux, AIX, Solaris, and Windows Server administration in enterprise environments.
  • Experience with ServiceNow, SailPoint, application inventory tools, and enterprise change management processes.
  • Demonstrated ability to lead major incident response efforts, coordinate cross-functional teams, and communicate effectively with business and technology stakeholders.

Applications will be accepted until August 28th, 2026.

Salary and Other Compensation

The annual salary for this position is between $ 63,000 to $ 99,500 depending on experience and other qualifications of the successful candidate.

This position is also eligible for Cognizant’s discretionary annual incentive program, based on performance and subject to the terms of Cognizant’s applicable plans.

Benefits
  • Medical/Dental/Vision/Life Insurance
  • Paid holidays plus Paid Time Off
  • 401(k) plan and contributions
  • Long-term/Short-term Disability
  • Paid Parental Leave
  • Employee Stock Purchase Plan

Disclaimer: The salary, other compensation, and benefits information is accurate as of the date of this posting. Cognizant reserves the right to modify this information at any time, subject to applicable law.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Sr. Site Reliability Engineer
Sr. Site Reliability Engineer

Cognizant • Bentonville (AR)

On-site
USD 90,000 - 101,000
Medical/Dental/Vision/Life Insurance
Paid holidays & PTO
401(k) plan
+3
Site Reliability Engineer
Site Reliability Engineer

Cognizant • Bentonville (AR)

On-site
USD 70,000 - 80,000
Medical/Dental/Vision/Life Insurance
Paid holidays + PTO
401(k) plan and contributions
+3
Senior DevOps Engineer with AWS
Senior DevOps Engineer with AWS

Cognizant • Town of Charlotte (NY)

Hybrid
USD 64,000 - 117,000
Medical/Dental/Vision/Life Insurance
401(k) plan and contributions
Paid holidays and PTO
+1
Sr. Java Backend Developer
Sr. Java Backend Developer

Cognizant • Arkansas

Hybrid
USD 60,000 - 94,000
Senior Java Backend Engineer (Hybrid)
Senior Java Backend Engineer (Hybrid)

Cognizant • Bentonville (AR)

Hybrid
USD 47,000 - 94,000
Health insurance
Paid time off
401(k) plan
Java Technical Lead (Onsite)
Java Technical Lead (Onsite)

Cognizant • Wilmington (DE)

On-site
USD 80,000 - 140,000
Medical/Dental/Vision/Life Insurance
401(k) plan
Paid parental leave
Senior Java Backend Engineer (Hybrid)
Senior Java Backend Engineer (Hybrid)

Cognizant • Arkansas

Hybrid
USD 46,000 - 94,000
Medical/Dental/Vision/Life Insurance
Paid holidays & PTO
401(k) plan and contributions
+3
Senior DevOps Engineer with AWS
Senior DevOps Engineer with AWS

Cognizant • Charlotte (NC)

Hybrid
USD 64,000 - 117,000
Medical/Dental/Vision/Life Insurance
Paid holidays plus PTO
401(k) plan and contributions
+3
Enterprise Compute Tower Lead (Windows & Linux Infrastructure)
Enterprise Compute Tower Lead (Windows & Linux Infrastructure)

Cognizant • Des Moines (IA)

Hybrid
USD 119,000 - 137,000
Medical/Dental/Vision/Life Insurance
401(k) plan and contribution
Paid holidays plus Paid Time Off
+3
Technical Project Manager
Technical Project Manager

Cognizant • Buffalo (NY)

On-site
USD 81,000 - 136,000
Discretionary annual incentive program
Medical, dental, and vision coverage
401(k) contributions
+2