Site Reliability Engineer - Vice President

Citi

New York (NY)

Hybrid

USD 37,000 - 66,000

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Citi is hiring a Site Reliability Engineer at Vice President level in a hybrid role based in Pune, India. You will lead SRE practices, ensure application resilience, and drive end-to-end observability and automation across services.

You will collaborate with cross-functional teams, implement cloud-native resiliency patterns, and mentor staff while shaping disaster recovery and SRE strategy in a regulated enterprise environment.

Qualifications

  • 13+ years of SRE experience with SLOs/SLIs, error budgets, and toil reduction.
  • Proven disaster recovery planning and resiliency testing for large systems.
  • Experience deploying and troubleshooting apps on OpenShift/Kubernetes.
  • Hands-on use of modern observability tools (Prometheus, Grafana, Loki, Mimir, Tempo, AppDynamics).

Responsibilities

  • Own daily operations, architecture resilience, and end-to-end SRE execution.
  • Drive observability strategy across metrics, logging, and tracing.
  • Collaborate with development for cloud-native resiliency patterns.
  • Lead recovery testing, automation, and One Touch Recovery initiatives.

Skills

SRE concepts (SLOs/SLIs)
Disaster Recovery
OpenShift
Kubernetes
Observability tools – Prometheus/Graf
Loki/Mimir/Tempo/AppDynamics
IaC (Ansible/Terraform)
Helm charts
Cloud platforms (GCP/AWS/Azure)
Java/Python/Go
Agile methodologies

Tools

OpenShift
Kubernetes
Prometheus
Grafana
Loki
Mimir
Tempo
AppDynamics
Ansible
Terraform
Helm

Job description

Site Reliability Engineer - Vice President

Location(s): Pune, Maharashtra, India

Job Type: Hybrid

Posted: Sep. 01, 2026

Discover your future at Citi

Working at Citi is far more than just a job. A career with us means joining a team of approximately 219,000 dedicated people from around the globe. At Citi, you’ll have the opportunity to grow your career, give back to your community and make a real impact.

Job Overview

The Site Reliability Engineer (SRE) is a strategic professional accountable for the daily operations, architectural resilience, and overall implementation of SRE principles in a complex, critical, and largescale multi-disciplinary environment. This role requires a comprehensive understanding of multiple technology domains and their interaction to achieve business objectives. As a recognized technical authority, you will apply an in depth understanding of the business impact of technical contributions and provide advice and counsel on strategic solutions.

We are seeking a passionate and experienced SRE to join our Production Management team. In this role, you will be instrumental in enhancing the reliability, performance, and efficiency of our Applications and Services. You will drive our strategy for end-to-end observability and resiliency, collaborating across the organization to ensure our services are stable, scalable, and fault tolerant. This is a key role that will influence strategic decisions and foster a culture of technical excellence and accountability.

Key Responsibilities
Culture & Strategy
  • Foster a culture of transparency, innovation, and accountability that encourages continuous improvement.
  • Communicate the progress and impact of SRE initiatives to stakeholders at all levels.
  • Operate effectively within a highly regulated environment, ensuring compliance with all relevant requirements.
Resiliency & Recovery
  • Ensure critical business applications meet stringent operational resilience requirements, including adherence to defined impact tolerances.
  • Oversee advanced recovery testing, including Production Swing Tests, Data Recovery Tests, and chaos engineering practices.
  • Drive the adoption and development of automation, such as One Touch Recovery solutions, to minimize recovery time.
  • Partner with development teams to leverage cloud native services and established resiliency patterns to enhance application reliability.
Observability & Performance
  • Collaborate across the organization to develop and scale observability solutions using modern tools for metrics, logging, and tracing.
  • Partner with development teams to effectively instrument applications, providing deep insights into system health and performance.
Essential Skills
  • 13 + years of deep understanding of SRE concepts, including SLOs, SLIs, error budgets, and toil reduction.
  • Demonstrable experience with Disaster Recovery planning, resiliency testing, and fault tolerant distributed system design.
  • Proficiency in deploying, managing, and troubleshooting applications on OpenShift/Kubernetes.
  • Hands on experience with modern observability tools (e.g., Prometheus, Grafana, Loki, Mimir, Tempo, AppDynamics).
  • Experience with Infrastructure as Code (IaC), configuration management, and automation tools (e.g., Ansible, Terraform).
  • Experience creating, modifying, and managing Helm charts for application deployment.
Desired Skills
  • Experience with major public cloud providers (e.g., Google Cloud, AWS, Azure).
  • Proven experience delivering software and infrastructure using Agile frameworks.
  • Experience presenting technical strategy to senior and executive level audiences.
  • Experience writing or maintaining code in Java, Python, Go, or similar languages.
Qualifications
  • Significant professional experience in production management, software development, or an equivalent field, with a strong focus on Site Reliability Engineering.
  • Expertise in analyzing complex application, database, network, and OS issues within large scale, customer facing systems.
  • A service-oriented attitude combined with excellent problem-solving and strategic thinking skills.
  • Strong communication and diplomacy skills, with a proven ability to work effectively across multiple business and technical teams.
Job Family Group:

Technology

Job Family:

Applications Support

Time Type:

Full time

Most Relevant Skills

Please see the requirements listed above.

Other Relevant Skills

For complementary skills, please see above and/or contact the recruiter.

Citi is an equal opportunity employer, and qualified candidates will receive consideration without regard to their race, color, religion, sex, sexual orientation, gender identity, national origin, disability, status as a protected veteran, or any other characteristic protected by law.

If you are a person with a disability and need a reasonable accommodation to use our search tools and/or apply for a career opportunity review Accessibility at Citi. View Citi’s EEO Policy Statement and the Know Your Rights poster.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

SRE Observability Technical Lead - Vice President
SRE Observability Technical Lead - Vice President

Citi • New York (NY)

Hybrid
USD 42,000 - 73,000
VP SRE — Lead Resilience, Observability & Reliability
VP SRE — Lead Resilience, Observability & Reliability

Citi • New York (NY)

Hybrid
USD 37,000 - 66,000
Lead Site Reliability Engineer (SRE) / Principal Site Reliability Engineer (SRE)
Lead Site Reliability Engineer (SRE) / Principal Site Reliability Engineer (SRE)

Mindlance • Irving (TX)

Hybrid
USD 120,000 - 160,000
Engineering - SRE Platforms - Site Reliability Engineer - Vice President - Dallas
Engineering - SRE Platforms - Site Reliability Engineer - Vice President - Dallas

The Goldman Sachs Group • Dallas (TX)

On-site
USD 180,000 - 280,000
Devops Engineer - Assistant Vice President
Devops Engineer - Assistant Vice President

Citi • New York (NY)

Hybrid
USD 3,700 - 5,800
Site Reliability Engineer
Site Reliability Engineer

Cognizant • Plano (TX)

On-site
USD 74,000 - 105,000
Medical/Dental/Vision/Life Insurance
Paid holidays plus Paid Time Off
401(k) plan
+3
Sr SRE Automation Engineer
Sr SRE Automation Engineer

Compunnel, Inc. • Austin (TX), Northern (KY)

Hybrid
USD 130,000 - 180,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Hobbsnews • Jersey City (NJ)

On-site
USD 153,000 - 192,000
Benefits eligible
Discretionary incentive plan
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Mission Staffing • New York (NY)

Hybrid
USD 140,000 - 200,000
Site Reliability Engineer
Site Reliability Engineer

Jobtailor • New Jersey

On-site
USD 120,000 - 180,000