Site Reliability Engineer-Vice President

Citi Bank

Pune District

On-site

INR 1,800,000 - 3,200,000

Full time

5 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Citi is seeking a senior Site Reliability Engineer in Pune, India, to strengthen production management and drive end-to-end observability and resiliency. You will partner with multiple teams to ensure services are stable, scalable, and fault-tolerant in a highly regulated environment.

The role emphasizes ownership of SRE practices, disaster recovery planning, and automation using IaC and modern tooling. A strong background in OpenShift/Kubernetes and telemetry is required.

Qualifications

  • 7+ years of experience in production systems.
  • Strong understanding of SLOs, SLIs, and error budgets.
  • Experience with disaster recovery planning and resiliency testing.
  • Hands-on with OpenShift/Kubernetes and modern observability tools.

Responsibilities

  • Foster transparency, innovation, and accountability in SRE culture.
  • Communicate progress and impact of SRE initiatives to stakeholders.
  • Ensure compliance within a regulated environment.
  • Oversee recovery testing and automation development to reduce downtime.

Skills

SRE concepts
Disaster recovery
Toil reduction
Communication

Education

Bachelor's or Master's degree

Tools

OpenShift/Kubernetes
Prometheus
Grafana
Ansible
Terraform

Job description

Discover your future at Citi Working at Citi is far more than just a job. A career with us means joining a team of more than 230,000 dedicated people from around the globe. At Citi, youll have the opportunity to grow your career, give back to your community and make a real impact.


Job Overview

The Site Reliability Engineer (SRE) is a strategic professional accountable for the daily operations, architectural resilience, and overall implementation of SRE principles in a complex, critical, and largescale multi-disciplinary environment. This role requires a comprehensive understanding of multiple technology domains and their interaction to achieve business objectives. As a recognized technical authority, you will apply an in depth understanding of the business impact of technical contributions and provide advice and counsel on strategic solutions. We are seeking a passionate and experienced SRE to join our Production Management team. In this role, you will be instrumental in enhancing the reliability, performance, and efficiency of our Applications and Services. You will drive our strategy for end-to-end observability and resiliency, collaborating across the organization to ensure our services are stable, scalable, and fault tolerant. This is a key role that will influence strategic decisions and foster a culture of technical excellence and accountability.


Key Responsibilities

Culture & Strategy

Foster a culture of transparency, innovation, and accountability that encourages continuous improvement.


Communicate the progress and impact of SRE initiatives to stakeholders at all levels.


Operate effectively within a highly regulated environment, ensuring compliance with all relevant requirements.


Resiliency & Recovery


  • Ensure critical business applications meet stringent operational resilience requirements, including adherence to defined impact tolerances.

  • Oversee advanced recovery testing, including Production Swing Tests, Data Recovery Tests, and chaos engineering practices.

  • Drive the adoption and development of automation, such as One Touch Recovery solutions, to minimize recovery time.

  • Partner with development teams to leverage cloud native services and established resiliency patterns to enhance application reliability.


Observability & Performance


  • Collaborate across the organization to develop and scale observability solutions using modern tools for metrics, logging, and tracing.

  • Partner with development teams to effectively instrument applications, providing deep insights into system health and performance.


Qualifications and Must Have Skills


  • 7+ Years of Experience is a must have.

  • Deep understanding of SRE concepts, including SLOs, SLIs, error budgets, and toil reduction.

  • Demonstrable experience with Disaster Recovery planning, resiliency testing, and fault tolerant distributed system design.

  • Proficiency in deploying, managing, and troubleshooting applications on OpenShift/Kubernetes.

  • Hands on experience with modern observability tools (e.g., Prometheus, Grafana, Loki, Mimir, Tempo, AppDynamics).

  • Experience with Infrastructure as Code (IaC), configuration management, and automation tools (e.g., Ansible, Terraform).

  • Experience creating, modifying, and managing Helm charts for application deployment.

  • Significant professional experience in production management, software development, or an equivalent field, with a strong focus on Site Reliability Engineering.

  • Expertise in analyzing complex application, database, network, and OS issues within large scale, customer facing systems.

  • A service-oriented attitude combined with excellent problem-solving and strategic thinking skills.

  • Strong communication and diplomacy skills, with a proven ability to work effectively across multiple business and technical teams.


Desired Skills


  • Experience with major public cloud providers (e.g., Google Cloud, AWS, Azure).

  • Proven experience delivering software and infrastructure using Agile frameworks.

  • Experience presenting technical strategy to senior and executive level audiences.

  • Experience writing or maintaining code in Java, Python, Go, or similar languages.


Education

Bachelors/University degree, Masters degree preferred .

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer - Vice President
Site Reliability Engineer - Vice President

Citi • Pune District

On-site
INR 4,000,000 - 5,500,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Cubic Transportation Systems • Hyderabad

On-site
INR 1,200,000 - 1,800,000
Site Reliability Engineer
Site Reliability Engineer

Lloyds Technology Centre • Hyderabad

On-site
INR 1,200,000 - 2,400,000
Site Reliability Engineer
Site Reliability Engineer

InOpTra Digital • Bengaluru

On-site
INR 1,200,000 - 2,000,000
Site Reliability Engineering Lead_Truist
Site Reliability Engineering Lead_Truist

Infosys • Bengaluru

On-site
INR 4,000,000 - 7,000,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Infosys • Hyderabad

On-site
INR 1,400,000 - 2,200,000
Lead Site Reliability Engineer
Lead Site Reliability Engineer

Technologies Pvt. Ltd. • Pune District

On-site
INR 900,000 - 1,400,000
SRE Engineer @ Investment Banking | Mumbai
SRE Engineer @ Investment Banking | Mumbai

Net Connect Global • Bengaluru, Mumbai

Hybrid
INR 1,800,000 - 2,400,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Falabella India • Bengaluru

On-site
INR 4,000,000 - 7,000,000
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

UST • Thiruvananthapuram

On-site
INR 1,200,000 - 1,600,000