SRE Observability Technical Lead - Vice President

Citi

New York (NY)

Hybrid

USD 42,000 - 73,000

Full time

14 days+
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Citi is seeking an experienced SRE Observability Technical Lead - VP to drive end-to-end observability for payments platforms. You will embed telemetry, implement SLOs, and build meaningful dashboards across on‑prem, cloud (AWS/GCP), and containerized environments.

You will bridge line‑of‑business needs with central infrastructure, guide SREs and developers, and shape the observability roadmap while adopting AI/ML‑driven insights and modern tooling to reduce MTTD and improve reliability.

Qualifications

  • 7+ years in SRE, Observability, or platform infrastructure focused on telemetry.
  • Hands-on with Grafana, Prometheus, OpenTelemetry, ELK or Splunk for monitoring.
  • Strong understanding of SLIs, SLOs, and telemetry in high-availability setups.

Responsibilities

  • Define the roadmap for engineering enablers for the Orion project aligned with reliability goals.
  • Translate strategy into actionable delivery plans with cross-functional teams.
  • Build scalable telemetry solutions and end-to-end monitoring for critical payments flows.
  • Develop dashboards and real-time visualizations for key client journeys across payments.
  • Guide teams on SLI/SLO implementation, golden signals, and alerting best practices.
  • Support observability tooling across on‑prem, cloud, and containerized environments.
  • Drive MTTD reduction and improve recovery outcomes through tooling improvements.

Skills

Grafana
Prometheus
OpenTelemetry
ELK
Splunk
SLOs/SLIs
Kubernetes
AWS/GCP
Dashboards
Telemetry

Education

Bachelor’s degree in Computer Science, Engineering, or a related technical field

Tools

Grafana
Prometheus
OpenTelemetry
ELK
Splunk
Kubernetes
AWS
GCP

Job description

SRE Observability Technical Lead - Vice President

Working at Citi is far more than just a job. A career with us means joining a team of approximately 219,000 dedicated people from around the globe. At Citi, you’ll have the opportunity to grow your career, give back to your community and make a real impact.

Job Req Id:

26971169

Location(s):

Chennai, Tamil Nadu, India, Pune, Maharashtra, India

Job Type:

Hybrid

Posted:

Aug. 18, 2026

Job Overview

The SRE Observability Specialist is a hands-on expert, delivering the future of Observability across Services Technology. This role is a part of a central SRE enablement team within Services Production, working closely with SREs, developers, and platform teams to embed telemetry, implement SLOs, and build meaningful visualizations for key production flows — particularly in critical Payments Business.

The ideal candidate will have deep technical knowledge, a collaborative mindset, and the ability to translate strategy into scalable engineering outcomes. You will also act as a bridge between Services Technology teams and central infrastructure/CTO teams, prioritizing observability needs from line-of-business teams and driving improvements. A strong understanding of observability tooling, evolving AI/ML capabilities, and enterprise tooling ecosystems will be essential.

This role requires providing technological Support solution for Function called Project Orion which provides End-to-End payment monitoring like Building an End-to-End payments Dashboard, Toil Reduction, Transformation of legacy monitoring into observability based monitoring solution, requires good understanding of different Payments Taxonomy (ACH, Wires, Instant Payments, etc.). Strong commercial awareness, technical credibility, and excellent communication skills are essential to negotiate internally, influence peers, and drive change. Some external communication may be necessary.

Key Responsibilities:

  • Define the roadmap for Engineering enablers for Project Orion team aligned with enterprise reliability and SRE Services organization goals.
  • Translate Organization strategy into an actionable delivery plan in partnership with Services Products, Operations & Engineering function, delivering incremental, high-value milestones.
  • Understand Critical Business Services functional scope and translate into End-to-End monitoring solutions.
  • Deliver against the observability roadmap for Services Technology by building scalable, reusable telemetry solutions.
  • Periodic review and analyze application monitoring TOIL and collaborate with stakeholders and remediate them as per organization goal.
  • Create and maintain dashboards and visualizations for critical client journeys, including real-time flows across Payments.
  • Guide line-of-business teams in implementing SLIs/SLOs, golden signals, and effective alerting to support operational excellence.
  • Support integration and adoption of observability tooling across on-prem, public cloud (AWS/GCP), and containerized environments (ECS, Kubernetes).
  • Customize shared dashboards and observability components in partnership with CTI and other central Engineering functions, ensuring usability and flexibility.
  • Provide technical support and implementation guidance to SREs and developers facing integration or tooling challenges.
  • Effectively manage the observability book of work for Services Technology and drive initiatives to reduce MTTD and improve recovery outcomes.
  • Serve as a key connection point between line-of-business SREs and central infrastructure functions by gathering tooling feedback, surfacing systemic issues, and influencing platform enhancements via the Services Observability Forum.
  • Stay current with observability trends, including AI/ML-driven insights, anomaly detection, and emerging OSS practices, and assess their applicability.
  • Maintain strong knowledge of observability platform features and vendor offerings to advise teams and maximize the value of tooling investments.
  • Foster AI adoption by building use cases performed by Orion L1 Functions and remediation using Citi AI tech stack.

Qualifications:

  • 7+ years of experience in SRE, Observability Engineering, or platform infrastructure roles focused on operational telemetry.
  • Hands-on experience in observability tools and stacks such as Grafana, Prometheus, OpenTelemetry, ELK, Splunk, and similar platforms.
  • Deep understanding of SLIs, SLOs, Error Budgets, and telemetry best practices in high-availability environments.
  • Proven ability to troubleshoot integration issues and support observability across hybrid platforms (on-prem, cloud, containers).
  • Experience building dashboards aligned to business outcomes and incident workflows, especially in critical flows like payments.
  • Familiarity with modern observability tooling ecosystems, including AI/ML capabilities, trace correlation, baselining, and alert tuning.
  • Strong interpersonal and collaboration skills — able to operate across federated engineering teams and central infrastructure groups.
  • Experience in enablement or platform teams with a track record of scaling best practices across diverse business units.

Education:

  • Bachelor’s degree in Computer Science, Engineering, or a related technical field, or equivalent practical experience.
Job Family Group:

Technology

Job Family:

Applications Support

Time Type:

Full time

Most Relevant Skills

Please see the requirements listed above.

Other Relevant Skills

For complementary skills, please see above and/or contact the recruiter.

Citi is an equal opportunity employer, and qualified candidates will receive consideration without regard to their race, color, religion, sex, sexual orientation, gender identity, national origin, disability, status as a protected veteran, or any other characteristic protected by law.

Accessibility at Citi.

View Citi’s EEO Policy Statement and the Know Your Rights poster.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Application Support Technology Lead Analyst - Vice President
Application Support Technology Lead Analyst - Vice President

Citi • New York (NY)

Hybrid
USD 130,000 - 170,000
Site Reliability Engineer - Vice President
Site Reliability Engineer - Vice President

Citi • New York (NY)

Hybrid
USD 37,000 - 66,000
SRE Observability Tech Lead, VP — Hybrid
SRE Observability Tech Lead, VP — Hybrid

Citi • New York (NY)

Hybrid
USD 42,000 - 73,000
Observability & SRE Lead (Payments) VP
Observability & SRE Lead (Payments) VP

Citi • New York (NY)

Hybrid
USD 130,000 - 170,000
Applications Support Senior Analyst, Assistant Vice President
Applications Support Senior Analyst, Assistant Vice President

Citi • Jacksonville (FL)

On-site
USD 87,000 - 131,000
Applications Support Senior Analyst, Assistant Vice President
Applications Support Senior Analyst, Assistant Vice President

Citigroup Inc. • Jacksonville (FL)

On-site
USD 87,000 - 131,000
Application Development Lead (Bigdata & Analytics) - Vice President
Application Development Lead (Bigdata & Analytics) - Vice President

Citi • New York (NY)

Hybrid
USD 37,000 - 74,000
Data & AI Transformation Lead - Senior Vice President
Data & AI Transformation Lead - Senior Vice President

Citi • New York (NY)

Hybrid
USD 37,000 - 84,000
Lead Site Reliability Engineer (SRE) / Principal Site Reliability Engineer (SRE)
Lead Site Reliability Engineer (SRE) / Principal Site Reliability Engineer (SRE)

Mindlance • Irving (TX)

Hybrid
USD 120,000 - 160,000
Senior Lead Java Development - Assistant Vice President
Senior Lead Java Development - Assistant Vice President

Citi • New York (NY)

Hybrid
USD 21,000 - 34,000