Sr. Manager, Data Reliability Engineering

Visa Consolidated Support Services India

Bengaluru

Hybrid

INR 3,600,000 - 6,000,000

Full time

12 days ago
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Visa seeks a Senior Manager for Site Reliability Engineering (SRE) Data Platforms in Bengaluru. You will lead teams to ensure data platforms are reliable, scalable, secure, and highly available, partnering with data engineering, analytics, ML, product, security, risk, and QA functions.

You will establish reliability strategies, automation, observability, incident response, and disaster recovery, while guiding cross-functional squads to deliver resilient data services.

Qualifications

  • 8+ years of relevant work experience and a Bachelor''s degree.
  • Experience in leading teams in the design, development, and deployment of large-scale engineering solutions.
  • Experience withcloud infrastructure, distributed systems, observability platforms, and reliability engineering practices.
  • Experience in building and maintaining scalable infrastructure, automation frameworks, incident response processes, and service reliability programs.

Responsibilities

  • Lead the design, development, deployment, and operation ofscalable, reliable, secure, and highly available SRE solutions supporting data platforms and data infrastructure.
  • Oversee automation, monitoring, alerting, observability, incident response, disaster recovery, and reliability practices for data systems across platforms and services.
  • Guide teams in troubleshooting and resolving complex production issues involving data pipelines, distributed processing systems, storage platforms, orchestration tools, and analytics environments.
  • Drive the adoption of automation, CI/CD, infrastructure as code, quality engineering, observability, reliability testing, and operational excellence practices across the Data area.
  • Coach and mentor SRE, infrastructure, platform, or operations engineering teams supporting Data, fostering continuous improvement, accountability, learning, reliability, data integrity, and technical excellence.
  • Collaborate with data engineering, analytics, machine learning, product, security, risk, compliance, QA, infrastructure, and business partners to ensure solutions meet priorities and regulatory requirements.

Skills

Leadership
Cloud infrastructure
Distributed systems
Observability
Automation
SRE
Python
Kubernetes
CI/CD
Infrastructure as Code

Education

Bachelor's degree

Tools

Terraform
Git
Jenkins

Job description

About Us
Visa is a world leader in payments technology, facilitating transactions between consumers, merchants, financial institutions and government entities across more than 200 countries and territories, dedicated to uplifting everyone, everywhere by being the best way to pay and be paid.

At Visa, you''ll have the opportunity to create impact at scale tackling meaningful challenges, growing your skills and seeing your contributions impact lives around the world.

Join Visa and do work that matters to you, to your community, and to the world. Progress starts with you.

Job Description

Job Summary

The Senior Manager, Site Reliability Engineering (SRE) Data Platforms leads teams responsible for ensuring thereliability, scalability, security, performance, and operational excellence of data platforms, data infrastructure, and data-intensive production systems. This role supports the Data area by partnering with data engineering, analytics, machine learning, product, risk, compliance, security, and infrastructure teams to operate and improve the systems that capture, process, store, move, and serve critical business data.

This role oversees the development and execution of reliability strategies, infrastructure automation, observability practices, incident management processes, and operational standards fordata pipelines, data platforms, distributed processing systems, cloud data services, streaming platforms, orchestration tools, and data storage environments. The Senior Manager ensures that data services meet expectations foravailability, latency, throughput, recoverability, security, compliance, and data integrity.

This role is responsible for establishingprocesses, automation, monitoring, alerting, service-level objectives, incident response practices, disaster recovery capabilities, capacity planning, and operational excellence frameworksbased on business and technical requirements. The Senior Manager guides teams in improving platform reliability, reducing toil, troubleshooting complex production issues, and supporting resilient data operations across high-volume, distributed environments.

The Senior Manager acts as a leader for SRE, infrastructure, platform, or operations engineering teams supporting the Data organization byremoving blockers, enabling productivity, improving operational maturity, and fostering an inclusive, high-performance engineering culture. The role also involves coaching and developing engineers to adapt to evolving cloud platforms, data technologies, automation tools, observability frameworks, reliability practices, and AI-driven workflows. This leader partners closely with software engineering, data engineering, product management, QA, risk, compliance, security, infrastructure, and business stakeholders to deliversafe, resilient, transparent, and customer‑centric data solutions.

All roles require digital fluency, including the ability to work with emerging technologies such asGenerative AI tools for example, ChatGPT and Claude Cod to support everyday work, improve operational efficiency, accelerate documentation, assist with troubleshooting, and enhance engineering productivity.


Key Responsibilities

  • Lead the design, development, deployment, and operation ofscalable, reliable, secure, and highly available SRE solutions supporting data platforms, data pipelines, and data infrastructure.
  • Oversee the creation and maintenance ofautomation, monitoring, alerting, observability, incident response, disaster recovery, and reliability engineering practicesfor data systems across multiple platforms and services.
  • Guide teams in troubleshooting and resolvingcomplex production issues involving data pipelines, distributed processing systems, storage platforms, orchestration tools, streaming systems, infrastructure services, performance bottlenecks, and service degradations.
  • Drive the adoption ofautomation, CI/CD, infrastructure as code, quality engineering, observability, reliability testing, and operational excellence practicesacross the Data area.
  • Coach and mentor SRE, infrastructure, platform, or operations engineering teams supporting Data, fostering a culture ofcontinuous improvement, accountability, learning, reliability, data integrity, and technical excellence.
  • Collaborate with data engineering, analytics, machine learning, product, security, risk, compliance, QA, infrastructure, and business partners to ensure solutions meetbusiness priorities, customer expectations, regulatory requirements, security standards, data governance expectations, and operational risk controls.
  • Ensure adherence to reliability, resilience, security, compliance, access management, disaster recovery, change management, and data protection standardsthroughout the technology and data lifecycle.
  • Lead the implementation ofservice-level indicators, service-level objectives, error budgets, capacity planning, performance optimization, incident metrics, reliability dashboards, and operational health checksfor data platforms and data services.
  • Partner with data engineering and platform teams to improve the reliability ofbatch pipelines, real-time streaming workloads, machine learning pipelines, data ingestion frameworks, transformation jobs, metadata processes, and reporting or analytics services.
  • Oversee the development and maintenance oftechnical documentation, runbooks, architecture diagrams, operational procedures, post-incident reviews, production-readiness checklists, and best practices for code quality, change review, and data platform supportability.

This is a hybrid position. Expectation of days in the office will be confirmed by your Hiring Manager.

Qualifications

Basic Qualifications

  • 8+ years of relevant work experience and a Bachelor''s degree, OR 11+ years of relevant work experience.
  • Experience in leading teams in the design, development, and deployment of large-scale engineering solutions.
  • Experience withcloud infrastructure, distributed systems, observability platforms, and reliability engineering practices.
  • Experience in building and maintainingscalable infrastructure, automation frameworks, incident response processes, and service reliability programs.
  • Experience with programming, scripting, and infrastructure tools such asPython, Go, Bash, Terraform, Kubernetes, CI/CD platforms, and version control systems such as Git.
  • Experience troubleshooting and resolving issues incomplex, high-volume, highly available production environments.
  • Experience implementingsecurity, compliance, access control, disaster recovery, and operational risk management standards.
  • Experience coaching, mentoring, and developingSRE, infrastructure, platform, or operations engineering teams.
  • Experience collaborating withengineering, product, security, infrastructure, and business stakeholdersto improve reliability, scalability, and operational excellence.
  • Experience developing and maintainingtechnical documentation, runbooks, post-incident reviews, service-level objectives, operational procedures, and reliability dashboards.


Preferred Qualifications

  • 9 or more years of relevant work experiencewith a Bachelors Degree, or7 or more years of relevant experiencewith an Advanced Degree, or3 or more years of experiencewith a PhD.
  • Experience supportinglarge-scale data platforms and data infrastructurein production environments.
  • Experience withcloud platforms and services, such as AWS, Azure, or GCP, used to operate reliable and scalable data systems.
  • Experience withSite Reliability Engineering practices, including monitoring, alerting, incident management, service-level objectives, error budgets, and post-incident reviews.
  • Experience improving the reliability ofdata pipelines, distributed processing systems, streaming platforms, orchestration tools, storage platforms, and analytics environments.
  • Experience troubleshootingcomplex production issuesacross data platforms, infrastructure, applications, and distributed systems.
  • Experience withautomation, CI/CD, infrastructure as code, observability tools, containerization, and orchestration platforms.
  • Experience supporting systems requiringhigh availability, low latency, scalability, resiliency, and disaster recovery.
  • Experience partnering withdata engineering, analytics, machine learning, product, security, risk, compliance, QA, and infrastructure teams.
  • Experience leadingcross-functional SRE, platform, infrastructure, operations, or data engineering teamsin a matrixed organization.

Visa is an EEO Employer

Qualified applicants will receive consideration for employment without regard to race, color, religion, sex, national origin, sexual orientation, gender identity, disability or protected veteran status. Visa will also consider for employment qualified applicants with criminal histories in a manner consistent with EEOC guidelines and applicable local law.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Sr. Manager, Data Reliability Engineering (8+ years)
Sr. Manager, Data Reliability Engineering (8+ years)

Pismo • Bengaluru

Hybrid
INR 4,500,000 - 7,000,000
Sr. Manager, Data Reliability Engineering (8+ years)
Sr. Manager, Data Reliability Engineering (8+ years)

Visa • Bengaluru

On-site
INR 4,200,000 - 5,600,000
Senior SW Engineer - SRE, Production Support, Linux, Networking, CI/CD
Senior SW Engineer - SRE, Production Support, Linux, Networking, CI/CD

PowerToFly • Bengaluru

Hybrid
INR 600,000 - 900,000
Hybrid work model
Senior SW Engineer - SRE, Production Support, Linux, Networking, CI/CD
Senior SW Engineer - SRE, Production Support, Linux, Networking, CI/CD

Visa Consolidated Support Services India • Bengaluru

Hybrid
INR 1,200,000 - 2,100,000
Site Reliability Engineer - Vice President
Site Reliability Engineer - Vice President

Citi • Pune District

On-site
INR 4,000,000 - 5,500,000
Software Engineer - Sr. Consultant Level (7+, SDET)
Software Engineer - Sr. Consultant Level (7+, SDET)

Visa Consolidated Support Services India • Bengaluru

On-site
INR 4,500,000 - 6,500,000
Sr. Manager - Software Engineering (11-14 years' experience, Java/Python, Kafka, Spark, Hive, Hadoop)
Sr. Manager - Software Engineering (11-14 years' experience, Java/Python, Kafka, Spark, Hive, Hadoop)

PowerToFly • Bengaluru

Hybrid
INR 3,800,000 - 6,500,000
Lead Software Engineer - (14-18 years - Java, Distributed Systems, Microservices, Cloud, Security, Platform Engineering, Gen AI)
Lead Software Engineer - (14-18 years - Java, Distributed Systems, Microservices, Cloud, Security, Platform Engineering, Gen AI)

Visa • Bengaluru

Hybrid
INR 3,000,000 - 6,000,000
Site Reliability Engineer-Vice President
Site Reliability Engineer-Vice President

Citi Bank • Pune District

On-site
INR 1,800,000 - 3,200,000
Staff Data Engineer
Staff Data Engineer

Visa Consolidated Support Services India • Bengaluru

Hybrid
INR 3,000,000 - 5,400,000