Director – Observability, Response & Reliability (ORR)

Ouro and Real Madrid Football Club

Dadri

On-site

INR 6,000,000 - 12,000,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Netspend Corporation seeks a visionary Director – Observability, Response & Reliability (ORR) to lead the India-based ORR organization. You will establish a world-class Observability and AIOps practice, institutionalize SRE principles, and govern reliability for critical FinTech systems, ACH processing, and payment rails across global platforms.

You will mentor a high-performing team, define career paths, and partner with security and cloud leaders to ensure PCI-DSS compliance while scaling to

Qualifications

  • Bachelor’s or Master’s Degree in Computer Science, Software Engineering, IT, or related field.
  • 15+ years of total experience in SRE, systems/platform engineering, or enterprise observability.
  • Strong expertise in enterprise telemetry and observability tooling.

Responsibilities

  • Define and drive an enterprise Observability strategy and end-to-end telemetry across microservices and cloud infrastructure.
  • Lead SRE practice with SLIs/SLOs, error budgets, and reliability reviews.
  • Champion AIOps, anomaly detection, and automated self-healing workflows.
  • Oversee major incident response, RCA, and executive reliability reporting.
  • Drive cloud transformation and IaC standards with GitLab CI/CD, Terraform, and Ansible.

Skills

Observability architecture
SRE leadership
AIOps & automation
Cloud architecture
Security & PCI-DSS

Education

Bachelor’s or Master’s Degree in CS/Software Eng or related

Tools

Splunk
Dynatrace
Datadog
Prometheus
Grafana
OpenTelemetry
Kafka
Cassandra

Job description

Ouro is a global, vertically-integrated financial services and technology company dedicated to the delivery of innovative financial empowerment solutions to consumers worldwide. Ouro's financial products and services span prepaid, debit, cross-border payments, and loyalty solutions for consumers and enterprise partners. Since its founding in 1999 by industry pioneers, Ouro products have processed billions of dollars in transaction volume and served millions of customers worldwide. The company is headquartered in Austin, Texas with regional offices around the world.
Director – Observability, Response & Reliability (ORR)
Location

India (Noida)

Employment Type

Full time

Department

About the Company:

Netspend Corporation is a global, vertically-integrated financial services and technology company dedicated to the delivery of innovative financial empowerment solutions to consumers worldwide. Netspend's financial products and services span prepaid, debit, cross-border payments, and loyalty solutions for consumers and enterprise partners.
Netspend provides prepaid and debit account solutions that connect customers with secure, convenient access to global payment networks so they can manage their money and make everyday purchases. With a nationwide U.S. retail network, customers can purchase and reload Netspend products at 130,000 reload points and over 100,000 distributing locations.
Since our founding in 1999 by industry pioneers, Netspend products have processed billions of dollars in transaction volume and served millions of customers worldwide. The company is headquartered in Austin, Texas with employees worldwide.

Role Overview

We are seeking a visionary Director – Observability, Response & Reliability (ORR) with 15+ years of overall technical experience, including 5+ years in engineering leadership, to serve as our primary authority on system resilience, full-stack observability, and enterprise incident management across Netspend's global financial ecosystem.

In this high-visibility role, you will lead our India-based ORR organization, driving the strategy that transforms how Netspend monitors, predicts, responds to, and resolves critical operational events. You will establish a world-class Observability and AIOps practice, institutionalize Site Reliability Engineering (SRE) principles across all product teams, and safeguard systems handling ACH processing, core payment rails, millions of daily card transactions, and bank-sensitive data.

Key Responsibilities
1. Observability Strategy & Telemetry Architecture

Unified Telemetry Vision: Design and execute the enterprise Observability roadmap, establishing full-stack, end-to-end visibility across microservices, legacy monoliths, cloud infrastructure, and data pipelines (Kafka, Cassandra).

MELT Standardizer: Define unified logging, metrics, traces, and synthetics standards (OpenTelemetry, Prometheus, Splunk, Dynatrace) across all engineering groups to enable granular transaction-level tracing for payment flows.

AIOps & Self-Healing Systems: Pioneer the adoption of AIOps, Machine Learning, and anomaly detection to shift the engineering culture from reactive alert firefighting to proactive noise reduction, predictive fault detection, and automated self-healing workflows.

Financial Data Observability: Build custom monitoring, alert thresholds, and real-time dashboards for critical FinTech protocols, payment rails, ACH file watching, and bank connectivity channels (Axway, SFTP, AS2).

2. Reliability Engineering, Resilience & SLO Management

SRE Practice Ownership: Institutionalize SRE paradigms (SLIs, SLOs, Error Budgets, Reliability Reviews) across all software product squads, embedding reliability directly into the software development lifecycle (SDLC).

Chaos Engineering & Testing: Establish proactive resilience practices, including failure injection, Chaos Engineering, and regular multi-AZ/multi-region Disaster Recovery (DR) simulations for critical financial services.

Cost-Effective Reliability (FinOps): Partner with cloud and finance leadership to balance extreme uptime demands with cost efficiency across Splunk, Dynatrace, AWS telemetry storage, and log ingestion limits.

3. Enterprise Incident Response & Operational Governance

Major Incident Command (P0/P1): Oversee the high-severity Incident Management framework, ensuring 24/7 incident readiness, rapid mean time to detect (MTTD), and swift mean time to resolve/recover (MTTR).

Blameless Post-Mortems & RCA: Champion a culture of psychological safety through rigorous, blameless post-mortems and Root Cause Analyses (RCAs) to drive systemic platform fixes and prevent recurring outages.

Executive Reliability Reporting: Establish transparent reporting frameworks to translate uptime metrics, availability SLAs, error budget consumption, and platform risks directly to C-suite leadership (CTO, CIO).

4. Strategic Cloud Transformation & Infrastructure Evolution

Strangler Pattern Execution: Provide reliability oversight for moving critical functions off legacy platforms (Xymon, Puppet, SVN) to cloud-native AWS architectures without risking live financial traffic or data loss.

CI/CD & IaC Standards: Enforce strict Infrastructure-as-Code (Terraform/Ansible) and GitLab CI/CD pipeline reliability, integrating automated canary deployments, security scans, and instant rollback mechanisms.

Security & Compliance Oversight: Partner with security teams to ensure all observability logging, tracing, and automation frameworks comply strictly with PCI-DSS, SOC 2, mTLS, and banking audit standards.

India ORR Center of Excellence: Build, mentor, and scale a high-performing team of Site Reliability Engineers, Observability Specialists, and Incident Response Leads in India.

Talent Development: Define technical career paths, performance benchmarks, and continuous learning opportunities in observability and SRE for mid-level and senior engineers.

Vendor Governance: Manage strategic relationships, licensing, contractual negotiations, and technical roadmaps with key observability and cloud partners (Splunk, Dynatrace, AWS, GitLab).

Required Qualifications & Experience
Education & Overall Experience

Bachelor’s or Master’s Degree in Computer Science, Software Engineering, Information Technology, or a related quantitative field.

15+ years of total experience in SRE, Systems/Platform Engineering, Infrastructure Architecture, and Enterprise Observability.

Core Observability & Reliability Expertise

Observability Mastery: Deep expertise architecting enterprise telemetry solutions using tools such as Splunk, Dynatrace, Datadog, Prometheus, Grafana, and OpenTelemetry.

AIOps & Automation: Proven track record implementing AIOps tools, automated incident remediation, and AI/ML-based anomaly detection engines.

SRE Leadership: Expert knowledge of SRE best practices, error budget policy enforcement, SLO modeling, and incident response frameworks.

Technical & Systems Foundation

Cloud & Modern Infrastructure: Strong hands-on architectural understanding of AWS services (EC2, ALB/NLB, Direct Connect, IAM, KMS, Transfer Family).

Legacy-to-Cloud Modernization: Practical experience executing "strangler fig" migration strategies off legacy tooling (Puppet, SVN, Xymon) to modern GitOps/IaC (GitLab CI/CD, Terraform, Ansible).

Data & Middleware: Experience monitoring and troubleshooting distributed middleware and database technologies (Apache Kafka, Cassandra) under heavy throughput.

FinTech & Secure Protocols: Understanding of secure file transfer (SFTP, AS2), mTLS, cryptographic key management (HSMs, Virtucrypt), and high-availability payment processing environments.

Preferred Certifications

Observability: Splunk Certified Architect, Dynatrace Master / Professional, or equivalent.

Cloud & DevOps: AWS Certified Solutions Architect – Professional

SRE & Methodologies: Certified Site Reliability Engineer (SRE), ITIL v4 (Incident/Problem Management focus).

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer
Site Reliability Engineer

Synechron • Bengaluru

On-site
INR 2,500,000 - 4,500,000
SRE New Relic
SRE New Relic

Tekskills • Hyderabad, Chennai District, Bengaluru

Hybrid
INR 400,000 - 700,000
Senior Security Incident Response Analyst
Senior Security Incident Response Analyst

Ouro and Real Madrid Football Club • Dadri

On-site
INR 1,500,000 - 2,300,000
Lead Engineer - Reliability Engineering
Lead Engineer - Reliability Engineering

StoneX Group Inc. • Bengaluru

Hybrid
INR 3,500,000 - 6,000,000
Application Support Technology Lead Analyst - Vice President
Application Support Technology Lead Analyst - Vice President

Citibank (Switzerland) AG • Pune District

On-site
Confidential
Site Reliability Engineer
Site Reliability Engineer

SourcingXPress • Maharashtra

On-site
INR 700,000 - 1,800,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Brillio • Bengaluru Urban

On-site
INR 1,200,000 - 2,000,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

AcquireX • Maharashtra

On-site
INR 1,800,000 - 2,500,000
Observability Engineer
Observability Engineer

Weekday (YC W21) • Mumbai

On-site
INR 4,000,000 - 6,000,000
SRE Observability Engineer
SRE Observability Engineer

Awign • Hyderabad

On-site
INR 4,200,000 - 6,500,000