Director of Data & Storage Reliability Engineering

ServiceNow

Santa Clara (CA)

On-site

USD 260,000 - 380,000

Full time

14 days+
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Benefits offered by this job

Generous family leave
Matched donations
Annual learning stipends
Flexible PTO
Competitive retirement plan
Paid volunteer time

Job summary

ServiceNow seeks a Director to lead the Data & Storage Reliability Engineering organization, improving reliability, observability, performance, scalability, and customer experience across database, storage, and platform infrastructure.

The role builds and scales engineering teams focused on reliability, observability, diagnostics, automation, and migration readiness, while driving systemic improvements and the adoption of AI-powered tooling.

Qualifications

  • 15+ years of software/ platform reliability experience in large-scale SaaS.
  • Strong product mindset with roadmaps, priorities, and measurable outcomes.
  • Experience leading cross-functional engineering teams and driving data-driven improvements.

Responsibilities

  • Lead Data & Storage Reliability Engineering org across database, storage, and platform infrastructure.
  • Build, develop, and scale high-performing reliability/observability teams.
  • Partner with Product, Cloud, Architecture, and Ops to embed reliability in SDLC.
  • Establish KPIs, governance, and operating models for platform reliability.
  • Drive AI-powered tooling to improve detection, insights, and productivity.

Skills

Reliability engineering
Observability
Performance engineering
Platform engineering
AI in engineering
Leadership
Stakeholder management
Cloud infrastructure
Data analytics

Education

Bachelor's degree in CS/Engineering

Tools

Observability platforms
Production analytics

Job description

  • The successful candidate will lead the Data & Storage Reliability Engineering organization responsible for improving reliability, resilience, performance, scalability, observability, and customer experience across ServiceNow’s database, storage, and supporting platform infrastructure
  • This leader will be responsible for building, developing, and scaling high-performing engineering teams focused on reliability engineering, observability, performance engineering, diagnostics, automation, production analytics, migration readiness, resilience engineering, and prevention engineering
  • Responsibilities include talent acquisition, performance management, career development, succession planning, objective setting, coaching, and prioritization of strategic initiatives
  • The role will establish a strong engineering-first culture centered on data-driven decision making, continuous improvement, operational excellence, customer experience, and systemic risk reduction
  • This position is accountable for identifying recurring failure patterns, reliability risks, performance bottlenecks, scalability constraints, migration challenges, and operational inefficiencies across database services, storage platforms, cloud infrastructure, and distributed application environments, and driving engineering improvements that eliminate entire classes of issues before they impact customers
  • The Director will partner closely with Product Engineering, Database Engineering, Cloud Infrastructure, Architecture, Storage Engineering, Support, and Operations teams to ensure reliability, observability, performance, and resilience considerations are incorporated throughout the software development lifecycle
  • The successful candidate will also partner closely with SWAT and Customer & Production Engineering teams to establish a continuous feedback loop between production operations and platform improvement. SWAT remains responsible for customer escalations, production operations, incident response, and service restoration, while this organization is responsible for identifying systemic opportunities, defining engineering priorities, and driving platform improvements that reduce future customer impact
  • The successful candidate will serve as the senior technical leader for complex reliability investigations, customer-critical escalation reviews, migration readiness assessments, and platform improvement initiatives, transforming production insights into long-term engineering outcomes
  • This role requires a strong product mindset. The leader will treat reliability, observability, resilience, performance, and automation capabilities as products with roadmaps, priorities, adoption goals, and measurable outcomes
  • They will be responsible for identifying the highest-value engineering opportunities, prioritizing investments, and driving adoption across multiple product and infrastructure organizations
  • The successful candidate will establish scalable reliability engineering practices, standards, governance processes, and operating models across the organization
  • They will drive adoption of observability standards, reliability engineering frameworks, resiliency assessments, migration readiness practices, diagnostics capabilities, engineering guardrails, and automation strategies
  • This leader will continuously evaluate incidents, customer escalations, migration outcomes, platform telemetry, performance trends, capacity signals, and operational data to identify systemic risks and drive long-term engineering improvements
  • The role will establish a formal review process with SWAT and Customer & Production Engineering teams to evaluate major incidents, recurring operational challenges, migration learnings, customer-impacting events, and emerging platform risks. These insights will be used to prioritize engineering investments and platform improvements
  • The successful candidate will establish meaningful KPIs and engineering metrics that provide visibility into platform reliability, resiliency, performance, operational efficiency, customer experience, engineering productivity, and risk reduction
  • The successful candidate will leverage AI-powered tools, analytics, automation frameworks, and production intelligence to identify emerging risks, improve detection coverage, accelerate engineering insights, reduce operational toil, and improve engineering productivity
  • They will use production telemetry, incident learnings, customer escalations, migration outcomes, observability data, and operational trends to drive architectural improvements, reliability investments, platform standards, and long-term engineering evolution
  • The Director will maintain a portfolio of reliability investments spanning observability, performance, diagnostics, resilience, automation, and prevention, balancing immediate customer needs with long-term platform strategy
  • The Director will champion a proactive reliability engineering model that shifts the organization from reactive issue response toward predictive analysis, prevention, resilience, and continuous optimization
Benefits
  • Generous family leave
  • Matched donations
  • Annual learning stipends
  • Flexible PTO
  • Competitive retirement plan
  • Paid volunteer time

8+ years of engineering leadership experience, including leading managers and globally distributed teamsExperience partnering closely with production operations, customer escalation teams, reliability organizations, and software engineering teams to drive systemic improvements based on operational learningsExtensive experience leading Reliability Engineering, Platform Engineering, Database Engineering, Infrastructure Engineering, Production Engineering, Performance Engineering, or related technical organizationsExperience driving engineering initiatives through data, metrics, customer impact analysis, and measurable business outcomes15+ years of experience in software engineering, platform engineering, reliability engineering, infrastructure engineering, database engineering, distributed systems, product management, or large-scale SaaS environmentsProven experience identifying systemic issues and converting operational insights into strategic engineering improvementsStrong product mindset with demonstrated experience treating technical capabilities as products with roadmaps, priorities, customers, adoption goals, and measurable business outcomesExperience in leveraging or critically thinking about how to integrate AI into work processes, decision-making, or problem-solving. This may include using AI-powered tools, automating workflows, analyzing AI-driven insights, or exploring AI’s potential impact on the function or industryExperience leveraging AI technologies to improve decision-making, analytics, engineering workflows, operational efficiency, reliability insights, automation, or customer outcomesStrong understanding of reliability engineering principles, observability, scalability, resiliency, operational excellence, and performance engineeringExperience translating production insights, customer pain points, operational challenges, reliability risks, and platform telemetry into prioritized engineering investments and long-term roadmapsExperience operating a portfolio of engineering investments, balancing short-term customer needs with long-term reliability, performance, scalability, and resilience objectivesDeep expertise in distributed systems, databases, storage technologies, cloud infrastructure, and large-scale SaaS architecturesExperience building and operating observability, telemetry, diagnostics, reliability, or performance capabilities at scaleBachelor’s degree in Computer Science, Engineering, or a related technical field, or equivalent practical experienceExceptional communication, stakeholder management, and leadership skillsExperience defining product strategies, developing roadmaps, prioritizing investments, and aligning stakeholders across multiple organizations without direct authorityExperience partnering closely with Product Management organizations to influence roadmaps and deliver customer-centric outcomesPrevious Product Management experience in a platform, infrastructure, cloud, database, storage, or SaaS environmentExperience applying product management disciplines such as roadmap planning, prioritization, customer-centric thinking, outcome measurement, and portfolio management to engineering organizationsExperience operating large-scale enterprise database and storage platforms supporting mission-critical workloadsExperience building and scaling Reliability Engineering, Performance Engineering, Platform Engineering, SRE, or Production Engineering organizationsExperience with observability platforms, telemetry systems, diagnostics frameworks, and production analyticsExperience with migration readiness, resiliency validation, reliability testing, operational risk reduction, and large-scale cloud transformationsExperience leveraging AI technologies to improve anomaly detection, forecasting, incident analysis, prioritization, and engineering productivityStrong understanding of distributed systems architecture, cloud platform operations, and hyperscale environmentsExperience developing executive-facing reliability scorecards, engineering metrics, and business impact reportingExperience influencing platform architecture, database strategy, storage strategy, and long-term engineering roadmapsExperience with Linux-based production environments and large-scale cloud infrastructureExperience supporting enterprise database technologies such as MySQL, MariaDB, PostgreSQL, Oracle, SQL Server, or cloud-native database platformsFamiliarity with ServiceNow platform architecture and large-scale SaaS operations

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Director of Site Reliability Engineering & Service Enablement
Director of Site Reliability Engineering & Service Enablement

ServiceNow • Santa Clara (CA)

On-site
USD 260,000 - 360,000
Generous family leave
Matched donations
Annual learning stipends
+3
Director, Data & Storage Reliability Engineering
Director, Data & Storage Reliability Engineering

SmartRecruiters, Inc. • Santa Clara (CA), Northern (KY)

Hybrid
USD 221,000 - 387,000
Director - Database Performance Engineering
Director - Database Performance Engineering

SmartRecruiters, Inc. • Santa Clara (CA), Northern (KY)

Hybrid
USD 221,000 - 387,000
Director - Database Performance Engineering
Director - Database Performance Engineering

ServiceNow • Santa Clara (CA)

On-site
USD 221,000 - 387,000
Director of Data & Storage Reliability & Platform
Director of Data & Storage Reliability & Platform

ServiceNow • Santa Clara (CA)

On-site
USD 260,000 - 380,000
Generous family leave
Matched donations
Annual learning stipends
+3
Director of Data & Storage Reliability (Remote)
Director of Data & Storage Reliability (Remote)

SmartRecruiters, Inc. • Santa Clara (CA), Northern (KY)

Hybrid
USD 221,000 - 387,000
Director, Site Reliability Engineering & Service Enablement
Director, Site Reliability Engineering & Service Enablement

ServiceNow • California (MO)

Hybrid
USD 221,000 - 387,000
Health plans
401(k) Plan with company match
ESPP
+3
Director, Data & Storage Reliability & Performance
Director, Data & Storage Reliability & Performance

SmartRecruiters, Inc. • Santa Clara (CA), Northern (KY)

Hybrid
USD 221,000 - 387,000
Director, Database Performance & Reliability
Director, Database Performance & Reliability

ServiceNow • Santa Clara (CA)

On-site
USD 221,000 - 387,000
Senior Staff Software Engineer – SRE & AIOps
Senior Staff Software Engineer – SRE & AIOps

ServiceNow • California (MO)

Hybrid
USD 191,000 - 334,000