Turn this role into an interview — a resume and cover letter built around what this employer wants.
ServiceNow in Santa Clara, CA is seeking a Director of Data & Storage Reliability Engineering to lead a global team focused on reliability, observability, and performance across databases, storage, and platform infrastructure.
You will drive a data-driven culture, partner with product engineering and SWAT, set roadmaps, mentor leaders, and push proactive risk reduction while balancing speed and stability.
It all started when engineer Fred Luddy wrote code that automated a tedious task for his coworker, Phyllis. She cried tears of joy. That moment inspired Fred to build a company that could do that for everyone—freeing people from busywork so they could focus on meaningful work. Today, ServiceNow is the AI control tower for business reinvention. Our ServiceNow AI platform brings together any AI, any data, and any workflow— helping 85% of the Fortune 500® work smarter, faster, and better. We're building an AI-native culture where technology and talent are unstoppable together. And we're just getting started.
It all started when engineer Fred Luddy wrote code that automated a tedious task for his coworker, Phyllis. She cried tears of joy. That moment inspired Fred to build a company that could do that for everyone—freeing people from busywork so they could focus on meaningful work. Today, ServiceNow is the AI control tower for business reinvention. Our ServiceNow AI platform brings together any AI, any data, and any workflow— helping 85% of the Fortune 500® work smarter, faster, and better. We're building an AI-native culture where technology and talent are unstoppable together. And we're just getting started. Join us to put AI to work for people.
The successful candidate will lead the Data & Storage Reliability Engineering organization responsible for improving reliability, resilience, performance, scalability, observability, and customer experience across ServiceNow's database, storage, and supporting platform infrastructure. This leader will be responsible for building, developing, and scaling high-performing engineering teams focused on reliability engineering, observability, performance engineering, diagnostics, automation, production analytics, migration readiness, resilience engineering, and prevention engineering. Responsibilities include talent acquisition, performance management, career development, succession planning, objective setting, coaching, and prioritization of strategic initiatives. The role will establish a strong engineering-first culture centered on data-driven decision making, continuous improvement, operational excellence, customer experience, and systemic risk reduction. This position is accountable for identifying recurring failure patterns, reliability risks, performance bottlenecks, scalability constraints, migration challenges, and operational inefficiencies across database services, storage platforms, cloud infrastructure, and distributed application environments, and driving engineering improvements that eliminate entire classes of issues before they impact customers. The Director will partner closely with Product Engineering, Database Engineering, Cloud Infrastructure, Architecture, Storage Engineering, Support, and Operations teams to ensure reliability, observability, performance, and resilience considerations are incorporated throughout the software development lifecycle. The successful candidate will also partner closely with SWAT and Customer & Production Engineering teams to establish a continuous feedback loop between production operations and platform improvement. SWAT remains responsible for customer escalations, production operations, incident response, and service restoration, while this organization is responsible for identifying systemic opportunities, defining engineering priorities, and driving platform improvements that reduce future customer impact. The successful candidate will serve as the senior technical leader for complex reliability investigations, customer-critical escalation reviews, migration readiness assessments, and platform improvement initiatives, transforming production insights into long-term engineering outcomes. They will influence architectural decisions and technology investments by providing reliability expertise, observability insights, performance guidance, and production-based evidence that improve platform resilience, scalability, efficiency, and customer outcomes. This role requires a strong product mindset. The leader will treat reliability, observability, resilience, performance, and automation capabilities as products with roadmaps, priorities, adoption goals, and measurable outcomes. They will be responsible for identifying the highest-value engineering opportunities, prioritizing investments, and driving adoption across multiple product and infrastructure organizations.
The successful candidate will establish scalable reliability engineering practices, standards, governance processes, and operating models across the organization. They will drive adoption of observability standards, reliability engineering frameworks, resiliency assessments, migration readiness practices, diagnostics capabilities, engineering guardrails, and automation strategies. This leader will continuously evaluate incidents, customer escalations, migration outcomes, platform telemetry, performance trends, capacity signals, and operational data to identify systemic risks and drive long-term engineering improvements. The role will establish a formal review process with SWAT and Customer & Production Engineering teams to evaluate major incidents, recurring operational challenges, migration learnings, customer-impacting events, and emerging platform risks. These insights will be used to prioritize engineering investments and platform improvements. The successful candidate will establish meaningful KPIs and engineering metrics that provide visibility into platform reliability, resiliency, performance, operational efficiency, customer experience, engineering productivity and risk reduction. The successful candidate will leverage AI-powered tools, analytics, automation frameworks, and production intelligence to identify emerging risks, improve detection coverage, accelerate engineering insights, reduce operational toil, and improve engineering productivity. They will use production telemetry, incident learnings, customer escalations, migration outcomes, observability data, and operational trends to drive architectural improvements, reliability investments, platform standards, and long-term engineering evolution. The Director will maintain a portfolio of reliability investments spanning observability, performance, diagnostics, resilience, automation, and prevention, balancing immediate customer needs with long-term platform strategy. The Director will champion a proactive reliability engineering model that shifts the organization from reactive issue response toward predictive analysis, prevention, resilience, and continuous optimization.
JV20
For positions in this location, we offer a base pay of $221,200 - $387,100, plus equity (when applicable), variable/incentive compensation and benefits. Sales positions generally offer a competitive On Target Earnings (OTE) incentive compensation structure. Please note that the base pay shown is a guideline, and individual total compensation will vary based on factors such as qualifications, skill level, competencies, and work location. We also offer health plans, including flexible spending accounts, a 401(k) Plan with company match, ESPP, matching donations, a flexible time away plan and family leave programs. Compensation is based on the geographic location in which the role is located and is subject to change based on work location.
We approach our distributed world of work with flexibility and trust. Work personas (flexible, remote, or required in office) are categories that are assigned to ServiceNow employees depending on the nature of their work and their assigned work location. Learn more here. To determine eligibility for a work persona, ServiceNow may confirm the distance between your primary residence and the closest ServiceNow office using a third-party service.
ServiceNow is an equal opportunity employer. All qualified applicants will receive consideration for employment without regard to race, color, creed, religion, sex, sexual orientation, national origin or nationality, ancestry, age, disability, gender identity or expression, marital status, veteran status, or any other category protected by law. In addition, all qualified applicants with arrest or conviction records will be considered for employment in accordance with legal requirements.
We strive to create an accessible and inclusive experience for all candidates. If you require a reasonable accommodation to complete any part of the application process, or are unable to use this online application and need an alternative method to apply, please contact globaltalentss@servicenow.com for assistance.
For positions requiring access to controlled technology subject to export control regulations, including the U.S. Export Administration Regulations (EAR), ServiceNow may be required to obtain export control approval from government authorities for certain individuals. All employment is contingent upon ServiceNow obtaining any export license or other approval that may be required by relevant export control authorities.
From Fortune. ©2026 Fortune Media IP Limited. All rights reserved. Used under license.