Turn this role into an interview — a resume and cover letter built around what this employer wants.
ServiceNow seeks a Director to lead the Data & Storage Reliability Engineering organization, improving reliability, observability, performance, scalability, and customer experience across database, storage, and platform infrastructure.
The role builds and scales engineering teams focused on reliability, observability, diagnostics, automation, and migration readiness, while driving systemic improvements and the adoption of AI-powered tooling.
8+ years of engineering leadership experience, including leading managers and globally distributed teamsExperience partnering closely with production operations, customer escalation teams, reliability organizations, and software engineering teams to drive systemic improvements based on operational learningsExtensive experience leading Reliability Engineering, Platform Engineering, Database Engineering, Infrastructure Engineering, Production Engineering, Performance Engineering, or related technical organizationsExperience driving engineering initiatives through data, metrics, customer impact analysis, and measurable business outcomes15+ years of experience in software engineering, platform engineering, reliability engineering, infrastructure engineering, database engineering, distributed systems, product management, or large-scale SaaS environmentsProven experience identifying systemic issues and converting operational insights into strategic engineering improvementsStrong product mindset with demonstrated experience treating technical capabilities as products with roadmaps, priorities, customers, adoption goals, and measurable business outcomesExperience in leveraging or critically thinking about how to integrate AI into work processes, decision-making, or problem-solving. This may include using AI-powered tools, automating workflows, analyzing AI-driven insights, or exploring AI’s potential impact on the function or industryExperience leveraging AI technologies to improve decision-making, analytics, engineering workflows, operational efficiency, reliability insights, automation, or customer outcomesStrong understanding of reliability engineering principles, observability, scalability, resiliency, operational excellence, and performance engineeringExperience translating production insights, customer pain points, operational challenges, reliability risks, and platform telemetry into prioritized engineering investments and long-term roadmapsExperience operating a portfolio of engineering investments, balancing short-term customer needs with long-term reliability, performance, scalability, and resilience objectivesDeep expertise in distributed systems, databases, storage technologies, cloud infrastructure, and large-scale SaaS architecturesExperience building and operating observability, telemetry, diagnostics, reliability, or performance capabilities at scaleBachelor’s degree in Computer Science, Engineering, or a related technical field, or equivalent practical experienceExceptional communication, stakeholder management, and leadership skillsExperience defining product strategies, developing roadmaps, prioritizing investments, and aligning stakeholders across multiple organizations without direct authorityExperience partnering closely with Product Management organizations to influence roadmaps and deliver customer-centric outcomesPrevious Product Management experience in a platform, infrastructure, cloud, database, storage, or SaaS environmentExperience applying product management disciplines such as roadmap planning, prioritization, customer-centric thinking, outcome measurement, and portfolio management to engineering organizationsExperience operating large-scale enterprise database and storage platforms supporting mission-critical workloadsExperience building and scaling Reliability Engineering, Performance Engineering, Platform Engineering, SRE, or Production Engineering organizationsExperience with observability platforms, telemetry systems, diagnostics frameworks, and production analyticsExperience with migration readiness, resiliency validation, reliability testing, operational risk reduction, and large-scale cloud transformationsExperience leveraging AI technologies to improve anomaly detection, forecasting, incident analysis, prioritization, and engineering productivityStrong understanding of distributed systems architecture, cloud platform operations, and hyperscale environmentsExperience developing executive-facing reliability scorecards, engineering metrics, and business impact reportingExperience influencing platform architecture, database strategy, storage strategy, and long-term engineering roadmapsExperience with Linux-based production environments and large-scale cloud infrastructureExperience supporting enterprise database technologies such as MySQL, MariaDB, PostgreSQL, Oracle, SQL Server, or cloud-native database platformsFamiliarity with ServiceNow platform architecture and large-scale SaaS operations