Staff Software Engineer – SRE & AIOps

Servicenow

Santa Clara (CA)

On-site

USD 180,000 - 240,000

Full time

6 days ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

ServiceNow in Santa Clara, CA, seeks a Staff Software Engineer – SRE & AIOps to drive infrastructure automation, resilience, and toil elimination across hybrid cloud and data center operations. You will design and implement automation-first systems that reduce manual intervention and accelerate incident remediation for global engineering teams.

Embedded within the Site Reliability & Database Engineering organization, you will architect SRE tooling, auto-remediation capabilities, and patterns

Responsibilities

  • Design, deploy, and operate enterprise-scale Kubernetes clusters across hybrid and multi-cloud environments.
  • Architect and implement closed-loop auto-remediation systems that detect, classify, and resolve transient infrastructure failures without human intervention, leveraging agentic AI and ML to predict failures, trigger preventive actions, and reduce MTTR.
  • Design and evolve the SRE tooling stack, including monitoring platforms, incident management systems, log aggregation, and observability integrations for global follow-the-sun on-call operations.
  • Establish SLO frameworks, error budgets, and alerting policies balancing rapid incident response with alert fatigue management, while automating runbooks and playbooks for on-call engineers.
  • Design and maintain Infrastructure-as-Code frameworks and GitOps pipelines for reproducible deployments with security and compliance guardrails.
  • Architect hybrid cloud and data center operations, including workload migration strategies, disaster recovery patterns, and cost optimization across multi-region deployments.

Job description

Company Description

It all started when engineer Fred Luddy wrote code that automated a tedious task for his coworker, Phyllis. She cried tears of joy. That moment inspired Fred to build a company that could do that for everyone—freeing people from busywork so they could focus on meaningful work. Today, ServiceNow is the AI control tower for business reinvention. Our ServiceNow AI platform brings together any AI, any data, and any workflow— helping 85% of the Fortune 500® work smarter, faster, and better. We're building an AI-native culture where technology and talent are unstoppable together. And we're just getting started.

Join us to put AI to work for people.

Job Description

About the Role

ServiceNow is seeking a Staff Software Engineer – SRE & AIOps to drive infrastructure automation, operational resilience, and toil elimination across our hybrid cloud and data center operations. Embedded within the Site Reliability & Database Engineering organization, you will design and implement automation‑first systems that reduce manual intervention, accelerate incident remediation, and enable our global engineering teams to operate reliably at scale.

This role combines strong hands-on technical expertise in Kubernetes, cloud platforms, and DevOps practices with technical leadership influence across infrastructure teams. You will architect SRE tooling, develop auto‑remediation capabilities, and establish patterns that allow ServiceNow's cloud platform to maintain high reliability while minimizing operational toil across follow-the-sun global teams.

What you get to do in this role:

  • Design, deploy, and operate enterprise‑scale Kubernetes clusters across hybrid and multi‑cloud environments, establishing governance, scaling policies, and operational practices that support high‑velocity application deployments at 99.99%+ availability targets.
  • Architect and implement closed-loop auto‑remediation systems that detect, classify, and resolve transient infrastructure failures without human intervention, leveraging agentic AI and machine learning frameworks to predict failures, trigger preventive actions, and continuously reduce MTTR and on-call burden.
  • Design and evolve the SRE tooling stack, including monitoring platforms, incident management systems, log aggregation, and observability integrations, that support global follow-the-sun on-call operations and enable data‑driven incident response.
  • Establish SLO frameworks, error budgets, and alerting policies that balance rapid incident response with alert fatigue management, while developing automated runbooks and playbooks that empower on-call engineers to resolve issues autonomously.
  • Design and maintain Infrastructure-as-Code frameworks and GitOps pipelines that enable reproducible, auditable infrastructure deployments across hybrid and multi‑cloud environments with consistent security and compliance guardrails.
  • Architect hybrid cloud and data center operations, spanning on‑premises infrastructure, public cloud environments, and edge computing, including workload migration strategies, disaster recovery patterns, and cost optimization practices across multi‑region deployments.
  • Drive adoption of containerization, microservices, and DevOps patterns across engineering teams, establishing CI/CD best practices, service mesh architectures, and network security controls that enable rapid, safe release cycles.
  • Design on‑call rotation schedules, escalation policies, and incident command systems that span across different time zones, ensuring 24/7 incident response while driving post‑incident review processes that capture learning and drive systemic improvements.
  • Mentor and guide junior SRE engineers and infrastructure teams on reliability patterns, incident investigation techniques, automation best practices, and agentic AI applications for infrastructure operations.
  • Champion a culture of blameless incident analysis, data‑driven decision‑making, continuous improvement, and experimentation across engineering teams, establishing knowledge‑sharing practices and technical documentation standards.
  • Reduce operational toil through systematic automation of repetitive tasks, from infrastructure provisioning to incident response to cost optimization, directly improving team capacity and job satisfaction across globally distributed operations.
Qualifications
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Software Engineer - SRE & AIOps
Senior Software Engineer - SRE & AIOps

Servicenow • Santa Clara (CA)

On-site
USD 143,000 - 243,000
Health plans
401(k) Plan with company match
ESPP
Senior SRE & AIOps Engineer — Cloud Reliability Lead
Senior SRE & AIOps Engineer — Cloud Reliability Lead

Servicenow • Santa Clara (CA)

On-site
USD 143,000 - 243,000
Health plans
401(k) Plan with company match
ESPP
Staff SRE & AIOps Engineer: Auto-Remediation in Hybrid Cloud
Staff SRE & AIOps Engineer: Auto-Remediation in Hybrid Cloud

Servicenow • Santa Clara (CA)

On-site
USD 180,000 - 240,000
Senior Reliability Engineer
Senior Reliability Engineer

Servicenow • Minneapolis (MN)

On-site
USD 115,000 - 195,000
Health plans
401(k) with company match
Equity
Senior Manager, Data Platform Engineering - Kubernetes - Distributed Systems
Senior Manager, Data Platform Engineering - Kubernetes - Distributed Systems

Servicenow • San Diego (CA)

On-site
USD 181,000 - 317,000
Health plans
401(k) Plan with company match
ESPP
+2
Principal Software Engineer - Multi-Cloud Control Plane - Kubernetes
Principal Software Engineer - Multi-Cloud Control Plane - Kubernetes

Servicenow • San Diego (CA)

Hybrid
USD 255,000 - 445,000
Health insurance
401(k) Plan
Flexible time off
Principal Software Engineer - Multi-Cloud - Control Plane - Kubernetes
Principal Software Engineer - Multi-Cloud - Control Plane - Kubernetes

Servicenow • Santa Clara (CA)

Hybrid
USD 255,000 - 445,000
Principal Software Engineer - Multi-Cloud - Control Plane - Kubernetes
Principal Software Engineer - Multi-Cloud - Control Plane - Kubernetes

ServiceNow • California (MO)

Hybrid
USD 255,000 - 445,000
Health plans
401(k) Plan with company match
ESPP
+3
Sr SRE Automation Engineer
Sr SRE Automation Engineer

Compunnel, Inc. • Austin (TX), Northern (KY)

Hybrid
USD 130,000 - 180,000
Senior Staff Software Engineer – SRE, Release & Test Platforms
Senior Staff Software Engineer – SRE, Release & Test Platforms

ServiceNow • California (MO)

On-site
USD 180,000 - 270,000
Health plans
401(k) Plan with company match
ESPP
+3