Staff Reliability Engineer

Servicenow

Kirkland (WA)

On-site

USD 150,000 - 210,000

Full time

5 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

ServiceNow in Kirkland, WA seeks an experienced SRE/Platform Engineer to design and operate cloud-native platforms for software validation, release validation, and production readiness. You will design production-like environments, integrate automated pipelines, and boost deployment confidence.

You will also build self-service tooling, drive automation, and mentor engineers while partnering with teams to improve platform reliability and cloud-native adoption.

Qualifications

  • 8+ years in SRE/DevOps/Platform/Software/Infrastructure with a related degree or equivalent experience.
  • Hands-on experience with Kubernetes across cluster operations, networking, storage, security, autoscaling, and multi-cluster environments.
  • Experience building and operating cloud-native platforms supporting scalable, highly available services.
  • Experience integrating Kubernetes with CI/CD, GitOps, automated test pipelines, deployment validation, and cloud-native deployment workflows.
  • Experience designing and implementing automation to improve developer productivity, release quality, and operational efficiency.
  • Experience with progressive delivery practices, including canary deployments, feature flags, automated rollback, and deployment verification.
  • Experience with chaos engineering, resilience testing, disaster recovery, and reliability validation.
  • Strong software engineering skills with hands-on experience in Python, Go, Java, or Ruby.

Responsibilities

  • Design, build, and operate cloud-native engineering platforms for software validation, release validation, and production readiness.
  • Design and maintain production-like release and test environments that improve release confidence and deployment readiness.
  • Build and integrate automated test pipelines, observability, reliability signals, and quality gates into CI/CD workflows.
  • Develop automation solutions that improve engineering productivity and reduce manual toil through shift-left practices.
  • Mentor engineers through code reviews, knowledge sharing, and engineering best practices.
  • Partner with engineering teams to improve platform reliability, release quality, and cloud-native adoption.

Skills

Kubernetes
Python
Go
Java
Ruby
AI-assisted engineering
Cloud-native platforms
Automation
CI/CD
GitOps
Observability

Education

Bachelor's degree or higher in a technical field

Tools

Playwright
Selenium
Cypress
REST Assured
PyTest
JUnit/TestNG
GitLab CI/CD
Argo CD
Flux

Job description

Job Description

Join us to build the next generation of cloud-native reliability, release, and test platforms that enable engineering excellence, developer productivity, and high-confidence ServiceNow releases through automation, observability, and AI-driven operations.

What you get to do in this role:
  • Design, build, and operate cloud-native engineering platforms for software validation, release validation, and production readiness
  • Design and maintain production-like release and test ServiceNow environments that improve release confidence and deployment readiness.
  • Build and integrate automated test pipelines, observability, reliability signals, deployment intelligence, and quality gates into CI/CD workflows.
  • Develop automation solutions that improve engineering productivity, streamline operations, and reduce manual toil through shift-left engineering practices.
  • Build reusable frameworks, self-service engineering environments, test data management, mock services, and developer productivity tooling.
  • Design and enhance Kubernetes-based platforms supporting scalable test infrastructure, release automation, cloud-native workloads, and developer self-service.
  • Implement automated validation for failure detection, deployment verification, policy enforcement, security checks, resilience testing, and operational health assessments.
  • Resolve complex platforms, infrastructure, and networking challenges through software engineering, systems design, and automation.
  • Partner closely with engineering teams to improve platform reliability, release quality, cloud-native adoption, and engineering best practices.
  • Participate in architecture reviews, technical design discussions, and implementation of scalable, automation-first engineering solutions.
  • Influence technical decisions through strong engineering execution, collaboration, and delivery of high-quality platform capabilities.
  • Mentor engineers through technical guidance, code reviews, knowledge sharing, and engineering best practices.
  • Foster a culture of reliability, automation, operational excellence, continuous improvement, and customer-focused engineering.
To be successful in this role you have:
  • Experience in leveraging or critically thinking about how to integrate AI into work processes, decision‑making, or problem‑solving. This may include using AI‑powered tools, automating workflows, analyzing AI‑driven insights, or exploring AI's potential impact on the function or industry.
  • 8+ years of experience in Site Reliability Engineering (SRE), DevOps, Platform Engineering, Software Engineering, or Infrastructure Engineering with a Bachelor's degree; or 6 years and a Master's degree; or a PhD with 3 years experience; or equivalent experience.
  • Hands‑on experience with Kubernetes across cluster operations, networking, storage, security, autoscaling, and multi‑cluster environments.
  • Experience building and operating cloud‑native platforms supporting scalable, highly available services.
  • Experience integrating Kubernetes with CI/CD, GitOps, automated test pipelines, deployment validation, and cloud‑native deployment workflows.
  • Experience designing and implementing automation to improve developer productivity, release quality, and operational efficiency.
  • Experience with progressive delivery practices, including canary deployments, feature flags, automated rollback, and deployment verification.
  • Experience with chaos engineering, resilience testing, disaster recovery, and reliability validation.
  • Strong software engineering skills with hands‑on experience designing, developing, testing, and debugging applications using Python, Go, Java, or Ruby.
  • Experience leveraging AI‑assisted engineering for intelligent testing, release risk analysis, incident diagnostics, or operational automation is a plus.
  • Strong understanding of observability, monitoring, SLI/SLOs, incident management, and production operations for distributed systems.
  • Demonstrated ability to solve complex technical problems, drive projects independently, and collaborate effectively across engineering teams.
  • Thrives in fast‑paced, ambiguous environments with a strong ownership mindset, bias for action, and a passion for continuous learning and automation.
  • Low ego, intellectually curious, and an effective collaborator who enjoys partnering with globally distributed teams to deliver reliable engineering solutions.
Good to have:
  • Experience with observability and monitoring platforms for applications, services, and distributed systems at scale.
  • Experience with DevOps automation, CI/CD pipelines, GitOps, and Agile development practices using tools such as GitLab CI/CD, Argo CD, or Flux.
  • Experience building and maintaining enterprise‑scale test automation frameworks using technologies such as Playwright, Selenium, Cypress, REST Assured, PyTest, JUnit/TestNG, or equivalent.
  • Experience with test orchestration, intelligent regression testing, test impact analysis, flaky test detection, parallel execution, and test data management.
  • Exper
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Staff Software Engineer – SRE, Release & Test Platforms
Senior Staff Software Engineer – SRE, Release & Test Platforms

ServiceNow • California (MO)

On-site
USD 180,000 - 270,000
Health plans
401(k) Plan with company match
ESPP
+3
Senior Staff Platform Engineer — AI-Driven Reliability
Senior Staff Platform Engineer — AI-Driven Reliability

ServiceNow • California (MO)

On-site
USD 180,000 - 270,000
Health plans
401(k) Plan with company match
ESPP
+3
Site Reliability Engineer
Site Reliability Engineer

SCIGON • Naperville (IL)

Hybrid
USD 110,000 - 170,000
Site Reliability Engineering Manager
Site Reliability Engineering Manager

O.C. Tanner • Salt Lake City (UT)

On-site
USD 180,000 - 260,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

O.C. Tanner • Salt Lake City (UT)

On-site
USD 130,000 - 180,000
Cloud Solutions Engineer
Cloud Solutions Engineer

Tyler Technologies • Lakewood (CO)

On-site
USD 120,000 - 150,000
Senior SRE (Contract/Hybrid)
Senior SRE (Contract/Hybrid)

Optomi • Orlando (FL)

Hybrid
USD 120,000 - 180,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Veloc Inc • Coppell (TX)

On-site
USD 140,000 - 190,000
Cloud Solutions Engineer
Cloud Solutions Engineer

Tyler-Technologies-29572f8 • Lakewood (CO)

On-site
USD 93,547 - 150,000
Cloud Solutions Engineer
Cloud Solutions Engineer

Tyler Technologies, Inc. • Plano (TX), Latham (NY), Lubbock (TX), Lakewood (CO)

On-site
USD 93,547 - 150,000