Principal Site Reliability Engineer

Saviynt

Vancouver (WA)

On-site

USD 260,000 - 275,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Medium is looking for an experienced Infrastructure Engineer in Vancouver, WA, focused on building and maintaining shared infrastructure services. The ideal candidate will have 9+ years of experience, particularly with Kubernetes, Go and Python programming, and a customer-centric mindset.

This role demands strong technical skills in multi-cloud environments, CI/CD pipelines, and observability tools. A competitive salary ranging from $260,000 to $275,000 annually is offered.

Qualifications

  • 9+ years of experience in Infrastructure Development, Platform Engineering, or Site Reliability Engineering.
  • Deep expertise with Kubernetes in production environments.
  • Strong programming skills in Go (Golang) and Python.
  • Proven experience designing and implementing event-driven architecture.

Responsibilities

  • Design and maintain shared infrastructure services and platforms.
  • Architect, implement, and manage scalable Kubernetes platforms.
  • Develop tools for infrastructure management using Go (Golang).
  • Establish observability and monitoring platforms for self-service insights.

Skills

Kubernetes expertise
Go (Golang) programming
Python programming
Cloud Provider experience
CI/CD pipeline tools
Event-driven architecture
Distributed systems design
Observability and monitoring

Education

Bachelor’s degree in Computer Science, Engineering, or related field

Tools

GitLab CI
AWS
Azure

Job description

Security & Compliance

This role requires compliance with Saviynt’s information security and privacy policies, including annual security training.

What You Will Be Doing
  • Design, build, and maintain shared infrastructure services and platforms that our product and application teams depend on.
  • Create reusable, reliable, and scalable solutions that abstract away complexity, enabling other teams to focus on core business logic and deliver features faster in a multi‑cloud environment.
  • Architect, implement, and manage highly available and scalable Kubernetes platforms as a service for internal consumers.
  • Develop robust internal‑facing tools and automation for infrastructure provisioning and management using Go (Golang).
  • Architect and optimize foundational solutions within Cloud environments (AWS, Azure, etc.), focusing on reusable patterns and modules for other teams.
  • Design and implement shared event‑driven architecture components and messaging platforms using technologies like Kafka or Google Pub/Sub.
  • Develop and maintain robust CI/CD pipelines (e.g., GitLab CI and ArgoCD) as a service, providing standardized and automated deployment workflows for various development teams.
  • Design and build resilient distributed systems components that serve as building blocks for other applications, focusing on reliability, fault tolerance, and performance.
  • Manage and optimize shared infrastructure across multi‑region cloud environments, ensuring platform services are globally available and performant for all consumers.
  • Establish and enhance centralized observability and monitoring platforms and tools that provide self‑service insights for consuming teams.
  • Define and implement clear, well‑documented RESTful API designs for the infrastructure services built.
  • Implement and manage Service Mesh (e.g., Envoy, Istio) capabilities for traffic management, security, and policy enforcement as a shared platform for services.
  • Design, implement, and optimize highly available relational database services or shared data platforms for broad organizational use.
  • Collaborate closely with product development teams to understand their infrastructure needs and pain points, providing technical guidance and support.
  • Participate in on‑call rotations to support the critical shared infrastructure built.
What You Bring
  • 9+ years of experience in an Infrastructure Development, Platform Engineering, or Site Reliability Engineering role, with a strong focus on building tools and services for other engineers.
  • Deep expertise with Kubernetes in production environments, particularly in providing it as a platform (single‑tenant and multi‑tenant deployment architectures).
  • Strong programming skills in Go (Golang) and Python, building robust, maintainable backend services and automation.
  • Extensive hands‑on experience with at least one major Cloud Provider (AWS, GCP, or Azure); multi‑cloud experience is a strong plus.
  • Proven experience designing and implementing event‑driven architecture and message queuing systems (e.g., Kafka, RMQ, NATS) as shared services.
  • Solid understanding and practical experience with CI/CD pipeline tools (especially GitLab CI) and establishing automated delivery processes for other teams.
  • Demonstrable experience designing and operating distributed systems, with patterns for reliable, shared components.
  • Familiarity with multi‑region cloud environments and strategies for building globally distributed and highly available platforms.
  • Proficiency in establishing and utilizing comprehensive observability and monitoring platforms (e.g., Prometheus, Grafana, ELK stack, Datadog) for shared infrastructure.
  • Strong experience with RESTful API design principles and building well‑documented, consumable APIs.
  • Knowledge of Service Mesh concepts and practical experience with solutions like Istio in a platform context.
  • Hands‑on experience with relational databases (e.g., MySQL, PostgreSQL), ideally managing them as a service.
  • Excellent communication skills and the ability to clearly articulate complex technical concepts to technical and non‑technical audiences.
  • A strong customer‑centric mindset, treating internal development teams as the primary customers.
  • Bachelor’s degree in Computer Science, Engineering, or a related field, or equivalent practical or military experience.

$260,000 - $275,000 a year

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior / Staff Site Reliability, Platform Engineering
Senior / Staff Site Reliability, Platform Engineering

Saviynt • Atlanta (GA)

On-site
USD 120,000 - 150,000
Competitive compensation
Benefits package
Career growth opportunities
Principal Engineer, Cloud Platforms
Principal Engineer, Cloud Platforms

Saviynt • Milpitas (CA)

On-site
USD 235,000 - 250,000
Competitive compensation
Benefits and growth opportunities
Work on a mission-critical platform
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Veloc Inc • Coppell (TX)

On-site
USD 140,000 - 190,000
Senior Engineer, Federal Cloud Platform
Senior Engineer, Federal Cloud Platform

Saviynt • Atlanta (GA)

On-site
USD 100,000 - 130,000
Competitive compensation
Growth opportunities
Mission-critical projects
Principal Engineer (Federal Cloud Platform)
Principal Engineer (Federal Cloud Platform)

Saviynt • San Francisco (CA)

On-site
USD 120,000 - 160,000
Competitive compensation
Benefits
Growth opportunities
Senior Engineer, Federal Cloud Platform
Senior Engineer, Federal Cloud Platform

Rival • Atlanta (OH)

On-site
USD 120,000 - 160,000
Competitive compensation
Growth opportunities
Mission-critical projects
Site Reliability Engineer
Site Reliability Engineer

Tata Consultancy Services • Scottsdale (AZ)

On-site
USD 100,000 - 110,000
Discretionary Annual Incentive
Comprehensive Medical Coverage
Family Support: Parental Leaves
+4
Site Reliability Engineer
Site Reliability Engineer

TalentDome Staffing • United States

On-site
USD 140,000 - 210,000
Site Reliability Engineer
Site Reliability Engineer

Govcio LLC • Arlington (TX)

Hybrid
USD 230,000 - 250,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Jobtailor • Arlington (VA)

On-site
USD 140,000 - 200,000