Platform Reliability Engineer II — Automate, Scale, Secure

Todyl, Inc.

Denver (CO)

On-site

USD 140,000 - 190,000

Full time

4 days ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Benefits offered by this job

Medical coverage
Dental coverage
Vision coverage
HSA/FSA
Life insurance
Disability insurance
Telehealth
Flexible PTO
401(k)
Parental leave

Job summary

Todyl, Inc. is seeking a Site Reliability Engineer to join the SRE team responsible for a reliable, secure platform built on Kubernetes, AWS, and IaC. You will own automation, tooling, and operational standards that enable developers to move fast without compromising safety and compliance.

The role emphasizes ownership, proactive reliability, and collaboration with product teams. Expect to modernize on-call, improve cost efficiency, and integrate security into daily operations while mentoring

Qualifications

  • Deep experience with Kubernetes-based platforms and cloud infrastructure.
  • Strong automation mindset and production reliability focus.
  • Proficient in IaC tools (Terraform, Salt) and CI/CD pipelines.
  • Experience with observability stacks (Grafana, Prometheus).
  • Proficient in Python or Bash; Git workflows.

Responsibilities

  • Build and operate production platform including Kubernetes, CI/CD, IaC, observability, secrets management, and AWS foundation.
  • Automate deployment path, enable self-service and guardrails to reduce manual work.
  • Drive cost visibility and efficiency across cloud footprint; tag resources, attribute COGs, right-size.
  • Modernize on-call: runbooks, trusted alerts, post-incident reviews.
  • Embed security into day-to-day operations with patching, access controls, rotation, and dependency hygiene.
  • Partner with product teams early on reliability for high-stakes projects.
  • Participate in weekly on-call rotation and document after incidents.
  • Mentor teammates through pairing and documentation.
  • Hand off mature components to self-manage.

Skills

Kubernetes
AWS
Terraform
Salt
CI/CD
Grafana
Prometheus
Linux
Python
Git

Tools

Terraform
Salt
Grafana
Prometheus

Job description

Todyl, Inc. is seeking a Site Reliability Engineer to join the SRE team responsible for a reliable, secure platform built on Kubernetes, AWS, and IaC. You will own automation, tooling, and operational standards that enable developers to move fast without compromising safety and compliance.

The role emphasizes ownership, proactive reliability, and collaboration with product teams. Expect to modernize on-call, improve cost efficiency, and integrate security into daily operations while mentoring

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Site Reliability Engineer
Site Reliability Engineer

Todyl • Denver (CO)

On-site
USD 90,000 - 130,000
Site Reliability Engineer II
Site Reliability Engineer II

Todyl, Inc. • Denver (CO)

On-site
USD 140,000 - 190,000
Medical coverage
Dental coverage
Vision coverage
+7
Site Reliability Engineer II — Scale & Resilience
Site Reliability Engineer II — Scale & Resilience

DAT Freight Solutions • Portland (OR)

Hybrid
USD 95,000 - 134,000
Medical, Dental, Vision
401k matching
Flexible vacation
+1
Site Reliability Engineer II — Hybrid, AWS & Kubernetes
Site Reliability Engineer II — Hybrid, AWS & Kubernetes

Xometry • Waltham (MA)

Hybrid
USD 135,000 - 155,000
401(k) match
Medical, dental and vision insurance
Life and disability insurance
+4
Site Reliability Engineer II: AWS, Kubernetes, CI/CD
Site Reliability Engineer II: AWS, Kubernetes, CI/CD

Xometry Europe GmbH • Lexington (VA)

Hybrid
USD 110,000 - 150,000
401(k) match
Medical insurance
Paid time off
Site Reliability Engineer II — Scale, Automate & Observe
Site Reliability Engineer II — Scale, Automate & Observe

Worky • Denver (CO)

Hybrid
USD 95,000 - 134,000
Medical, Dental, Vision
401k matching
Employee Stock Purchase Plan
+3
Senior Site Reliability Engineer: Scale, Automate, Observe
Senior Site Reliability Engineer: Scale, Automate, Observe

TalentDome Staffing • United States

On-site
USD 140,000 - 210,000
SRE II: Scale Systems with Automation & Observability
SRE II: Scale Systems with Automation & Observability

DAT Freight & Analytics • Seattle (WA)

Hybrid
USD 95,000 - 134,000
Medical insurance
Dental insurance
Vision insurance
+5
SRE Platform Engineer II: Scale, Automate Cloud Systems
SRE Platform Engineer II: Scale, Automate Cloud Systems

DAT Freight & Analytics • Portland (OR)

Hybrid
USD 95,000 - 134,000
Medical, Dental, Vision
Parental Leave
Flexible Vacation Time
+2
Site Reliability Engineer II: Build Scalable Infra
Site Reliability Engineer II: Build Scalable Infra

Xometry • Boston (MA)

Hybrid
USD 135,000 - 165,000
401(k) matching
Medical, dental and vision insurance
Paid time off
+1