SRE II: Platform Reliability & Self-Service Automation

Todyl

Atlanta (GA)

Hybrid

USD 130,000 - 160,000

Full time

14 days+
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

Medical, dental, vision coverage
Health savings and flexible spending
Life insurance
Disability protection
Telehealth services
Employee Assistance Program (EAP)
Flexible PTO + 13 holidays
401(k)
Parental leave

Job summary

Todyl is seeking a Site Reliability Engineer to build, operate, and automate a secure cloud platform. You will partner with developers to deploy, scale, and configure services while maintaining reliability and security.

In this role you will own automation-first platform capabilities, drive cost efficiency, and embed security into daily operations. The team values ownership, proactive work, and collaboration across engineering groups.

Qualifications

  • Experience building and operating production platforms with Kubernetes, CI/CD, and IaC.
  • Familiarity with AWS-based cloud infrastructure and observability tools.
  • Proficiency in Python or Bash for tooling and Git workflows.
  • Ability to work with developers and balance security with speed.

Responsibilities

  • Build and operate the production platform including Kubernetes, CI/CD, IaC, observability, secrets management, and AWS foundation on which services run.
  • Automate the path to production with self-service capabilities so engineering teams can deploy and scale without routine help.
  • Drive cost visibility and efficiency across the cloud footprint, including AWS resource tagging and right-sizing.
  • Modernize on-call: living runbooks, trusted alerting, post-incident reviews as a normal part of operations.
  • Embed security into day-to-day operations through patching, access controls, secrets rotation, and dependency hygiene.
  • Partner with product teams early on reliability for high-stakes projects, shaping design rather than last-minute review.
  • Participate in weekly on-call rotation, resolve most issues independently, and document after incidents.
  • Plan and estimate honestly; break work into increments, and write tests for automation that runs in production.
  • Treat code review as a quality lever; push back on debt and monitor dashboards and logs after changes.
  • Mentor less-tenured teammates through pairing and documentation; knowledge flows across the team.
  • When something is mature, hand it off or make it self-managing rather than holding onto it.

Skills

Kubernetes
AWS
Infrastructure-as-code
CI/CD pipelines
Observability
Linux
Python
Git
Bash

Tools

Terraform
Salt

Job description

Todyl is seeking a Site Reliability Engineer to build, operate, and automate a secure cloud platform. You will partner with developers to deploy, scale, and configure services while maintaining reliability and security.

In this role you will own automation-first platform capabilities, drive cost efficiency, and embed security into daily operations. The team values ownership, proactive work, and collaboration across engineering groups.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

SRE II: Cloud Reliability & Automation
SRE II: Cloud Reliability & Automation

Xometry • North Bethesda (MD)

Hybrid
USD 135,000 - 155,000
401(k) match
Medical insurance
Dental & Vision insurance
+1
Hybrid SRE II: Cloud Reliability & Automation
Hybrid SRE II: Cloud Reliability & Automation

Socket.dev • Lexington (KY)

Hybrid
USD 100,000 - 140,000
401(k) match
Medical, dental, and vision insurance
Generous paid time off
Site Reliability Engineer
Site Reliability Engineer

Todyl • Denver (CO)

On-site
USD 90,000 - 130,000
SRE II: Build Reliable, Automated Systems
SRE II: Build Reliable, Automated Systems

WEX, Inc. • Seattle (WA)

On-site
USD 85,000 - 102,000
Senior Cloud SRE: Secure, Scalable Infra & Automation
Senior Cloud SRE: Secure, Scalable Infra & Automation

Okta • Bellevue (WA)

On-site
USD 147,000 - 202,000
Amazing Benefits
Making Social Impact
Fostering Diversity, Equity, Inclusion
Remote SRE II: Cloud-Native Reliability & Automation
Remote SRE II: Cloud-Native Reliability & Automation

NationsBenefits, LLC • Plantation (FL)

On-site
USD 110,000 - 160,000
Unlimited PTO
Competitive compensation & benefits
Career growth opportunities
+1
SRE II: Cloud Reliability, Automation & Incidents
SRE II: Cloud Reliability, Automation & Incidents

Artha Nexgen • O’Fallon (MO), Northern (KY)

Hybrid
USD 100,000 - 150,000
Senior SRE: Automation, Reliability & Self-Healing Expert
Senior SRE: Automation, Reliability & Self-Healing Expert

TechDigital Group • Pittsburgh

On-site
USD 120,000 - 160,000
SRE for AI Platform: Reliability at Scale
SRE for AI Platform: Reliability at Scale

Mosaic.tech • San Francisco (CA)

On-site
USD 350,000 - 475,000
Health, dental, and vision benefits
Unlimited PTO
Parental leave
+1
Remote SRE II: Kubernetes & Automation
Remote SRE II: Kubernetes & Automation

NationsBenefits, LLC • United States

On-site
USD 110,000 - 160,000
Unlimited PTO
Competitive benefits
Career growth