Site Reliability Engineer II

Todyl

Atlanta (GA)

Hybrid

USD 130,000 - 160,000

Full time

7 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Medical, dental, vision coverage
Health savings and flexible spending
Life insurance
Disability protection
Telehealth services
Employee Assistance Program (EAP)
Flexible PTO + 13 holidays
401(k)
Parental leave

Job summary

Todyl is seeking a Site Reliability Engineer to build, operate, and automate a secure cloud platform. You will partner with developers to deploy, scale, and configure services while maintaining reliability and security.

In this role you will own automation-first platform capabilities, drive cost efficiency, and embed security into daily operations. The team values ownership, proactive work, and collaboration across engineering groups.

Qualifications

  • Experience building and operating production platforms with Kubernetes, CI/CD, and IaC.
  • Familiarity with AWS-based cloud infrastructure and observability tools.
  • Proficiency in Python or Bash for tooling and Git workflows.
  • Ability to work with developers and balance security with speed.

Responsibilities

  • Build and operate the production platform including Kubernetes, CI/CD, IaC, observability, secrets management, and AWS foundation on which services run.
  • Automate the path to production with self-service capabilities so engineering teams can deploy and scale without routine help.
  • Drive cost visibility and efficiency across the cloud footprint, including AWS resource tagging and right-sizing.
  • Modernize on-call: living runbooks, trusted alerting, post-incident reviews as a normal part of operations.
  • Embed security into day-to-day operations through patching, access controls, secrets rotation, and dependency hygiene.
  • Partner with product teams early on reliability for high-stakes projects, shaping design rather than last-minute review.
  • Participate in weekly on-call rotation, resolve most issues independently, and document after incidents.
  • Plan and estimate honestly; break work into increments, and write tests for automation that runs in production.
  • Treat code review as a quality lever; push back on debt and monitor dashboards and logs after changes.
  • Mentor less-tenured teammates through pairing and documentation; knowledge flows across the team.
  • When something is mature, hand it off or make it self-managing rather than holding onto it.

Skills

Kubernetes
AWS
Infrastructure-as-code
CI/CD pipelines
Observability
Linux
Python
Git
Bash

Tools

Terraform
Salt

Job description

At Todyl, we are on a mission to protect small and medium-sized businesses from ever-changing cyber threats. The Todyl platform fully integrates threat, risk, and compliance management to provide exceptional, affordable, unified cybersecurity solutions to MSPs (Managed Service Providers) and their end customers.


At the end of the day, we're here to keep our partners and customers safe and help them manage the risks and comply with regulations. Protecting others requires a team that works together with trust and cares deeply about carrying out our mission.


About the Role

The Site Reliability Engineering team at Todyl exists to make our platform reliable, secure, and easy for engineering teams to ship to. We do that by building automation, self-service tooling, and operational standards that let developers move fast without putting customers at risk. Our success is measured by how much production reliability and developer velocity we enable, not by how much work flows through us.


You’ll spend your time building the tooling and platform capabilities that let engineering teams deploy, scale, and configure their services without having to file a ticket with us. You’ll partner closely with developers, take operational reliability seriously, and bring an automation-first mindset to a platform that handles security workloads at the heart of our product.


In this role, we’re looking for someone who:


  • Has a bias for action and a strong sense of ownership. They finish what they start and stay with the work through stabilization, not just through a successful deploy.

  • Sees SRE as a service to the engineering organization, not a gate. They build trust with developers and make other teams' jobs easier.

  • Treats security as a normal part of platform operations, not an afterthought, and brings a growth mindset to security regardless of starting expertise.

  • Gets energized by eliminating toil. They look at repetitive work and ask, "How do we make this go away?"

  • Actively uses AI tooling in their day-to-day work and is curious about where it goes next.

  • Can communicate technical decisions clearly to engineers and non-engineers, and is comfortable saying no or pushing back constructively when it matters.


What you’ll do:


  • You’ll build and operate the production platform, including Kubernetes, CI/CD pipelines, infrastructure-as-code, observability, secrets management, and the AWS foundation on which our services run.

  • You’ll automate the path to production, investing in self-service capability so engineering teams can deploy and scale without depending on you for routine work. We’re shifting from reactive to proactive, and we’d rather build guardrails than approve every deploy.

  • You’ll drive cost visibility and efficiency across our cloud footprint, including AWS resource tagging, COGs attribution, and right-sizing across the platform.

  • You’ll modernize how we run on-call: living runbooks, alerting we trust, and post-incident reviews as a normal part of how the team operates.

  • You’ll embed security into day-to-day operations through patching, access controls, secrets rotation, and dependency hygiene, as part of the platform you operate rather than a separate workstream.

  • You’ll partner with product teams early on reliability for high-stakes projects, helping shape the design rather than reviewing it the week before launch.

  • You’ll participate in a weekly on-call rotation, resolve most issues independently, and update documentation after incidents.

  • You’ll plan and estimate honestly. Break work into smaller increments, communicate delays early, and write tests for the automation you build because it runs in production.

  • You’ll treat code review as a quality lever, not a checkbox. Catch missing tests, push back on tech debt, and watch dashboards and logs to verify your own changes after they ship.

  • You’ll mentor less-tenured teammates through pairing, documentation, and the example you set. The team has engineers at different stages, and we expect knowledge to flow across them.

  • When something you’ve built is mature and stable, you’ll look for ways to hand it off or make it self-managing rather than holding onto it forever.


Important note:

We expect the person in this role to actively use AI tools, including Claude, to accelerate automation development, reduce toil, and solve infrastructure problems more quickly. Pairing strong SRE fundamentals with AI-assisted development is increasingly how modern platform teams move at the speed the market requires, and we want a teammate who is comfortable working this way. We’ll talk about how you use AI in your work during the interview, and we expect your fluency with these tools to grow as part of your professional development here.


We don’t expect deep knowledge across every item below, but familiarity with several of these will help you ramp quickly. Most importantly, we’re looking for a strong technical background and the willingness to learn what you don’t already know.



  • Kubernetes and containerization

  • AWS and cloud-native infrastructure

  • Infrastructure-as-code (Terraform, Salt)

  • CI/CD pipelines and automation

  • Observability stack (Grafana, Prometheus)

  • Linux at scale

  • Python or Bash for tooling

  • Git and modern development workflows


What We Offer:

For full-time employees, Todyl offers comprehensive benefits including:



  • Medical, dental, and vision coverage

  • Health savings and flexible spending accounts (HSA/FSA)

  • Life insurance

  • Short- and long-term disability

  • Access to on-demand healthcare and telehealth services

  • Employee Assistance Program (EAP)

  • Flexible PTO in addition to 13 company holidays

  • 401(k)

  • Generous parental leave programs


This role is based in Denver, CO, or Atlanta, GA, with 3 days per week in our office.


The salary range for this position is $130,000–$160,000, plus equity. The actual annual salary for this role will depend on each candidate's experience, qualifications, and work location, with most new hires placed near the midpoint of the posted range to ensure fairness and consistency across our team.


Todyl provides equal employment opportunities to all employees and applicants for employment without regard to race, color, religion, gender, sexual orientation, transgender status, gender identity or expression, national origin, age, disability, marital status, genetic information, military status, or any other status protected by applicable federal, state, or local laws.


We encourage you to apply even if you don’t meet all the requirements listed. We’re looking for the best person for the job, someone who brings a unique combination of skills and experience that makes them exceptional, even if they don’t check every box.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer
Site Reliability Engineer

Todyl • Denver (CO)

On-site
USD 90,000 - 130,000
Senior Software Engineer, DevSecOps
Senior Software Engineer, DevSecOps

Takt • Reston (VA)

On-site
USD 140,000 - 190,000
Health Insurance
401(k) with employer match
Unlimited PTO
+2
Software/Site Reliability Engineer - FedRAMP
Software/Site Reliability Engineer - FedRAMP

Embedded Shishya • San Francisco (CA)

Hybrid
USD 131,500 - 175,500
Staff TDI Site Reliability Engineer, Okta Federal
Staff TDI Site Reliability Engineer, Okta Federal

United States Digital Space LLC • San Francisco (CA)

Hybrid
USD 174,000 - 239,000
Customer Success Manager II
Customer Success Manager II

Todyl • Denver (CO)

On-site
USD 80,000 - 85,000
Medical, dental, and vision coverage
HSA/FSA options
Competitive 401(k)
+1
Senior Site Reliability Engineer II
Senior Site Reliability Engineer II

Juniper Square • United States

On-site
USD 165,000 - 195,000
Health, dental, and vision care
Life insurance
Mental wellness coverage
+3
Senior Backend Engineer
Senior Backend Engineer

In Tandem • Minnesota

On-site
USD 150,000 - 210,000
Medical premium coverage for employees
401k match
Paid parental leave
+3
Staff TDI Site Reliability Engineer, Okta Federal
Staff TDI Site Reliability Engineer, Okta Federal

United States Digital Space LLC • Washington

Hybrid
USD 174,000 - 239,000
Equity
Bonus
Health insurance
+5
Operations Lead, Strategic Growth Initiatives
Operations Lead, Strategic Growth Initiatives

Cozi • Indiana (PA)

On-site
USD 120,000 - 180,000
Medical coverage
401k match
Paid parental leave
+3
Senior Site Reliability Engineer
Senior Site Reliability Engineer

United States Digital Space LLC • Bellevue (CA)

On-site
USD 180,000 - 230,000
Amazing Benefits
Making Social Impact
Fostering Diversity, Equity, Inclusion