Site Reliability Engineer II

Todyl

Atlanta (GA)

On-site

USD 130,000 - 160,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Medical, dental, and vision coverage
Competitive 401(k)
Flexible PTO and 13 company holidays
Generous parental leave

Job summary

A cybersecurity firm in Atlanta is seeking a Platform SRE to develop and maintain services for its application hosting infrastructure, primarily utilizing Kubernetes. The successful candidate will have strong skills in scripting, cloud security, and production environments. Responsibilities include automation, security policy enforcement, and operational reliability. The firm offers competitive salaries, flexible PTO, and excellent health benefits, with a compensation range of $130K - $160K.

Qualifications

  • Experience managing Kubernetes and applications running on Kubernetes.
  • General competency in one or more scripting or programming languages, including Python or Bash.
  • Demonstrated experience identifying and remediating vulnerabilities in production infrastructure.
  • Experience managing production Linux systems at scale.
  • Working knowledge of REST APIs.
  • Comfort with cloud security tooling.
  • Comfort with cloud cost management concepts.
  • Production experience using CI/CD for code deployment.

Responsibilities

  • Develop tools and services for application hosting infrastructure, including Kubernetes environments.
  • Build automation to improve reliability and reduce human interaction for Day-2 operations.
  • Implement and enforce security policies and system patching.
  • Own attack-surface management and drive CVE resolution.
  • Collaborate with product and engineering teams to improve monitoring and alerting.
  • Participate in a weekly on-call rotation and resolve issues independently.

Skills

Kubernetes management
Scripting languages (Python, Bash)
Vulnerability remediation
Production Linux systems
REST APIs
Networking fundamentals
Cloud security tooling
Cloud cost management
CI/CD practices

Job description

About The Role

At Todyl, our Application Platform Engineering team is dedicated to building infrastructure, services and patterns that enable our application development teams to quickly and safely deploy services at the core of our security offering. As a member of this innovative team, you will play a pivotal role in designing and engineering cutting‑edge solutions that are highly performant, highly resilient and low‑maintenance. Your work will not only directly impact the reliability and security of our platform but also empower the engineering team to continuously push the boundaries of what’s possible in the security space.

Responsibilities
  • As a Platform SRE (Site Reliability Engineer) at Todyl, you will develop tools and services that support Todyl’s application hosting infrastructure, including but not limited to Kubernetes‑based environments.
  • Build automation to improve reliability and reduce human interaction for Day‑2 operations, with an emphasis on infrastructure‑as‑code practices.
  • Implement and enforce security policies, access controls, and system patching—treating security hygiene as a first‑class operational responsibility.
  • Own attack‑surface management for production infrastructure: identify exposure, prioritize remediation, and drive CVE resolution to completion rather than leaving findings unactioned.
  • Operationalize security tooling by building integrations, establishing remediation workflows, and ensuring findings are consistently acted upon.
  • Own features and services through deployment and stabilization—work isn’t done until it’s stable in production and documented.
  • Collaborate with product and engineering teams to deliver solutions that meet the needs of stakeholders and the business; improve application monitoring and alerting to minimize time to detect and time to restore; review dashboards and logs to verify deployments succeeded.
  • Identify and drive cost‑optimization opportunities, including resource labeling, right‑sizing, and efficiency improvements, to reduce COGs.
  • Participate in a weekly on‑call rotation, resolve most issues independently, and update runbooks and documentation after incidents.
Requirements
  • Must have: experience managing Kubernetes and applications running on Kubernetes.
  • Must have: general competency in one or more scripting or programming languages, including Python or Bash.
  • Must have: demonstrated experience identifying and remediating vulnerabilities in production infrastructure, including CVE triage and remediation workflows.
  • Experience managing production Linux systems at scale.
  • Working knowledge of REST APIs.
  • Familiarity with networking fundamentals and common attack‑surface concepts (exposed services, misconfigured access controls, unpatched dependencies).
  • Comfort with cloud security tooling and the ability to operationalize findings into actionable remediation work.
  • Comfort with cloud cost management concepts, including resource tagging and cost attribution strategies.
  • Breaks work into incremental deliverables; communicates delays early and tracks progress against estimates.
  • Writes and maintains tests for the automation and tooling you build; proactively considers edge cases and failure conditions.
  • Ability to quickly learn new concepts, frameworks, and technologies, including AI‑assisted tools to accelerate development and reduce toil.
  • Comfortable building and maintaining production services with a strong sense of ownership from build through stabilization.
  • Production experience using CI/CD for code deployment.
  • Experience with on‑call rotations and incident response processes.
What we Offer
  • Health & Wellbeing
    • Medical, dental, and vision coverage for you and your family
    • HSA/FSA options
    • Life insurance and short‑and‑long‑term disability coverage
  • Financial & Future
    • Competitive 401(k) to invest in your future
    • Short- and long-term disability coverage for when life gets unpredictable
  • Flexibility & Time Off
    • Hybrid work schedule
    • Flexible PTO + 13 company holidays
    • Generous parental leave

Todyl provides equal employment opportunities to all employees and applicants for employment without regard to race, color, religion, gender, sexual orientation, transgender status, gender identity or expression, national origin, age, disability, marital status, genetic information, military status or any other status protected by applicable federal, state or local laws.

Compensation Range: $130K - $160K

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer
Site Reliability Engineer

Todyl • Denver (CO)

On-site
USD 90,000 - 130,000
Site Reliability Engineer
Site Reliability Engineer

ProdataKey • Draper (UT)

On-site
USD 75,000 - 125,000
Comprehensive medical coverage
Dental and vision coverage
401(k) with company match
+2
Senior Site Reliability Engineer (In-Office Required)
Senior Site Reliability Engineer (In-Office Required)

Tavily • United States

Hybrid
USD 147,000 - 224,000
Health insurance
401(k) plan
Parental leave
+2
Software/Site Reliability Engineer - FedRAMP
Software/Site Reliability Engineer - FedRAMP

Tenable, Inc. • San Francisco (CA)

Hybrid
USD 131,000 - 176,000
Medical, dental, vision
401(k) with company match
Employee stock purchase plan
+2
Site Reliability Engineer
Site Reliability Engineer

VantageScore® • San Francisco (CA)

On-site
USD 150,000
Medical insurance
Dental insurance
401(k) plan
+1
Software/Site Reliability Engineer - FedRAMP
Software/Site Reliability Engineer - FedRAMP

Embedded Shishya • San Francisco (CA)

Hybrid
USD 131,500 - 175,500
Senior Software Engineer - SRE
Senior Software Engineer - SRE

Socure • New York (NY)

On-site
USD 160,000 - 180,000
DevOps / Site Reliability Engineer ID70127
DevOps / Site Reliability Engineer ID70127

AgileEngine • Atlanta (GA)

Hybrid
USD 100,000 - 130,000
Professional growth
Competitive compensation
Exciting projects
+1
DevOps / Site Reliability Engineer ID70127
DevOps / Site Reliability Engineer ID70127

AgileEngine • Chicago (IL)

Hybrid
USD 120,000 - 150,000
Professional growth
Competitive compensation
Exciting projects
+1
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

Thinking Machines Lab • San Francisco (CA)

On-site
USD 350,000 - 475,000
Health benefits
Unlimited PTO
Paid parental leave
+1