Site Reliability Engineer

Todyl

Denver (CO)

On-site

USD 90,000 - 130,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

A technology company is looking for two Site Reliability Engineers to join their Application Platform Engineering team in Denver, Colorado. The ideal candidates will have experience managing Kubernetes and Linux systems at scale while developing tools and services to enhance application hosting infrastructure. Responsibilities include implementing security policies and improving application monitoring. This is a unique opportunity to make a significant impact on the company's security offering in a collaborative and innovative environment.

Qualifications

  • Experience managing production Linux systems at scale.
  • Experience managing Kubernetes and applications running on Kubernetes.
  • General competency in one or more scripting languages including Python, Perl, or Bash.

Responsibilities

  • Develop tools and services for Application hosting infrastructure.
  • Implement and enforce security policies and system patching.
  • Build automation to improve reliability for Day 2 Operations.
  • Collaborate with product and engineering teams.
  • Improve application monitoring and alerting.
  • Participate in on-call rotation for emergency pages.

Skills

Managing production Linux systems
Managing Kubernetes (k8s)
Scripting languages (Python, Perl, Bash)
Working knowledge of REST APIs
Networking fundamentals
Building production services
Incident response processes

Job description

At Todyl, our Application Platform Engineering team is dedicated to building infrastructure, services and patterns that enable our application development teams to quickly and safely deploy services at the core of our security offering. As a member of this innovative team, you will play a pivotal role in designing and engineering cutting‑edge solutions that are highly performant, highly resilient and low maintenance. Your work will not only directly impact the reliability and security of our platform but also empower the engineering team to continuously push the boundaries of what’s possible in the security space.

Todyl is excited to grow our team and is hiring two Site Reliability Engineers (SRE I and SRE II) to help build and scale our platform!

Responsibilities
  • As an SRE (Site Reliability Engineer), you will be responsible for developing tools and services that support Todyl’s Application hosting infrastructure, including but not limited to K8s and baremetal.
  • Implement and enforce security policies, access control and system patching.
  • Build automation to improve the reliability and reduce human interaction for Day 2 Operations.
  • You will collaborate with product and engineering and deliver solutions that meet the needs of stakeholders and the business.
  • Improve Application monitoring and alerting to minimize time to detect and time to restore.
  • Participate in a weekly on‑call rotation with the team and be available during off‑hours for emergency pages.
Requirements
  • Experience managing production Linux systems at scale
  • MUST HAVE: Experience managing k8s and applications running on k8s.
  • MUST HAVE: General competency in one or more scripting languages including Python, Pearl, or Bash.
  • Working knowledge of REST APIs.
  • Familiarity with building custom Linux ISOs and AMIs.
  • Familiarity with networking fundamentals.
  • Ability to quickly learn new concepts, frameworks, and technologies.
  • Comfortable building and maintaining production services.
  • Experience with on‑call rotations and incident response processes.

Todyl provides equal employment opportunities to all employees and applicants for employment without regard to race, color, religion, gender, sexual orientation, transgender status, gender identity or expression, national origin, age, disability, marital status, genetic information, military status or any other status protected by applicable federal, state or local laws.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer II
Site Reliability Engineer II

Todyl • Atlanta (GA)

On-site
USD 130,000 - 160,000
Medical, dental, and vision coverage
Competitive 401(k)
Flexible PTO and 13 company holidays
+1
Site Reliability Engineer
Site Reliability Engineer

SRE • Puerto Rico

Hybrid
USD 120,000 - 180,000
Site Reliability Engineer
Site Reliability Engineer

TalentDome Staffing • United States

On-site
USD 140,000 - 210,000
Staff TDI Site Reliability Engineer, Okta Federal
Staff TDI Site Reliability Engineer, Okta Federal

Okta • San Francisco (CA)

On-site
USD 140,000 - 200,000
Site Reliability Engineer
Site Reliability Engineer

ProdataKey • Draper (UT)

On-site
USD 75,000 - 125,000
Comprehensive medical coverage
Dental and vision coverage
401(k) with company match
+2
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

Thinking Machines Lab • San Francisco (CA)

On-site
USD 350,000 - 475,000
Health benefits
Unlimited PTO
Paid parental leave
+1
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

Thinkingmachines • San Francisco (CA)

On-site
USD 350,000 - 475,000
Generous health benefits
Unlimited PTO
Paid parental leave
+1
Site Reliability Engineer
Site Reliability Engineer

NextGen | GTA: A Kelly Telecom Company • Mount Laurel Township (NJ)

On-site
Site Reliability Engineer (Secret Clearance)
Site Reliability Engineer (Secret Clearance)

ROI Services LLC • Huntsville (AL)

On-site
USD 110,000 - 150,000
Technical Support Engineer III
Technical Support Engineer III

Todyl • Denver (CO)

On-site
USD 75,000 - 95,000