Senior Site Reliability Engineer

Drata

San Francisco (CA)

On-site

USD 166,900 - 225,900

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Stock equity
Up to 100% employer-paid medical coverage
401(k) plan
Paid parental leave
Flexible vacation policy

Job summary

Drata is looking for a Site Reliability Engineer to join their SRE team in San Francisco. In this role, you'll operate at the intersection of software and systems engineering, focusing on reliability through automation. You'll lead production readiness reviews, define service-level objectives, and contribute to cross-team infrastructure standards. This position requires at least 6 years of experience in SRE or cloud engineering, with strong skills in Terraform and Datadog. The compensation range is $166,900 - $225,900 annually along with equity options.

Qualifications

  • 6+ years of experience in Site Reliability Engineering or building and maintaining scalable services.
  • Hands-on experience with Datadog for monitoring and alerting.
  • Strong understanding of observability concepts.

Responsibilities

  • Lead Production Readiness Reviews before new services launch.
  • Partner with product engineering to define SLOs and SLIs.
  • Build reusable artifacts like SLO templates and alerting standards.

Skills

Site Reliability Engineering
Cloud Engineering
Terraform
Docker
Monitoring with Datadog
Automation in Python
CI/CD pipeline automation
Incident management
Observability concepts
Container orchestration

Tools

GitHub Actions
AWS ECS
Kubernetes

Job description

Job Summary

Drata's SRE team operates as both a central engineering function and an embedded reliability practice. You'll be part of a close-knit SRE team where you grow your career, shape standards, and collaborate with peers—while also serving as the dedicated reliability partner for one of Drata's product engineering teams across the full lifecycle of their work.

This is a highly technical role at the intersection of software engineering and systems engineering. The best SREs at Drata are engineers first—they solve problems by building solutions, not by executing manual processes. Automation is a core value, and nowhere is that more visible than in how we approach reliability.

Our infrastructure runs on AWS across multiple accounts, defined entirely in Terraform. You'll work across a modern cloud‑native stack to help Drata scale reliably for a rapidly growing customer base.

Reliability Architecture for Your Product Team

You are the reliability expert for your aligned product team. You engage early—during architecture reviews and design discussions—to surface risks before they become incidents.

  • Lead Production Readiness Reviews (PRRs) before new services launch, with authority to flag gaps and gate launches when critical reliability standards aren't met.
  • Partner with product engineering leads and staff engineers to define SLOs and SLIs for critical services, turning reliability from a vague goal into a measurable commitment.
  • Participate in team planning and architecture reviews to provide proactive reliability guidance.
  • Build reusable artifacts—SLO templates, observability checklists, alerting standards, reference dashboards—to raise the reliability floor across the team, not just the services you touch directly.
Eliminating Toil Through Engineering

You handle operational needs from your product team, but your job isn't to be a help desk. Your goal is to make each request the last of its kind. When an engineer needs something, your priority is to automate it so anyone can do it, document it so the team can self‑serve, and execute it manually only as a last resort.

  • Build and maintain Datadog monitors, dashboards, and alert routing—enforcing infrastructure‑as‑code standards via Terraform so those resources are owned, versioned, and auditable.
  • Handle infrastructure requests: ECS task management, secret rotations, Terraform changes, capacity adjustments.
  • Identify repeated manual work and convert it into self‑service tooling or runbooks.
  • Audit existing services for reliability anti‑patterns and surface top risks before they cause incidents.
Central SRE Platform Work

Beyond your product team, you contribute to cross‑cutting infrastructure, tooling, and standards that benefit every team at Drata. Recent examples include automated Datadog governance workflows, dynamic AWS account provisioning, and disaster recovery exercises.

  • Design and build shared platform infrastructure—Reusable Terraform modules, standardized observability stacks, service templates—to compound reliability improvements across the organization.
  • Participate in the on‑call rotation and lead incident response when needed; conduct thorough post‑incident reviews to drive lasting fixes.
  • Design and manage CI/CD pipelines using GitHub Actions.
  • Contribute to evolving SRE standards, tooling, and practices across the organization.
What you’ll bring
  • 6+ years of experience in Site Reliability Engineering, Cloud Engineering, or building and maintaining scalable, resilient services.
  • Robust knowledge of cloud computing technologies: Terraform, Docker, Git, and Linux.
  • Hands‑on experience with Datadog for monitoring, alerting, dashboards, SLO tracking, and distributed tracing.
  • Experience building software systems as a software engineer.
  • Experience developing tooling and automation in Python and/or Bash.
  • Experience with CI/CD pipeline automation, specifically GitHub Actions.
  • Experience with disaster recovery practices and incident management.
  • Strong understanding of observability concepts—monitoring, logging, distributed tracing, and metrics—and how to apply them to production systems.
  • Experience with container orchestration and deployment technologies including AWS ECS Fargate and/or Kubernetes.
  • Experience working with relational databases (MySQL proficiency is a plus).
  • Ability to take ownership of problems and act on them independently in a constantly evolving environment.
Nice to Have
  • Experience with AIOps—using AI/ML‑based tooling for anomaly detection, predictive alerting, or automated incident triage.
  • Familiarity with the reliability characteristics of AI/ML‑backed services (e.g., LLM inference latency, non‑determinism, prompt pipeline observability).
  • Experience with the JavaScript/Node.js ecosystem.
  • Certified Kubernetes Administrator (CKA) certification.
  • Familiarity with compliance frameworks like SOC 2, ISO 27001, or NIST.
AI Experience (required – at least one of the following)
  • Hands‑on experience using AI‑assisted development tools (e.g., GitHub Copilot, Cursor, or similar) to accelerate automation, scripting, or infrastructure work.
  • Demonstrated use of AI/AIOps capabilities for reliability tasks—anomaly detection, incident triage, runbook generation, or alert noise reduction.
  • Familiarity with the operational characteristics of AI/ML‑backed services and what it means to make them observable and reliable in production.
  • Demonstrated passion for AI through personal projects, contributions, or continuous learning in the context of infrastructure or reliability engineering.
How we support you

We offer a comprehensive total rewards package designed to power well‑being, accelerate growth, and maintain work‑life balance.

  • Shared Success: Stock equity to give every employee ownership and a share in company growth.
  • Health & Wellness: Up to 100% employer‑paid premiums for medical, dental, and vision coverage for employees and their dependents, along with comprehensive wellness benefits and healthcare concierge services.
  • Financial Well‑being: 401(k) plan, company‑paid life and disability insurance, tax‑advantaged spending accounts, and discounted voluntary offerings.
  • Family Support: Paid parental leave policy after six months of employment; fertility and family‑building benefits and dedicated leave specialists.
  • Growth & Development: Annual stipends for professional and personal development, internal learning opportunities, and support to advance career.
  • Time Off & Flexibility: Flexible vacation policy, paid holidays, and other perks to recharge.

Compensation

This role will receive a competitive base salary, benefits, and stock, typically in the form of Restricted Stock Units (RSUs). The applicable salary range for this role is: $166,900 - $225,900.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Software Engineer II
Senior Software Engineer II

Drata • San Francisco (CA)

On-site
USD 174,000 - 216,000
Stock equity
100% employer-paid medical, dental, and vision premiums
401(k) plan
+3
Senior Software Engineer 2, IAM
Senior Software Engineer 2, IAM

Drata • San Francisco (CA)

Hybrid
USD 174,000 - 237,000
Stock equity
100% employer-paid health premiums
Flexible vacation policy
+1
Staff Software Engineer, Core GRC
Staff Software Engineer, Core GRC

Cacheflow • San Francisco (CA)

Hybrid
USD 200,000 - 272,000
100% employer-paid premiums for medical, dental, and vision coverage
401(k) plan
Paid Parental Leave policy
+1
Staff Platform Engineer, Interoperability
Staff Platform Engineer, Interoperability

Cacheflow • San Francisco (CA)

Hybrid
USD 200,000 - 272,000
Stock equity
Health and wellness benefits
Time off and flexibility
Staff Software Engineer, Core GRC
Staff Software Engineer, Core GRC

Drata • San Francisco (CA)

On-site
USD 200,000 - 272,000
Equity ownership through stock awards
Employer-paid medical, dental, and vision coverage
401(k) and life insurance benefits
+3
Staff Platform Engineer, Interoperability
Staff Platform Engineer, Interoperability

Drata • San Francisco (CA)

Hybrid
USD 200,000 - 272,000
Stock equity
100% employer-paid health premiums
Flexible vacation policy
+1
Head of Product, Assurance
Head of Product, Assurance

Drata • San Francisco (CA)

On-site
USD 207,000 - 281,000
Health and wellness benefits
401(k) with company match
Flexible vacation policy
Senior IT Engineer
Senior IT Engineer

Drata • San Francisco (CA)

Hybrid
USD 136,000 - 169,000
Stock equity opportunities
Up to 100% employer-paid medical, dental, and vision coverage
401(k) plan with financial benefits
+3
Site Reliability Engineer
Site Reliability Engineer

ProdataKey • Draper (UT)

On-site
USD 75,000 - 125,000
Comprehensive medical coverage
Dental and vision coverage
401(k) with company match
+2
Manager, Technical Support - US
Manager, Technical Support - US

jobr.pro • United States

Hybrid
USD 100,000 - 156,000
Stock equity
Comprehensive health benefits
401(k) plan
+2