Staff Site Reliability Engineer

Jobtailor

California (MO)

On-site

USD 140,000 - 210,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Jobtailor is seeking an experienced SRE/DevOps leader to design, build, and scale a secure AI-native platform. Own reliability, automate deployments, and drive incident response across cloud environments (AWS/GCP).

Candidates should have deep cloud experience, strong IaC (Terraform/Pulumi), observability expertise, and leadership across multi-stakeholder teams; cybersecurity familiarity is a plus in a high-growth enterprise SaaS setting.

Qualifications

  • SRE/DevOps or software/systems engineering with 10+ years managing production systems at scale.
  • Deep expertise in AWS/GCP/Azure and Kubernetes/EKS/GKE, networks, and storage.
  • Extensive IaC experience with Terraform, Pulumi, or similar tools.
  • Hands-on observability stack knowledge with Prometheus, Grafana, ELK, Datadog.
  • Solid understanding of microservices, distributed databases, and event-driven systems.

Responsibilities

  • Design, build, and maintain highly available, scalable infrastructure.
  • Develop internal tooling to streamline deployments, incident response, and capacity planning.
  • Monitor performance and optimize for low-latency AI workloads.
  • Lead incident response, post-mortems, and implement long-term solutions.
  • Collaborate with Product and Security to bake reliability into the development lifecycle.

Skills

SRE/DevOps/Software Eng
Cloud Infra (AWS/GCP/Azure)
Terraform/Pulumi
Observability (Prometheus/Grafana/ELK/
Distributed Systems
Clear Communication
Leadership in multi-team initiatives
Analytical problem-solving
Cybersecurity domain knowledge
AI/ML infra familiarity
Startup mentality
Agentic Workflows

Tools

Kubernetes (EKS/GKE)
Prometheus
Grafana
ELK
Datadog
Terraform
Pulumi
Agentic Workflows
SIEM
EDR

Job description

  • System Resilience: Design, build, and maintain highly available, scalable, and secure infrastructure to support our AI-native cybersecurity platform.
  • Automation & Tooling: Develop internal tooling and automation to streamline deployment processes, incident response, and capacity planning.
  • Performance Engineering: Monitor system performance and proactively identify bottlenecks, optimizing infrastructure for low-latency, high-throughput AI workloads.
  • Incident Management: Lead incident response efforts, conduct post-mortems, and implement long-term solutions to prevent recurring reliability issues.
  • Infrastructure as Code (IaC): Manage infrastructure via code, driving consistency, auditability, and scalability across our cloud environments (e.g., AWS, GCP).
  • Cross-Functional Collaboration: Partner with sibling Engineering teams, Product, and Security teams to ensure reliability is baked into our development lifecycle from concept to production.
Requirements
  • Core Engineering: 10+ years of experience in SRE, DevOps, or Software/Systems Engineering, particularly in managing production systems at scale.
  • Cloud Infrastructure: Deep expertise in public cloud environments (AWS, GCP, or Azure) and managing services such as Kubernetes (EKS/GKE), networking, and storage.
  • Infrastructure as Code: Extensive experience with tools like Terraform, Pulumi, or similar technologies to manage complex infrastructure deployments.
  • Observability: Hands-on experience with monitoring, logging, and tracing stacks (e.g., Prometheus, Grafana, ELK, Datadog) to drive data-informed reliability decisions.
  • Distributed Systems: Solid understanding of microservices architecture, distributed databases, and event-driven systems.
  • Communication: Clear, concise communication skills and a bias for collaborative problem-solving.
  • Leadership Alignment: Proven track record of guiding multi-stakeholder initiatives and influencing engineering practices across teams.
  • Analytical Rigor: Strong problem-solving, debugging, and analytical skills, especially in high-pressure environments.
  • Domain Background: Prior work in cybersecurity, specifically regarding SIEM, EDR, or SOAR infrastructure is nice-to-have.
  • AI/ML Infrastructure: Experience supporting infrastructure for large-scale AI/ML workloads (e.g., GPU scheduling, LLM serving optimization) is nice-to-have.
  • Startup Mentality: Background driving high-impact engineering initiatives in high-growth startups or enterprise SaaS is nice-to-have.
  • Strong familiarity with Agentic Workflows such as Agno, Temporal, etc. is nice-to-have.
Core Competencies

Demonstrates expertise in designing and maintaining scalable, secure cloud infrastructure, with a strong focus on automation, performance engineering, and incident management. Proven ability to collaborate across teams and drive reliability in AI-native cybersecurity platforms.

Highest-signal resume keywords
  • 10+ Years Experience in SRE, DevOps, or Software Engineering
  • Deep Expertise in AWS, GCP, or Azure
  • Extensive Experience with Terraform or Pulumi
  • Hands-On Experience with Prometheus, Grafana, or ELK
  • Strong Problem-Solving and Analytical Skills
ATS Optimization Keywords
Hard Skills
  • Infrastructure as Code
  • Performance Engineering
  • Incident Management
  • Cloud Infrastructure Management
  • Distributed Systems Understanding
  • AI/ML Infrastructure Support
  • Automation Development
  • Monitoring and Observability
  • Capacity Planning
  • Microservices Architecture
Soft Skills
  • Clear Communication Skills
  • Collaborative Problem-Solving
  • Leadership in Multi-Stakeholder Initiatives
Industry Keywords
  • Cybersecurity
  • AI-Native Platforms
  • High-Impact Engineering
  • Startup Mentality
  • Enterprise SaaS
Tools & Technologies
  • Kubernetes (EKS/GKE)
  • Prometheus
  • Grafana
  • ELK
  • Datadog
  • Terraform
  • Pulumi
  • Agentic Workflows
  • SIEM
  • EDR
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Site Reliability Engineer
Senior Site Reliability Engineer

Jobtailor • Town of Florida (NY)

Hybrid
USD 150,000 - 190,000
Lead AI Security Automation Engineer
Lead AI Security Automation Engineer

Jobtailor • Massachusetts

On-site
USD 200,000 - 260,000
Senior AI Security Automation Engineer
Senior AI Security Automation Engineer

Jobtailor • Massachusetts

On-site
USD 180,000 - 260,000
Software Engineer
Software Engineer

Jobtailor • New Jersey

On-site
USD 110,000 - 170,000
Principal Platform Engineer, AI – Automation
Principal Platform Engineer, AI – Automation

Jobtailor • Phoenix (AZ)

On-site
USD 170,000 - 210,000
Principal Cloud Engineer – AI
Principal Cloud Engineer – AI

Jobtailor • West Chester

On-site
USD 150,000 - 210,000
Senior Engineer, Applications Systems
Senior Engineer, Applications Systems

Jobtailor • Town of Florida (NY)

On-site
USD 120,000 - 160,000
Senior Manager, Staff Engineering – Software Development, Microservices
Senior Manager, Staff Engineering – Software Development, Microservices

Jobtailor • Maryland

On-site
USD 180,000 - 240,000
Senior Engineer, IT Infrastructure Engineering – Data Center
Senior Engineer, IT Infrastructure Engineering – Data Center

Jobtailor • Atlanta (GA)

Hybrid
USD 140,000 - 180,000
DevOps Software Engineer
DevOps Software Engineer

Jobtailor • Connecticut

On-site
USD 120,000 - 160,000