Senior Staff Site Reliability Engineer

Archer

San Jose (CA)

On-site

USD 160,000 - 210,000

Full time

7 days ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Archer is seeking a highly experienced Sr. Staff Site Reliability Engineer to join our growing team in California. You will own the reliability, scalability, performance, and security of core systems and services, designing robust cloud-native infrastructure on AWS and EKS.

You will drive observability, CI/CD automation, data pipelines, and secure operations, collaborating with development teams and handling on-call rotations to maintain production systems' health and resilience.

Qualifications

  • 10+ years in Site Reliability Engineering, DevOps, or similar roles focusing on operational excellence.
  • Strong expertise in AWS, Kubernetes, CI/CD, and security best practices.
  • Hands-on experience with observability stacks (Prometheus, Grafana, ELK).
  • Excellent problem solving, analytics, and communication skills.

Responsibilities

  • Design, implement, and maintain cloud-native infrastructure (AWS/EKS).
  • Develop observability strategies with monitoring, logging, and alerting.
  • Improve CI/CD pipelines and release management practices.
  • Build and maintain data pipelines (Kafka, Airflow, Spark).
  • Collaborate with development teams to bake reliability into software lifecycle.
  • Handle on-call rotations and troubleshoot production issues.

Skills

SRE
DevOps
CI/CD
Observability
Python
Bash
PowerShell
Docker/Kubernetes
Security
Networking
Distributed systems

Education

Bachelor's degree in CS/Engineering

Tools

Jenkins
GitLab CI
ArgoCD
Prometheus
Grafana
ELK
Kafka
Airflow
Spark
Terraform
OpenRouter
Kubernetes
AWS

Job description

  • We are seeking a highly experienced and passionate Sr. Staff Site Reliability Engineer (SRE) to join our growing team. In this critical role, you will be responsible for the reliability, scalability, performance, and security of our core systems and services
  • You will leverage your extensive expertise in various technologies to design, implement, and maintain robust infrastructure and automation solutions
  • Implement and maintain the infrastructure and pipeline required for an internal LLM-powered chat service, potentially leveraging platforms like OpenRouter or similar alternatives
  • Implement and maintain highly available, scalable, and secure cloud-native infrastructure on Amazon Elastic Kubernetes Service (EKS)
  • Develop and implement comprehensive observability strategies, including monitoring, logging, and alerting, to ensure the health and performance of our systems
  • Architect and optimize data pipelines to ensure efficient and reliable data flow across various platforms
  • Drive the continuous improvement of our CI/CD pipelines, promoting best practices for automated testing, deployment, and release management
  • Champion cloud-first strategies, leveraging the full capabilities of cloud platforms for infrastructure, services, and operations
  • Implement and enforce robust security practices across our infrastructure, applications, and data
  • Design and maintain Docker-based containerization solutions for our applications
  • Develop and maintain automation scripts and tools using Python, Bash, and PowerShell
  • Collaborate with development teams to ensure reliability is built into the software development lifecycle from inception
  • Troubleshoot complex production issues across various layers of the stack, identifying root causes and implementing preventative measures
  • Participate in on-call rotations to support production systems

Proven track record in designing and implementing robust data pipelines (e.g., Kafka, Airflow, Spark)Ability to work independently and as part of a highly collaborative teamExpert-level knowledge of cloud platforms (AWS preferred), including infrastructure-as-code principlesStrong background in CI/CD methodologies and tools (e.g., Jenkins, GitLab CI, ArgoCD)Comprehensive understanding of security best practices for cloud environments, applications, and dataSolid understanding of networking concepts, distributed systems, and operating systems12+ years of experience in Site Reliability Engineering, DevOps, or a similar role with a strong focus on operational excellenceExcellent problem-solving, analytical, and communication skillsAdvanced scripting and programming skills in Python, Bash, and PowerShellProficiency in Docker for containerization and orchestrationExtensive experience with observability tools and practices, including Prometheus, Grafana, ELK stack, or similarBachelor’s degree in Computer Science, Engineering, or a related field, or equivalent practical experienceDeep expertise in Amazon EKS, including cluster provisioning, management, and troubleshootingSuccessful candidates must be able to demonstrate U.S. citizenship, permanent residency, or status as a protected individual to satisfy ITAR, contractual, and/or regulatory requirementsCertifications in AWS, Kubernetes, or other relevant technologiesExperience with other Kubernetes distributions or cloud providersFamiliarity with compliance frameworks (e.g., SOC 2, HIPAA, GDPR)

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer
Site Reliability Engineer

TalentDome Staffing • United States

On-site
USD 140,000 - 210,000
Site Reliability Engineer
Site Reliability Engineer

Evlo AI • Minneapolis (MN)

On-site
USD 120,000 - 180,000
Site Reliability Engineer
Site Reliability Engineer

Harrison Clarke • New York (NY)

On-site
USD 120,000 - 160,000
Site Reliability Engineering Manager
Site Reliability Engineering Manager

Calance • United States

Hybrid
USD 150,000 - 200,000
Sr Staff Site Reliability Engineer
Sr Staff Site Reliability Engineer

Archer • San Jose (CA)

On-site
USD 140,000 - 180,000
Site Reliability Engineer
Site Reliability Engineer

JobCubby • Barrington (RI), Northern (KY)

On-site
USD 110,000 - 170,000
Site Reliability Engineer
Site Reliability Engineer

ASCENDING • United States

Remote
USD 150,000 - 210,000
Senior/Staff Cloud Reliability Engineer
Senior/Staff Cloud Reliability Engineer

Cerebras • Mountain View (CA)

On-site
USD 190,000 - 240,000
Senior/Staff Cloud Reliability Engineer
Senior/Staff Cloud Reliability Engineer

ThoughtSpot • Mountain View (CA)

On-site
USD 180,000 - 240,000
Site Reliability Engineer II
Site Reliability Engineer II

Jobtailor • Arlington (VA)

On-site
USD 120,000 - 180,000