Staff Software Engineer (Site Reliability Engineer)

Harvey

San Francisco (CA)

On-site

USD 130,000 - 180,000

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Deepstreamtech is looking for a Staff Software Engineer for the Site Reliability team in San Francisco. In this role, you will ensure the reliability, scalability, and performance of our legal AI platform. Responsibilities include leading incident management, automating operational workflows, and establishing best practices for security and compliance. The ideal candidate has over 10 years of experience in Site Reliability Engineering and is proficient with cloud infrastructure and automation tools. Join us to make our systems fast, secure, and resilient.

Qualifications

  • 10+ years of experience in Site Reliability Engineering or similar roles.
  • Proven ability to mentor and guide technical teams.
  • Expertise in infrastructure as code tools like Pulumi, Terraform, CloudFormation.

Responsibilities

  • Ensure reliability, scalability, and performance of the legal AI platform.
  • Design and manage monitoring, alerting, and infrastructure resources across global regions.
  • Lead incident management processes and drive actionable improvements.
  • Automate operational tasks and workflows for high reliability.

Skills

Site Reliability Engineering
Infrastructure as Code (IaC)
Observability Tools
Cloud Infrastructure Platforms
Programming Skills (Python, Bash, Go)
CI/CD
Kubernetes
Cloud Security Principles
Problem-Solving Skills

Job description

Requirements
  • If you’re passionate about building robust systems and reducing complexity through automation, we’d love to work with you
  • 10+ years of experience in Site Reliability Engineering or similar roles supporting production environments, with proven ability to mentor and guide technical teams
  • Expertise in infrastructure as code (IaC) tools (Pulumi, Terraform, CloudFormation, etc.)
  • Deep familiarity with observability tools (Datadog, Sentry, etc.) and incident response practices (PagerDuty, IncidentIO, etc.)
  • Proficiency with cloud infrastructure platforms (Azure, GCP, AWS, etc.)
  • Strong programming skills (Python, Bash, Go, or similar languages)
  • Proven track record of diagnosing complex system problems and implementing durable solutions
  • Solid understanding of CI/CD, Kubernetes, containerization, networking, and cloud security principles
  • Excellent problem-solving skills, meticulous attention to detail, and a commitment to operational excellence
What the job involves
  • As a Staff Software Engineer on the Site Reliability team at Harvey, you will ensure the reliability, scalability, and performance of our legal AI platform. You’ll join a high-leverage team that sits at the intersection of infrastructure and product, owning the systems that keep our platform fast, secure, and always on
  • From scaling across 50+ regions to automating mission-critical operations, your work will ensure that Harvey remains resilient as we grow
  • Design, implement, and manage monitoring, alerting, and infrastructure resources (compute, storage, networking) across 50+ global regions
  • Lead incident management processes, including postmortems, root cause analyses, and driving actionable improvements
  • Automate operational tasks and workflows, building tools and processes for capacity planning, graceful rollouts, and safe data access to maintain high reliability and reduce manual intervention
  • Establish best practices for security, compliance, and reliability and collaborate across teams to drive these principles throughout the software lifecycle
  • Optimize infrastructure costs through strategic capacity planning and build-versus-buy decisions while maintaining system performance, reliability, and functionality
  • Provide technical mentorship and leadership, promoting best practices and fostering team growth
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Software Engineer, Site Reliability Engineer
Senior Software Engineer, Site Reliability Engineer

Harvey • San Francisco (CA)

On-site
USD 200,000 - 260,000
Staff Software Engineer, Site Reliability Engineer
Staff Software Engineer, Site Reliability Engineer

Harvey, Inc. • San Francisco (CA), Northern (KY)

On-site
USD 238,000 - 290,000
Relocation assistance
In-person work model
Senior Software Engineer, Site Reliability Engineer
Senior Software Engineer, Site Reliability Engineer

Harvey, Inc. • San Francisco (CA), Northern (KY)

On-site
USD 200,000 - 260,000
Staff Software Engineer, Core Infrastructure
Staff Software Engineer, Core Infrastructure

Harvey • San Francisco (CA)

On-site
USD 236,000 - 290,000
Staff Software Engineer, Core Infrastructure
Staff Software Engineer, Core Infrastructure

Harvey • New York (NY)

On-site
USD 201,000 - 264,000
Staff Software Engineer, Developer Experience
Staff Software Engineer, Developer Experience

Harvey, Inc. • San Francisco (CA), Northern (KY)

On-site
USD 238,000 - 290,000
Senior Software Engineer, Core Infrastructure
Senior Software Engineer, Core Infrastructure

Harvey • San Francisco (CA)

On-site
USD 200,000 - 250,000
Staff/Sr. Staff Software Engineer, Product Engineering
Staff/Sr. Staff Software Engineer, Product Engineering

Harvey • New York (NY)

Hybrid
USD 231,000 - 340,000
Staff/Sr. Staff Software Engineer, Product Engineering
Staff/Sr. Staff Software Engineer, Product Engineering

Harvey • San Francisco (CA)

Hybrid
USD 231,000 - 340,000
Staff Software Engineer, Frontend
Staff Software Engineer, Frontend

Harvey • New York (NY)

On-site
USD 231,000 - 340,000