Site Reliability Engineer

Latent

San Francisco (CA)

On-site

USD 140,000 - 200,000

Full time

14 days+
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Latent in San Francisco, CA is seeking an SRE to own production environments and ensure 99.9%+ stability for our clinical AI platform used by health systems. You will drive Kubernetes/Helm, IaC with Terraform, and optimize CI/CD pipelines for TypeScript and Python/ML, while improving DevX and supporting scalable deployments.

You will contribute to maintainability, reliability, and performance of distributed systems in a five-day in-office setting, collaborating with a cross-functional team in a

Qualifications

  • Extensive experience designing, implementing, and maintaining production-grade infrastructure.
  • Hands-on deployment optimization for TypeScript and Python/ML pipelines.
  • Experience with high-availability, distributed systems, and large-scale deployments.
  • Office-based in San Francisco, five days per week.

Responsibilities

  • Infrastructure Ownership: Design, implement, and maintain the production environment, having previously handled 500+ machine deployments.
  • Kubernetes Mastery: Own our containerized infrastructure, leveraging deep expertise in Kubernetes and Helm to manage deployment, scaling, and operational health.
  • CI/CD & Deployment Optimization: Optimize and streamline both the TypeScript and Python/ML deployment pipelines to support high-velocity feature release while maintaining the highest reliability.
  • DevX Support: Support Developer Experience (DevX) work to streamline developer workflows, enhance tool proficiency, and improve CI/CD systems.
  • Infrastructure as Code (IaC): Manage and maintain infrastructure definitions using Terraform.

Skills

Automation
Ownership
Problem solving
DevX support

Tools

Kubernetes
Helm
Terraform
TypeScript
Python/ML
PostgreSQL
Redis
Kafka
CI/CD pipelines

Job description

SRE

Location: San Francisco, CA (5 Days In-Office)

You are the infrastructure expert who enables our rapid product development and guarantees 99.9%+ stability and performance of our clinical AI platform for major health systems. Your focus on operational excellence is directly tied to a patient's access to life-saving treatment.

What We Look for in a Great Engineer
  • Tool Proficiency: You are highly proficient with your tools—you speak command line fluently and have mastered keyboard shortcuts.
  • Ownership: You thrive on owning complex systems and have a proven track record of scaling mission-critical deployments.
  • Automation Drive: You love automating things, always finding new ways to increase your own leverage, and defining standards for operational excellence.
  • Problem Solver: You won't wait for someone else to solve a problem that you're in a position to solve; you are willing to jump into whatever needs to get done.
What You'll Work On (Responsibilities)

As our SRE, you will own the entire production environment and improve the development experience:

  • Infrastructure Ownership: Design, implement, and maintain the production environment, having previously handled 500+ machine deployments.
  • Kubernetes Mastery: Own our containerized infrastructure, leveraging deep expertise in Kubernetes and Helm to manage deployment, scaling, and operational health.
  • CI/CD & Deployment Optimization: Optimize and streamline both the TypeScript and Python/ML deployment pipelines to support high-velocity feature release while maintaining the highest reliability.
  • DevX Support: Support Developer Experience (DevX) work to streamline developer workflows, enhance tool proficiency, and improve CI/CD systems.
  • Infrastructure as Code (IaC): Manage and maintain infrastructure definitions using Terraform.
Technical Qualifications & Environment
  • IaC & Orchestration: Deep, demonstrable experience with Kubernetes, Helm, and Terraform.
  • Scaling Systems: Proven ability to architect and maintain complex, distributed systems with high-availability requirements.
  • Deployment Experience: Hands‑on experience optimizing deployment pipelines for both application code (TypeScript) and machine learning models (Python/ML). Also PostgreSQL, Redis, Kakfa.
  • Core Team Member: Excitement about working five days per week in our San Francisco office.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Site Reliability Engineer
Site Reliability Engineer

Knack Solutions • Richmond (VA)

On-site
USD 100,000 - 130,000
Site Reliability Engineer — 99.9% Uptime, Kubernetes
Site Reliability Engineer — 99.9% Uptime, Kubernetes

Latent • San Francisco (CA)

On-site
USD 140,000 - 200,000
Site Reliability Engineer
Site Reliability Engineer

Amiri Recruiting • Mountain View (CA)

On-site
USD 130,000 - 160,000
Site Reliability Engineer
Site Reliability Engineer

Compunnel, Inc. • Greenwood Village (CO)

On-site
USD 120,000 - 150,000
Site Reliability Engineer
Site Reliability Engineer

TalentDome Staffing • United States

On-site
USD 140,000 - 210,000
Site Reliability Engineer
Site Reliability Engineer

Harvey Nash • United States

Remote
USD 120,000 - 150,000
Site Reliability Engineer
Site Reliability Engineer

Veritas Search Group • Tustin (CA)

On-site
USD 140,000 - 190,000
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

Methodic • San Francisco (CA)

On-site
USD 140,000 - 210,000
Member of Technical Staff, DevOps
Member of Technical Staff, DevOps

Reactor • San Francisco (CA)

On-site
USD 100,000 - 160,000
Competitive salary and early equity
Visa sponsorship
Generous health, dental, and vision coverage
Sr SRE Automation Engineer
Sr SRE Automation Engineer

Compunnel, Inc. • Austin (TX), Northern (KY)

On-site
USD 130,000 - 180,000