Site Reliability Engineer

Latent

San Francisco (CA)

On-site

USD 140,000 - 200,000

Full time

4 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Latent in San Francisco, CA is seeking an SRE to own production environments and ensure 99.9%+ stability for our clinical AI platform used by health systems. You will drive Kubernetes/Helm, IaC with Terraform, and optimize CI/CD pipelines for TypeScript and Python/ML, while improving DevX and supporting scalable deployments.

You will contribute to maintainability, reliability, and performance of distributed systems in a five-day in-office setting, collaborating with a cross-functional team in a

Qualifications

  • Extensive experience designing, implementing, and maintaining production-grade infrastructure.
  • Hands-on deployment optimization for TypeScript and Python/ML pipelines.
  • Experience with high-availability, distributed systems, and large-scale deployments.
  • Office-based in San Francisco, five days per week.

Responsibilities

  • Infrastructure Ownership: Design, implement, and maintain the production environment, having previously handled 500+ machine deployments.
  • Kubernetes Mastery: Own our containerized infrastructure, leveraging deep expertise in Kubernetes and Helm to manage deployment, scaling, and operational health.
  • CI/CD & Deployment Optimization: Optimize and streamline both the TypeScript and Python/ML deployment pipelines to support high-velocity feature release while maintaining the highest reliability.
  • DevX Support: Support Developer Experience (DevX) work to streamline developer workflows, enhance tool proficiency, and improve CI/CD systems.
  • Infrastructure as Code (IaC): Manage and maintain infrastructure definitions using Terraform.

Skills

Automation
Ownership
Problem solving
DevX support

Tools

Kubernetes
Helm
Terraform
TypeScript
Python/ML
PostgreSQL
Redis
Kafka
CI/CD pipelines

Job description

SRE

Location: San Francisco, CA (5 Days In-Office)

You are the infrastructure expert who enables our rapid product development and guarantees 99.9%+ stability and performance of our clinical AI platform for major health systems. Your focus on operational excellence is directly tied to a patient's access to life-saving treatment.

What We Look for in a Great Engineer
  • Tool Proficiency: You are highly proficient with your tools—you speak command line fluently and have mastered keyboard shortcuts.
  • Ownership: You thrive on owning complex systems and have a proven track record of scaling mission-critical deployments.
  • Automation Drive: You love automating things, always finding new ways to increase your own leverage, and defining standards for operational excellence.
  • Problem Solver: You won't wait for someone else to solve a problem that you're in a position to solve; you are willing to jump into whatever needs to get done.
What You'll Work On (Responsibilities)

As our SRE, you will own the entire production environment and improve the development experience:

  • Infrastructure Ownership: Design, implement, and maintain the production environment, having previously handled 500+ machine deployments.
  • Kubernetes Mastery: Own our containerized infrastructure, leveraging deep expertise in Kubernetes and Helm to manage deployment, scaling, and operational health.
  • CI/CD & Deployment Optimization: Optimize and streamline both the TypeScript and Python/ML deployment pipelines to support high-velocity feature release while maintaining the highest reliability.
  • DevX Support: Support Developer Experience (DevX) work to streamline developer workflows, enhance tool proficiency, and improve CI/CD systems.
  • Infrastructure as Code (IaC): Manage and maintain infrastructure definitions using Terraform.
Technical Qualifications & Environment
  • IaC & Orchestration: Deep, demonstrable experience with Kubernetes, Helm, and Terraform.
  • Scaling Systems: Proven ability to architect and maintain complex, distributed systems with high-availability requirements.
  • Deployment Experience: Hands‑on experience optimizing deployment pipelines for both application code (TypeScript) and machine learning models (Python/ML). Also PostgreSQL, Redis, Kakfa.
  • Core Team Member: Excitement about working five days per week in our San Francisco office.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer
Site Reliability Engineer

SRE • Puerto Rico

Hybrid
USD 120,000 - 180,000
Site Reliability Engineer — 99.9% Uptime, Kubernetes
Site Reliability Engineer — 99.9% Uptime, Kubernetes

Latent • San Francisco (CA)

On-site
USD 140,000 - 200,000
Site Reliability Engineer
Site Reliability Engineer

Knack Solutions • Richmond (VA)

On-site
USD 100,000 - 130,000
SRE Leader
SRE Leader

Kontakt Micro-Location Sp. Z.o.o. • New York (NY)

Hybrid
USD 180,000 - 260,000
Equity in a high-growth company
Health, dental, and vision coverage
401k
+3
Site Reliability Engineer
Site Reliability Engineer

Amiri Recruiting • Mountain View (CA)

On-site
USD 130,000 - 160,000
Site Reliability Engineer
Site Reliability Engineer

TalentDome Staffing • United States

On-site
USD 140,000 - 210,000
Site Reliability Engineer
Site Reliability Engineer

Compunnel, Inc. • Greenwood Village (CO)

On-site
USD 120,000 - 150,000
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

Virtual Tech Gurus • Puerto Rico

On-site
USD 140,000 - 210,000
Site Reliability Engineer
Site Reliability Engineer

Jobtailor • New Hampshire

On-site
USD 110,000 - 160,000
Site Reliability Engineer (SRE) | Cognitive Minds | Washington, DC
Site Reliability Engineer (SRE) | Cognitive Minds | Washington, DC

Cognitive Minds • Washington

Hybrid
USD 120,000 - 150,000