SRE: Scalable ML Infra & CI/CD Architect

Baseten

San Francisco (CA)

On-site

USD 165,000 - 330,000

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

Competitive compensation with equity
100% medical, dental, and vision coverage
Generous PTO including Winter Break
Paid parental leave
Company-facilitated 401(k)

Job summary

A leading AI firm in San Francisco seeks a Site Reliability Engineer to ensure scalable and reliable infrastructure for machine learning models. Responsibilities include maintaining infrastructure, automating CI/CD processes, and mentoring team members. Ideal candidates hold a degree in Computer Science or similar and have extensive Kubernetes experience. The position offers competitive compensation, full medical coverage, generous PTO, and a collaborative team environment. Join us to shape the future of AI!

Qualifications

  • Bachelor’s, Master’s, or Ph.D. degree in a relevant field.
  • Extensive experience with Kubernetes.
  • Experience in building and maintaining scalable infrastructure.

Responsibilities

  • Build and maintain scalable infrastructure to support machine learning models.
  • Establish standards and best practices for reliability across infrastructure.
  • Automate processes for managing CI/CD pipelines.

Skills

Experience with Kubernetes
Scalable infrastructure development
Infrastructure as code tools
CI/CD tooling experience
Observability tools knowledge

Education

Bachelor’s, Master’s, or Ph.D. in Computer Science, Engineering, Mathematics

Tools

Terraform
GitHub Actions
Prometheus
Grafana stack

Job description

A leading AI firm in San Francisco seeks a Site Reliability Engineer to ensure scalable and reliable infrastructure for machine learning models. Responsibilities include maintaining infrastructure, automating CI/CD processes, and mentoring team members. Ideal candidates hold a degree in Computer Science or similar and have extensive Kubernetes experience. The position offers competitive compensation, full medical coverage, generous PTO, and a collaborative team environment. Join us to shape the future of AI!
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

SRE: AI Inference Platform & ML Systems
SRE: AI Inference Platform & ML Systems

Cohere • San Francisco (CA)

Hybrid
USD 120,000 - 150,000
Site Reliability Engineer — ML Infra, Scale & Equity
Site Reliability Engineer — ML Infra, Scale & Equity

Baseten • New York (NY)

On-site
USD 165,000 - 330,000
SRE: AI Infra & ML Platforms in Hybrid Cloud - Equity
SRE: AI Infra & ML Platforms in Hybrid Cloud - Equity

FLUIX • Palo Alto (CA)

On-site
USD 120,000 - 150,000
Attractive compensation package including equity options
Comprehensive health, dental, and vision insurance
Opportunities for professional growth
SRE: AI/ML Infra on Kubernetes, AWS & Terraform
SRE: AI/ML Infra on Kubernetes, AWS & Terraform

Deepgram • United States

Hybrid
USD 120,000 - 150,000
Senior SRE: ML Infra at Scale, Multi-Cloud K8s
Senior SRE: ML Infra at Scale, Multi-Cloud K8s

The Consensus • New York (NY)

On-site
USD 110,000 - 140,000
Competitive compensation
100% insurance coverage
Flexible PTO policy
+3
Site Reliability Engineer — ML Infra & Observability
Site Reliability Engineer — ML Infra & Observability

Baseten • San Francisco (CA)

On-site
USD 135,000 - 285,000
Competitive compensation including equity
100% coverage of medical, dental, and vision insurance
Flexible PTO policy
+3
Senior SRE: Kubernetes, GPU Infra & ML Ops Leader
Senior SRE: Kubernetes, GPU Infra & ML Ops Leader

Gruve • Redwood City (CA)

On-site
USD 120,000 - 150,000
SRE for AI Training Pipelines & RL Runs
SRE for AI Training Pipelines & RL Runs

Mosaic.tech • San Francisco (CA)

On-site
USD 350,000 - 475,000
Visa sponsorship
Relocation support
Unlimited PTO
+1
Global Remote SRE for AI Infrastructure & Kubernetes
Global Remote SRE for AI Infrastructure & Kubernetes

Andromeda Cluster • San Francisco (CA)

Hybrid
USD 120,000 - 160,000
Site Reliability Engineer
Site Reliability Engineer

Latent • San Francisco (CA)

On-site
USD 140,000 - 200,000