Site Reliability Engineer — Scale an AI‑Powered SaaS Platform

Instrumental Inc.

Palo Alto (CA)

On-site

USD 140,000 - 165,000

Full time

4 days ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Benefits offered by this job

Health insurance
Vision insurance
Dental plan
Commuter plans
Parental leave

Job summary

Instrumental Inc. is seeking a Site Reliability Engineer to operate, improve, and scale our AWS-based SaaS platform in the Bay Area.

You’ll join a hands-on team focused on reliability, automation, observability, and operational excellence while participating in a bi-weekly on-call rotation. You’ll own production issues from detection through remediation and long-term resolution, collaborating with software engineers to deliver scalable, production-ready services.

Qualifications

  • 3–4 years of SRE/DevOps experience in production SaaS.
  • Strong hands-on AWS experience: EC2, VPC, IAM, RDS, ECS, S3.
  • Experience with Terraform or IaC tooling.
  • Proficient with CI/CD pipelines (GitHub Actions/Jenkins/GitLab).
  • Familiar with Datadog for dashboards, alerts, and APM.
  • Experience with Docker and Kubernetes.
  • On-call experience including incident response and RCA.
  • Ownership mindset to drive production issues to resolution.

Responsibilities

  • Operate, improve, and scale the AWS-based SaaS platform.
  • Design, implement, and uphold reliability and observability.
  • Automate repetitive tasks to reduce toil and errors.
  • Participate in on-call rotations and incident response.
  • Collaborate with software engineers to ship production-ready services.
  • Lead long-term remediation of production issues.

Skills

AWS
Terraform
CI/CD
Datadog
Docker
Kubernetes
Python
On-call
Ownership

Tools

GitHub Actions
Jenkins
GitLab CI/CD
Terraform
Datadog
Docker
Kubernetes
Python

Job description

Instrumental Inc. is seeking a Site Reliability Engineer to operate, improve, and scale our AWS-based SaaS platform in the Bay Area.

You’ll join a hands-on team focused on reliability, automation, observability, and operational excellence while participating in a bi-weekly on-call rotation. You’ll own production issues from detection through remediation and long-term resolution, collaborating with software engineers to deliver scalable, production-ready services.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior SRE: Scale & Reliability for AI-Driven SaaS Platform
Senior SRE: Scale & Reliability for AI-Driven SaaS Platform

Instrumental Inc. • Palo Alto (CA)

On-site
USD 175,000 - 229,000
Health benefits
Commuter plans
Parental leave
Site Reliability Engineer — Scale & Resilience for AI Ops
Site Reliability Engineer — Scale & Resilience for AI Ops

HappyRobot • San Francisco (CA)

On-site
USD 120,000 - 160,000
Senior Site Reliability Engineer — AI Platform Scale
Senior Site Reliability Engineer — AI Platform Scale

Future Secure AI • Austin (TX)

On-site
USD 140,000 - 190,000
Site Reliability Engineer — Scale, Observability & Cloud
Site Reliability Engineer — Scale, Observability & Cloud

Air Apps • San Francisco (CA)

On-site
USD 140,000 - 180,000
Apple hardware ecosystem for work.
Annual Bonus.
Medical Insurance (vision & dental)
+2
Site Reliability Engineer - Scale & Observability
Site Reliability Engineer - Scale & Observability

gamma.app • San Francisco (CA)

On-site
USD 120,000 - 160,000
Flexible work-from-home options
Creative and collaborative team environment
Site Reliability Engineer: Production-Scale AI Infra
Site Reliability Engineer: Production-Scale AI Infra

Better Tomorrow Ventures • New York (NY)

On-site
USD 180,000 - 260,000
Health & Wellness benefits
Unlimited PTO
Office meals stipend
+3
Senior SRE — Scale, Reliability & AI-Driven Ops Leader
Senior SRE — Scale, Reliability & AI-Driven Ops Leader

Instrumental • Palo Alto (CA)

On-site
USD 175,000 - 229,000
Health insurance
Vision insurance
Dental insurance
+2
Site Reliability Engineer
Site Reliability Engineer

Instrumental Inc. • Palo Alto (CA)

On-site
USD 140,000 - 165,000
Health insurance
Vision insurance
Dental plan
+2
Site Reliability Engineer – AI Platform Infra (Onsite NYC)
Site Reliability Engineer – AI Platform Infra (Onsite NYC)

getbasis.ai • New York (NY)

On-site
USD 140,000 - 190,000
Health & Wellness benefits
Time off — unlimited PTO + holidays
In-Office perks — meals, kitchen, desk
Site Reliability Engineer — Scale AI Infra with Ownership
Site Reliability Engineer — Scale AI Infra with Ownership

Happyrobot Inc. • San Francisco (CA)

On-site
USD 100,000 - 140,000
Competitive salary + equity
Ownership & autonomy in projects
Opportunity to work with top-tier engineers