Site Reliability Engineer

Instrumental Inc.

Palo Alto (CA)

On-site

USD 140,000 - 165,000

Full time

20 hours ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Benefits offered by this job

Health insurance
Vision insurance
Dental plan
Commuter plans
Parental leave

Job summary

Instrumental Inc. is seeking a Site Reliability Engineer to operate, improve, and scale our AWS-based SaaS platform in the Bay Area.

You’ll join a hands-on team focused on reliability, automation, observability, and operational excellence while participating in a bi-weekly on-call rotation. You’ll own production issues from detection through remediation and long-term resolution, collaborating with software engineers to deliver scalable, production-ready services.

Qualifications

  • 3–4 years of SRE/DevOps experience in production SaaS.
  • Strong hands-on AWS experience: EC2, VPC, IAM, RDS, ECS, S3.
  • Experience with Terraform or IaC tooling.
  • Proficient with CI/CD pipelines (GitHub Actions/Jenkins/GitLab).
  • Familiar with Datadog for dashboards, alerts, and APM.
  • Experience with Docker and Kubernetes.
  • On-call experience including incident response and RCA.
  • Ownership mindset to drive production issues to resolution.

Responsibilities

  • Operate, improve, and scale the AWS-based SaaS platform.
  • Design, implement, and uphold reliability and observability.
  • Automate repetitive tasks to reduce toil and errors.
  • Participate in on-call rotations and incident response.
  • Collaborate with software engineers to ship production-ready services.
  • Lead long-term remediation of production issues.

Skills

AWS
Terraform
CI/CD
Datadog
Docker
Kubernetes
Python
On-call
Ownership

Tools

GitHub Actions
Jenkins
GitLab CI/CD
Terraform
Datadog
Docker
Kubernetes
Python

Job description

Instrumental builds the manufacturing acceleration platform behind the world’s most complex electronics. We capture digital exhaust and engineering context from assembly lines—images, test logs, BOM data, performance, repair cycles—and our AI engines identify insights that are difficult or impossible for human engineers to find. We accelerate the companies building the AI era by improving manufacturing yield, throughput, and ramp. NVIDIA, Meta, Cisco, and their manufacturing partners rely on Instrumental to accelerate new product introduction and production.

The Instrumental platform collects, intelligently transforms, and contextually presents manufacturing data to technical end-users, enabling them to optimize their manufacturing process in real-time. Our core technology is proprietary ML algorithms, packaged in an accessible, user-centric user interface—we believe we must have both the best technology and the best access to that technology to win.

As a Site Reliability Engineer, you’ll operate, improve, and scale our AWS-based SaaS platform. You’ll combine hands-on production operations with engineering, focusing on reliability, automation, observability, and operational excellence. You’ll participate in a bi-weekly on-call rotation, but the goal isn’t simply to keep systems running—it’s to continuously engineer away the operational complexity that comes with scaling our platform and customer base.

Requirements
  • 3–4 years of experience in Site Reliability Engineering, DevOps, Cloud Operations, Platform Engineering, or Systems Engineering supporting production SaaS environments.
  • Strong hands‑on experience with AWS, including EC2, VPC, IAM, RDS, ECS, and S3.
  • Experience managing infrastructure using Terraform or other Infrastructure as Code technologies.
  • Experience designing and supporting CI/CD pipelines using GitHub Actions, Jenkins, GitLab CI/CD, or similar platforms.
  • Strong experience with monitoring and observability tools, preferably Datadog, including dashboards, alerting, logging, and APM.
  • Experience with Docker and Kubernetes.
  • Scripting experience with Python and/or Bash.
  • Experience supporting production environments through an on‑call rotation, including incident response and root cause analysis.
  • Proven ability to take ownership of production issues and drive them through investigation, remediation, and long‑term resolution.
Who You Are
  • Dead serious about performance, scalability, and reliability (PSR): You care deeply about how systems behave in the real world and continuously look for ways to make them more reliable, scalable, observable, and supportable.
  • Automation, automation, automation: If something is repetitive, manual, or error‑prone, your first instinct is to automate it and make it disappear.
  • An engineer at heart: You don’t want to repeatedly fight the same fires. You look for the underlying cause and build durable engineering solutions that reduce operational toil and technical debt.
  • Strong systems thinker: You understand how infrastructure, applications, networks, deployments, monitoring, and people interact—and can troubleshoot complex production issues across those boundaries.
  • Collaborative and reliable: You partner closely with software engineers to make services production‑ready, improve operational workflows, and build reliability into systems before they become problems.
  • Comfortable with growth and ambiguity: You’re comfortable making good decisions without perfect information and adapting as the platform, customer base, and company scale quickly.
Nice To Have
  • Experience working in a high‑growth B2B SaaS environment.
  • Experience implementing SRE practices such as SLIs, SLOs, and error budgets.
  • Experience building internal tooling and automation to eliminate operational toil.
  • Experience supporting multi‑region AWS environments.
  • AWS cost optimization or FinOps experience.
  • Network, application security, and compliance experience.
  • Experience introducing AI tools or processes into engineering and operational workflows.

This position requires access to items and data that are developed under U.S. government contracts and subject to dissemination controls that limit access to U.S. citizens only.

We’re a growing team that works collaboratively, is supportive of each other, and is highly energized by the opportunity for a large impact. We actively work to promote an inclusive environment, valuing passion and the ability to learn. You’re encouraged to apply even if your experience doesn’t precisely match the job description!

The following is a representative annual base salary range for this position within the Bay Area: $140,000-$165,000. Job level and salary opportunities are evaluated through our interview process – we review the experience, knowledge, skills, and abilities of each applicant.

Instrumental is proud to offer a highly‑rated variety of benefits, including health, vision, dental, commuter plans, and parental leave.

At Instrumental, protecting company and customer information is a shared responsibility. Employees are expected to comply with company engineering, security, access control, and privacy policies, and promptly report suspected security incidents or policy violations.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Site Reliability Engineer (SRE)
Senior Site Reliability Engineer (SRE)

Instrumental • Palo Alto (CA)

On-site
USD 175,000 - 229,000
Health insurance
Vision insurance
Dental insurance
+2
Senior Site Reliability Engineer (SRE)
Senior Site Reliability Engineer (SRE)

Instrumental Inc. • Palo Alto (CA)

On-site
USD 175,000 - 229,000
Health benefits
Commuter plans
Parental leave
Senior Solutions Engineer
Senior Solutions Engineer

Instrumental Inc. • Palo Alto (CA)

On-site
USD 170,000 - 198,000
Health
Vision
Dental
+2
Forward Deployed Manufacturing Engineer
Forward Deployed Manufacturing Engineer

Instrumental Inc. • Palo Alto (CA)

On-site
USD 155,000 - 179,000
Health insurance
Vision insurance
Dental insurance
+2
Forward-Deployed Manufacturing Engineer
Forward-Deployed Manufacturing Engineer

Instrumental Inc. • Dallas (TX)

On-site
USD 116,000 - 158,000
Health
Vision
Dental
+2
Principal Solutions Architect
Principal Solutions Architect

Instrumental Inc. • Palo Alto (CA)

On-site
USD 210,000 - 300,000
Health
Vision
Dental
+2
Solutions Engineer
Solutions Engineer

Instrumental Inc. • Palo Alto (CA)

On-site
USD 134,000 - 175,000
Health, Vision, Dental
Public Transit/Commuter Plans
Maternity/Paternity Leave
Lead Recruiter
Lead Recruiter

Instrumental Inc. • Palo Alto (CA)

On-site
USD 170,000 - 221,000
Health
Vision
Dental
+2
Forward-Deployed Manufacturing Engineer
Forward-Deployed Manufacturing Engineer

Clutch Canada • Cincinnati (OH)

On-site
USD 116,000 - 176,000
Health, Vision, Dental
Public Transit/Commuter Plans
Parental Leave
Lead Recruiter
Lead Recruiter

Instrumental • Palo Alto (CA), Northern (KY)

Hybrid
USD 170,000 - 221,000
Health insurance
Vision care
Dental care
+2