Staff Site Reliability Engineer

Careers

San Francisco (CA)

Hybrid

USD 180,000 - 240,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Sight Machine is seeking a senior Cloud Infrastructure Engineer to drive reliability, automation, and scale across our AI-driven platform. You will lead IaC, CI/CD, and observability, mentor engineers, and own high‑impact reliability improvements.

You’ll design and operate the infrastructure for agentic AI workloads, LLM gateways, and cross-team reliability reviews, while remaining hands-on with code and incidents.

Qualifications

  • 10+ years designing, building, or operating Kubernetes/Docker in multi-tenant cloud environments.
  • Strong Linux and networking foundations (TCP/IP, app-layer).
  • Hands-on with IaC, CI/CD tooling (Terraform/OpenTofu, FluxCD, Jenkins/GitHub Actions).
  • Experience with agentic AI/LLM-based systems in production is a plus.

Responsibilities

  • Lead reliability practices and IaC across teams; mentor engineers; own critical incidents.
  • Design and operate infrastructure for AI workloads, LLM gateways, and orchestration layers.
  • Develop monitoring, alerting, and observability for AI-driven services.
  • Contribute to runbooks, ADRs, and technical documentation for engineers.

Skills

Kubernetes
Docker
Python
Go
IaC
CI/CD
Linux fundamentals
Networking
Cloud platforms
SRE practices
Documentation

Tools

Terraform
FluxCD
Jenkins
GitHub Actions

Job description

About the team

Sight Machine is built on the shoulders of a unique, robust and highly scalable Infrastructure as Code model. This enables the creation and operation of customer instances in our ecosystem in a standardized and simplified manner. We are looking for team members to help us build, maintain, and improve the infrastructure that makes Sight Machine the leading provider of Manufacturing Data Pipelines and Analytics.

Great things happen when people can bring their authentic selves to work. We empower all of our team members to share their perspectives, passions and experiences because collectively we make a better, stronger team through always “open communications” mind.

Our team collaborates closely with peers & cross functional stakeholders throughout the business, our clients on the forefront of digital transformation, and the cutting edge of digital manufacturing thought leadership.

Sight Machine has offices in San Francisco, CA and Ann Arbor, Mi. We do have a remote-friendly culture with people based all around the US and the rest of the world. For this role in particular, the ideal candidate is located near either of our offices and willing to work in a hybrid capacity. We would still consider 100% remote for exceptional candidates if they aren’t located near an office.

About the role

Join the Cloud Infrastructure Team as a technical leader driving reliability, automation, and scalability across the systems running Sight Machine’s platform. You’ll operate at the intersection of classic SRE discipline which include IaC, CI/CD, observability, incident response and the emerging demands of running agentic AI systems in production: LLM gateways, agent orchestration, and the operational patterns that come with non-deterministic workloads.

This is a senior level IC role. You’ll help set and drive technical direction for infrastructure and reliability practices across teams, mentor senior engineers, and be a primary escalation point for the org’s hardest systems problems while still being hands‑on with code, infrastructure, and incidents.

Success requires deep technical range, sound judgment on risk vs. customer impact, and the ability to influence architecture decisions across Development Engineering without formal authority.

What You’ll Actually Work On
  • Champion an agentic‑AI‑first engineering mindset: identify where AI‑driven automation and agent‑based tooling can replace manual toil, and hold that work to the same quality, testing, and reliability bar as any other production system
  • Evolve reliability practices for meeting reliability SLO’s, error budgets, drive incident postmortems to systemic (not just symptomatic) fixes, and lead reliability reviews for new services before they hit production
  • Troubleshoot and resolve the org’s most complex, cross‑layer systems problems CI/CD, container orchestration, networking, OS, cloud resources, databases, and increasingly, agentic AI/LLM orchestration layers
  • Design, build, and operate the infrastructure supporting agentic AI workloads, LLM gateway routing, agent orchestration frameworks, monitoring of non‑deterministic/AI‑driven services, and the operational tooling needed to run them reliably at scale
  • Architect and instrument monitoring, alerting, and observability infrastructure for critical services, with an eye toward what “critical” means for AI‑driven systems specifically
  • Author and continuously improve operational runbooks and automation, increasingly incorporating agentic/AI‑assisted tooling (e.g., automated triage, AI‑assisted incident response) where it measurably reduces toil
  • Design and build internal platforms and developer tooling that other engineers build on top of
  • Participate in on‑call coverage and help evolve the program as we scale including escalation paths and reducing avoidable pages through better automation
  • Bring a startup mindset of daily engagement: staying close to what’s breaking, what customers are hitting, and where the team needs help, even outside a formal ticket or rotation
  • Mentor senior and mid‑level engineers; act as a technical sounding board across teams
  • Proactively identify and drive cross‑team initiatives that improve stability, reliability, and availability, this is expected to be self‑directed, not assigned
What We’re Looking For
  • Demonstrated experience designing, building, or operating agentic AI/LLM‑based systems in production, held to the same quality‑first, test‑driven rigor as traditional infrastructure code, not just prototype‑grade work
  • Embody a quality‑first and security‑first culture in all that you do
  • 10+ years of experience with Kubernetes/Docker in at least one top‑tier cloud provider (Azure, GCP, AWS), including production‑scale multi‑tenant or multi‑cluster environments
  • 10+ years coding experience (Python, Go, Java, or similar) with a track record of building tools/platforms used by other engineers, not just scripts
  • 10+ years with IaC and CI/CD tooling (Terraform/OpenTofu, FluxCD or similar GitOps tooling, Jenkins/GitHub Actions)
  • Strong, provable Linux and networking fundamentals (TCP/IP and application‑layer)
  • Practical experience integrating or operating LLM/agentic AI systems in a production context this can be API‑based orchestration, LLM gateways, or agent frameworks
  • A track record of authoring technical documentation (design docs, ADRs, runbooks) that other engineers actually use
  • Demonstrated mentorship of other engineers, without needing formal management authority to do it
  • Strong bias for action over endless planning, hands‑on, has made mistakes, learned from them, and can weigh risk vs. customer impact under pressure
  • Clear, empathetic communicator, comfortable pushing back on architecture decisions across teams
  • Operational experience with monitoring/alerting systems (Prometheus, Grafana, Loki, Sentry, Signoz or equivalents)
  • Deep understanding of cloud performance, able to diagnose and resolve bottlenecks others can’t
Nice to Have
  • Experience with elements of our current tech stack are a plus: Kubernetes, FluxCD, Terraform, Helm Charts, Prometheus, Elasticsearch, Python, Java, Kafka, Postgres, and Jenkins
  • Previous experience or a keen interest in industrial IoT, analytics, or manufacturing a plus
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Staff Site Reliability Engineer
Staff Site Reliability Engineer

Jobless • Ann Arbor (MI)

Hybrid
USD 180,000 - 240,000
Health Care Coverage
Life Insurance
Health Savings Account
+3
Senior Site Reliability Engineer
Senior Site Reliability Engineer

hardrockdigital • United States

Hybrid
USD 120,000 - 160,000
Competitive pay and benefits
Flexible vacation allowance
Startup culture with global brand support
+1
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Veloc Inc • Coppell (TX)

On-site
USD 140,000 - 190,000
Site Reliability Engineer
Site Reliability Engineer

TalentDome Staffing • United States

On-site
USD 140,000 - 210,000
Head of SRE
Head of SRE

Wand AI • Palo Alto (CA)

On-site
USD 130,000 - 180,000
Platform Site Reliability Engineer
Platform Site Reliability Engineer

Specter • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 250,000
Staff Site Reliability Engineer
Staff Site Reliability Engineer

Jobtailor • California (MO)

On-site
USD 140,000 - 210,000
Site Reliability Engineer
Site Reliability Engineer

Motion Recruitment Partners LLC • Lake Buena Vista (FL)

On-site
USD 120,000 - 150,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

SDI International • Chicago (IL)

Hybrid
USD 130,000 - 180,000
Site Reliability Engineer
Site Reliability Engineer

United States Digital Space LLC • San Francisco (CA)

On-site
USD 120,000 - 150,000