Site Reliability Engineer

United States Digital Space LLC

San Francisco (CA)

On-site

USD 120,000 - 150,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

United States Digital Space LLC is looking for a passionate Site Reliability Engineer to enhance our monitoring, alerting infrastructure, and incident response process. Located in San Francisco, CA, this hybrid role involves closely collaborating with engineering to ensure system health and visibility.

Ideal candidates will possess hands-on experience with observability tools like Prometheus and Grafana and familiarity with AWS. Join our small team to use LLMs effectively as part of your workflow.

Qualifications

  • Hands-on experience with Prometheus and Grafana (or similar).
  • Strong instincts for what to instrument and good alerting.
  • Comfort debugging distributed systems across the entire stack.

Responsibilities

  • Manage observability stack including Prometheus and Grafana.
  • Drive incident response and build postmortems.
  • Define SLIs, SLOs, and error budgets for reliability.

Skills

Hands-on experience with Prometheus
Experience with Grafana or similar tools
Strong instincts for monitoring
Debugging distributed systems
AWS familiarity
LLM awareness for automation

Tools

AWS
Prometheus
Grafana
CDK or Terraform

Job description

the company | Site Reliability Engineer | San Francisco, CA (Hybrid) | Full-time

the company is a no-code data workflow automation tool that helps operations teams move, transform, and automate their data without writing code. LLMs are a core part of our product — we use them to help users build and reason about their workflows — and they're increasingly part of how we run infrastructure too. We're a small, product-focused team and our infrastructure runs on AWS. We're looking for an SRE that's passionate about observability and keeping systems healthy and understandable. You'll own our monitoring and alerting infrastructure, drive incident response, and work closely with engineering to make sure we have deep visibility into everything that matters. We expect you to use LLMs heavily in your work — writing runbooks, generating alert configs, drafting postmortems, building dashboards — and we want someone who's already figured out how to make that feel natural.

Responsibilities
  • Observability stack — Prometheus, Grafana, dashboards, alerting, and on-call workflows
  • Incident response and postmortems — building a culture of learning from failures
  • SLIs, SLOs, and error budgets — helping the team make data-driven reliability decisions
  • Monitoring LLM-specific infrastructure: latency, token throughput, model error rates, cost attribution
  • AWS infrastructure across our stack (Lambda, ECS, RDS, OpenSearch, CloudFront, etc.)
  • CDK-based IaC and CI/CD pipelines as needed
Qualifications
  • Hands-on experience with Prometheus and Grafana (or similar — Datadog, Honeycomb, etc.)
  • Strong instincts for what to instrument and what good alerting actually looks like
  • Comfort debugging distributed systems across the full stack
  • Experience owning on-call and incident response end to end
  • AWS familiarity and enough IaC experience to get things done (CDK or Terraform)
  • Someone who reaches for an LLM before writing boilerplate from scratch — and knows when not to
Nice to have
  • Experience instrumenting LLM pipelines
  • Familiarity with TypeScript/Node.js
  • Startup experience
  • Background in security and compliance
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer
Site Reliability Engineer

TalentDome Staffing • United States

On-site
USD 140,000 - 210,000
Staff Site Reliability Engineer
Staff Site Reliability Engineer

Careers • San Francisco (CA)

Hybrid
USD 180,000 - 240,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

The ReWork Group • New York (NY)

On-site
USD 120,000 - 160,000
Lead Site Reliability Engineer
Lead Site Reliability Engineer

Selby Jennings • Wilmington (NC)

On-site
USD 140,000 - 200,000
Observability SRE — Incident Response & Dashboards
Observability SRE — Incident Response & Dashboards

United States Digital Space LLC • San Francisco (CA)

Hybrid
USD 120,000 - 150,000
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

Methodic • San Francisco (CA)

On-site
USD 140,000 - 210,000
Site Reliability Engineer
Site Reliability Engineer

Harrison Clarke • New York (NY)

On-site
USD 120,000 - 160,000
Sr. Site Reliability Engineer
Sr. Site Reliability Engineer

MeridianLink • United States

Remote
USD 140,000 - 190,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Jobtailor • Arlington (VA)

On-site
USD 140,000 - 200,000
Staff Site Reliability Engineer
Staff Site Reliability Engineer

Sight Machine • Ann Arbor (MI)

Hybrid
USD 150,000 - 210,000
Stock Options
Health Care Coverage
Life Insurance
+8