Get more replies from employers
Send a job-specific resume in minutes.
United States Digital Space LLC is looking for a passionate Site Reliability Engineer to enhance our monitoring, alerting infrastructure, and incident response process. Located in San Francisco, CA, this hybrid role involves closely collaborating with engineering to ensure system health and visibility.
Ideal candidates will possess hands-on experience with observability tools like Prometheus and Grafana and familiarity with AWS. Join our small team to use LLMs effectively as part of your workflow.
the company | Site Reliability Engineer | San Francisco, CA (Hybrid) | Full-time
the company is a no-code data workflow automation tool that helps operations teams move, transform, and automate their data without writing code. LLMs are a core part of our product — we use them to help users build and reason about their workflows — and they're increasingly part of how we run infrastructure too. We're a small, product-focused team and our infrastructure runs on AWS. We're looking for an SRE that's passionate about observability and keeping systems healthy and understandable. You'll own our monitoring and alerting infrastructure, drive incident response, and work closely with engineering to make sure we have deep visibility into everything that matters. We expect you to use LLMs heavily in your work — writing runbooks, generating alert configs, drafting postmortems, building dashboards — and we want someone who's already figured out how to make that feel natural.