Remote SRE: Cloud Reliability for AI-Driven Logistics

outpost

United States

Remote

USD 140,000 - 180,000

Full time

14 days+
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Outpost is hiring a Senior SRE to own uptime and incident response as the platform scales. You will lead reliability targets across backend, API, and ML pipelines, enhance monitoring, and build auto-remediation.

You’ll partner with engineering to triage alerts and improve system resilience, with on-call duties and blameless postmortems. You will collaborate across a small, mission-critical team building a real-time, scalable logistics platform that supports critical freight operations globally.

Qualifications

  • 4+ years in an SRE, infrastructure, or backend role with production on-call ownership.
  • Deep experience with a major cloud provider (GCP preferred).
  • Experience building monitoring/alerting/observability stacks (Grafana, Prometheus, Datadog, etc.).
  • Strong scripting/automation skills (Python, Bash, or similar).
  • Comfortable with containerized workloads (Docker) and CI/CD pipelines.
  • Track record of reducing incident volume and improving reliability metrics.
  • Strong English communication, able to engage with technical and non-technical stakeholders.

Responsibilities

  • Own reliability targets across backend/API, worker services, and pipelines.
  • Level up monitoring/alerting and implement auto-remediation at scale.
  • Collaborate to build agents for alert triage and routine remediation.
  • Harden and optimize GCP infrastructure for cost and performance.
  • Own database scale, queries, read replicas, and capacity planning.
  • Improve ML training/monitoring infrastructure reliability with CV/ML teams.
  • Run blameless postmortems and drive root-cause fixes.
  • Participate in on-call rotation.

Skills

SRE experience
GCP expertise
Monitoring stack
Scripting: Python/Bash
Docker & CI/CD
Reliability metrics
English communication

Tools

Grafana
Prometheus
Zabbix
Datadog

Job description

Outpost is hiring a Senior SRE to own uptime and incident response as the platform scales. You will lead reliability targets across backend, API, and ML pipelines, enhance monitoring, and build auto-remediation.

You’ll partner with engineering to triage alerts and improve system resilience, with on-call duties and blameless postmortems. You will collaborate across a small, mission-critical team building a real-time, scalable logistics platform that supports critical freight operations globally.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Site Reliability Engineer (Contract)
Site Reliability Engineer (Contract)

outpost • United States

Remote
USD 140,000 - 180,000
Remote SRE: AI Platform Reliability & Automation
Remote SRE: AI Platform Reliability & Automation

Runpod • United States

On-site
USD 150,000 - 200,000
Remote work first
Competitive base salary
Stock options equity
+2
Senior SRE: AI-Driven Cloud Reliability & Automation
Senior SRE: AI-Driven Cloud Reliability & Automation

Hidden Jobs • United States

Remote
USD 191,000 - 226,000
Equity incentive
Flexible PTO
Health insurance
+2
Senior SRE: Scale Resilient AI Platforms & Automation
Senior SRE: Scale Resilient AI Platforms & Automation

Relx Plc • Philadelphia

Hybrid
USD 95,000 - 159,000
Remote SRE Lead: Reliability & CloudOps Champion
Remote SRE Lead: Reliability & CloudOps Champion

Loadsmart • United States

Remote
USD 58,000 - 82,000
Equity package
PTO with no limit
Remote Brazil
+1
Remote SRE Manager: Lead AI-Driven Reliability & Cloud Ops
Remote SRE Manager: Lead AI-Driven Reliability & Cloud Ops

Arcoro Holdings Corp • Phoenix (AZ), Northern (KY)

Hybrid
USD 200,000 - 220,000
Remote Work
401(k) with Company match
Flexible PTO and Company-paid holidays
Senior SRE: Scale & Reliability for AI-Driven SaaS Platform
Senior SRE: Scale & Reliability for AI-Driven SaaS Platform

Instrumental Inc. • Palo Alto (CA)

On-site
USD 175,000 - 229,000
Health benefits
Commuter plans
Parental leave
Remote SRE for High-Impact AI Platform
Remote SRE for High-Impact AI Platform

Vannevar Labs • San Diego (CA)

Hybrid
USD 140,000 - 200,000
Health insurance
Dental insurance
Vision insurance
+7
Staff SRE Architect: Cloud, Kubernetes & Resilience
Staff SRE Architect: Cloud, Kubernetes & Resilience

Shipt, Inc. • California (MO)

Hybrid
USD 96,000 - 185,000
Medical, dental, vision
401(k)
Paid time off
Senior Production SRE: Cloud & On-Prem Reliability
Senior Production SRE: Cloud & On-Prem Reliability

Weights & Biases • New York (NY)

On-site
USD 140,000 - 180,000
Medical Insurance
Dental Insurance
Vision Insurance
+15