Site Reliability Engineer (Contract)

outpost

United States

Remote

USD 140,000 - 180,000

Full time

14 days+
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Outpost is hiring a Senior SRE to own uptime and incident response as the platform scales. You will lead reliability targets across backend, API, and ML pipelines, enhance monitoring, and build auto-remediation.

You’ll partner with engineering to triage alerts and improve system resilience, with on-call duties and blameless postmortems. You will collaborate across a small, mission-critical team building a real-time, scalable logistics platform that supports critical freight operations globally.

Qualifications

  • 4+ years in an SRE, infrastructure, or backend role with production on-call ownership.
  • Deep experience with a major cloud provider (GCP preferred).
  • Experience building monitoring/alerting/observability stacks (Grafana, Prometheus, Datadog, etc.).
  • Strong scripting/automation skills (Python, Bash, or similar).
  • Comfortable with containerized workloads (Docker) and CI/CD pipelines.
  • Track record of reducing incident volume and improving reliability metrics.
  • Strong English communication, able to engage with technical and non-technical stakeholders.

Responsibilities

  • Own reliability targets across backend/API, worker services, and pipelines.
  • Level up monitoring/alerting and implement auto-remediation at scale.
  • Collaborate to build agents for alert triage and routine remediation.
  • Harden and optimize GCP infrastructure for cost and performance.
  • Own database scale, queries, read replicas, and capacity planning.
  • Improve ML training/monitoring infrastructure reliability with CV/ML teams.
  • Run blameless postmortems and drive root-cause fixes.
  • Participate in on-call rotation.

Skills

SRE experience
GCP expertise
Monitoring stack
Scripting: Python/Bash
Docker & CI/CD
Reliability metrics
English communication

Tools

Grafana
Prometheus
Zabbix
Datadog

Job description

About Us:

Outpost is building the backbone of freight. We’re reinventing how supply chain infrastructure works in America with carrier agnostic truck terminals. As a vertically integrated real estate, operations, and technology company, we acquire and operate mission-critical real estate across the country to serve the largest logistics providers in the world. Backed by $1B from Greenpoint Partners, we’re scaling and building the most valuable logistics network in the country.

We thrive on accountability, integrity, and a shared drive to raise the bar. If you’re excited to reshape the industry alongside a high-performance team with a championship mindset that executes relentlessly, welcome aboard.

Role Summary:

Our platform combines AI-powered gate automation, computer vision, and operational software to help logistics operators run smarter, faster facilities. We're a small, high-conviction team shipping real software that ends up in real yards, at real gates, moving real freight; if our system goes down, trucks stop moving and a customer's yard stops running. What we build is mission-critical to the people who depend on it.

We’re scaling fast, with load expected to 10X over the next 18 months, and reliability is now core to whether customers trust us to run their gates. We need an SRE to own uptime and incident response as the system grows, and to help the team get proactive about issues instead of reactive.

Project Details:
  • Own reliability targets across our backend/API, worker services, applications and CV pipeline; MTD, MTM, MTR, and follow-through on root causes.

  • Level up our monitoring and alerting, and build out auto-remediation, so on-call load scales with automation, not headcount.

  • Partner with our agentic engineering work to build agents that triage alerts and handle routine remediation.

  • Harden and optimize our GCP infrastructure (Cloud Run, Cloud SQL, GCS) for cost and performance as load scales.

  • Own database scale and performance; connection pooling, query optimization and indexing, read replicas, and capacity planning, so Postgres doesn't become the bottleneck as data volume grows.

  • Improve the reliability of our ML training and monitoring infrastructure, in partnership with the CV/ML team.

  • Run blameless postmortems and drive fixes for root causes, not just symptoms.

  • Participate in on-call rotation.

Qualifications:
  • 4+ years in an SRE, infrastructure, or backend engineering role with production on-call ownership.

  • Deep experience with a major cloud provider (GCP preferred); compute, managed databases, object storage, networking.

  • Experience building monitoring/alerting/observability stacks (Grafana, Prometheus, Zabbix, Datadog, or similar).

  • Strong scripting/automation skills (Python, Bash, or similar).

  • Comfortable with containerized workloads (Docker) and CI/CD pipelines.

  • Track record of reducing incident volume or improving reliability metrics — not just responding to incidents.

  • Strong communication skills, comfortable working with both technical and non-technical stakeholders, know when and how to elevate urgency, and build strong working relationships across teams.

  • Strong communication skills in English — you write clearly and engage well async.

Preferred Qualifications:
  • Experience with ML/data infrastructure — training pipelines, model monitoring, feature stores.

  • Experience building or integrating AI agents for operational automation (alert triage, auto-remediation).

  • Infrastructure-as-code experience (Terraform or similar).

  • PostgreSQL performance tuning at scale.

  • Background supporting physical/IoT systems (edge devices, cameras, on-site hardware).

  • Experience with bare-metal infrastructure in colocation environments, hardware monitoring, redundancy, and failover/high-availability configuration.

Our Stack:

TypeScript / Node.js · Next.js · React · Apollo Server · Express · PostgreSQL · GCP · Docker · Python (ML/data workloads). We’re pragmatic — the right tool matters more than the familiar one.

Location:

Remote — Latin America preferred

Working Arrangements:

This role is a full-time contract position. You’ll work closely with our core engineering team — embedded in our sprints, standups, and Slack channels — but employment is managed through the agency. We’ve built this model successfully with engineers in Latin America, and it’s been a great fit for both sides.

Outpost is an Equal Opportunity Employer and Prohibits Discrimination of Any Kind.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Site Reliability Engineer Engineer
Site Reliability Engineer Engineer

Modus Create • Aurora (IL)

Remote
USD 120,000 - 160,000
Software Engineer (Generalist)
Software Engineer (Generalist)

Outpost • United States

Remote
USD 120,000 - 160,000
Member of Technical Staff, DevOps
Member of Technical Staff, DevOps

Reactor • San Francisco (CA)

On-site
USD 100,000 - 160,000
Competitive salary and early equity
Visa sponsorship
Generous health, dental, and vision coverage
Remote SRE: Cloud Reliability for AI-Driven Logistics
Remote SRE: Cloud Reliability for AI-Driven Logistics

outpost • United States

Remote
USD 140,000 - 180,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Alembic Technologies • Dunwoody (GA)

On-site
USD 200,000 - 225,000
Site Reliability Engineer
Site Reliability Engineer

Seek Now • Atlanta (GA)

Hybrid
USD 120,000 - 180,000
Competitive salary
Health, dental, and vision coverage
401(k) with company match
+1
Senior DevOps/SRE Engineer
Senior DevOps/SRE Engineer

VITG • Ellicott City (MD)

On-site
USD 90,000 - 120,000
401(k) with employer contribution
Medical/Dental/Vision insurance
Paid vacation (PTO)
Senior Forward Deployed Engineer (DevOps/SRE)
Senior Forward Deployed Engineer (DevOps/SRE)

LeoForce • Pleasanton (CA)

On-site
USD 300,000 - 350,000
Medical benefits
401(k) plan
Free meals and snacks
+2
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Alembic • Dunwoody (GA)

On-site
USD 150,000 - 190,000
Site Reliability Engineer
Site Reliability Engineer

Harrison Clarke • New York (NY)

On-site
USD 120,000 - 160,000