Site Reliability Engineer

Specter

San Francisco (CA)

On-site

USD 120,000 - 160,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Specter is seeking a Site Reliability Engineer to manage the operational health of our sensor platform, ensuring the reliability of both edge hardware deployed at customer sites and cloud infrastructure.

This role requires strong Linux systems administration skills, as well as experience with networking and cloud services. You will debug issues, build fleet management systems, and design observability strategies. Join us in automating the physical world with cutting-edge AI technology.

Qualifications

  • Strong Linux systems administration, comfortable working over SSH in production.
  • Experience with edge or on-prem hardware alongside cloud infrastructure.
  • Solid networking fundamentals including DNS, firewalls, and VPNs.

Responsibilities

  • Debug and triage issues across diverse Linux-based sensor nodes.
  • Build and maintain fleet management systems for updates and diagnostics.
  • Design and implement observability across edge devices and cloud infrastructure.

Skills

Linux systems administration
Scripting in Python, Go, or Bash
Networking fundamentals
Experience with containerization
Embedded systems experience
AWS experience

Tools

Docker
Kubernetes

Job description

Company Background

Specter's mission is to help automate the physical world.

Today, we build video sensors with state‑of‑the‑art AI agents that answer any question, anywhere in their environments. Our systems can automatically detect and reason about any physical activity captured on camera, from security incidents (e.g. perimeter intrusion, theft, LPR), to safety monitoring (e.g. PPE detection, injured people), to operational efficiency (e.g. material tracking, congestion monitoring). We offer both long‑range wireless (1km range) and wired sensor variants to suit any deployment.

Our co‑founders Xerxes and Philip are passionate about empowering our partners in the fast approaching world of physical AI and robotics. We are a small, fast growing team who hail from Anduril, Tesla, Uber, and the U.S. Special Forces.

The Role

We're hiring a Site Reliability Engineer to own the operational health of our connected sensor platform — spanning a live fleet of edge hardware deployed at customer sites and the cloud infrastructure behind it.

This is a high‑ownership role at the intersection of ops and platform engineering. You'll drive reliability across our sensor fleet — triaging issues in the field, building the systems that prevent them from recurring, and owning the observability that keeps us ahead of problems as we scale.

You set your own priorities across all three:

Responsibilities
Reactive — Triage & Recovery
  • Debug and triage issues across a live fleet of diverse Linux‑based sensor nodes and edge appliances deployed at customer sites.
  • SSH into field hardware to diagnose, patch, and recover systems — often with limited remote access and incomplete information.
  • Own site bring‑ups end to end; be the person who gets things back online.
Systems Builder — Close the Loop
  • Build and maintain fleet management systems: OTA update pipelines, device health tracking, remote diagnostics, and lifecycle tooling.
  • Identify repeat fires and eliminate them — build tooling, pre‑deployment checks, and root cause processes that prevent recurrence.
  • Automate toil relentlessly: if you're doing something twice, you should be scripting it.
  • Collaborate with embedded systems, and platform teams to define reliability and deployment requirements.
Observability Owner — Fleet Visibility
  • Design and implement observability (logging, metrics, alerting) across edge devices and cloud infrastructure (AWS).
  • Surface and close telemetry gaps; build fleet‑wide visibility that enables data‑driven reliability decisions.
  • Develop runbooks, incident response procedures, and participate in on‑call rotations.
Qualifications
  • Strong Linux systems administration — comfortable working over SSH in production, not just dev environments.
  • Experience with edge or on‑prem hardware alongside cloud infrastructure.
  • Solid networking fundamentals: DNS, firewalls, VPNs, subnets, secure remote access.
  • Scripting or programming in Python, Go, or Bash for operational tooling.
  • Familiarity with containerization (Docker, Kubernetes a plus).
  • Embedded systems experience — reading firmware logs, understanding hardware‑software boundaries, and reasoning about what's happening below the OS is a meaningful edge in this role.
  • Deeper cloud experience (AWS infrastructure, IAM, networking, observability tooling) is a strong plus for owning the cloud side of the fleet.
  • Rust or C experience — we have firmware in both; being able to read and reason about low‑level code accelerates triage significantly.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Site Reliability Engineer
Senior Site Reliability Engineer

SDI International • Chicago (IL)

Hybrid
USD 130,000 - 180,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Veloc Inc • Coppell (TX)

On-site
USD 140,000 - 190,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Jobtailor • Arlington (VA)

On-site
USD 140,000 - 200,000
Edge & Cloud SRE: Fleet Reliability & Observability
Edge & Cloud SRE: Fleet Reliability & Observability

Specter • San Francisco (CA)

On-site
USD 120,000 - 160,000
Product Reliability Engineer
Product Reliability Engineer

Scope • Town of Texas (WI)

On-site
USD 80,000 - 110,000
Site Reliability Engineer
Site Reliability Engineer

TalentDome Staffing • United States

On-site
USD 140,000 - 210,000
Staff Site Reliability Engineer
Staff Site Reliability Engineer

Careers • San Francisco (CA)

Hybrid
USD 180,000 - 240,000
Site Reliability Engineer
Site Reliability Engineer

Jobtailor • New Hampshire

On-site
USD 110,000 - 160,000
Site Reliability Engineer
Site Reliability Engineer

asobbi • California (MO)

On-site
USD 170,000 - 220,000
Fully remote (US timezone)
Senior Site Reliability Engineer
Senior Site Reliability Engineer

synthesia • United States

On-site
USD 110,000 - 150,000