SRE: Cloud & Edge Reliability Lead

Picogrid, Inc.

El Segundo (CA)

On-site

USD 170,000 - 195,000

Full time

14 days+
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Benefits offered by this job

Stock options
401(k) with employer matching
Full health coverage (medical, dental,
Vision insurance
Relocation assistance
Unlimited PTO
11 paid holidays per year
Paid parental leave
Lunch provided in-office
Free EV charging at HQ
Unique office in El Segundo

Job summary

Picogrid, Inc. is looking for its first Site Reliability Engineer to own production reliability across cloud and edge, including observability, incident response, and edge device management. You will help build practices to ensure mission-critical systems remain available in challenging environments.

You will work with DevSecOps to stand up new infrastructure and support scrappy, fast-paced development with a strong on-call culture and robust guardrails.

Qualifications

  • 3+ years of experience as an SRE or related role.
  • Deep Kubernetes operations experience including node lifecycle and live cluster debugging.
  • Experience designing comprehensive observability dashboards and high signal-to-noise ratio alerting rules.
  • Competent incident responder with evidence-first triage and blameless postmortems.
  • Production Terraform or OpenTofu experience.
  • Fluent in AWS including IAM, networking, multi-account environments, and hardening.
  • Experience managing high availability database deployments.
  • IoT or edge fleet operation experience.
  • Comfortable operating in scrappy, fast-paced environments and turning ambiguous requirements into concrete solutions.

Responsibilities

  • Own, define and drive reliability SLIs and SLOs for cloud deployments.
  • Own, define and drive reliability SLIs and SLOs for edge devices in remote/contested areas.
  • Own the observability stack and dashboards, with versioned configs and guarded alert rules.
  • Participate in on-call and incident response with blameless postmortems.
  • Encode reliability into infrastructure as code.

Skills

Kubernetes operations
Observability design
Incident response
Edge/IoT fleet operations
Cloud security best practices

Tools

Grafana
Prometheus
Loki
OpenTelemetry
Terraform
OpenTofu
AWS IAM & networking

Job description

Picogrid, Inc. is looking for its first Site Reliability Engineer to own production reliability across cloud and edge, including observability, incident response, and edge device management. You will help build practices to ensure mission-critical systems remain available in challenging environments.

You will work with DevSecOps to stand up new infrastructure and support scrappy, fast-paced development with a strong on-call culture and robust guardrails.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

SRE: Cloud & Edge Reliability for Defense Tech
SRE: Cloud & Edge Reliability for Defense Tech

Picogrid • El Segundo (CA)

On-site
USD 170,000 - 195,000
Stock options
401(k) matching
Health coverage
+4
Site Reliability Engineer
Site Reliability Engineer

Picogrid • El Segundo (CA)

On-site
USD 170,000 - 195,000
Stock options
401(k) matching
Health coverage
+4
Remote Senior Network Reliability Engineer (SRE)
Remote Senior Network Reliability Engineer (SRE)

Gainbridge • Zionsville (IN), Northern (KY)

On-site
USD 135,000 - 190,000
Health Insurance
Dental Insurance
Vision Insurance
+4
Site Reliability Engineer
Site Reliability Engineer

Picogrid, Inc. • El Segundo (CA)

On-site
USD 170,000 - 195,000
Stock options
401(k) with employer matching
Full health coverage (medical, dental,
+8
Senior Production SRE: Cloud & On-Prem Reliability
Senior Production SRE: Cloud & On-Prem Reliability

Weights & Biases • New York (NY)

On-site
USD 140,000 - 180,000
Medical Insurance
Dental Insurance
Vision Insurance
+15
Edge & Cloud SRE: Fleet Reliability & Observability
Edge & Cloud SRE: Fleet Reliability & Observability

Specter • San Francisco (CA)

On-site
USD 120,000 - 160,000
SRE Engineering Manager — Lead Reliability, Remote Flexible
SRE Engineering Manager — Lead Reliability, Remote Flexible

DevOpsChat • California (MO)

Hybrid
USD 140,000 - 220,000
Health insurance
Professional development opportunities
Senior SRE & Cloud Reliability Architect
Senior SRE & Cloud Reliability Architect

United States Digital Space LLC • United States

Remote
USD 180,000 - 250,000
SRE Lead: Reliability & Cloud Observability Architect
SRE Lead: Reliability & Cloud Observability Architect

BlackCube Labs • San Diego (CA)

On-site
USD 190,000 - 280,000
Senior Site Reliability Engineer — Cloud Observability
Senior Site Reliability Engineer — Cloud Observability

Guidehouse • San Antonio (TX)

On-site
USD 106,000 - 176,000
Medical, Rx, Dental & Vision Insurance
401(k) Retirement Plan
Tuition Reimbursement
+2