Senior SRE: AI-Ready Energy Grid Reliability

GridCARE

Redwood City (CA)

Hybrid

USD 180,000 - 230,000

Full time

10 days ago
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

Competitive salary
Equity
Health insurance
Lunch provided
Hybrid schedule
Access to partners
Mission-driven team

Job summary

GridCARE seeks a Senior SRE to own reliability, scalability, and observability of production systems. You’ll collaborate with platform and data engineering to keep high-throughput, data-intensive services available for utilities and data center operators.

You will design infrastructure on AWS with Terraform and Kubernetes, build robust monitoring with Prometheus, Grafana, and Datadog, and automate pipelines and self-healing systems. Hybrid work is available in a fast-moving startup.

Qualifications

  • 5+ years in SRE, DevOps, or infrastructure engineering roles.
  • Deep experience with Kubernetes, Terraform/IaC, and cloud platforms (AWS Preferred).
  • Strong scripting/programming ability (Python, Bash).
  • Observability Experience (Prometheus, Grafana, Datadog).
  • Track record of running on-call for production systems and leading incident response.
  • Experience with CI/CD pipelines (Github Actions) and infrastructure automation.
  • Experience with Gitops concepts and tooling (ArgoCD/Flux).
  • Solid understanding of networking, distributed systems, and database reliability.
  • Comfortable operating in a fast-moving startup environment with ambiguity.

Responsibilities

  • Design and operate infrastructure on AWS using Terraform and Kubernetes.
  • Build monitoring, alerting, and observability (Prometheus, Grafana, Datadog, or similar) with meaningful SLOs/SLIs.
  • Automate away toil — deployment pipelines, capacity management, self-healing systems.
  • Partner with engineering on architecture reviews to catch reliability and scalability risks before they ship.
  • Manage database and data pipeline reliability for large-scale, real-time grid data processing.
  • Drive security and compliance best practices across infrastructure.

Skills

Kubernetes
Terraform/IaC
AWS
Python
Bash
Prometheus
Grafana
Datadog
GitOps (ArgoCD/Flux)
GitHub Actions
On-call / incident response
CI/CD pipelines
Networking & distributed systems

Tools

ArgoCD
Flux
GitHub Actions
Terraform

Job description

GridCARE seeks a Senior SRE to own reliability, scalability, and observability of production systems. You’ll collaborate with platform and data engineering to keep high-throughput, data-intensive services available for utilities and data center operators.

You will design infrastructure on AWS with Terraform and Kubernetes, build robust monitoring with Prometheus, Grafana, and Datadog, and automate pipelines and self-healing systems. Hybrid work is available in a fast-moving startup.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Site Reliability Engineer — AI-Scale Energy Infra
Senior Site Reliability Engineer — AI-Scale Energy Infra

Artha Nexgen • Irvine (CA)

Hybrid
USD 180,000 - 230,000
Salary bonus
Equity
HealthCoverage
+4
Senior SRE - AI-Driven Energy Infra (Hybrid)
Senior SRE - AI-Driven Energy Infra (Hybrid)

GridCARE, Inc. • Redwood City (CA), Northern (KY)

Hybrid
USD 180,000 - 230,000
Hybrid schedule
Lunch provided three days a week in?
Health, dental, and vision coverage
+1
Senior Site Reliability Engineer
Senior Site Reliability Engineer

GridCARE, Inc. • Redwood City (CA), Northern (KY)

Hybrid
USD 180,000 - 230,000
Hybrid schedule
Lunch provided three days a week in?
Health, dental, and vision coverage
+1
T Senior Site Reliability Engineer TP-Link Systems Irvine, California, US
T Senior Site Reliability Engineer TP-Link Systems Irvine, California, US

Artha Nexgen • Irvine (CA)

Hybrid
USD 180,000 - 230,000
Salary bonus
Equity
HealthCoverage
+4
Senior Site Reliability Engineer
Senior Site Reliability Engineer

GridCARE • Redwood City (CA)

Hybrid
USD 180,000 - 230,000
Competitive salary
Equity
Health insurance
+4
Senior SRE: AI-Driven Cloud Reliability
Senior SRE: AI-Driven Cloud Reliability

SupportFinity™ • San Francisco (CA)

Hybrid
USD 164,000 - 205,000
BetterUp coaching
Competitive pay
Medical, dental, and vision insurance
+7
Hybrid Senior Full-Stack Engineer: Grid & AI Platform
Hybrid Senior Full-Stack Engineer: Grid & AI Platform

GridCARE • Redwood City (CA)

Hybrid
USD 170,000 - 195,000
Hybrid schedule
Catered lunches
Equity
+1
Senior SRE: Automation, AI-Driven Reliability (Enterprise)
Senior SRE: Automation, AI-Driven Reliability (Enterprise)

ManpowerGroup Global, Inc. • Austin (TX)

On-site
USD 66,000 - 90,000
Senior SRE: AI-Driven Reliability & Cloud Automation
Senior SRE: AI-Driven Reliability & Cloud Automation

NDEAVOUR CONSULTING • United States

Hybrid
USD 120,000 - 150,000
Remote Office
Parking Space
Fun Office Space
+7
Senior SRE: Observability, Automation & Hybrid Cloud
Senior SRE: Observability, Automation & Hybrid Cloud

Colorado-Public-Employees • Denver (CO)

Hybrid
USD 140,000 - 165,000
Hybrid work option
On-call rotation
Work from home eligibility