Senior Site Reliability Engineer — AI-Scale Energy Infra

Artha Nexgen

Irvine (CA)

Hybrid

USD 180,000 - 230,000

Full time

2 days ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Benefits offered by this job

Salary bonus
Equity
HealthCoverage
Lunch provided
Hybrid schedule
Industry partners
Mission-driven team

Job summary

GridCARE is hiring a Senior SRE to own reliability, scalability, and observability of production systems. You will collaborate with platform and data engineering to keep high-throughput, real-time grid services running at the availability required by utilities and data centers.

Role emphasizes design and operation of AWS-based infrastructure using Terraform and Kubernetes, with strong emphasis on monitoring, incident response, and secure, scalable deployments in a fast-paced startup.

Qualifications

  • 5+ years in SRE, DevOps, or infrastructure engineering roles.
  • Deep experience with Kubernetes, Terraform/IaC, and cloud platforms (AWS Preferred).
  • Strong scripting/programming ability (Python, Bash).
  • Observability Experience (Prometheus, Grafana, Datadog).
  • Track record of running on-call for production systems and leading incident response.
  • Experience with CI/CD pipelines (Github Actions) and infrastructure automation.
  • Solid understanding of networking, distributed systems, and database reliability.
  • Comfortable operating in a fast-moving startup environment with ambiguity.

Responsibilities

  • Design and operate infrastructure on AWS using Terraform and Kubernetes.
  • Build monitoring, alerting, and observability with meaningful SLOs/SLIs.
  • Automate away toil — deployment pipelines, capacity management, self-healing systems.
  • Partner with engineering on architecture reviews to catch reliability and scalability risks before they ship.
  • Manage database and data pipeline reliability for large-scale, real-time grid data processing.
  • Drive security and compliance best practices across infrastructure.

Skills

Kubernetes
Terraform/IaC
Python
Bash
Observability
Incident response
CI/CD (Github Actions)
Networking
Distributed systems

Tools

AWS

Job description

GridCARE is hiring a Senior SRE to own reliability, scalability, and observability of production systems. You will collaborate with platform and data engineering to keep high-throughput, real-time grid services running at the availability required by utilities and data centers.

Role emphasizes design and operation of AWS-based infrastructure using Terraform and Kubernetes, with strong emphasis on monitoring, incident response, and secure, scalable deployments in a fast-paced startup.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

T Senior Site Reliability Engineer TP-Link Systems Irvine, California, US
T Senior Site Reliability Engineer TP-Link Systems Irvine, California, US

Artha Nexgen • Irvine (CA)

Hybrid
USD 180,000 - 230,000
Salary bonus
Equity
HealthCoverage
+4
Senior Site Reliability Engineer — AI Platform Scale
Senior Site Reliability Engineer — AI Platform Scale

Future Secure AI • Austin (TX)

On-site
USD 140,000 - 190,000
Senior Backend Engineer - Real-Time AI Cloud Infra
Senior Backend Engineer - Real-Time AI Cloud Infra

GridCARE, Inc. • Redwood City (CA), Northern (KY)

Hybrid
USD 184,000 - 284,000
Hybrid schedule
Competitive salary
Equity
+4
Senior Backend Platform Engineer – Real-Time Cloud-Native
Senior Backend Platform Engineer – Real-Time Cloud-Native

GridCARE • Redwood City (CA)

Hybrid
USD 180,000 - 240,000
Competitive compensation
Equity
Health coverage
+2
Senior AI Platform SRE: Scale Cloud Infra & Kubernetes
Senior AI Platform SRE: Scale Cloud Infra & Kubernetes

GCS Recruitment • Mount Laurel Township (NJ)

On-site
USD 110,000 - 170,000
Site Reliability Engineer
Site Reliability Engineer

Motion Recruitment Partners LLC • Chicago (IL), Northern (KY)

Hybrid
USD 140,000 - 170,000
Senior SRE: AI-Driven Cloud Reliability
Senior SRE: AI-Driven Cloud Reliability

SupportFinity™ • San Francisco (CA)

Hybrid
USD 164,000 - 205,000
BetterUp coaching
Competitive pay
Medical, dental, and vision insurance
+7
Senior/Staff Backend Software Engineer
Senior/Staff Backend Software Engineer

GridCARE • Redwood City (CA)

On-site
USD 184,000 - 284,000
Competitive salary
Performance bonus
Equity
+5
Senior SRE – AI Infrastructure Reliability Leader
Senior SRE – AI Infrastructure Reliability Leader

Nscale • San Francisco (CA), Seattle (WA), Houston (TX)

On-site
USD 170,000 - 265,000
Equity
Ownership from start
Flexible schedule
Senior SRE - AI-Powered Cloud Reliability (Remote)
Senior SRE - AI-Powered Cloud Reliability (Remote)

ServiceTitan, Inc. • Northern (KY)

Hybrid
USD 148,000 - 221,000