T Senior Site Reliability Engineer TP-Link Systems Irvine, California, US

Artha Nexgen

Irvine (CA)

Hybrid

USD 180,000 - 230,000

Full time

2 days ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Benefits offered by this job

Salary bonus
Equity
HealthCoverage
Lunch provided
Hybrid schedule
Industry partners
Mission-driven team

Job summary

GridCARE is hiring a Senior SRE to own reliability, scalability, and observability of production systems. You will collaborate with platform and data engineering to keep high-throughput, real-time grid services running at the availability required by utilities and data centers.

Role emphasizes design and operation of AWS-based infrastructure using Terraform and Kubernetes, with strong emphasis on monitoring, incident response, and secure, scalable deployments in a fast-paced startup.

Qualifications

  • 5+ years in SRE, DevOps, or infrastructure engineering roles.
  • Deep experience with Kubernetes, Terraform/IaC, and cloud platforms (AWS Preferred).
  • Strong scripting/programming ability (Python, Bash).
  • Observability Experience (Prometheus, Grafana, Datadog).
  • Track record of running on-call for production systems and leading incident response.
  • Experience with CI/CD pipelines (Github Actions) and infrastructure automation.
  • Solid understanding of networking, distributed systems, and database reliability.
  • Comfortable operating in a fast-moving startup environment with ambiguity.

Responsibilities

  • Design and operate infrastructure on AWS using Terraform and Kubernetes.
  • Build monitoring, alerting, and observability with meaningful SLOs/SLIs.
  • Automate away toil — deployment pipelines, capacity management, self-healing systems.
  • Partner with engineering on architecture reviews to catch reliability and scalability risks before they ship.
  • Manage database and data pipeline reliability for large-scale, real-time grid data processing.
  • Drive security and compliance best practices across infrastructure.

Skills

Kubernetes
Terraform/IaC
Python
Bash
Observability
Incident response
CI/CD (Github Actions)
Networking
Distributed systems

Tools

AWS

Job description

GridCARE is a leading venture-backed startup solving the most critical constraint in AI’s growth trajectory: immediate access to power. As demand for computing skyrockets, access to energy has become the defining bottleneck in the AI infrastructure race. While leading tech companies invest billions in speculative, long-term solutions that may take decades to arrive, GridCARE’s pioneering physics-based generative AI platform unlocks gigawatts of hidden capacity in today’s electric grid — enabling hyperscalers, data center developers, and utilities to power AI infrastructure years sooner than conventional approaches and without costly upgrades.

Founded at Stanford’s Doerr School of Sustainability and backed by leading climate-tech and deep-tech investors, GridCARE has assembled a world-class team spanning power systems, AI, and infrastructure.

At GridCARE, you will:

Work at the intersection of AI, energy, and infrastructure — the foundation of the next industrial revolution.

Partner with hyperscalers, developers, and utilities on high-impact, real-world deployments.

Help shape a more abundant, efficient, and resilient energy future for the digital era.

Join a company defining a new category — capacity acceleration for AI.

Receive competitive compensation, equity, and benefits in a fast-growth, mission-driven environment.

Learn more about GridCARE:

TechCrunch: GridCARE thinks more than 100 GW of data-center capacity is hiding in the grid

Utility Dive: Portland General Electric invests in AI-powered flexibility to speed data-center connection

Data Center Dynamics: From Years to Months — Creating an AI Fast Lane for Data Centers

GridCARE Raises $64 Million Series A to Create a New Category: Power Acceleration

Job Description

We're looking for a Senior SRE to own the reliability, scalability, and observability of our production systems. You'll work closely with platform and data engineering to keep high-throughput, data-intensive services running at the availability our customers (utilities, data center operators) require.

Responsibilities
  • Design and operate infrastructure on AWS using Terraform and Kubernetes
  • Build monitoring, alerting, and observability (Prometheus, Grafana, Datadog, or similar) with meaningful SLOs/SLIs
  • Automate away toil — deployment pipelines, capacity management, self-healing systems
  • Partner with engineering on architecture reviews to catch reliability and scalability risks before they ship
  • Manage database and data pipeline reliability for large-scale, real-time grid data processing
  • Drive security and compliance best practices across infrastructure
Qualifications
Required
  • 5+ years in SRE, DevOps, or infrastructure engineering roles
  • Deep experience with Kubernetes, Terraform/IaC, and cloud platforms (AWS Preferred)
  • Strong scripting/programming ability (Python, Bash)
  • Observability Experience (Prometheus, Grafana, Datadog)
  • Track record of running on-call for production systems and leading incident response
  • Experience with CI/CD pipelines (Github Actions) and infrastructure automation
  • Solid understanding of networking, distributed systems, and database reliability
  • Comfortable operating in a fast-moving startup environment with ambiguity
Preferred
  • Experience with data-intensive or real-time processing systems
  • Background in energy, climate tech, or critical infrastructure
  • Experience scaling infrastructure through hypergrowth
What We Offer
  • Competitive salary, performance bonus, and equity.
  • Comprehensive health, dental, and vision coverage.
  • Lunch provided three days a week in office.
  • Hybrid schedule for local employees: 3 days in office for collaboration, 2 days remote for focused work.
  • Access to leading academic, industry, and government partners in the AI-energy ecosystem.
  • A mission-driven team focused on shaping the future of the energy transition.
Salary Range

$180,000-$230,000 Total

Join us in tackling one of the most important infrastructure challenges of our time — enabling the energy foundation for the age of AI.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior/Staff Backend Software Engineer
Senior/Staff Backend Software Engineer

GridCARE • Redwood City (CA)

On-site
USD 184,000 - 284,000
Competitive salary
Performance bonus
Equity
+5
Senior/Staff Backend Software Engineer
Senior/Staff Backend Software Engineer

GridCARE, Inc. • Redwood City (CA), Northern (KY)

Hybrid
USD 184,000 - 284,000
Hybrid schedule
Competitive salary
Equity
+4
Senior Software Engineer
Senior Software Engineer

GridCARE • Redwood City (CA)

Hybrid
USD 180,000 - 240,000
Competitive compensation
Equity
Health coverage
+2
Director, Utility Partnerships
Director, Utility Partnerships

GridCARE • Redwood City (CA)

Hybrid
USD 185,000 - 285,000
Lunch provided
Hybrid work schedule
Health coverage
+1
Senior Product Manager
Senior Product Manager

GridCARE • Redwood City (CA)

Hybrid
USD 150,000 - 210,000
Competitive compensation
Equity
Health, dental, and vision coverage
+4
Modeling Engineer
Modeling Engineer

GridCARE • Redwood City (CA)

Hybrid
USD 120,000 - 190,000
Hybrid work in Redwood City
Remote flexibility
Competitive salary
+2
Optimization Engineer
Optimization Engineer

GridCARE • Redwood City (CA)

Hybrid
USD 110,000 - 165,000
Competitive salary
Equity
Health + dental + vision
+3
Optimization Engineer, Grid Systems
Optimization Engineer, Grid Systems

GridCARE • Redwood City (CA)

Hybrid
USD 140,000 - 220,000
Competitive salary
Performance bonus
Equity
+3
Senior Site Reliability Engineer — AI-Scale Energy Infra
Senior Site Reliability Engineer — AI-Scale Energy Infra

Artha Nexgen • Irvine (CA)

Hybrid
USD 180,000 - 230,000
Salary bonus
Equity
HealthCoverage
+4
Physical Systems Modeling Engineer
Physical Systems Modeling Engineer

GridCARE, Inc. • Redwood City (CA)

Hybrid
USD 140,000 - 250,000
Hybrid in Redwood City with remoteflex
Competitive salary and equity