Senior DevOps Engineer

Gridware

San Francisco (CA)

On-site

USD 190,000 - 215,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Health, Dental & Vision benefits
Paid parental leave
Off the Grid - 2 weeks paid break

Job summary

Gridware in San Francisco is seeking a seasoned DevOps/Platform Engineer to own AWS infrastructure, Kubernetes (EKS), CI/CD, and observability end-to-end. You will work with Cloud Security, backend, firmware, and data teams to scale production systems and build reliable, secure, and scalable deployments.

As an early member of the team, you will design and operate pipelines, implement IaC with Terraform/Terragrunt, manage IAM and security controls, and optimize costs across environments.

Qualifications

  • 5+ years in DevOps, SRE, or Platform Engineering with production AWS experience.
  • Hands-on Kubernetes administration (EKS or equivalent) and GitOps (Argo CD/ Flux).
  • Proficiency with Terraform; Terragrunt or similar wrapper.
  • Experience designing and maintaining CI/CD pipelines, preferably GitHub Actions.
  • Production experience with Kafka (MSK).
  • Strong networking, DNS, TLS, and security fundamentals, IdP-driven access control.
  • Experience with monitoring/logging stacks (Grafana, Prometheus, Loki, Mimir).
  • Scripting in Python or Bash for automation.

Responsibilities

  • Design, build, and maintain scalable, secure, and highly available infrastructure on AWS (EKS, EC2, RDS, MSK, S3, IAM).
  • Manage and optimize Kubernetes clusters (EKS) across environments with Argo CD GitOps.
  • Implement and maintain CI/CD pipelines using GitHub Actions and reusable workflows.
  • Operate and tune Kafka-based event streams (MSK) for high-throughput data pipelines.
  • Define and manage IaC with Terraform/Terragrunt with environment separation.
  • Manage IAM across platforms with Auth0/EntraID and service accounts.
  • Build observability with Grafana, Loki, Prometheus, Mimir for on-call fault-finding.
  • Cost optimization across environments in partnership with engineering teams.

Skills

DevOps experience
Kubernetes (EKS)
GitOps
CI/CD
Kafka
Networking & security
Linux scripting
Auth0/EntraID
Monitoring & observability

Tools

Terraform
Terragrunt
Argo CD
GitHub Actions
AWS (EKS, MSK, RDS)
Grafana / Prometheus / Loki / Mimir

Job description

About Gridware

Gridware is a San Francisco-based technology company dedicated to protecting and enhancing the electrical grid. We pioneered a groundbreaking new class of grid management called active grid response (AGR), focused on monitoring the electrical, physical, and environmental aspects of the grid that affect reliability and safety. Gridware’s advanced Active Grid Response platform uses high-precision sensors to detect potential issues early, enabling proactive maintenance and fault mitigation. This comprehensive approach helps improve safety, reduce outages, and ensure the grid operates efficiently. The company is backed by climate-tech and Silicon Valley investors. For more information, please visit www.Gridware.io.

Role Description

We’re scaling the deployment of critical infrastructure monitoring devices to detect real-world fault events that lead to wildfires. The platform you’ll build and operate ingests millions of events per day from devices in the field, powers customer-facing dashboards and alerting, and supports the data science work that turns raw signals into grid intelligence.

You will own AWS infrastructure, Kubernetes (EKS), CI/CD, and observability end-to-end, partnering with our Cloud Security team to keep the platform safe and compliant, and with backend, firmware, and data teams to keep them shipping fast. As an early member of the DevOps team, you’ll have a direct hand in shaping how Gridware builds, deploys, and runs production systems for years to come.

Responsibilities
  • Design, build, and maintain scalable, secure, and highly available infrastructure on AWS (EKS, EC2, RDS / Aurora Postgres, MSK, S3, VPC, IAM).
  • Manage and optimize Kubernetes clusters (EKS) across multiple environments, and deploy applications using Argo CD with GitOps best practices.
  • Implement and maintain CI/CD pipelines using GitHub Actions, including reusable workflows, build/push/scan flows for ECR, and frontend deployment pipelines.
  • Operate and tune Kafka-based event streaming on Amazon MSK for high-throughput, low-latency device data pipelines.
  • Define and manage Infrastructure as Code with Terraform and Terragrunt, with reusable modules, sensible environment separation, and review-friendly plans.
  • Manage identity and access across platforms with Auth0 / EntraID integrations, IAM roles for service accounts (IRSA), and short-lived credentials.
  • Build and maintain observability with Grafana, Loki, Prometheus / Mimir, and related tooling so on-call engineers can quickly find and fix issues.
  • Monitor and optimize infrastructure cost across environments, partnering with engineering teams on right-sizing, capacity planning, and waste reduction.
  • Partner with our Cloud Security team to enforce security standards, integrate with SIEM tooling, and respond to vulnerabilities and incidents.
  • Debug complex production issues across infrastructure, deployment, and networking layers, and turn the lessons learned into automation and runbooks.
Required Skills
  • 5+ years in DevOps, SRE, or Platform Engineering with production experience operating AWS infrastructure.
  • Deep hands-on experience administering Kubernetes (EKS or equivalent) and deploying via GitOps (Argo CD or Flux).
  • Proficiency with Infrastructure as Code using Terraform; comfort with Terragrunt or a similar wrapper.
  • Hands-on experience designing and maintaining CI/CD pipelines, preferably with GitHub Actions and reusable workflows.
  • Production experience operating distributed systems such as Kafka (MSK).
  • Strong understanding of networking, DNS, TLS, and security best practices, including IdP-driven access control (Auth0, EntraID, or similar).
  • Solid experience with monitoring and logging stacks such as Grafana, Loki, Prometheus, Mimir, or equivalents.
  • Ability to debug complex production issues across infrastructure, deployment, and networking layers.
  • Comfortable working in Linux environments with strong scripting skills (Python or Bash preferred for automation).
  • Knowledge of version control workflows, automated testing, and release management.
Bonus Skills
  • Experience operating Apollo Router / GraphQL federation gateways in production.
  • Experience operating Argo Workflows or similar Kubernetes-native job / pipeline runners in production.
  • Familiarity with Databricks or ML Ops pipelines for data and model deployment.
  • Experience designing, operating, and exercising Disaster Recovery (DR) environments, including cross-region replication, backups, and tested failover runbooks.
  • Experience with Tailscale or other zero-trust networking tools.
  • Experience supporting IoT / embedded fleets at scale, including secure device-to-cloud connectivity.
  • Experience in high-growth startup environments where you must wear many hats.

$190,000 - $215,000 a year

This describes the ideal candidate; many of us have picked up this expertise along the way. Even if you meet only part of this list, we encourage you to apply!

Benefits

Health, Dental & Vision (Gold and Platinum with some providers plans fully covered)

Paid parental leave

Alternating day off (every other Monday)

“Off the Grid”, a two week per year paid break for all employees.

Commuter allowance

Company-paid training

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Cloud Engineer
Senior Cloud Engineer

Gridware Technologies Inc. • San Francisco (CA)

On-site
USD 190,000 - 215,000
Health, Dental & Vision
Paid parental leave
Alternating day off
+3
Senior Platform Engineer
Senior Platform Engineer

Gridware Technologies Inc. • San Francisco (CA)

On-site
USD 190,000 - 210,000
Health, Dental & Vision
Paid parental leave
Alternating day off
+3
Senior Software Engineer, Grid Communications & Platform
Senior Software Engineer, Grid Communications & Platform

Gridware • San Francisco (CA)

On-site
USD 190,000 - 2,150,000
Health, Dental & Vision
Parental leave
Off the Grid breaks
+2
Senior Cloud Engineer
Senior Cloud Engineer

Gridware • San Francisco (CA)

On-site
USD 120,000 - 160,000
Health, Dental & Vision Plans
Paid parental leave
Alternating day off
+3
Senior Research Engineer, Electrical
Senior Research Engineer, Electrical

Gridware Technologies Inc. • San Francisco (CA)

On-site
USD 185,000 - 200,000
Health, Dental & Vision coverage
Paid parental leave
Alternating day off every other Monday
+3
Senior Fleet Intelligence Analyst
Senior Fleet Intelligence Analyst

Gridware • San Francisco (CA)

On-site
USD 130,000 - 150,000
Health, Dental & Vision
Paid parental leave
Alternating day off
+3
Senior ML Engineer, Multi-Sensor Modeling
Senior ML Engineer, Multi-Sensor Modeling

Gridware • San Francisco (CA)

On-site
USD 190,000 - 205,000
Health, Dental & Vision insurance
Paid parental leave
Alternating day off
+3
Research Engineer, Mechanical
Research Engineer, Mechanical

Gridware • San Francisco (CA)

Hybrid
USD 140,000 - 165,000
Health, Dental & Vision
Paid parental leave
Alternating day off
+3
Senior Software Engineer, Grid Communications
Senior Software Engineer, Grid Communications

Gridware • San Francisco (CA)

On-site
USD 140,000 - 190,000
Health, Dental & Vision (Gold and-Plat
Paid parental leave
Alternating day off
+3
Senior Director, Full Stack Product
Senior Director, Full Stack Product

Gridware • San Francisco (CA), Northern (KY)

Hybrid
USD 260,000 - 290,000
Health, Dental & Vision
Parental leave
Alternating day off
+3