Site Reliability Engineer: AI-Driven Kubernetes Automation

Kindredventures

United States

Remote

USD 8,504 - 14,882

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

fal is seeking a senior platform engineer to own and operate Kubernetes infrastructure, including cluster lifecycle, upgrades, networking, and multi-tenant isolation for customer workloads. You will build and maintain CI/CD pipelines, leverage AI to automate production analysis, and drive reliability improvements with automation, SLOs, and incident processes, collaborating across teams in Turkey.

Strong focus on observability with Prometheus, Grafana, Loki, and Datadog, and experience with

Qualifications

  • 5+ years of experience managing critical production systems and workflows.
  • Experience operating Kubernetes at scale with IaC (Terraform, Ansible).
  • Deep knowledge of Linux networking, container networking, DNS.
  • Experience building CI/CD and GitOps (FluxCD, ArgoCD).
  • Proficiency in Python and Go or Bash for tooling.
  • Strong logging, monitoring, and alerting (Prometheus, Grafana, Loki, Datadog).

Responsibilities

  • Own and operate Kubernetes infrastructure including cluster lifecycle and multi-tenant isolation.
  • Build and maintain CI/CD pipelines and deployment infrastructure.
  • Leverage AI to automate analysis and resolution of production issues.
  • Build dashboards, alerting, and anomaly detection across systems.
  • Define and enforce SLOs and incident response processes.
  • Manage networking, load balancing, and service mesh configurations.
  • Drive reliability improvements through automation and chaos engineering.

Skills

Production systems
Kubernetes at scale
Python
Go
Bash
CI/CD
GitOps
Logging & Monitoring
Linux networking
Communication

Tools

Terraform
Ansible
FluxCD
ArgoCD
Prometheus
Grafana
Loki
Thanos
Datadog

Job description

fal is seeking a senior platform engineer to own and operate Kubernetes infrastructure, including cluster lifecycle, upgrades, networking, and multi-tenant isolation for customer workloads. You will build and maintain CI/CD pipelines, leverage AI to automate production analysis, and drive reliability improvements with automation, SLOs, and incident processes, collaborating across teams in Turkey.

Strong focus on observability with Prometheus, Grafana, Loki, and Datadog, and experience with

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Site Reliability Engineer — Kubernetes & AI-Driven Ops
Senior Site Reliability Engineer — Kubernetes & AI-Driven Ops

Fal • San Francisco (CA)

On-site
USD 180,000 - 250,000
Relocation assistance to San Francisco
Visa sponsorship
Competitive salary and equity
+1
Senior Site Reliability Engineer — AI Platform Scale
Senior Site Reliability Engineer — AI Platform Scale

Future Secure AI • Austin (TX)

On-site
USD 140,000 - 190,000
Senior Platform Reliability Engineer (Kubernetes & CI/CD)
Senior Platform Reliability Engineer (Kubernetes & CI/CD)

Optomi • Orlando (FL)

Hybrid
USD 120,000 - 180,000
Platform Reliability Engineer - AI, Kubernetes & GitOps
Platform Reliability Engineer - AI, Kubernetes & GitOps

Nasdaq • Philadelphia

Hybrid
USD 140,000 - 200,000
401(k) with company match
Employee stock purchase program (ESPP)
Health, dental, vision coverage
+2
Senior SRE: AI Cloud Platform & Kubernetes Expert
Senior SRE: AI Cloud Platform & Kubernetes Expert

Lambda • Bellevue (WA)

On-site
USD 180,000 - 260,000
Health insurance
Dental insurance
Vision insurance
+3
Senior Site Reliability Engineer - Kubernetes & AI Platform
Senior Site Reliability Engineer - Kubernetes & AI Platform

Optomi • Seattle (WA)

On-site
USD 140,000 - 180,000
Senior Kubernetes SRE — Scale AI Clusters & Automation
Senior Kubernetes SRE — Scale AI Clusters & Automation

Neura Market • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 240,000
Equity compensation
Health, dental and vision coverage
Wellness and commuter stipends
+2
Site Reliability Engineer
Site Reliability Engineer

Flanksource Inc. • Mission (KS)

Remote
USD 90,000 - 130,000
100% remote work
Flexible hours
Opportunity to work with cutting-edge technology
Senior SRE (Contract/Hybrid)
Senior SRE (Contract/Hybrid)

Optomi • Orlando (FL)

Hybrid
USD 120,000 - 180,000
Senior Kubernetes SRE - Platform Automation & Istio Expert
Senior Kubernetes SRE - Platform Automation & Istio Expert

Talanto • Bellevue (WA), Northern (KY)

Hybrid
USD 174,000 - 267,000