Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.
Level AI is seeking a Senior SRE at the intersection of backend engineering, infrastructure operations, and FinOps. The role focuses on reducing Kubernetes overprovisioning, driving right-sizing, and maintaining cost telemetry to guide backend teams.
Responsibilities include building tooling, dashboards, and processes that empower backend groups to own cost and reliability budgets, and ensuring robust surface area coverage for cost-at-scale and reliability.
Level AI is on a mission to turn every customer interaction into a strategic advantage. Our AI-native platform helps enterprises transform contact centers from cost centers into engines of customer intelligence, operational efficiency, and business growth. By combining advanced AI with deep domain understanding of customer experience, Level AI empowers teams to unlock actionable insights, automate workflows, and deliver more consistent, higher-quality support across the customer journey.
Level AI is on a mission to turn every customer interaction into a strategic advantage. Our AI-native platform helps enterprises transform contact centers from cost centers into engines of customer intelligence, operational efficiency, and business growth. By combining advanced AI with deep domain understanding of customer experience, Level AI empowers teams to unlock actionable insights, automate workflows, and deliver more consistent, higher-quality support across the customer journey.
Headquartered in Mountain View, California, Level AI is a Series C company backed by leading investors including Battery Ventures and ENIAC. Our platform leverages Large Language Models and Custom Small Language Models (SLMs) to power AI Agents across the entire CX journey—customer-facing agents, agent‑assist, and backend automation—along with deep conversation analytics for QA, coaching, and insights.
The Senior SRE will be positioned at the intersection of backend engineering, infrastructure operations, and FinOps. The role is explicitly broader than a traditional DevOps engineer and explicitly more hands‑on than a pure architect.
Backend engineering depth: production experience in Python, Go/Rust, comfortable owning services end to end, able to read and reason about backend code across teams.
Kubernetes at scale: scheduler behavior, resource requests/limits, HPA/VPA, node pool design, cost‑aware autoscaling (Cast AI, Karpenter, or equivalent).
Cloud and on‑premise infrastructure: GCP fluency, IaC (Terraform), CI/CD, and comfort operating in hy brid setups including on‑prem GPU clusters.
GPU workload understanding: familiarity with throughput profiling, batching, KV‑cache behaviour, inference server tuning, and GPU utilisation metrics.
Observability and reliability: metrics, traces, logs, SLOs, and the discipline to instrument systems properly rather than reactively.
FinOps mindset: demonstrated history of converting infrastructure choices into measurable cost outcomes.
Security baseline: able to take on platform‑security workstreams without requiring constant handoff to the DevOps team.