Senior Site Reliability Engineer (Noida, BLR, India)

Level AI

Mountain View (CA)

On-site

USD 180,000 - 240,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Level AI, headquartered in Mountain View, California, seeks a Senior SRE to join the infrastructure and FinOps focus. The role bridges backend engineering, operations, and cost optimization, delivering tooling and dashboards that empower backend teams to own costs and reliability while advancing GPU performance and security readiness.

Ideal candidates bring 4–5 years of hands-on systems experience, deep Kubernetes expertise at scale, and strong cloud/on-prem infrastructure skills.

Qualifications

  • 4–5 years of hands-on systems experience.
  • Backend engineering depth with Python, Go/Rust, and end-to-end ownership.
  • Kubernetes at scale: scheduler behavior, resources, HPA/VPA, cost-aware autoscaling.
  • Cloud and on-prem infrastructure: GCP fluency, IaC (Terraform), CI/CD, and hybrid setups with on-prem GPU clusters.
  • GPU workload understanding: throughput profiling, batching, GPU utilization metrics.
  • Observability and reliability: metrics, traces, logs, SLOs, proper instrumentation.
  • FinOps mindset: history of turning infra choices into cost savings.
  • Security baseline: ability to lead platform-security workstreams.

Responsibilities

  • Infrastructure cost efficiency and FinOps. Own reduction of Kubernetes overprovisioning and cost telemetry for backend teams.
  • GPU throughput optimization. Run structured experiments on on-prem GPU clusters in partnership with AI service owners.
  • Backend enablement, not ownership absorption. Build tooling, dashboards, and processes for teams to own cost and reliability budgets.
  • Reliability instrumentation. Ensure surface area is captured for cost-at-scale and reliability.
  • Selective security workstreams. Take on security-adjacent platform changes without overloading DevOps.

Skills

Python
Go/Rust
Kubernetes
Cloud/GCP
Terraform
CI/CD
GPU workloads
Observability
FinOps
Security baseline

Tools

Cast AI
Karpenter
Terraform
CI/CD
GCP
GPU clusters

Job description

About Level AI

Level AI is on a mission to turn every customer interaction into a strategic advantage. Our AI-native platform helps enterprises transform contact centers from cost centers into engines of customer intelligence, operational efficiency, and business growth. By combining advanced AI with deep domain understanding of customer experience, Level AI empowers teams to unlock actionable insights, automate workflows, and deliver more consistent, higher-quality support across the customer journey.

Headquartered in Mountain View, California, Level AI is a Series C company backed by leading investors including Battery Ventures and ENIAC. Our platform leverages Large Language Models and Custom Small Language Models (SLMs) to power AI Agents across the entire CX journey—customer-facing agents, agent-assist, and backend automation—along with deep conversation analytics for QA, coaching, and insights.

About the role

The Senior SRE will be positioned at the intersection of backend engineering, infrastructure operations, and FinOps. The role is explicitly broader than a traditional DevOps engineer and explicitly more hands-on than a pure architect.

What you'll be liable for:
  • Infrastructure cost efficiency and FinOps. Own the continued reduction of Kubernetes overprovisioning, drive right-sizing programs, and maintain the cost telemetry that backend teams use to make decisions.

    GPU throughput optimization.Run a structured experimentation program on on-premise GPU clusters, partnering with AI service owners. Lead by the Engineering leadership, with this role providing the experimental bandwidth.

    Backend enablement, not ownership absorption.Build the tooling, dashboards, and processes that let backend teams from other groups own their own cost and reliability budgets. The deliverable is leverage, not headcount-shaped work.

    Reliability instrumentation.As the infra team owns most of the instrumentation across new and offline flows, this role takes a central seat in making sure that surface area is captured properly for both cost-at-scale and reliability.

    Selective security workstreams.Take on a defined slice of the active security work so that senior DevOps engineers are not the single point of execution for security-adjacent platform changes.

We'll love to explore more about you if you have:
  • This role explicitly requires 4-5 years of hands-on systems experience. We are not looking for someone who will lean entirely on AI tooling to discover what to do; we are looking for someone who already knows what to ask, and can use AI tooling as a force multiplier on top of that judgement.

    Backend engineering depth:production experience in Python, Go/Rust, comfortable owning services end to end, able to read and reason about backend code across teams.

    Kubernetes at scale:scheduler behavior, resource requests/limits, HPA/VPA, node pool design, cost-aware autoscaling (Cast AI, Karpenter, or equivalent).

    Cloud and on-premise infrastructure:GCP fluency, IaC (Terraform), CI/CD, and comfort operating in hy brid setups including on-prem GPU clusters.

    GPU workload understanding:familiarity with throughput profiling, batching, KV-cache behavior, inference server tuning, and GPU utilization metrics.

    Observability and reliability:metrics, traces, logs, SLOs, and the discipline to instrument systems properly rather than reactively.

    FinOps mindset:demonstrated history of converting infrastructure choices into measurable cost outcomes.

    Security baseline:able to take on platform-security workstreams without requiring constant handoff to the DevOps team.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Site Reliability Engineer -AI Infrastructure Operations
Senior Site Reliability Engineer -AI Infrastructure Operations

Nscale • San Francisco (CA), Seattle (WA), Houston (TX)

On-site
USD 170,000 - 265,000
Equity
Ownership from start
Flexible schedule
Senior Site Reliability Engineer (SRE) - AI Inftastructure
Senior Site Reliability Engineer (SRE) - AI Inftastructure

Hamilton Barnes Associates Limited • San Francisco (CA)

On-site
USD 270,000 - 330,000
Equity
Senior / Lead Infrastructure & Operations Engineer
Senior / Lead Infrastructure & Operations Engineer

Austin Werner • Boston (MA)

On-site
USD 120,000 - 150,000
Site Reliability Engineer
Site Reliability Engineer

Amiri Recruiting • Mountain View (CA)

On-site
USD 130,000 - 160,000
Sr. Site Reliability Engineer
Sr. Site Reliability Engineer

Tiger Analytics, LLC • Washington

Hybrid
USD 120,000 - 160,000
Career development opportunities
Sr. Site Reliability Engineer
Sr. Site Reliability Engineer

Tiger Analytics • Washington

Hybrid
USD 100,000 - 140,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

SDI International • Chicago (IL)

Hybrid
USD 130,000 - 180,000
Staff+ Software Engineer - Backend
Staff+ Software Engineer - Backend

Resolve AI • San Francisco (CA)

On-site
USD 120,000 - 160,000
Comprehensive Medical, Dental, and Vision Insurance
Monthly Housing Stipend
Flexible (Unlimited) Paid Time Off
+5
Staff Site Reliability Engineer – Automation and Platform
Staff Site Reliability Engineer – Automation and Platform

Cerebras • Sunnyvale (CA)

On-site
USD 150,000 - 200,000
Site Reliability Engineer Austin, TX
Site Reliability Engineer Austin, TX

Future Secure AI Pty • Austin (TX)

On-site
USD 100,000 - 140,000
Flexible work environment
Competitive salary
Growth trajectory