Senior SRE: FinOps-Driven Infra, GPU & Scale

Level AI

Mountain View (CA)

On-site

USD 180,000 - 240,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Level AI, headquartered in Mountain View, California, seeks a Senior SRE to join the infrastructure and FinOps focus. The role bridges backend engineering, operations, and cost optimization, delivering tooling and dashboards that empower backend teams to own costs and reliability while advancing GPU performance and security readiness.

Ideal candidates bring 4–5 years of hands-on systems experience, deep Kubernetes expertise at scale, and strong cloud/on-prem infrastructure skills.

Qualifications

  • 4–5 years of hands-on systems experience.
  • Backend engineering depth with Python, Go/Rust, and end-to-end ownership.
  • Kubernetes at scale: scheduler behavior, resources, HPA/VPA, cost-aware autoscaling.
  • Cloud and on-prem infrastructure: GCP fluency, IaC (Terraform), CI/CD, and hybrid setups with on-prem GPU clusters.
  • GPU workload understanding: throughput profiling, batching, GPU utilization metrics.
  • Observability and reliability: metrics, traces, logs, SLOs, proper instrumentation.
  • FinOps mindset: history of turning infra choices into cost savings.
  • Security baseline: ability to lead platform-security workstreams.

Responsibilities

  • Infrastructure cost efficiency and FinOps. Own reduction of Kubernetes overprovisioning and cost telemetry for backend teams.
  • GPU throughput optimization. Run structured experiments on on-prem GPU clusters in partnership with AI service owners.
  • Backend enablement, not ownership absorption. Build tooling, dashboards, and processes for teams to own cost and reliability budgets.
  • Reliability instrumentation. Ensure surface area is captured for cost-at-scale and reliability.
  • Selective security workstreams. Take on security-adjacent platform changes without overloading DevOps.

Skills

Python
Go/Rust
Kubernetes
Cloud/GCP
Terraform
CI/CD
GPU workloads
Observability
FinOps
Security baseline

Tools

Cast AI
Karpenter
Terraform
CI/CD
GCP
GPU clusters

Job description

Level AI, headquartered in Mountain View, California, seeks a Senior SRE to join the infrastructure and FinOps focus. The role bridges backend engineering, operations, and cost optimization, delivering tooling and dashboards that empower backend teams to own costs and reliability while advancing GPU performance and security readiness.

Ideal candidates bring 4–5 years of hands-on systems experience, deep Kubernetes expertise at scale, and strong cloud/on-prem infrastructure skills.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior GPU Infra SRE — Remote, Low-Level Linux & Scale
Senior GPU Infra SRE — Remote, Low-Level Linux & Scale

Luma • Redwood City (CA)

Remote
USD 180,000 - 240,000
Senior SRE: GPU-Driven, Global Scale & Causal AI
Senior SRE: GPU-Driven, Global Scale & Causal AI

Crossing Hurdles • San Francisco (CA)

On-site
USD 180,000 - 240,000
Senior SRE: 24/7 GPU & Kubernetes Reliability
Senior SRE: 24/7 GPU & Kubernetes Reliability

Nvidia Corporation in • Austin (TX)

On-site
USD 208,000 - 334,000
Equity
Benefits
Senior SRE, BCM/DGX Cloud - Scale GPU Clusters
Senior SRE, BCM/DGX Cloud - Scale GPU Clusters

NVIDIA • Santa Clara (CA)

On-site
USD 168,000 - 334,000
Equity
Benefits
Senior SRE: AI Infra on-Site in SF, GPU & Cloud
Senior SRE: AI Infra on-Site in SF, GPU & Cloud

The Recruiting Guy • Arlington (VA)

On-site
USD 175,000 - 250,000
Senior AI GPU Infra SRE - Scale, Automation & Equity
Senior AI GPU Infra SRE - Scale, Automation & Equity

Hamilton Barnes Associates Limited • San Francisco (CA)

On-site
USD 270,000 - 330,000
Equity
Senior Staff SRE - Global Infra (On-Prem/Cloud), Equity
Senior Staff SRE - Global Infra (On-Prem/Cloud), Equity

NVIDIA Corporation • California (MO)

Hybrid
USD 200,000 - 322,000
Equity
Benefits package
Senior Site Reliability Engineer (SRE) - AI Inftastructure
Senior Site Reliability Engineer (SRE) - AI Inftastructure

Hamilton Barnes Associates Limited • San Francisco (CA)

On-site
USD 270,000 - 330,000
Equity
Senior Site Reliability Engineer (Noida, BLR, India)
Senior Site Reliability Engineer (Noida, BLR, India)

Level AI • Mountain View (CA)

On-site
USD 180,000 - 240,000
Senior SRE - Scalable Infra & Reliability (Equity)
Senior SRE - Scalable Infra & Reliability (Equity)

NVIDIA • Durham (NC)

On-site
USD 224,000 - 431,250
Equity
Benefits