Staff DevOps Engineer/SRE

FlexAI

Bengaluru

On-site

INR 2,000,000 - 3,500,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

A leading AI infrastructure firm is seeking a Staff DevOps / SRE Engineer to architect and scale their next-generation AI and PaaS platform. This role involves leading the design of cloud-native systems, establishing SRE best practices, and mentoring engineers. Ideal candidates have a strong background in Kubernetes, Terraform, and CI/CD, with over 8 years of relevant experience, including leadership roles. Join an innovative team to shape the future of AI infrastructure in Bengaluru.

Qualifications

  • 8+ years of experience in DevOps, SRE, or Infrastructure Engineering.
  • At least 2 years in a Staff/Lead-level capacity.
  • Strong hands-on experience with Terraform and IaC practices.

Responsibilities

  • Architect and evolve infrastructure for multi-cloud environments.
  • Design highly available and fault-tolerant systems.
  • Establish SRE principles including SLIs, SLOs, and error budgets.

Skills

Kubernetes
Terraform
Python
Observability (Prometheus, Grafana)
CI/CD systems

Education

Bachelor’s or Master’s degree in Computer Science, Engineering, or related field

Tools

Docker
GitOps

Job description

Build and Deploy AI the right way, anywhere.

The FlexAI Compute Infrastructure Platform provides an "end-to-end AI compute layer" for running and managing workloads across any cloud, any GPU, and any deployment model (public, hybrid, or on-prem). It brings together "1-click simplicity" for users with "enterprise-grade orchestration, security, and automation" under the hood.

Founded by Brijesh Tripathi, who brings experience from Nvidia, Apple, Tesla, Intel and Zoox, FlexAI is not just building a product – we’re shaping the future of AI. Our teams are strategically distributed across Paris, Silicon Valley, and Bangalore, united by a shared mission: to deliver more compute with less complexity.

If you're passionate about shaping the future of artificial intelligence, driving innovation, and contributing to a sustainable and inclusive AI ecosystem, FlexAI is the place for you !

Position Overview

FlexAI is looking for a highly experienced Staff DevOps / SRE Engineer to architect, scale, and operate the core infrastructure powering our next-generation AI and PaaS platform. Building on the foundation of our current DevOps/SRE function supporting a multi-architecture PaaS product in Beta, this role will operate at a strategic and technical leadership level, driving reliability, scalability, and performance of a globally distributed AI compute platform.

As a Staff Engineer, you will go beyond execution to define infrastructure strategy, establish SRE best practices, and design resilient systems capable of supporting large-scale AI workloads, distributed runtimes, and high-performance cloud-native environments. You will work closely with Runtime, Backend, and Platform teams to ensure our platform delivers enterprise-grade reliability while maintaining the agility of a fast-scaling deep-tech startup.

This is a high-impact role suited for engineers who thrive in complex, multi-cloud, high-throughput environments and enjoy solving hard infrastructure problems at scale.

What You’ll Do
  • Architect and evolve the infrastructure backbone for FlexAI’s PaaS and AI runtime platform across multi-cloud and multi-architecture environments
  • Design highly available, fault-tolerant, and scalable systems that support mission-critical AI and compute workloads
  • Establish and own SRE principles including SLIs, SLOs, error budgets, and reliability frameworks

Infrastructure at Scale

  • Lead the design and implementation of Infrastructure as Code (IaC) using Terraform and other modern tooling to ensure repeatable and scalable infrastructure
  • Own and optimize Kubernetes clusters, container runtimes, and orchestration frameworks across heterogeneous architectures
  • Drive infrastructure standardization and automation to support global deployments and rapid product iteration
  • Define and scale advanced CI/CD pipelines to support rapid, reliable, and secure releases of platform components
  • Build self-healing systems, automated remediation workflows, and reliability-focused automation
  • Champion GitOps and platform engineering best practices across the organization

Observability & Performance Engineering

  • Design and implement end-to-end observability (metrics, logs, traces) for distributed systems and AI workloads
  • Proactively identify performance bottlenecks and drive system optimization for latency, throughput, and cost efficiency
  • Lead incident management, root cause analysis, and postmortem culture across the infrastructure stack

Cross-Functional Collaboration

  • Partner with Runtime, AI, Backend, and Security teams to ensure seamless integration and high system reliability
  • Act as a technical advisor to leadership on infrastructure scaling, reliability risks, and architectural decisions
  • Mentor and guide senior and mid-level DevOps/SRE engineers, setting technical standards and best practices

Security, Compliance & Resilience

  • Embed security best practices into infrastructure design and deployment workflows
  • Ensure platform resilience through chaos engineering, disaster recovery planning, and capacity forecasting
  • Support compliance and governance requirements in global cloud environments
What You’ll Need
  • Bachelor’s or Master’s degree in Computer Science, Engineering, or a related field
  • 8+ years of experience in DevOps, SRE, or Infrastructure Engineering, with at least 2+ years in a Staff/Lead-level capacity
  • Proven experience designing and operating large-scale, cloud-native distributed systems
  • Deep expertise in Kubernetes, Docker, and container orchestration in production environments
  • Strong hands‑on experience with Terraform and modern IaC practices
  • Advanced knowledge of CI/CD systems, release engineering, and automation frameworks
  • Strong proficiency in scripting and programming (Python, Go, Bash; Rust is a plus)
  • Extensive experience with cloud platforms (AWS, GCP, Azure) and multi‑cloud architectures
  • Strong understanding of observability stacks (Prometheus, Grafana, OpenTelemetry, etc.)
  • Experience building highly reliable platforms with defined SLOs, SLAs, and incident response frameworks
  • Excellent troubleshooting and systems thinking skills in high‑scale environments
Nice to Have
  • Experience supporting AI/ML infrastructure, GPU workloads, or model training pipelines
  • Familiarity with distributed runtimes and high-performance compute environments
  • Exposure to platform engineering and internal developer platform (IDP) design
  • Experience in startup or deep-tech environments building products from Beta to scale
What Makes You a Great Fit
  • You think in systems, not just tools
  • You enjoy solving reliability and scale challenges in cutting‑edge platforms
  • You bring a strong ownership mindset and thrive in fast‑paced, ambiguous environments
  • You can balance long‑term architecture with immediate operational excellence
  • You are entrepreneurial, pragmatic, and excited about building foundational infrastructure for AI‑native platforms

At FlexAI, you’ll play a critical role in shaping the infrastructure that powers next-generation AI and cloud platforms, working at the intersection of reliability, scale, and innovation.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Lead/Senior Backend Engineer
Lead/Senior Backend Engineer

FlexAI • Bengaluru

On-site
INR 1,200,000 - 1,800,000
Senior FullStack Engineer
Senior FullStack Engineer

FlexAI • Bengaluru

On-site
INR 1,200,000 - 1,800,000
Staff DevOps Engineer
Staff DevOps Engineer

Sia • Mumbai

On-site
INR 1,800,000 - 3,000,000
Opportunity to lead AI projects
Collaborative team environment
Equal opportunity employer
Staff DevOps Engineer
Staff DevOps Engineer

Sia Partners' • Mumbai

On-site
INR 1,500,000 - 2,500,000
Opportunity to lead cutting-edge AI projects
Dynamic and collaborative team environment
Lead SDE - DevOps
Lead SDE - DevOps

Flourish Ventures • Chennai District

On-site
INR 2,000,000 - 3,000,000
Inclusive and people-first culture
Health & wellness programs
Comprehensive medical insurance
+2
Site Reliability Engineer
Site Reliability Engineer

United States Digital Space LLC • Karnataka

On-site
INR 900,000 - 1,200,000
Significant equity in a venture-backed company
Opportunity to work with modern tech stack
Cloud Infrastructure & Platform Engineer
Cloud Infrastructure & Platform Engineer

Virallens • Bengaluru

On-site
INR 3,000,000 - 4,200,000
Principal Site Reliability Engineer
Principal Site Reliability Engineer

Namely • India

On-site
INR 1,500,000 - 2,500,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

SourcingXPress • Hyderabad

On-site
INR 3,000,000 - 5,000,000
Member of Technical Staff (MTS) - DevOps & Infrastructure
Member of Technical Staff (MTS) - DevOps & Infrastructure

Fi Money • Bengaluru

On-site
INR 1,000,000 - 1,500,000