Build and Deploy AI the right way, anywhere.
The FlexAI Compute Infrastructure Platform provides an "end-to-end AI compute layer" for running and managing workloads across any cloud, any GPU, and any deployment model (public, hybrid, or on-prem). It brings together "1-click simplicity" for users with "enterprise-grade orchestration, security, and automation" under the hood.
Founded by Brijesh Tripathi, who brings experience from Nvidia, Apple, Tesla, Intel and Zoox, FlexAI is not just building a product – we’re shaping the future of AI. Our teams are strategically distributed across Paris, Silicon Valley, and Bangalore, united by a shared mission: to deliver more compute with less complexity.
If you're passionate about shaping the future of artificial intelligence, driving innovation, and contributing to a sustainable and inclusive AI ecosystem, FlexAI is the place for you !
Position Overview
FlexAI is looking for a highly experienced Staff DevOps / SRE Engineer to architect, scale, and operate the core infrastructure powering our next-generation AI and PaaS platform. Building on the foundation of our current DevOps/SRE function supporting a multi-architecture PaaS product in Beta, this role will operate at a strategic and technical leadership level, driving reliability, scalability, and performance of a globally distributed AI compute platform.
As a Staff Engineer, you will go beyond execution to define infrastructure strategy, establish SRE best practices, and design resilient systems capable of supporting large-scale AI workloads, distributed runtimes, and high-performance cloud-native environments. You will work closely with Runtime, Backend, and Platform teams to ensure our platform delivers enterprise-grade reliability while maintaining the agility of a fast-scaling deep-tech startup.
This is a high-impact role suited for engineers who thrive in complex, multi-cloud, high-throughput environments and enjoy solving hard infrastructure problems at scale.
What You’ll Do
- Architect and evolve the infrastructure backbone for FlexAI’s PaaS and AI runtime platform across multi-cloud and multi-architecture environments
- Design highly available, fault-tolerant, and scalable systems that support mission-critical AI and compute workloads
- Establish and own SRE principles including SLIs, SLOs, error budgets, and reliability frameworks
Infrastructure at Scale
- Lead the design and implementation of Infrastructure as Code (IaC) using Terraform and other modern tooling to ensure repeatable and scalable infrastructure
- Own and optimize Kubernetes clusters, container runtimes, and orchestration frameworks across heterogeneous architectures
- Drive infrastructure standardization and automation to support global deployments and rapid product iteration
- Define and scale advanced CI/CD pipelines to support rapid, reliable, and secure releases of platform components
- Build self-healing systems, automated remediation workflows, and reliability-focused automation
- Champion GitOps and platform engineering best practices across the organization
Observability & Performance Engineering
- Design and implement end-to-end observability (metrics, logs, traces) for distributed systems and AI workloads
- Proactively identify performance bottlenecks and drive system optimization for latency, throughput, and cost efficiency
- Lead incident management, root cause analysis, and postmortem culture across the infrastructure stack
Cross-Functional Collaboration
- Partner with Runtime, AI, Backend, and Security teams to ensure seamless integration and high system reliability
- Act as a technical advisor to leadership on infrastructure scaling, reliability risks, and architectural decisions
- Mentor and guide senior and mid-level DevOps/SRE engineers, setting technical standards and best practices
Security, Compliance & Resilience
- Embed security best practices into infrastructure design and deployment workflows
- Ensure platform resilience through chaos engineering, disaster recovery planning, and capacity forecasting
- Support compliance and governance requirements in global cloud environments
What You’ll Need
- Bachelor’s or Master’s degree in Computer Science, Engineering, or a related field
- 8+ years of experience in DevOps, SRE, or Infrastructure Engineering, with at least 2+ years in a Staff/Lead-level capacity
- Proven experience designing and operating large-scale, cloud-native distributed systems
- Deep expertise in Kubernetes, Docker, and container orchestration in production environments
- Strong hands‑on experience with Terraform and modern IaC practices
- Advanced knowledge of CI/CD systems, release engineering, and automation frameworks
- Strong proficiency in scripting and programming (Python, Go, Bash; Rust is a plus)
- Extensive experience with cloud platforms (AWS, GCP, Azure) and multi‑cloud architectures
- Strong understanding of observability stacks (Prometheus, Grafana, OpenTelemetry, etc.)
- Experience building highly reliable platforms with defined SLOs, SLAs, and incident response frameworks
- Excellent troubleshooting and systems thinking skills in high‑scale environments
Nice to Have
- Experience supporting AI/ML infrastructure, GPU workloads, or model training pipelines
- Familiarity with distributed runtimes and high-performance compute environments
- Exposure to platform engineering and internal developer platform (IDP) design
- Experience in startup or deep-tech environments building products from Beta to scale
What Makes You a Great Fit
- You think in systems, not just tools
- You enjoy solving reliability and scale challenges in cutting‑edge platforms
- You bring a strong ownership mindset and thrive in fast‑paced, ambiguous environments
- You can balance long‑term architecture with immediate operational excellence
- You are entrepreneurial, pragmatic, and excited about building foundational infrastructure for AI‑native platforms
At FlexAI, you’ll play a critical role in shaping the infrastructure that powers next-generation AI and cloud platforms, working at the intersection of reliability, scale, and innovation.