Staff AI Runtime Engineer

FlexAI

Santa Clara (CA)

On-site

USD 180,000 - 240,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Competitive salary and benefits package
Work on cutting-edge AI infrastructure
High ownership and fast execution

Job summary

A forward-thinking AI infrastructure company is seeking a Staff AI Runtime Engineer to lead the design and optimization of their AI compute platform. In this leadership role, you'll enhance AI training and inference capabilities. Successful candidates will have over 8 years of experience in systems engineering, expertise with PyTorch and TensorFlow, and strong programming skills in Python and C++. This role is based in Santa Clara, CA, and offers a competitive salary along with the chance to work on cutting-edge technology.

Qualifications

  • 8+ years of experience in systems/software engineering, especially in AI runtime and distributed systems.
  • Proven experience optimizing and scaling deep learning runtimes for large-scale applications.
  • Strong programming skills in Python and C++, with knowledge of Go or Rust being a plus.
  • Familiarity with distributed training frameworks and resource orchestration.
  • Experience with multi-GPU, multi-node, or cloud-native AI workloads.
  • Solid understanding of containerized workloads and job scheduling in production.

Responsibilities

  • Own the core runtime architecture supporting AI training and inference at scale.
  • Design resilient runtime features within our custom PyTorch stack.
  • Profile and enhance system performance across training and inference pipelines.
  • Improve packaging, deployment, and integration of customer models in production.
  • Mentor junior engineers and collaborate with Research, Infrastructure, and Product teams.
  • Champion CI/CD and software quality across the runtime stack.

Skills

Systems/software engineering
AI runtime optimization
Python
C++
Distributed systems
Containerization
Multi-GPU/Multi-node

Tools

PyTorch
TensorFlow
Kubernetes

Job description

Build and Deploy AI the right way, anywhere.

The FlexAI Compute Infrastructure Platform provides an "end-to-end AI compute layer" for running and managing workloads across any cloud, any GPU, and any deployment model (public, hybrid, or on-prem). It brings together "1-click simplicity" for users with "enterprise-grade orchestration, security, and automation" under the hood.

Founded by Brijesh Tripathi, who brings experience from Nvidia, Apple, Tesla, Intel and Zoox, FlexAI is not just building a product – we’re shaping the future of AI. Our teams are strategically distributed across Silicon Valley and Bengaluru, united by a shared mission: to deliver more compute with less complexity.

If you're passionate about shaping the future of artificial intelligence, driving innovation, and contributing to a sustainable and inclusive AI ecosystem, FlexAI is the place for you !

Role Overview

At FlexAI, we’re building a high-performance, cloud-agnostic AI compute platform designed for next-generation training and inference workloads. As a Staff AI Runtime Engineer, you’ll play a pivotal role in the design, development, and optimization of the core runtime infrastructure that powers distributed training and deployment of large AI models (LLMs and beyond).

This is a hands‑on leadership role – perfect for a systems‑minded software engineer who thrives at the intersection of AI workloads, runtimes, and performance‑critical infrastructure. You’ll own critical components of our PyTorch-based stack, lead technical direction, and collaborate across engineering, research, and product to push the boundaries of elastic, fault‑tolerant, high‑performance model execution.

What You’ll Do
Lead Runtime Design & Development:
  • Own the core runtime architecture supporting AI training and inference at scale.
  • Design resilient and elastic runtime features (e.g. dynamic node scaling, job recovery) within our custom PyTorch stack.
  • Optimize distributed training reliability, orchestration, and job-level fault tolerance.
Drive Performance at Scale:
  • Profile and enhance low‑level system performance across training and inference pipelines.
  • Improve packaging, deployment, and integration of customer models in production environments.
  • Ensure consistent throughput, latency, and reliability metrics across multi‑node, multi‑GPU setups.
Build Internal Tooling & Frameworks:
  • Design and maintain libraries and services that support model lifecycle: training, checkpointing, fault recovery, packaging, and deployment.
  • Implement observability hooks, diagnostics, and resilience mechanisms for deep learning workloads.
  • Champion best practices in CI/CD, testing, and software quality across the AI Runtime stack.
  • Work cross-functionally with Research, Infrastructure, and Product teams to align runtime development with customer and platform needs.
  • Guide technical discussions, mentor junior engineers, and help scale the AI Runtime team’s capabilities.
What You’ll Need to Be Successful
  • 8+ years of experience in systems/software engineering, with deep exposure to AI runtime, distributed systems, or compiler/runtime interaction.
  • Proven experience optimizing and scaling deep learning runtimes (e.g. PyTorch, TensorFlow, JAX) for large‑scale training and/or inference.
  • Strong programming skills in Python and C++ (Go or Rust is a plus).
  • Familiarity with distributed training frameworks, low‑level performance tuning, and resource orchestration.
  • Experience working with multi‑GPU, multi‑node, or cloud-native AI workloads.
  • Solid understanding of containerized workloads, job scheduling, and failure recovery in production environments.
Nice to Have
  • Contributions to PyTorch internals or open-source DL infrastructure projects.
  • Familiarity with LLM training pipelines, checkpointing, or elastic training orchestration.
  • Experience with Kubernetes, Ray, TorchElastic, or custom AI job orchestrators.
  • Background in systems research, compilers, or runtime architecture for HPC or ML.
  • Start‑up previous experience.

This position is In‑Person and located at our Santa Clara, CA Office.

What We Offer
  • A competitive salary and benefits package
  • Work on cutting‑edge AI infrastructure
  • Build products used by developers and enterprises
  • High ownership, fast execution, real impact
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Backend Engineer
Senior Backend Engineer

FlexAI • San Jose (CA)

On-site
USD 180,000 - 240,000
Competitive salary
Cutting-edge AI infra
High ownership
+2
Senior Backend Engineer
Senior Backend Engineer

FlexAI • Santa Clara (CA)

On-site
USD 120,000 - 150,000
Competitive salary and benefits
Work on cutting-edge AI infrastructure
High ownership and fast execution
Senior Full Stack Engineer
Senior Full Stack Engineer

FlexAI • Santa Clara (CA)

On-site
USD 120,000 - 160,000
Competitive salary and benefits package
Opportunity to work on cutting-edge AI infrastructure
High ownership and real impact
Presales Engineer (AI Workload/ Cloud Infrastructure)
Presales Engineer (AI Workload/ Cloud Infrastructure)

FlexAI • Santa Clara (CA)

Hybrid
USD 120,000 - 160,000
Competitive Compensation
Performance-Based Incentives
Comprehensive Benefits
+2
Runtime Engineer
Runtime Engineer

Amadeus Search • San Francisco (CA)

Hybrid
USD 120,000 - 160,000
Equity opportunities
Medical, dental, and vision coverage
Retirement savings plan
+2
Software Engineer, Inference Runtime
Software Engineer, Inference Runtime

Lm-Studio • New York (NY)

Hybrid
USD 140,000 - 210,000
Competitive salary and equity grants
Excellent medical, vision, dental care
Catered team lunch / expensed dinners
+3
AI Training Infrastructure Engineer
AI Training Infrastructure Engineer

Designworks Talent • Bellevue (WA)

Hybrid
USD 180,000 - 240,000
Medical insurance
401(k) with company match
Paid holidays
Staff Software Engineer (AI Infrastructure)
Staff Software Engineer (AI Infrastructure)

DeepRec.ai • Palo Alto (CA)

On-site
USD 180,000 - 320,000
Training Performance Engineer
Training Performance Engineer

Slope • San Francisco (CA)

On-site
USD 250,000 - 460,000
Relocation assistance
Flexible working hours
Collaborative work environment
Staff Product Manager AI
Staff Product Manager AI

FlexAI • New York (NY)

Remote
USD 100,000 - 180,000
Competitive salary
Continuous growth opportunities
Support for personal and professional development