Staff Software Engineer (AI Infrastructure)

DeepRec.ai

Palo Alto (CA)

On-site

USD 180,000 - 320,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

DeepRec.ai in the Bay Area seeks a Staff/Lead Software Engineer, AI Infrastructure to design and scale GPU infrastructure, model serving APIs, and production deployment for high-performance AI features.

You will mentor engineers, drive the technical vision, and partner with ML and backend teams to deliver reliable, scalable AI capabilities at scale in California.

Qualifications

  • 5+ years as a software engineer on systems infrastructure, including hands-on ML serving and GPU orchestration.
  • Strong knowledge of distributed systems, Kubernetes, and cloud-native infrastructure (AWS/GCP/Azure).
  • Experience building and optimizing APIs for large-scale AI model serving (TensorFlow Serving, Triton, TorchServe).
  • Experience with high-throughput, scalable GPU fleet management, scheduling, and efficient model execution.
  • Backend programming in Python, Go, or C++ with a focus on performance and reliability.
  • Ownership mentality with ability to solve complex problems in ambiguous, high-growth environments.
  • Excellent communication, collaboration, and mentorship skills.

Responsibilities

  • Design, develop, and maintain scalable GPU infrastructure for training and serving state-of-the-art AI models.
  • Architect and optimize high-throughput, low-latency APIs for AI model serving and inference.
  • Lead orchestration and scheduling of heterogeneous GPU resources across clusters.
  • Build and support robust systems for model deployment, monitoring, scaling, and reliability in production.
  • Collaborate with ML, backend, and platform teams to deliver AI-powered features.
  • Drive technical direction, code reviews, and mentorship across the AI Infrastructure team.

Skills

Ownership mindset
Mentorship
Communication
Python
Go
C++
Distributed systems
System design

Tools

Kubernetes
AWS
GCP
Azure
TensorFlow Serving
Triton
TorchServe

Job description

Staff/Lead Software Engineer, AI Infrastructure

About the Company

A well-funded Bay Area AI startup operating at the frontier of generative media, with a product shipping to users at scale. The company is building the core infrastructure that powers its AI capabilities, and this is a senior, high-ownership hire on that team.

About the Role

This is a critical hire to build and scale the infrastructure behind the company's AI capabilities. You'll lead the design and implementation of GPU infrastructure, AI model serving APIs, and general AI infrastructure execution, enabling the machine learning features that drive the product.

You'll architect robust, distributed systems optimized for high-performance AI workloads, large-scale GPU orchestration, and low-latency, reliable API serving. Your work will directly shape how users experience generative AI at scale. As a senior technical leader, you'll also mentor engineers, drive best practices, and set the technical vision for AI infrastructure.

What You'll Do

  • Design, develop, and maintain scalable GPU infrastructure for training and serving state-of-the-art AI models.
  • Architect and optimize high-throughput, low-latency APIs for AI model serving and inference.
  • Lead the orchestration, scheduling, and efficient utilization of heterogeneous GPU resources across clusters.
  • Build and support robust systems for model deployment, monitoring, scaling, and reliability in production.
  • Collaborate with ML, backend, and platform engineering teams to deliver seamless AI-powered product features.
  • Drive technical direction, code reviews, and mentorship across the AI Infrastructure team.

What We're Looking For

  • 5+ years as a software engineer working on systems infrastructure, including hands-on ML serving and GPU orchestration.
  • Deep knowledge of distributed systems, Kubernetes (or similar orchestration frameworks), and cloud-native infrastructure (AWS/GCP/Azure).
  • Proven expertise building and optimizing APIs for large-scale AI model serving (TensorFlow Serving, Triton, TorchServe, or similar).
  • Familiarity with the challenges of high-throughput, scalable GPU fleet management, scheduling, and efficient model execution.
  • Proficiency in backend languages such as Python, Go, or C++, with experience optimizing for performance and reliability.
  • Ownership mentality and the drive to solve complex problems independently in ambiguous, high-growth environments.
  • Excellent communication, collaboration, and mentorship skills.

Nice to Have

  • Experience with multi-modal AI model infrastructure (LLMs, generative models, video/image/speech models).
  • Background building infra for multi-tenant SaaS, enterprise AI/ML platforms, or operational automation at scale.
  • Previous startup experience, or a track record leading high-impact projects through ambiguity and rapid iteration.
  • Experience with competitive coding or large-scale distributed computing environments.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Software Engineer, AI Infra
Software Engineer, AI Infra

Makers Fund • Palo Alto (CA)

On-site
USD 180,000 - 240,000
Equity
Health benefits
Monthly stipends
+1
AI Training Infrastructure Engineer
AI Training Infrastructure Engineer

Designworks Talent • Bellevue (WA)

Hybrid
USD 180,000 - 240,000
Medical insurance
401(k) with company match
Paid holidays
Software Engineer, AI Infrastructure
Software Engineer, AI Infrastructure

Harell Data • Palo Alto (CA)

On-site
USD 180,000 - 260,000
Senior Software Engineer, Infrastructure Software for AI (Centralized AI Data Centers & Distrib[...]
Senior Software Engineer, Infrastructure Software for AI (Centralized AI Data Centers & Distrib[...]

Intelliswift - An LTTS Company • Sunnyvale (CA)

On-site
USD 120,000 - 150,000
Competitive salary
Health insurance
Flexible work hours
Senior Member of Technical Staff
Senior Member of Technical Staff

DeepRec.ai • Palo Alto (CA)

Hybrid
USD 180,000 - 240,000
Competitive salary + meaningful equity
Flexible hybrid working model
High-performing, collaborative team environment
Software Engineer, AI Infra
Software Engineer, AI Infra

Sarah Smith Fund • Palo Alto (CA)

Hybrid
USD 140,000 - 180,000
Competitive salary
Equity
Health benefits
+1
Senior Software Engineer - Infrastructure
Senior Software Engineer - Infrastructure

InCommon • San Francisco (CA)

On-site
USD 180,000 - 280,000
System Software Engineer - AI
System Software Engineer - AI

Delos Data • Palo Alto (CA)

Hybrid
USD 140,000 - 200,000
Equity
401k
Benefits
Software Engineer - AI Infrastructure
Software Engineer - AI Infrastructure

Hamilton Barnes Associates Limited • San Francisco (CA)

Hybrid
USD 300,000 - 500,000
Early-stage equity
Founding engineer role
Equity package
Member of Technical Staff- Distributed Systems
Member of Technical Staff- Distributed Systems

Acceler8 Talent • San Francisco (CA), Northern (KY)

Hybrid
USD 150,000 - 350,000
Equity