Staff Infrastructure Engineer — Real-Time AI Systems
Strativ Group
Menlo Park (CA)
On-site
USD 250,000 - 320,000
Full time
14 days+
Get more replies from employers
Send a job-specific resume in minutes.
Start fresh or import an existing resume
Job summary
A leading AI lab in Menlo Park is seeking a Staff Infrastructure Engineer to architect and build the compute substrate for next-generation AI systems. The ideal candidate has 5+ years in Software / ML Infrastructure Engineering, with deep experience in distributed systems and GPU orchestration. Join a team pushing the boundaries of real-time generative models with a focus on low-latency performance and scalability.
Qualifications
5+ years of experience in Software / ML Infrastructure Engineering.
Deep experience with distributed systems and GPU orchestration for high‑performance ML workloads.
Proficiency in Python, Go, or similar, and strong grasp of software engineering best practices.
Hands‑on expertise with Kubernetes, Docker, and IaC (Terraform).
Experience optimizing model serving and data pipelines for latency and scalability.
Responsibilities
Work directly with founders to architect and build the compute substrate.
Design and optimize the inference platform and GPU-based training clusters.
Play a key role in scaling systems for research and production.
Skills
Distributed systems
GPU orchestration
Python
Go
Kubernetes
Docker
Infrastructure as Code (Terraform)
Job description
A leading AI lab in Menlo Park is seeking a Staff Infrastructure Engineer to architect and build the compute substrate for next-generation AI systems. The ideal candidate has 5+ years in Software / ML Infrastructure Engineering, with deep experience in distributed systems and GPU orchestration. Join a team pushing the boundaries of real-time generative models with a focus on low-latency performance and scalability.