An application made for this job — a tailored resume and cover letter that speak straight to the posting.
Strativ Group in Palo Alto is building AI infrastructure for the next generation of agentic AI. You will design and optimize production infrastructure for large-scale AI workloads, focusing on TPU and GPU inference, distributed serving, and runtime performance.
As part of a small founding team, you will work directly with the founders, own significant parts of the infrastructure, and help shape the architecture and direction toward autonomous, self-improving models with minimal human supervision.
We are working with an early-stage AI infrastructure company in Palo Alto building the systems that will power the next generation of agentic AI.
The company is currently focused on high-performance TPU and GPU inference serving, with a broader vision to build an agent cloud for self-improving models and agents.
This is a genuinely early founding-team opportunity. You will work directly with the founders and have significant influence over both the technical architecture and the direction of the company.
You will build and optimize production infrastructure for large-scale AI workloads, with a focus on:
The work sits at the intersection of ML systems, distributed systems and AI infrastructure. You could be working anywhere from low-level serving and runtime performance through to the infrastructure required for long-running, self-improving agent systems.
You do not need experience across the entire stack, but you should have a strong background in one or more of:
TPU experience is particularly valuable, although candidates with strong GPU infrastructure or inference backgrounds are equally relevant.
The company is still only a small founding team, so this is an opportunity to join before the engineering organization scales significantly.
You will have unusually high ownership, work directly alongside the founders and help define the infrastructure behind a much bigger ambition than simply building another inference platform.
The longer-term goal is infrastructure where models, agents and the systems they run on can continuously improve together with minimal human supervision.