Get more replies from employers
Send a job-specific resume in minutes.
Acceler8 Talent is seeking a Member of Technical Staff to join our San Francisco onsite team building a next-gen AI inference platform. You will design and implement production-grade ML inference and model serving systems, optimizing latency and throughput for large-scale workloads.
You'll collaborate with compiler, kernel, networking, and distributed systems engineers to push performance and efficiency, supporting diverse hardware and heterogeneous compute in production environments.
Member of Technical Staff – ML Systems & Inference
San Francisco, CA (Onsite)
I am seeking a Member of Technical Staff to join one of the most exciting AI infrastructure companies building the next generation of inference systems.
As AI models continue to grow in size and complexity, the challenge is no longer simply adding more GPUs, it's about making diverse hardware work together efficiently. This team is building the infrastructure that intelligently executes AI workloads across heterogeneous compute, delivering significant improvements in performance, efficiency, and scalability for production AI applications.
You'll join a small, highly technical engineering team solving some of the hardest problems in AI systems, working across inference runtimes, scheduling, memory management, and distributed infrastructure.
What You'll Do:
What We're Looking For:
This is an opportunity to join an exceptionally well-funded AI infrastructure company with a small, world-class engineering team already supporting production deployments for Fortune 500 and AI-native organisations.
You'll work across compiler systems, GPU kernels, distributed scheduling, inference optimisation, and heterogeneous compute, solving difficult engineering challenges that directly impact how modern AI workloads are executed in production.