A complete application in a minute — tailored resume and cover letter, ready to send.
Acceler8 Talent is seeking a Member of Technical Staff, Inference Performance to accelerate AI model inference across heterogeneous hardware. You will drive optimizations in latency, throughput, and memory usage, and push the serving stack with batching, caching, and quantization techniques.
The role involves developing CUDA/HIP/Triton kernels, deploying on accelerator clusters, and benchmarking models under real production workloads to improve efficiency and cost effectiveness.
Member of Technical Staff, Inference Performance
200k base + equity
I’m working with a AI infrastructure startup building a high-performance inference cloud for open models.
The team optimizes the full path from model and serving engine through kernels, accelerators, and production infrastructure. Its platform is already processing trillions of tokens per month, and the company is expanding due to customer demand growing faster than its current capacity.
This role will focus on making inference faster, more reliable, and more cost-efficient across different models and hardware architectures.
You’ll work on problems such as:
Looking for engineers who have strong evidence in one or more of:
Strong candidates will be able to explain what they personally optimized, how they measured it, and the production impact it created.
This is an intense, highly hands-on environment with direct founder access and broad ownership. It will suit engineers who want to move across models, kernels, hardware, and infrastructure rather than remain within a narrowly defined area.