Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.
Orbifold AI in Palo Alto is seeking a Member of Technical Staff, ML Engineer (Inference & Performance) to own the end-to-end serving layer for production models on Ray Serve and PyTorch, across a heterogeneous GPU fleet. This role focuses on throughput, latency, memory fit, and cost efficiency.
You will push quantization, memory layout, and kernel-level optimizations, with opportunities to develop custom CUDA/Triton code, benchmarks, and production-grade observability for scalable, reliable
Orbifold AI in Palo Alto is seeking a Member of Technical Staff, ML Engineer (Inference & Performance) to own the end-to-end serving layer for production models on Ray Serve and PyTorch, across a heterogeneous GPU fleet. This role focuses on throughput, latency, memory fit, and cost efficiency.
You will push quantization, memory layout, and kernel-level optimizations, with opportunities to develop custom CUDA/Triton code, benchmarks, and production-grade observability for scalable, reliable