Stand out for this role — generate a tailored resume and cover letter in about a minute.
Zoox in Foster City is seeking a Model Optimization & Deployment Engineer to bring production-ready, large-scale models to our on-vehicle stack. You will design low-latency C++/CUDA code for real-time perception on edge devices and optimize models with quantization and mixed-precision inference.
You will build TensorRT pipelines for edge deployment, perform parity checks against PyTorch, and develop custom ML ops and TensorRT plugins to maximize memory bandwidth and minimize latency on AI
Zoox in Foster City is seeking a Model Optimization & Deployment Engineer to bring production-ready, large-scale models to our on-vehicle stack. You will design low-latency C++/CUDA code for real-time perception on edge devices and optimize models with quantization and mixed-precision inference.
You will build TensorRT pipelines for edge deployment, perform parity checks against PyTorch, and develop custom ML ops and TensorRT plugins to maximize memory bandwidth and minimize latency on AI