Get more replies from employers
Send a job-specific resume in minutes.
Nava is seeking an AI/ML deployment engineer to design and deploy low-latency inference pipelines for LLMs, diffusion models, and vision transformers. You will optimize serving stacks across CPU, GPU, and NPU with Triton Inference Server and ONNX Runtime, and you will containerize services using Docker and Kubernetes to ensure high availability and auto-scaling.
You will implement monitoring, health checks, and A/B testing to validate performance and drift in production, and collaborate with ML