Get more replies from employers
Send a job-specific resume in minutes.
Nebius in Palo Alto is seeking a Senior MLE to own end-to-end model and endpoint optimization. You will work at the intersection of distributed systems, GPU performance, and production ML engineering, debugging serving problems and delivering measurable improvements with minimal supervision.
You will deploy and optimize LLM/VLM backends, quantify latency and cost, and collaborate with researchers, kernel engineers, and platform teams to push the performance frontier of AI workloads.
Nebius in Palo Alto is seeking a Senior MLE to own end-to-end model and endpoint optimization. You will work at the intersection of distributed systems, GPU performance, and production ML engineering, debugging serving problems and delivering measurable improvements with minimal supervision.
You will deploy and optimize LLM/VLM backends, quantify latency and cost, and collaborate with researchers, kernel engineers, and platform teams to push the performance frontier of AI workloads.