Turn this role into an interview — a resume and cover letter built around what this employer wants.
Cohere is seeking an engineer to build reliable ML systems and optimize LLM inference for enterprise-grade AI workloads. You will dive into the inference stack to reduce latency, improve throughput, and ensure quality across diverse tasks.
The role involves collaborating with modeling and systems teams, exploring GPU/CUDA optimizations, and implementing kernel-level improvements for large-scale architectures.
Cohere is seeking an engineer to build reliable ML systems and optimize LLM inference for enterprise-grade AI workloads. You will dive into the inference stack to reduce latency, improve throughput, and ensure quality across diverse tasks.
The role involves collaborating with modeling and systems teams, exploring GPU/CUDA optimizations, and implementing kernel-level improvements for large-scale architectures.