Stand out for this role — generate a tailored resume and cover letter in about a minute.
Emploive is seeking an experienced ML systems engineer to own the cost and performance of our inference stack, shaping how we serve models as workloads, traffic, and hardware evolve. You will partner with the engineers operating the serving fleet to optimize caching, batching, quantization, decoding, and kernel-level tuning.
You will implement profiling systems to show where time and memory are spent and tune routing between internal infrastructure and external providers using cost and capacity
Emploive is seeking an experienced ML systems engineer to own the cost and performance of our inference stack, shaping how we serve models as workloads, traffic, and hardware evolve. You will partner with the engineers operating the serving fleet to optimize caching, batching, quantization, decoding, and kernel-level tuning.
You will implement profiling systems to show where time and memory are spent and tune routing between internal infrastructure and external providers using cost and capacity