An application made for this job — a tailored resume and cover letter that speak straight to the posting.
Inception is seeking engineers and scientists to design, optimize, and scale the diffusion LLM serving systems powering production inference. Your work will help make inference faster, more cost-efficient, and more reliable.
You will extend orchestration frameworks (Kubernetes, Ray, SLURM) for distributed inference, evaluation, and large-batch serving, and implement load balancing, autoscaling, and traffic routing for model endpoints.
We\'re looking for engineers and scientists to design, optimize, and scale the systems that power our diffusion LLMs in production. Your work will make inference faster, more cost-effective, and more reliable.