Stand out for this role — generate a tailored resume and cover letter in about a minute.
Scale AI is hiring to build and optimize the foundation platform powering our ML research and development. You will collaborate with research teams to enable scalable LLM training, inference, and data curation.
The role emphasizes building high-performance distributed ML systems and integrating cutting-edge techniques to optimize our AI stack across multiple nodes and models.
If you are excited about shaping the future AI via fundamental innovations, we would love to hear from you!
Strong software engineering skills, proficient in frameworks and tools such as CUDA, Pytorch, transformers, flash attention, etcStrong written and verbal communication skills and the ability to operate in a cross functional team environmentExperience with multi-node LLM training and inferenceStrong excitement about system optimizationExperience with developing large-scale distributed ML systemsDemonstrated expertise in post-training methods &/or next generation use cases for large language models including instruction tuning, RLHF, tool use, reasoning, agents, and multimodal, etc