Get more replies from employers
Send a job-specific resume in minutes.
Scale AI is building the foundation platform that powers all ML research and development. You’ll collaborate with Scale’s ML teams to optimize training and inference, enabling the next generation of LLMs, data curation, and research timelines.
You will profile, tune, and integrate cutting-edge technologies to push performance and scalability of large-scale distributed ML systems, working across researchers and engineers to deliver robust solutions.
Strong software engineering skills, proficient in frameworks and tools such as CUDA, Pytorch, transformers, flash attention, etcStrong written and verbal communication skills and the ability to operate in a cross functional team environmentExperience with multi-node LLM training and inferenceStrong excitement about system optimizationExperience with developing large-scale distributed ML systemsDemonstrated expertise in post-training methods &/or next generation use cases for large language models including instruction tuning, RLHF, tool use, reasoning, agents, and multimodal, etc