Get more replies from employers
Send a job-specific resume in minutes.
OpenAI in San Francisco is seeking an experienced engineer to architect, build, and scale high-performance distributed training and inference infrastructure for massive LLMs. You will optimize GPU cluster utilization and memory management for multi-node training, partnering with research teams to accelerate experiments.
The role requires deep expertise in distributed systems, networking, and high-throughput computing, plus strong coding in C++, Python, and CUDA, with hands-on experience on large
OpenAI is an AI research and deployment company dedicated to ensuring that artificial general intelligence benefits all of humanity. Our infrastructure team builds the massive distributed systems required to train and serve frontier models.