Destaca para este puesto — genera un currículum y una carta de presentación adaptados en cuestión de un minuto.
Proxima is seeking a Principal ML Performance Engineer to optimize training and inference for structural and generative models, with a focus on CUDA, Triton, and distributed training across large GPU clusters. You will profile bottlenecks, write efficient kernels, and drive multi-node scalability.
As a senior leader, you will scale training across 32–64 nodes, reduce inference costs on thousands of GPUs, and mentor engineers while defining technical direction for the team.
Proxima is seeking a Principal ML Performance Engineer to optimize training and inference for structural and generative models, with a focus on CUDA, Triton, and distributed training across large GPU clusters. You will profile bottlenecks, write efficient kernels, and drive multi-node scalability.
As a senior leader, you will scale training across 32–64 nodes, reduce inference costs on thousands of GPUs, and mentor engineers while defining technical direction for the team.