Get more replies from employers
Send a job-specific resume in minutes.
Unknown Company in San Francisco, CA seeks a lead ML infrastructure engineer to own the distributed training and inference backbone for a foundation model trained from scratch. You will stand up clusters, build data pipelines at petabyte scale, and optimize GPU performance across model scales.
You will work with FSDP/DeepSpeed, NVIDIA GPUs, Linux, Python and C++, in a distributed cloud environment across GCP/AWS/Azure, with relocation supported.
Unknown Company in San Francisco, CA seeks a lead ML infrastructure engineer to own the distributed training and inference backbone for a foundation model trained from scratch. You will stand up clusters, build data pipelines at petabyte scale, and optimize GPU performance across model scales.
You will work with FSDP/DeepSpeed, NVIDIA GPUs, Linux, Python and C++, in a distributed cloud environment across GCP/AWS/Azure, with relocation supported.