An application made for this job — a tailored resume and cover letter that speak straight to the posting.
Tower Research Capital in New York seeks a Training Performance and Distributed Training Engineer to accelerate ML model training at scale. You will optimize end-to-end training pipelines from data ingestion to kernel performance, enabling researchers to iterate on increasingly complex models.
You will benchmark across CPUs/GPUs, design distribution strategies, and collaborate with ML researchers, HPC engineers, and infrastructure teams to improve throughput, stability, and cost efficiency.
Tower Research Capital in New York seeks a Training Performance and Distributed Training Engineer to accelerate ML model training at scale. You will optimize end-to-end training pipelines from data ingestion to kernel performance, enabling researchers to iterate on increasingly complex models.
You will benchmark across CPUs/GPUs, design distribution strategies, and collaborate with ML researchers, HPC engineers, and infrastructure teams to improve throughput, stability, and cost efficiency.