Get more replies from employers
Send a job-specific resume in minutes.
Jobtailor is seeking an experienced data infrastructure engineer to build and optimize large-scale batch and streaming data pipelines for ML training in a US setting. You will work with Flink, Spark, and Ray to generate datasets and integrate with Flyte and Airflow for reliable multi-stage workflows.
You will drive reproducibility, observability, and cost-aware optimization across distributed systems while collaborating with ML teams on experimentation and model iteration.
Demonstrates expertise in developing and optimizing large-scale data pipelines for machine learning, utilizing distributed computing frameworks like Flink, Spark, and Ray. Proficient in integrating orchestration systems and ensuring pipeline reliability and cost efficiency.