Stand out for this role — generate a tailored resume and cover letter in about a minute.
Get past ATS filters
Job summary
Reka is seeking a Data Engineer to ensure high-quality data production at petabyte scale. You will collaborate with model researchers and build algorithms for automated data quality assessment while tracking datasets for reproducibility. The ideal candidate has strong machine learning and deep learning fundamentals, practical experience with distributed processing tools, and solid Python skills. This position requires a blend of research and production engineering, offering a dynamic work environment in the United States.
Qualifications
Strong fundamentals in ML and experience with large-scale systems.
Comfortable with both research and production engineering.
Demonstrated experience with data quality and dataset releases.
Ability to design experiments with unbiased outcomes.
Practical experience with distributed processing tools.
Responsibilities
Define quality metrics, validation checks, and acceptance thresholds for data.
Create internal datasets for building fundamental World Models.
Build algorithms for automated data quality assessment.
Track datasets, metadata, and ensure experiments are reproducible.
Own CI/CD for the data stack and automate workflows.
Skills
Machine Learning fundamentals
Deep learning experience
Data quality assessment
Python skills
Distributed processing
GitHub experience
Tools
PyTorch
Spark
Airflow
Job description
Reka is seeking a Data Engineer to ensure high-quality data production at petabyte scale. You will collaborate with model researchers and build algorithms for automated data quality assessment while tracking datasets for reproducibility. The ideal candidate has strong machine learning and deep learning fundamentals, practical experience with distributed processing tools, and solid Python skills. This position requires a blend of research and production engineering, offering a dynamic work environment in the United States.