Data Ingestion Engineer for Scalable AI Training Pipelines
Reflection
New York (NY)
Hybrid
USD 100,000 - 130,000
Full time
14 days+
Get more replies from employers
Send a job-specific resume in minutes.
Start fresh or import an existing resume
Benefits offered by this job
Top-tier compensation
Comprehensive health insurance
Fully paid parental leave
Paid time off
Daily lunch and dinner
Regular team celebrations
Job summary
Reflection, located in New York, is searching for a Data Engineer to build robust data ingestion systems essential for AI training. The ideal candidate will be skilled in web crawling and data acquisition, comfortable working with large datasets, and have excellent communication abilities. The role emphasizes collaboration with researchers and iterative processes based on measurable impact. Benefits include top-tier compensation, health insurance, paid parental leave, and opportunities for team engagement.
Qualifications
Experience building systems using Ray, Beam, Spark, or similar.
Familiarity with LLM training and evaluation.
Ability to work with large datasets (multi-TB to PB).
Excellent communication skills.
Responsibilities
Build and operate data ingestion systems for pre-training.
Run experiments on strategies and methods.
Analyze ingested data for gaps and improvements.
Develop specialized crawlers for high-priority sources.
Skills
Web crawling
Data ingestion
Large-scale data acquisition
Ray
Beam
Spark
Communication skills
Job description
Reflection, located in New York, is searching for a Data Engineer to build robust data ingestion systems essential for AI training. The ideal candidate will be skilled in web crawling and data acquisition, comfortable working with large datasets, and have excellent communication abilities. The role emphasizes collaboration with researchers and iterative processes based on measurable impact. Benefits include top-tier compensation, health insurance, paid parental leave, and opportunities for team engagement.