Pre-Training Data Engineer: Scalable ML Data Pipelines
Cohere
City Of London
On-site
GBP 60,000 - 90,000
Full time
14 days+
Get more replies from employers
Send a job-specific resume in minutes.
Start fresh or import an existing resume
Benefits offered by this job
Open and inclusive culture
Weekly lunch stipend
Full health and dental benefits
6 weeks of vacation
Job summary
A leading AI technology company in London seeks a Machine Learning Engineer (Pre-Training Data) to design and implement data pipelines for advanced language models. The role requires strong software engineering expertise in Python and familiarity with data processing frameworks. Join a diverse team to transform data into critical components of innovative AI systems. Employees benefit from a remote-friendly environment and generous perks.
Qualifications
Strong software engineering skills, proficient in Python.
Familiarity with data processing frameworks like Apache Spark.
Experience working with large-scale datasets across multiple domains.
Responsibilities
Design and build scalable data pipelines for diverse datasets.
Conduct data ablations to assess and improve data quality.
Collaborate with teams to ensure data meets model demands.
Skills
Software engineering skills
Proficiency in Python
Data processing frameworks (Apache Spark, Pandas)
Experience with large-scale datasets
Knowledge of data quality assessment techniques
Job description
A leading AI technology company in London seeks a Machine Learning Engineer (Pre-Training Data) to design and implement data pipelines for advanced language models. The role requires strong software engineering expertise in Python and familiarity with data processing frameworks. Join a diverse team to transform data into critical components of innovative AI systems. Employees benefit from a remote-friendly environment and generous perks.