Data Engineer

Institute of Foundation Models

United States

On-site

USD 120,000 - 180,000

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

Comprehensive medical, dental, and视觉
Bonus
401K Plan
Generous paid time off, sick leave and
Paid Parental Leave
Employee Assistance Program
Life insurance and disability

Job summary

Institute of Foundation Models in the United States seeks a Data Engineer focused on NLP and large-scale data processing to rapidly gather, curate, and prepare high-quality datasets for cutting-edge NLP research. You will enable researchers by delivering essential data through scalable pipelines, web crawling, and LLM refinement using Python and related technologies.

Collaborate with researchers and engineers to ensure data meets quality criteria, document collection methodologies and pipeline

Qualifications

  • Bachelor's degree in Computer Science, Data Science, Engineering, or a related technical field.
  • Experience in data engineering, data processing, and automation using Python.
  • Strong understanding of data structures, databases, SQL, and performance optimization.
  • Experience with cloud infrastructure and distributed data processing frameworks (AWS, Spark, Kafka, Kubernetes).
  • Excellent problem-solving and collaboration skills.

Responsibilities

  • Rapidly collect, curate, and preprocess datasets per NLPresearchers' specs, delivering data within tight timelines.
  • Develop and maintain efficient web crawling solutions, APIs, and automated workflows to continuously improve data collection processes.
  • Refine and evaluate outputs from Large Language Models (LLMs) to generate structured datasets suitable for model training and benchmarking.
  • Implement scalable data pipelines, ensuring efficient data processing, storage, retrieval, and distribution to research teams.
  • Collaborate closely with researchers and engineers to ensure collected data meets specified quality and relevance criteria.
  • Document data collection methodologies, dataset characteristics, and pipeline architecture clearly and effectively.
  • Engage with peer teams and participate in technical reviews to uphold best practices and data quality standards.
  • Represent MBZUAI at industry and research forums, showcasing technical capabilities in large-scale data processing and AI data infrastructure.

Skills

Python
Data engineering
Web crawling
Cloud infrastructure
SQL
Distributed data processing

Education

Bachelor's degree in Computer Science, Data Science, Engineering, or related field
Master's degree or PhD preferred

Tools

AWS
Spark
Kafka
Kubernetes
Airflow

Job description

About the Institute of Foundation Models

We are a dedicated research lab for building, understanding, using, and risk-managing foundation models. Our mandate is to advance research, nurture the next generation of AI builders, and drive transformative contributions to a knowledge-driven economy.

As part of our team, you’ll have the opportunity to work on the core of cutting-edge foundation model training, alongside world-class researchers, data scientists, and engineers, tackling the most fundamental and impactful challenges in AI development. You will participate in the development of groundbreaking AI solutions that have the potential to reshape entire industries.Strategic and innovative problem-solving skills will be instrumental in establishing MBZUAI as a global hub forhigh-performance computing in deep learning, driving impactful discoveries that inspire the next generation of AIpioneers.

The Role

As a Data Engineer specializing in Natural Language Processing (NLP) and large-scale data processing, you will quickly and effectively gather, curate, and prepare high-quality datasets to support cutting‑edge NLP research. Your role will be instrumental in enabling researchers by delivering essential data through efficient and scalable engineering practices, including web crawling, LLM-generated content refinement, and robust data pipelines, primarily leveraging Python and related technologies.

Key Responsibilities
  • Rapidly collect, curate, and preprocess datasets based on detailed specifications provided by NLPresearchers,delivering data within tight timelines.
  • Develop and maintain efficient web crawling solutions, APIs, and automated workflows to continuously improve data collection processes.
  • Refine and evaluate outputs from Large Language Models (LLMs) to generate structured datasets suitable for model training and benchmarking.
  • Implement scalable data pipelines, ensuring efficient data processing, storage, retrieval, and distribution to research teams.
  • Collaborate closely with researchers and engineers to ensure collected data meets specified quality and relevance criteria.
  • Document data collection methodologies, dataset characteristics, and pipeline architecture clearly and effectively.
  • Engage with peer teams and participate in technical reviews to uphold best practices and data quality standards.
  • Represent MBZUAI at industry and research forums, showcasing technical capabilities in large-scale data processing and AI data infrastructure.
Academic Qualifications
  • Bachelor’s degree in Computer Science, Data Science, Engineering, or a related technical field required
  • Master’s degree or PhD degree or equivalent experience in Computer Science, Data Engineering, or related technical fields preferred.
Professional Experience - Required
  • Extensive experience in data engineering, data processing, and automation using Python.
  • Demonstrated proficiency in designing and deploying web crawling solutions, automated data extraction, and processing pipelines.
  • Strong understanding of data structures, algorithms, databases, SQL, and performance optimization.
  • Experience working with cloud infrastructure and distributed data processing frameworks (e.g., AWS, Spark, Kafka, Kubernetes).
  • Excellent problem-solving abilities, attention to detail, and the capability to rapidly address technical challenges.
  • Strong communication and collaboration skills with cross-functional teams.
Professional Experience - Preferred
  • Proven track record of supporting NLP or AI research teams with rapid and reliable data delivery.
  • Experience working with large language models, including evaluation, efficient inference, and prompt engineering.
  • Experience with refining outputs from large-scale AI models, such as LLM-generated data.
  • Contributions to open-source projects, coding competitions, or high visibility in coding communities (e.g., GitHub, Stack Overflow).
  • Familiarity with the latest advancements in NLP data processing and large language model technologies.

Visa Sponsorship

This position is eligible for visa sponsorship.

Benefits Include

  • Comprehensive medical, dental, and vision benefits
  • Bonus
  • 401K Plan
  • Generous paid time off, sick leave and holidays
  • Paid Parental Leave
  • Employee Assistance Program
  • Life insurance and disability
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Data Engineer
Data Engineer

Institute of Foundation Models • Sunnyvale (CA)

On-site
USD 150,000 - 450,000
Comprehensive medical, dental, and vision benefits
401K Plan
Generous paid time off
+2
Research Scientist - NLP
Research Scientist - NLP

Institute of Foundation Models • Sunnyvale (CA)

On-site
USD 150,000 - 450,000
Comprehensive medical, dental, and vision benefits
Bonus
401K Plan
+4
Distributed Machine Learning Engineer
Distributed Machine Learning Engineer

Institute of Foundation Models • Sunnyvale (CA)

On-site
USD 150,000 - 450,000
Comprehensive medical, dental, and vision benefits
Bonus
401K Plan
+4
Research Scientist – World Modeling, Data
Research Scientist – World Modeling, Data

Institute of Foundation Models • Sunnyvale (CA)

On-site
USD 150,000 - 400,000
Medical benefits
Dental/vision
401K
+5
Machine Learning Engineer – World Model
Machine Learning Engineer – World Model

Institute of Foundation Models • Sunnyvale (CA)

On-site
USD 150,000 - 450,000
Comprehensive medical, dental, and vision benefits
Bonus
401K Plan
+4
Research Scientist – World Modeling, Data
Research Scientist – World Modeling, Data

Ifm Us • Sunnyvale (CA)

On-site
USD 140,000 - 210,000
Comprehensive medical benefits
Dental benefits
Vision benefits
+6
Research Scientist - Reinforcement Learning
Research Scientist - Reinforcement Learning

Institute of Foundation Models • Sunnyvale (CA)

On-site
USD 150,000 - 450,000
Comprehensive medical, dental, and vision benefits
401K Plan
Generous paid time off
+2
Research Scientist - Distributed Machine Learning
Research Scientist - Distributed Machine Learning

Institute of Foundation Models • Sunnyvale (CA)

On-site
USD 300,000 - 600,000
Comprehensive medical, dental, and vision benefits
401K Plan
Generous paid time off
+3
AI Research Internship - LLM
AI Research Internship - LLM

Institute of Foundation Models • Sunnyvale (CA)

On-site
USD 100,000 - 140,000
Machine Learning Infrastructure Engineer
Machine Learning Infrastructure Engineer

Institute of Foundation Models • Sunnyvale (CA)

On-site
USD 150,000 - 450,000
Comprehensive medical, dental, and vision
401(k) program
Generous PTO
+3