NLP Data Engineer: Scalable Pipelines & LLM Data

Institute of Foundation Models

United States

On-site

USD 120,000 - 180,000

Full time

14 days+
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Benefits offered by this job

Comprehensive medical, dental, and视觉
Bonus
401K Plan
Generous paid time off, sick leave and
Paid Parental Leave
Employee Assistance Program
Life insurance and disability

Job summary

Institute of Foundation Models in the United States seeks a Data Engineer focused on NLP and large-scale data processing to rapidly gather, curate, and prepare high-quality datasets for cutting-edge NLP research. You will enable researchers by delivering essential data through scalable pipelines, web crawling, and LLM refinement using Python and related technologies.

Collaborate with researchers and engineers to ensure data meets quality criteria, document collection methodologies and pipeline

Qualifications

  • Bachelor's degree in Computer Science, Data Science, Engineering, or a related technical field.
  • Experience in data engineering, data processing, and automation using Python.
  • Strong understanding of data structures, databases, SQL, and performance optimization.
  • Experience with cloud infrastructure and distributed data processing frameworks (AWS, Spark, Kafka, Kubernetes).
  • Excellent problem-solving and collaboration skills.

Responsibilities

  • Rapidly collect, curate, and preprocess datasets per NLPresearchers' specs, delivering data within tight timelines.
  • Develop and maintain efficient web crawling solutions, APIs, and automated workflows to continuously improve data collection processes.
  • Refine and evaluate outputs from Large Language Models (LLMs) to generate structured datasets suitable for model training and benchmarking.
  • Implement scalable data pipelines, ensuring efficient data processing, storage, retrieval, and distribution to research teams.
  • Collaborate closely with researchers and engineers to ensure collected data meets specified quality and relevance criteria.
  • Document data collection methodologies, dataset characteristics, and pipeline architecture clearly and effectively.
  • Engage with peer teams and participate in technical reviews to uphold best practices and data quality standards.
  • Represent MBZUAI at industry and research forums, showcasing technical capabilities in large-scale data processing and AI data infrastructure.

Skills

Python
Data engineering
Web crawling
Cloud infrastructure
SQL
Distributed data processing

Education

Bachelor's degree in Computer Science, Data Science, Engineering, or related field
Master's degree or PhD preferred

Tools

AWS
Spark
Kafka
Kubernetes
Airflow

Job description

Institute of Foundation Models in the United States seeks a Data Engineer focused on NLP and large-scale data processing to rapidly gather, curate, and prepare high-quality datasets for cutting-edge NLP research. You will enable researchers by delivering essential data through scalable pipelines, web crawling, and LLM refinement using Python and related technologies.

Collaborate with researchers and engineers to ensure data meets quality criteria, document collection methodologies and pipeline

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

NLP Data Engineer for Foundation Models
NLP Data Engineer for Foundation Models

Institute of Foundation Models • Sunnyvale (CA)

On-site
USD 150,000 - 450,000
Comprehensive medical, dental, and vision benefits
401K Plan
Generous paid time off
+2
Senior ML Engineer: Data Pipelines & Scalable AI Systems
Senior ML Engineer: Data Pipelines & Scalable AI Systems

NLP PEOPLE • Dublin (CA)

On-site
USD 120,000 - 150,000
GenAI NLP Data Engineer – Python, LLMs, Pipelines (Contract)
GenAI NLP Data Engineer – Python, LLMs, Pipelines (Contract)

Blue Chip Talent • Charlotte (NC)

Hybrid
USD 80,000 - 120,000
LLM Automation Engineer for Scalable AI Pipelines
LLM Automation Engineer for Scalable AI Pipelines

Basis Research Institute • United States

Remote
USD 80,000 - 120,000
Member of Technical Staff, Data
Member of Technical Staff, Data

Inception • San Francisco (CA)

On-site
USD 180,000 - 240,000
NLP Data Scientist – LLM
NLP Data Scientist – LLM

Yakamconsulting • Northern (KY)

Hybrid
USD 120,000 - 190,000
Member of Technical Staff, Data Infrastructure
Member of Technical Staff, Data Infrastructure

Inception • San Francisco (CA)

On-site
USD 140,000 - 190,000
Remote LLM & Data Pipeline Engineer
Remote LLM & Data Pipeline Engineer

Crossing Hurdles • United States

On-site
USD 110,000 - 180,000
Data Infrastructure Engineer — Scalable ML Data Pipelines | Flexible Hours
Data Infrastructure Engineer — Scalable ML Data Pipelines | Flexible Hours

JobCubby • San Francisco (CA), Northern (KY)

Hybrid
USD 500,000 - 850,000
Visa sponsorship
Office in San Francisco
Flexible working hours
Staff Data Infrastructure Engineer: Scale AI Pipelines
Staff Data Infrastructure Engineer: Scale AI Pipelines

Inception • San Francisco (CA)

On-site
USD 140,000 - 190,000