NLP Engineer

HiredBuddy

New York (NY)

On-site

USD 70,000 - 120,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Hands-on experience with LLM datasets
Mentorship from senior AI engineers
Competitive salary

Job summary

HiredBuddy is seeking a Junior to Intermediate NLP Engineer in New York to develop text preprocessing pipelines and design tokenization workflows for LLM fine-tuning. You will automate quality evaluation for text annotations and optimize RLHF datasets for client ML teams.

You will work on transforming raw annotation outputs into production-ready dataset formats, leveraging Python and leading NLP libraries, with mentorship from senior engineers.

Qualifications

  • 1–3 years of hands-on experience in Natural Language Processing, Computational Linguistics, or Machine Learning Data Engineering.
  • Proficiency in Python and essential NLP libraries (Hugging Face Transformers, spaCy, NLTK, LangChain/LlamaIndex).
  • Deep familiarity with handling text data formats (JSONL, CSV, Parquet) and regular expressions (Regex).
  • Solid understanding of modern NLP architectures (Transformers, Tokenization, Embeddings, LLM fine-tuning concepts).
  • Experience with Git, Docker, REST APIs, and SQL.

Responsibilities

  • Build and maintain automated text data pipelines to clean, tokenize, format, and validate large text datasets (JSON, JSONL, Parquet).
  • Develop programmatic validation scripts for NER, intent classification, sentiment analysis, and multi-turn prompt-response labeling.
  • Implement statistical metrics to evaluate Inter-Annotator Agreement (IAA) and semantic consistency across datasets.
  • Benchmark AI model outputs using evaluation frameworks (ROUGE, BLEU, RAG) and support RLHF workflows.
  • Convert raw annotation outputs into structured dataset formats for fine-tuning transformer models.

Skills

Python
NLP Concepts
Regular Expressions

Tools

Hugging Face Transformers
spaCy
NLTK
LangChain/LlamaIndex
Git
Docker
SQL

Job description

About the Role

We are an AI data annotation and labeling contracting company delivering high-quality text, conversation, and multilingual datasets for Large Language Models (LLMs) and NLP applications. We are seeking a detail-oriented Junior to Intermediate NLP Engineer to build text pre-processing pipelines, design tokenization and parsing workflows, automate quality evaluation for text annotations, and optimize RLHF datasets for client ML teams.

Key Responsibilities
  • Text Data Pipelines: Build and maintain automated pipelines to clean, tokenize, format, and validate large text datasets (JSON, JSONL, Parquet) for LLM fine-tuning and evaluation.
  • Annotation Guidelines & Automation: Develop programmatic validation scripts for named entity recognition (NER), intent classification, sentiment analysis, and multi-turn prompt-response labeling.
  • Quality Assurance & IAA: Implement statistical metrics to evaluate Inter-Annotator Agreement (IAA) and semantic consistency across complex linguistic datasets.
  • LLM & Prompt Evaluation: Benchmark AI model outputs using evaluation frameworks (e.g., ROUGE, BLEU, RAG metrics) and support Reinforcement Learning from Human Feedback (RLHF) workflows.
  • Client Dataset Delivery: Convert raw annotation outputs into structured dataset formats ready for fine-tuning transformer models.
Qualifications & Skills
  • Experience: 1-3 years of hands-on experience in Natural Language Processing, Computational Linguistics, or Machine Learning Data Engineering.
  • Core Programming: Proficiency in Python and essential NLP libraries (Hugging Face Transformers, spaCy, NLTK, LangChain/LlamaIndex).
  • Data Handling: Deep familiarity with handling text data formats (JSONL, CSV, Parquet) and regular expressions (Regex).
  • ML Foundations: Solid understanding of modern NLP architectures (Transformers, Tokenization, Embeddings, LLM fine-tuning concepts).
  • DevOps Basics: Experience with Git, Docker, REST APIs, and SQL.
What We Offer
  • Direct hands-on experience building text datasets for industry-leading Large Language Models.
  • Mentorship from senior AI engineers and clear career progression.
  • Competitive salary package.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

NLP Engineer: Build LLM Data Pipelines & RLHF
NLP Engineer: Build LLM Data Pipelines & RLHF

HiredBuddy • New York (NY)

On-site
USD 70,000 - 120,000
Hands-on experience with LLM datasets
Mentorship from senior AI engineers
Competitive salary
Senior Data Scientist - NLP/LLM Specialist
Senior Data Scientist - NLP/LLM Specialist

Scismic • San Diego (CA)

On-site
USD 155,000 - 240,000
Unlimited PTO
401k program
Month-long sabbatical
+2
Linguist II at Infotech Sourcing United States
Linguist II at Infotech Sourcing United States

Fairweather, LLC • United States

On-site
USD 85,000 - 120,000
AI Engineer
AI Engineer

LeoTechnologies • Boca Raton (FL), Northern (KY)

Hybrid
USD 140,000 - 190,000
AI/ML Engineer
AI/ML Engineer

RiskForce • Northern (KY)

Hybrid
USD 120,000 - 155,000
Senior Data Engineer / AI ML Engineer with Python, AI/ML & LLMs
Senior Data Engineer / AI ML Engineer with Python, AI/ML & LLMs

Ampcus Inc • Reston (VA)

On-site
USD 100,000 - 130,000
Junior AI/ML Engineer
Junior AI/ML Engineer

Drivenintelligence • United States

Hybrid
USD 75,000 - 120,000
AWS AI/ML Engineer
AWS AI/ML Engineer

Tanisha Systems, Inc • Cincinnati (OH)

On-site
USD 90,000 - 150,000
LLM Training & Model Development Engineer
LLM Training & Model Development Engineer

InOpTra Digital • United States

Remote
USD 90,000 - 120,000
Competitive salary
Opportunity for remote work
Health benefits
AI Developer
AI Developer

Salvo Software LLC • Northern (KY)

Hybrid
USD 120,000 - 190,000