About the Role
We are seeking aSenior Machine Learning Engineer AI-Assisted Data Annotationto own the automated annotation track within ABBYY’sDocument AI Data team.
This role sits at the intersection oflarge model capabilities and production data engineering, leveraging LLMs and vision-language models to generate high-quality training data at scale. You will design and buildAI-assisted annotation pipelines, ensuring outputs areaccurate, measurable, and reliable for downstream model training.
This is an ideal role for engineers who combinedeep modelexpertisewith strong system-building instinctsand thrive in fast-moving, experimental environments.
Key Responsibilities
Technical Development & Innovation
- Design and implementAI-powered annotation pipelinesusing large models to generate ground truth labels at scale
- Develop and refineprompting strategies, few-shot examples, and fine-tuning approachesto improve accuracy and consistency
- Build systems forlabel verification, confidence scoring, and quality validation
- Evaluate which tasks are suitable forautomated annotation vs. human review, and define decision criteria
- Createevaluation frameworksto benchmark automated annotations against human-labeled data
- Continuously improve annotation quality using feedback from human review workflows
Project Ownership & Leadership
- Own the automated annotation trackend-to-end, from architecture through production monitoring
- Drive technical decisions acrossmodel selection, pipeline design, and validation strategies
- Define integration points withplatform infrastructure and model serving systems
- Collaborate with Data Operations to designhuman-in-the-loop workflowsfor efficient review
- Contribute to roadmap planning with Principal-level technical leadership
Infrastructure & Scale
- Build andoptimizelarge-scale inference pipelinesfor processing millions of documents
- Implement monitoring and alerting forquality degradation and system failures
- Design batching, caching, and fallback mechanisms to balancecost, throughput, and accuracy
- Collaborate with Platform teams onmodel serving, APIs, and infrastructure scaling
- Maintain clear documentation ofannotation strategies, metrics, and known limitations
Qualifications
Education & Experience
- MS or PhD in Computer Science, Engineering, Mathematics, or related field
- 5+ years of experience inMachine Learning / AI, with focus on:
- Large Language Models (LLMs)
- Vision-Language Models (VLMs)
- Data annotation or labeling systems
- Demonstrated success usinglarge AI models to automate annotation at production scale
- Strong background inevaluation design and quality measurement
Technical Expertise
- DeepexpertiseinLLMs and VLMs, including prompting, instruction tuning, and output evaluation
- Strong understanding ofdocument understanding tasks(classification, extraction, layout analysis, semantic parsing)
- Experience designinglabel quality metrics, confidence scoring, and agreement analysis
- Strong programming skills inPythonandproficiencywithPyTorchor similar frameworks
- Experience withlarge-scale inference pipelines and model serving systems
- Familiarity withhuman-in-the-loop annotation systemsand automation trade-offs