Overview
Thomson Reuters | Posted Mar 6 | Full-time | New York | Advanced (5-10 yrs)
This role sits within the applied science function. You will design, build, and deploy document understanding systems that directly power Westlaw, PracticalLaw, and CoCounsel. The problems are real, the scale is large, and the expectation is shipped, reliable, measurable impact.
You will work across semantic chunking, document enrichment, knowledge graph construction, and synthetic data generation for complex legal, tax, and accounting content. Multiple product teams depend on what this function delivers.
About You
You hold a PhD or Master\'s in Computer Science, AI, NLP, or a related field, with 5+ years of post-degree industry experience taking NLP and document understanding systems from development to production at scale. You have hands-on depth across model development, distillation, evaluation, and deployment. You publish, you work independently, lead through influence in an applied research setting, and measure success by what ships and performs in production.
What You'll Do
- Design and deploy semantic chunking models for lengthy, non-uniformly structured legal documents with adjustable granularity across use cases
- Build document enrichment systems using legal and customer-defined taxonomies
- Develop LLM-based knowledge graph construction pipelines that extract and link citations, entities, and legal concepts across diverse legal content
- Build scalable synthetic data generation systems for model training, multi-hop query simulation, and hallucination-free answer generation
- Apply knowledge distillation techniques to compress large models into latency-constrained, production-ready SLMs
- Design evaluation frameworks — component-level and end-to-end — using expert annotation and synthetic data
- Drive technical decisions on architecture, chunking strategy, classification approach, and knowledge extraction methods
- Partner with engineering on delivery, reliability, and scale across multiple product lines
- Contribute to published research at venues such as ACL, EMNLP, ICLR, NeurIPS, SIGIR, and KDD, and to intellectual property
Required Qualifications
- PhD or Master\'s in Computer Science, AI, NLP, or a related field
- 5+ years of post-degree industry experience shipping document understanding, information extraction, or knowledge graph systems into production — not research-only experience
- Publications at ACL, EMNLP, ICLR, NeurIPS, SIGIR, KDD, or equivalent
- Experience leading through influence in an applied research setting
- Production Python and experience with PyTorch, Hugging Face Transformers, and DeepSpeed
Hands-on Production Depth
- Document layout analysis and semantic chunking beyond fixed-size or paragraph-based methods
- Hierarchical, multi-label document classification with domain-specific and customer-defined schemas
- Entity recognition and linking, relation extraction, citation parsing, and knowledge graph construction from unstructured text
- LLM-based information extraction, few-shot and multi-task learning, and post-training
- Knowledge distillation, model compression, and SLM deployment under latency constraints
- Synthetic data generation and annotation workflow design
- End-to-end evaluation framework design for document understanding
Preferred Qualifications
- Legal document understanding, legal IE, or legal AI experience
- Complex document structures: nested hierarchies, cross-references, non-uniform formatting
- Retrieval or QA systems over large document collections
- RAG and agentic workflows in enterprise settings
- Knowledge graph frameworks for legal or enterprise applications
- AzureML or AWS SageMaker
What’s in It For You?
- Flexibility & Work-Life Balance: Flex My Way is a set of supportive workplace policies designed to help manage personal and professional responsibilities, including work from anywhere for up to 8 weeks per year.
- Career Development and Growth: By fostering a culture of continuous learning and skill development, we prepare our talent to tackle tomorrow’s challenges and deliver real-world solutions.
- Industry Competitive Benefits: Comprehensive benefit plans including flexible vacation, mental health days, Headspace, retirement savings, tuition reimbursement, incentive programs, and wellbeing resources.
- Culture: Global recognition for inclusion and belonging, flexibility, work-life balance, and more. Our values: Obsess over our Customers, Compete to Win, Challenge (Y)our Thinking, Act Fast / Learn Fast, and Stronger Together.
- Social Impact: Two paid volunteer days annually and opportunities for pro-bono and ESG initiatives.
- Making a Real-World Impact: We help customers pursue justice, truth, and transparency, upholding the rule of law and providing trusted information globally.
Thomson Reuters complies with local laws that require upfront disclosure of the expected pay range for a position. The base compensation range varies across locations. For eligible US locations, the base range is $127,400 USD - $236,600 USD. Base pay is one part of a Total Reward program with benefits and wellbeing programs. This role may be eligible for an Annual Bonus based on performance.
This job posting will close.
About Us
Thomson Reuters informs the way forward by bringing together trusted content and technology for professionals across legal, tax, accounting, compliance, government, and media. Reuters, part of Thomson Reuters, is a world-leading provider of trusted journalism and news. We are powered by 26,000 employees across 70+ countries, and we value objectivity, accuracy, fairness, and transparency. We are an Equal Employment Opportunity Employer. We make reasonable accommodations for qualified individuals with disabilities and for sincerely held religious beliefs in accordance with applicable law. Learn more on how to protect yourself from fraudulent job postings here. More information about Thomson Reuters can be found on thomsonreuters.com