Applied Scientist, Document Understanding

TempWorks Software Incorporated

Frisco (TX)

On-site

USD 120,000 - 160,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

TempWorks Software Incorporated is seeking a talented AI Engineer specialized in developing systems for legal document understanding. Ideal candidates will hold a PhD or Master's in a related field, bringing at least 3 years of production-based experience. You will design semantic chunking models and build document enrichment systems that categorize legal documents based on customer-defined taxonomies.

The role requires proficiency in Python, PyTorch, and a solid publication record. Join a cutting-edge team and contribute to impactful projects in the legal AI space.

Qualifications

  • 3+ years of post-degree industry experience shipping document understanding and information extraction into production.
  • Publication record at ACL, EMNLP, or similar venues.
  • Hands-on production depth across model development and evaluation.

Responsibilities

  • Design and deploy semantic chunking models for legal documents.
  • Build document enrichment systems for classification and metadata extraction.
  • Drive independent technical decisions on classification and extraction methods.

Skills

Document understanding
Information extraction
Knowledge graph systems
Production Python
PyTorch
Hugging Face Transformers
DeepSpeed
Hierarchical document classification
Entity recognition
Knowledge distillation

Education

PhD or Master's in Computer Science, AI, NLP or related field

Tools

AzureML
AWS SageMaker

Job description

You hold a PhD or Master's in Computer Science, AI, NLP, or a related field, with 3+ years of post-degree industry experience shipping document understanding, information extraction, or knowledge graph systems into production. You have hands‑on depth across model development, distillation, evaluation, and deployment. You work independently and measure success by what ships and performs in production.

What You’ll Do
  • Design and deploy semantic chunking models for lengthy, non-uniformly structured legal documents with adjustable granularity across use cases.
  • Build document enrichment systems that classify documents according to legal and customer-defined taxonomies and extract rich metadata.
  • Develop LLM-based knowledge graph construction pipelines that extract and link citations, entities, and legal concepts across diverse legal content.
  • Build scalable synthetic data generation systems for model training, multi-hop query simulation, and hallucination-free answer generation.
  • Apply knowledge distillation techniques to compress large models into latency-constrained, production-ready SLMs.
  • Design evaluation frameworks — component-level and end-to-end — using expert annotation and synthetic data.
  • Drive independent technical decisions on chunking strategy, classification approach, knowledge extraction methods, and multi-document reasoning architecture.
  • Partner with engineering on delivery, reliability, and scale across multiple product lines.
  • Contribute to published research at venues such as ACL, EMNLP, ICLR, NeurIPS, SIGIR, and KDD, and to intellectual property.
Required Qualifications
  • PhD or Master's in Computer Science, AI, NLP, or a related field.
  • 3+ years of post-degree industry experience shipping document understanding, information extraction, or knowledge graph systems into production — not research-only experience.
  • Publication record at ACL, EMNLP, ICLR, NeurIPS, SIGIR, KDD, or equivalent.
  • Production Python and experience with PyTorch, Hugging Face Transformers, and DeepSpeed.
  • Hands‑on production depth required in:
  • Document layout analysis and semantic chunking beyond fixed-size or paragraph-based methods.
  • Hierarchical, multi-label document classification with domain-specific and customer-defined schemas.
  • Entity recognition and linking, relation extraction, citation parsing, and knowledge graph construction from unstructured text.
  • LLM-based information extraction, few-shot and multi-task learning, and post-training.
  • Knowledge distillation, model compression, and SLM deployment under latency constraints.
  • Synthetic data generation for NLP: query-answer generation with verification and scalable data augmentation.
  • Annotation workflow design and evaluation framework development for document understanding tasks.
Preferred Qualifications
  • Legal document understanding, legal information extraction, or legal AI applications.
  • Complex document structures common in legal content: nested hierarchies, cross-references, non-uniform formatting, and embedded elements.
  • Retrieval, QA, or analysis systems over large document collections.
  • Knowledge graph frameworks for legal or enterprise applications.
  • RAG and agentic workflows for enterprise knowledge systems.
  • AzureML or AWS SageMaker.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Lead Applied Scientist, Document Understanding
Lead Applied Scientist, Document Understanding

Refinitiv • New York (NY)

On-site
USD 180,000 - 260,000
Flexible vacation
Mental Health Days
Headspace app
+4
Senior Applied Scientist, Document Understanding
Senior Applied Scientist, Document Understanding

Thomson Reuters • Ann Arbor (MI)

On-site
USD 140,000 - 210,000
Flexible work options
Tuition reimbursement
Mental health and well‑being programs
+1
Senior Applied Scientist, Document Understanding
Senior Applied Scientist, Document Understanding

Refinitiv • New York (NY)

Hybrid
USD 127,400 - 236,600
Flexible vacation
Mental Health Days
Tuition reimbursement
+2
Senior AI Scientist: Document Understanding & KG (Remote)
Senior AI Scientist: Document Understanding & KG (Remote)

Thomson Reuters • Ann Arbor (MI)

On-site
USD 140,000 - 210,000
Flexible work options
Tuition reimbursement
Mental health and well‑being programs
+1
Lead Applied Scientist, Document Understanding (Legal AI)
Lead Applied Scientist, Document Understanding (Legal AI)

Refinitiv • New York (NY)

On-site
USD 180,000 - 260,000
Flexible vacation
Mental Health Days
Headspace app
+4
Applied Scientist, Document Understanding & Graphs
Applied Scientist, Document Understanding & Graphs

TempWorks Software Incorporated • Frisco (TX)

On-site
USD 120,000 - 160,000
Senior Data Scientist
Senior Data Scientist

LexisNexis • United States

On-site
USD 120,000 - 160,000
Senior Applied Scientist — Production-Scale Legal NLP
Senior Applied Scientist — Production-Scale Legal NLP

SupportFinity™ • New York (NY)

Hybrid
USD 127,000 - 237,000
Flexible vacation
Mental health days
Tuition reimbursement
+2
Senior Applied Scientist, Document Understanding
Senior Applied Scientist, Document Understanding

Thomson Reuters • Eagan (MN)

On-site
USD 127,000 - 237,000
Flex My Way: work from anywhere up to
Grow My Way: continuous learning
Two paid volunteer days off annually
+1
Senior Document Understanding Scientist — Remote-Ready
Senior Document Understanding Scientist — Remote-Ready

Thomson Reuters • Eagan (MN)

On-site
USD 127,000 - 237,000
Flex My Way: work from anywhere up to
Grow My Way: continuous learning
Two paid volunteer days off annually
+1