Senior Data Scientist

LexisNexis

United States

On-site

USD 120,000 - 160,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

A global provider of information-based analytics is seeking a candidate for the design and continuous evolution of a multimodal document understanding platform. The selected candidate will lead the multimodal model strategy, focusing on tasks such as complex layout analysis and data extraction. The role requires a strong foundation in machine learning and data science, with a focus on multimodal representations. A Master’s degree and at least 5 years of experience in relevant fields are essential. This position offers an opportunity to innovate and deliver effective solutions in the legal market.

Qualifications

  • 5+ years of hands-on machine learning or data science experience.
  • Proven delivery experience in multimodal document understanding.
  • Strong foundation in statistical analysis and experimental design.

Responsibilities

  • Design and iterate the multimodal document parsing pipeline.
  • Build and optimize multi-agent collaboration mechanisms.
  • Define model selection and composition strategies.

Skills

Machine learning
Data science
Multimodal representations
Cross-modal alignment
Statistical analysis

Education

Master’s degree in Statistics or related field

Tools

Python
Deep learning frameworks

Job description

LexisNexis Legal & Professional, which serves customers in more than 150 countries with 11,800 employees worldwide, is part of RELX (www.relx.com), a global provider of information-based analytics and decision tools for professional and business customers. Our company has been a long-time leader in deploying AI and advanced technologies to the legal market to improve productivity and transform the overall business and practice of law, deploying ethical and powerful generative AI solutions with a flexible, multi-model approach that prioritizes using the best model from today’s top model creators for each individual legal use case.

The company employs over 2,000 technologists, data scientists, and experts to develop, test, and validate solutions in line with RELX Responsible AI Principles (https://stories.relx.com/responsible-ai-principles/index.html).

About the Role:

This role is responsible for the end‑to‑end design and continuous evolution of a multimodal document understanding and structured data extraction platform: complex PDF / scanned page layout analysis, semantic extraction, structural reconstruction, quality validation, and business integration. Leads multimodal model strategy (vision + language + layout) and multi‑agent collaboration (task decomposition, verification, conflict reconciliation, feedback loops) and plans future customized training and ongoing optimization of models.

Key Responsibilities

  • Design and iterate the multimodal document parsing pipeline: layout / structural modeling, semantic extraction, cross‑modal alignment, structural reconstruction.
  • Build and optimize a multi‑agent collaboration mechanism: task splitting, parallel / sequential scheduling, peer review, iterative quality improvement loops.
  • Define model selection / composition / routing strategies (dynamic dispatch by document type, structural patterns, quality signals).
  • Plan and execute model fine‑tuning, domain adaptation, continual learning, active learning, and data feedback loops.
  • Establish end‑to‑end metrics: extraction accuracy, structural consistency, agent collaboration effectiveness, latency, stability, and cost.
  • Build quality assurance and risk controls: drift & anomaly monitoring, confidence estimation, fallback strategies, alignment / compliance checks.
  • Drive mapping and consistency between agent / model outputs and business knowledge field standards.

Required Qualifications

  • Education: Master’s degree or above in a quantitative or technical field (Statistics, Computer Science, Mathematics, Data Science, etc.).
  • Experience:5+ years of hands‑on machine learning / data science experience. Proven delivery experience in multimodal (vision + text) or complex document understanding. Practical cases of orchestrating agents (or modular processing logic) in production workflows.
  • Capabilities: Solid foundation in machine learning / deep learning fundamentals, multimodal representations, and cross‑modal alignment concepts. Deep understanding of core principles and common algorithms for multimodal large models: cross‑modal attention & representation alignment, vision/text embedding fusion, hierarchical & layout structure modeling, instruction & contrastive paradigms, long‑context and retrieval‑augmented mechanisms, evaluation and failure mode dissection. Familiar with classic image and signal processing methods: edge & contour detection, filtering & denoising, morphological operations, segmentation & keypoint feature extraction, frequency / time‑frequency analysis, image enhancement & quality assessment; understands trade‑offs and complementarity with deep features. Knowledge of multi‑agent collaboration patterns: role assignment, task routing, feedback loops, redundancy & cross‑checks. Strong in statistical analysis & experimental design: hypothesis testing, factorial design, power analysis, A/B and multivariate evaluation. Able to decompose complex problems and build metric‑driven optimization paths. Rigorous in data quality & error analysis; rapid bottleneck identification. Ability to translate research pseudo‑code into maintainable, testable Python modules with benchmarking & regression harnesses.

Preferred / Nice to Have

  • Designed customization / fine‑tuning of multimodal foundation models, representation learning, or structural understanding subsystems.
  • Built an agent orchestration platform: task decomposition, iterative self‑checks, consensus or voting mechanisms.
  • Experience solving robustness & generalization challenges in large‑scale long documents / heterogeneous layouts.
  • Demonstrated results in cost optimization (model pruning, parameter‑efficient tuning, inference acceleration) or adaptive load scheduling.
  • Publications / patents or open‑source contributions.
  • Demonstrated Python systems optimization (e.g., custom Cython / CUDA kernels, vectorization replacing Python loops, latency reductions in inference pipelines).
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Data Scientist - NLP/LLM Specialist
Senior Data Scientist - NLP/LLM Specialist

Scismic • San Diego (CA)

On-site
USD 155,000 - 240,000
Unlimited PTO
401k program
Month-long sabbatical
+2
Senior Data Scientist / Machine Learning Scientist
Senior Data Scientist / Machine Learning Scientist

NLP PEOPLE • Raleigh (NC)

On-site
USD 100,000 - 140,000
Applied Scientist, Document Understanding
Applied Scientist, Document Understanding

TempWorks Software Incorporated • Frisco (TX)

On-site
USD 120,000 - 160,000
(Team and People) Lead Data Scientist
(Team and People) Lead Data Scientist

Lexis Nexis • Raleigh (NC)

On-site
USD 100,000 - 140,000
Disability insurance
Dependent care and commuter spending accounts
Life and accident insurance
+2
Principal Machine Learning Engineer I
Principal Machine Learning Engineer I

RELX • Raleigh (NC)

On-site
USD 136,000 - 253,000
Country-specific benefits
Senior Data Engineer — AI/ML Data Platforms
Senior Data Engineer — AI/ML Data Platforms

Akoncagua AI • Lakeland (FL)

On-site
USD 120,000 - 180,000
Machine Learning Engineer Lead
Machine Learning Engineer Lead

RELX • Raleigh (NC)

On-site
USD 115,000 - 192,000
Annual incentive bonus
Country-specific benefits
Lead Data Scientist
Lead Data Scientist

LexisNexis • Raleigh (NC)

On-site
USD 104,000 - 175,000
Lead Data Scientist
Lead Data Scientist

RELX • Raleigh (NC)

On-site
USD 105,000 - 175,000
Annual incentive bonus
Country-specific benefits
Principal Machine Learning Engineer I
Principal Machine Learning Engineer I

LexisNexis • Raleigh (NC)

On-site
USD 136,000 - 253,000