Senior Machine Learning Engineer - Intelligent Document Processing / Production AI Systems

Pantheon-Data

Reston (VA)

On-site

USD 140,000 - 200,000

Full time

8 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Pantheon Data is seeking a Senior Machine Learning Engineer in Reston, VA with deep experience in NLP, OCR, and LLM-based document understanding. You will design production AI components, build data pipelines for unstructured sources, and deploy services in cloud/container environments, while mentoring peers and collaborating with cross-functional teams.

This role emphasizes building observable, scalable systems that move beyond notebook work, delivering reliable solutions for complex documents

Qualifications

  • Bachelor's degree in Computer Science, Engineering, or a related technical field from an ABET accredited university.
  • 5+ years of hands-on ML/AI/DS engineering experience; plus 5 years in software engineering.
  • Experience shipping production ML/AI pipelines, APIs, or services.
  • Strong Python engineering skills with readable, maintainable code.
  • Hands-on NLP, OCR, CV, LLMs, embeddings, or retrieval techniques.

Responsibilities

  • Design and build ML/AI components for IDP and Generative AI systems.
  • Build data pipelines for unstructured data (PDFs, forms, tables, images).
  • Develop NLP, OCR, CV, and LLM-based approaches for document understanding.
  • Create APIs, internal tools, dashboards, validation workflows.
  • Deploy and support ML/AI services in cloud or containers.
  • Mentor engineers and communicate with stakeholders.

Skills

Python
NLP
OCR
LLMs
MLOps
APIs
Cloud
Data pipelines
Debugging

Education

ABET accredited degree
Bachelor's in CS/Engineering

Tools

PyTorch
TensorFlow
scikit-learn
Hugging Face
Git
Docker
Kubernetes

Job description

Company Overview

Pantheon Data (a Kenific Holding company) is a private, small business based in the Washington, DC, area. Pantheon Data was founded in 2011, initially providing acquisition and supply chain management services to the US Coast Guard. Our service offerings have grown in the past ten years, including infrastructure resiliency, contact center operations, information technology, software engineering, program management, strategic communications, engineering, and cybersecurity. We have also grown our customer base to include commercial clients. The company has used this experience to expand our service offerings to other agencies within the Department of Homeland Security (DHS), the Department of Defense (DoD), and other Federal Civilian Agencies.

Position Overview

We are seeking a Senior Machine Learning Engineer to help build production-oriented AI and Intelligent Document Processing (IDP) systems. This role is for a hands-on engineer who can move beyond experiments and build working software that ingests, processes, analyzes, retrieves, and explains information from complex unstructured and semi-structured sources.

The ideal candidate has real depth in machine learning, NLP, OCR, computer vision, LLMs, and retrieval-based systems, but also has the broader engineering judgment to understand the system around the model: data pipelines, APIs, databases, cloud infrastructure, containers, testing, evaluation, observability, and production failure modes.

This is not a notebook-only, prompt-only, or research-only role. A successful candidate should be prepared to discuss specific systems they have built, including the data flow, model or inference architecture, deployment approach, evaluation strategy, failure modes, and what they personally implemented.

What This Role Will Work On
  • Design and build AI/ML capabilities for Intelligent Document Processing, including OCR post-processing, document parsing, NLP/LLM extraction, semantic search, retrieval, evidence grounding, and structured data generation.
  • Develop production-quality Python services, pipelines, and tooling that turn messy source documents into reliable, traceable, usable information.
  • Work across the full lifecycle of AI systems: data ingestion, preprocessing, model or LLM integration, evaluation, deployment, monitoring, and iterative improvement.
  • Build and improve systems that process PDFs, scanned documents, tables, forms, drawings, images, technical manuals, and other complex document types.
  • Collaborate with software engineers, data engineers, cloud engineers, product leads, customers, and leadership to turn ambiguous technical problems into working solutions.
  • Make practical engineering decisions about when to use deterministic logic, classical NLP, OCR, embeddings, LLMs, fine-tuned models, or human review workflows.
  • Help establish engineering standards for evaluation, reproducibility, model behavior, data quality, traceability, and responsible use of AI in customer-facing systems.
Responsibilities
  • Design, implement, and maintain ML/AI software components for IDP and Generative AI systems.
  • Build data pipelines for unstructured and semi-structured data, including document ingestion, extraction, cleaning, enrichment, validation, and storage.
  • Develop and evaluate NLP, OCR, computer vision, embedding, retrieval, and LLM-based approaches for document understanding use cases.
  • Create APIs, internal tools, review interfaces, dashboards, or validation workflows that allow engineers and users to inspect, correct, and trust system output.
  • Contribute production-quality code with clear structure, tests, logging, error handling, and documentation.
  • Deploy and support ML/AI services in cloud or containerized environments, including model serving, batch processing, and workflow automation.
  • Design evaluation approaches for extraction quality, retrieval quality, model behavior, hallucination risk, and end-to-end system performance.
  • Troubleshoot system behavior across model output, data quality, retrieval, schema design, infrastructure, latency, cost, and user workflow issues.
  • Mentor other engineers and help raise the technical quality of the team.
  • Communicate clearly with both technical and non-technical stakeholders, including project managers, customers, and executive leadership.
Required Skills and Experience
  • Bachelor's degree in Computer Science, Engineering, or a related technical field from an ABET accredited university.
  • 5+ years of professional hands-on experience in machine learning engineering, AI engineering, data science engineering, or a closely related software engineering role. Plus an additional 5 years of experience in technology and software engineering.
  • Demonstrated experience building AI/ML systems beyond notebooks, prototypes, or demos. Candidates should have shipped or supported pipelines, services, APIs, inference endpoints, evaluation harnesses, or production-facing tools.
  • Strong Python engineering experience, including readable, maintainable code; debugging; testing; packaging; and integration with other systems.
  • Hands-on experience with NLP, OCR, computer vision, LLMs, embeddings, semantic search, RAG, or other document-understanding techniques.
  • Experience working with unstructured or semi-structured data such as PDFs, scanned documents, forms, tables, images, logs, contracts, technical manuals, or engineering documentation.
  • Ability to design and reason about end-to-end data flow: source data, preprocessing, transformation, model/inference step, persistence, API/service boundary, evaluation, and user-facing output.
  • Familiarity with common ML frameworks and tooling such as PyTorch, TensorFlow, scikit-learn, Hugging Face, MLflow, or similar technologies.
  • Experience with databases and data stores, including SQL and at least one relational or non-relational data platform.
  • Experience using Git-based development workflows, code review, issue tracking, and team-based software delivery practices.
  • Clear written and verbal communication skills, including the ability to explain technical tradeoffs, limitations, and failure modes.
  • Ability to work effectively remotely in cross-functional teams.
  • Ability to meet deadlines and produce quality work.
  • Proficient in Microsoft Suite software including Outlook, Word, Excel, SharePoint, and PowerPoint.
Preferred Skills and Experience
  • Direct experience with Intelligent Document Processing, document AI, OCR pipelines, table extraction, form extraction, layout-aware processing, or evidence-grounded retrieval.
  • Experience building, deploying, or operating LLM-backed systems, including inference serving, prompt/version management, model evaluation, retrieval, observability, or cost/latency management.
  • Experience with cloud platforms such as AWS or Azure, including storage, compute, serverless, networking basics, IAM, monitoring, or managed ML/AI services.
  • Experience with containers and deployment workflows, including Docker, Kubernetes, CI/CD pipelines, automated tests, and environment promotion.
  • Experience building user-facing or internal tools such as validation interfaces, review workflows, dashboards, admin tools, or lightweight full-stack applications.
  • Experience with data engineering tools such as pandas, NumPy, Spark/PySpark, Databricks, Airflow, or similar workflow/data platforms.
  • Experience with observability, performance profiling, or debugging tools such as Grafana, CloudWatch, TensorBoard, tracing tools, GPU profiling tools, or application logs.
  • Experience with evaluation design, benchmarking, reproducibility, statistical analysis, error analysis, or human-in-the-loop validation.
  • Bachelor's or advanced degree in Computer Science, Engineering, Mathematics, Physics, Statistics, Data Science, or another technical discipline. Equivalent professional experience will also be considered.
  • Demonstrated ability to mentor junior developers or contribute to team technical direction.
Clearance Requirements

Applicants selected will be subject to a security investigation and may need to meet eligibility requirements. Secret Clearance is required for continued employment.

Work Location

Reston, VA - Remote

  • Our company prioritizes the benefits of flexibility and collaboration, whether that happens in person or remotely.
  • If the position is remote or hybrid, you may periodically work from a Pantheon Data office location or client site.
  • If this position is assigned to a Pantheon Data office location or client site, you'll work with colleagues and clients in person, as needed for specific client requirements.
Interview Requirement

Candidates who are local to the area should be prepared to participate in an in-person interview as part of the selection process. Candidates outside the local area may be considered for a virtual interview.

Compensation

The salary range for this position is $140,000 - $200,000. This is not, however, a guarantee of compensation or salary. Rather, salary will be set based on experience, geographic location and possibly contractual requirements and could fall outside of this range.

Benefits Overview

We are always looking for good people! Pantheon Data is committed to providing its employees with competitive salaries and benefits in order to increase employee satisfaction and productivity. In addition to our benefits, we also offer SmartBenefits through the Washington Metro Area Transportation Authority, where you specify an amount of your pre-tax wages be paid directly to your SmarTrip account. In some cases, tuition assistance may be available for continuing education expenses and certifications related to their position. Additional details may be found at https://pantheon-data.com/careers/

Pantheon Data Important Information

All qualified applicants will be considered for employment without regard to disability, status as a protected veteran, or any other status protected by applicable federal, state, local, or international law.

As part of the application process, you are expected to be on camera during interviews and assessments. We reserve the right to take your picture to verify your identity and prevent fraud.

If you require reasonable accommodation in completing this application, interviewing, completing any pre-employment testing, or otherwise participating in the employee selection process, please direct your inquiries to our Talent Team at Recruiting@pantheon-data.com or by phone (571) 363-4020.

This company uses E-Verify to confirm each employee's work authorization. For more information, click here E-Verify Participation Poster

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Data Engineer
Senior Data Engineer

Pantheon-Data • Reston (VA)

Hybrid
USD 140,000 - 160,000
SmartBenefits program
Transportation benefits
Tuition assistance may be available
Senior Full Stack Developer
Senior Full Stack Developer

Pantheon-Data • Reston (VA)

On-site
USD 145,000 - 200,000
SmartBenefits (SmarTrip)
Tuition assistance
E-Verify participation
AI/ML Specialist
AI/ML Specialist

Pantheon Data • Washington

Hybrid
USD 80,000 - 140,000
SmartBenefits through the Washington-M
SeniorDevSecOpsEngineer
SeniorDevSecOpsEngineer

Pantheon Data • Reston (VA)

On-site
USD 140,000 - 200,000
SmartBenefits through WMATA
Tuition assistance
SeniorDevSecOpsEngineer
SeniorDevSecOpsEngineer

Pantheon-Data • Reston (VA)

Hybrid
USD 140,000 - 200,000
SmartBenefits via SmarTrip
Tuition assistance
Pre-tax SmarTrip payments
Senior Applied AI Engineer
Senior Applied AI Engineer

AgentGraph • Northern (KY)

Hybrid
USD 150,000 - 190,000
401(k) match
Company HSA contributions
Wellness reimbursement
+6
Senior Cloud Engineer
Senior Cloud Engineer

Pantheon Data • Charlotte (NC)

Remote
USD 100,000 - 150,000
Chief Product Officer (CPO)
Chief Product Officer (CPO)

Pantheon Data • Reston (VA)

On-site
USD 150,000 - 250,000
Competitive salary
Tuition assistance
Flexibility in work location
Software Development Engineer - Document Intelligence
Software Development Engineer - Document Intelligence

Workday • Pleasanton (CA)

Hybrid
USD 148,000 - 222,000
Flexible work arrangements
Competitive benefits
Senior Cloud Engineer
Senior Cloud Engineer

Pantheon-Data • Reston (VA)

On-site
USD 100,000 - 150,000