Senior Data Engineer — AI/ML Data Platforms

Akoncagua AI

Lakeland (FL)

On-site

USD 120,000 - 180,000

Full time

48 hours ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Akoncagua AI is seeking a Senior Data Engineer with 7+ years of production data engineering experience to own the data foundation behind our AI systems. This role requires someone who can design highly reliable pipelines for collecting, processing, validating, curating, versioning, and serving large-scale multimodal and technical datasets.

The engineer will work closely with Agentic AI, ML, research, and domain teams to ensure models and agents operate on accurate, traceable, production-quality

Qualifications

  • 7+ years of production Data Engineering, Data Platform, ML Data Infrastructure, or closely related experience.
  • Expert-level Python and SQL skills with strong software-engineering fundamentals.
  • Proven experience designing and operating large-scale batch and streaming data pipelines in production.
  • Deep experience with AWS, including technologies such as S3, ECS/EKS, Lambda, Step Functions, Glue, EMR, Athena, SQS/SNS, EventBridge, RDS/Aurora, DynamoDB, CloudWatch, and IAM as appropriate to the architecture.
  • Production experience with GCP and/or Azure is strongly preferred.

Responsibilities

  • Design and operate production-scale end-to-end data pipelines for AI/ML and agentic systems.
  • Build ingestion pipelines for PDFs, scanned documents, high-resolution technical images/drawings, structured data, APIs, object stores, databases, and other enterprise sources.
  • Develop robust document-processing pipelines including parsing, layout extraction, OCR when appropriate, image processing, metadata extraction, chunking, classification, and quality-control.
  • Design automated data validation and quality-control frameworks covering completeness, correctness, consistency, duplication, corruption, schema validity, provenance, and extraction confidence.
  • Lead architecture reviews, code reviews, data-quality standards, and technical direction for other data engineers.

Skills

Python
SQL
Data pipelines
Software engineering

Education

BS/MS in CS/CE/Data Eng/AI

Tools

AWS
GCP
Azure
Spark
Airflow
Dagster
Databricks
Kafka

Job description

Akoncagua AI is building domain-specialized AI systems for technically demanding industries.

Our work spans large language models, multimodal AI, agentic systems, model adaptation,

synthetic data, evaluation, RAG, and production AI infrastructure.

We are seeking a Senior Data Engineer with 7+ years of production experience to own the data

foundation behind our AI systems. This role requires someone who can design highly reliable

pipelines for collecting, processing, validating, curating, versioning, and serving large-scale

multimodal and technical datasets.

The engineer will work closely with Agentic AI, ML, research, and domain teams to ensure models

and agents operate on accurate, traceable, production-quality data.

Role Overview

This role sits at the intersection of Data Engineering, AI/ML infrastructure, document intelligence,

multimodal data processing, and RAG.

You will build end-to-end data platforms capable of processing complex data including PDFs,

scanned documents, high-resolution technical drawings, images, geospatial data, structured

datasets, and enterprise documents.

A major responsibility will be establishing rigorous data validation, curation, provenance,

benchmarking, and quality-control systems so downstream AI systems can reliably trust the

information they receive.

You will also provide technical leadership to engineers working on ingestion, processing, retrieval,

and data-quality systems.

Responsibilities
  • Design and operate production-scale end-to-end data pipelines for AI/ML and agentic systems.
  • Build ingestion pipelines for PDFs, scanned documents, high-resolution technical images/drawings, structured data, APIs, object stores, databases, and other enterprise sources.
  • Develop robust document-processing pipelines including parsing, layout extraction, OCR when appropriate, image processing, metadata extraction, chunking, classification, and
  • Design automated data validation and quality-control frameworks covering completeness, correctness, consistency, duplication, corruption, schema validity, provenance, and extraction confidence.
  • Build data curation pipelines for model training, fine-tuning, RAG, evaluation, and refresh pipelines.
  • Establish dataset lineage, provenance, versioning, reproducibility, auditability, and quality metrics.
  • Design benchmarks to compare document parsers, OCR/extraction approaches, embedding models, retrieval systems, vector databases, storage architectures, and processing pipelines.
  • Build large-scale processing infrastructure primarily on AWS, with experience supporting GCP and Azure environments.
  • Optimize pipelines for throughput, reliability, latency, storage efficiency, compute cost, and
  • Design human-in-the-loop data review and annotation workflows for high-accuracy datasets.
  • Build monitoring, alerting, replay, retry, dead-letter, and failure-recovery systems.
  • Partner with AI engineers to transform agent/model failures into data-quality improvements, new datasets, benchmarks, and evaluation cases.
  • Lead architecture reviews, code reviews, data-quality standards, and technical direction for other data engineers.
Required Qualifications
  • 7+ years of production Data Engineering, Data Platform, ML Data Infrastructure, or closely related experience.
  • BS, MS, or equivalent industry experience in Computer Science, Computer Engineering, Data Engineering, Artificial Intelligence, or a related technical discipline.
  • Expert-level Python and SQL skills with strong software-engineering fundamentals.
  • Proven experience designing and operating large-scale batch and streaming data pipelines in production.
  • Deep experience with AWS, including technologies such as S3, ECS/EKS, Lambda, Step Functions, Glue, EMR, Athena, SQS/SNS, EventBridge, RDS/Aurora, DynamoDB, CloudWatch, and IAM as appropriate to the architecture.
  • Production experience with GCP and/or Azure is strongly preferred.
  • Experience designing distributed, fault-tolerant processing systems for large document and
  • Strong experience processing PDFs, scanned documents, high-resolution technical images, drawings, tables, and complex document layouts.
  • Strong experience with data validation, cleansing, normalization, deduplication, schema enforcement, lineage, provenance, and dataset versioning.
  • Experience designing data-quality benchmarks and automated validation harnesses with measurable acceptance criteria.
  • Experience preparing and curating datasets for LLM training, fine-tuning, evaluation, RAG, and agentic AI systems.
  • Experience with distributed processing technologies such as Spark, Ray, Databricks, Kafka, Airflow, Dagster, or equivalent systems.
  • Strong experience with relational databases, object storage, data warehouses/lakes/lakehouses, and modern data architectures.
  • Strong understanding of testing, observability, profiling, performance optimization, CI/CD, Docker, infrastructure-as-code, and production reliability.
  • Experience designing systems where data quality, provenance, and confidence are critical
  • Demonstrated technical leadership, including mentoring engineers, reviewing architecture, establishing standards, and driving teams toward high-quality production outcomes.
Preferred Qualifications
  • Experience building data infrastructure specifically for LLMs, multimodal models, AI agents, fine-tuning, synthetic data, and evaluation systems.
  • Experience with large-scale technical-document intelligence and multimodal extraction.
  • Experience benchmarking OCR/document understanding models, parsers, vision-language models, embeddings, rerankers, and retrieval architectures.
  • Experience with human annotation, expert review, active learning, and data-curation platforms.
  • Experience building golden datasets and ground-truth benchmarks for AI/ML evaluation.
  • Experience with vector databases and search technologies such as OpenSearch/Elasticsearch, pgvector, Pinecone, Weaviate, Milvus, or similar systems.
  • Experience with data catalogs, governance, lineage, and observability platforms.
  • Experience with Terraform, Kubernetes, distributed compute, GPU workloads, and cost optimization.
  • Experience handling geospatial, CAD, engineering, scientific, or other highly technical datasets.
  • Familiarity with GIS, AutoCAD/Civil 3D, DXF/DWG, engineering drawings, maps, and geospatial data is highly preferred.
  • Experience working closely with research scientists and Agentic AI/ML engineers to create
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Data AI Engineer
Data AI Engineer

Compunnel, Inc. • Columbus (OH)

On-site
USD 100,000 - 130,000
AI & Data Architect
AI & Data Architect

Insaito Software • United States

Remote
USD 150,000 - 230,000
Senior Data Engineer / AI ML Engineer with Python, AI/ML & LLMs
Senior Data Engineer / AI ML Engineer with Python, AI/ML & LLMs

Ampcus Inc • Reston (VA)

On-site
USD 100,000 - 130,000
Senior AI Engineer - GenAI + Data Platform - AWS
Senior AI Engineer - GenAI + Data Platform - AWS

Compunnel, Inc. • Los Angeles (CA)

On-site
USD 120,000 - 160,000
Senior Software Engineering Lead
Senior Software Engineering Lead

Vanigent • Lead (SD)

On-site
USD 180,000 - 240,000
AI/ML Engineer
AI/ML Engineer

Be The Match in • Minneapolis (MN)

On-site
USD 120,000 - 180,000
AI Engineer
AI Engineer

Kaleidoscope Innovation • Fort Worth (TX)

On-site
USD 140,000 - 190,000
Senior Data Scientist
Senior Data Scientist

Flexjet LLC • Cleveland (OH)

On-site
USD 120,000 - 190,000
Lead Data AI Scientist
Lead Data AI Scientist

Nityo Infotech • Santa Clara (CA)

Hybrid
USD 180,000 - 240,000
AI Engineer
AI Engineer

INFOSYS NOVA HOLDINGS LLC • Fort Worth (TX)

On-site
USD 150,000 - 230,000