Akoncagua AI is building domain-specialized AI systems for technically demanding industries.
Our work spans large language models, multimodal AI, agentic systems, model adaptation,
synthetic data, evaluation, RAG, and production AI infrastructure.
We are seeking a Senior Data Engineer with 7+ years of production experience to own the data
foundation behind our AI systems. This role requires someone who can design highly reliable
pipelines for collecting, processing, validating, curating, versioning, and serving large-scale
multimodal and technical datasets.
The engineer will work closely with Agentic AI, ML, research, and domain teams to ensure models
and agents operate on accurate, traceable, production-quality data.
Role Overview
This role sits at the intersection of Data Engineering, AI/ML infrastructure, document intelligence,
multimodal data processing, and RAG.
You will build end-to-end data platforms capable of processing complex data including PDFs,
scanned documents, high-resolution technical drawings, images, geospatial data, structured
datasets, and enterprise documents.
A major responsibility will be establishing rigorous data validation, curation, provenance,
benchmarking, and quality-control systems so downstream AI systems can reliably trust the
information they receive.
You will also provide technical leadership to engineers working on ingestion, processing, retrieval,
and data-quality systems.
Responsibilities
- Design and operate production-scale end-to-end data pipelines for AI/ML and agentic systems.
- Build ingestion pipelines for PDFs, scanned documents, high-resolution technical images/drawings, structured data, APIs, object stores, databases, and other enterprise sources.
- Develop robust document-processing pipelines including parsing, layout extraction, OCR when appropriate, image processing, metadata extraction, chunking, classification, and
- Design automated data validation and quality-control frameworks covering completeness, correctness, consistency, duplication, corruption, schema validity, provenance, and extraction confidence.
- Build data curation pipelines for model training, fine-tuning, RAG, evaluation, and refresh pipelines.
- Establish dataset lineage, provenance, versioning, reproducibility, auditability, and quality metrics.
- Design benchmarks to compare document parsers, OCR/extraction approaches, embedding models, retrieval systems, vector databases, storage architectures, and processing pipelines.
- Build large-scale processing infrastructure primarily on AWS, with experience supporting GCP and Azure environments.
- Optimize pipelines for throughput, reliability, latency, storage efficiency, compute cost, and
- Design human-in-the-loop data review and annotation workflows for high-accuracy datasets.
- Build monitoring, alerting, replay, retry, dead-letter, and failure-recovery systems.
- Partner with AI engineers to transform agent/model failures into data-quality improvements, new datasets, benchmarks, and evaluation cases.
- Lead architecture reviews, code reviews, data-quality standards, and technical direction for other data engineers.
Required Qualifications
- 7+ years of production Data Engineering, Data Platform, ML Data Infrastructure, or closely related experience.
- BS, MS, or equivalent industry experience in Computer Science, Computer Engineering, Data Engineering, Artificial Intelligence, or a related technical discipline.
- Expert-level Python and SQL skills with strong software-engineering fundamentals.
- Proven experience designing and operating large-scale batch and streaming data pipelines in production.
- Deep experience with AWS, including technologies such as S3, ECS/EKS, Lambda, Step Functions, Glue, EMR, Athena, SQS/SNS, EventBridge, RDS/Aurora, DynamoDB, CloudWatch, and IAM as appropriate to the architecture.
- Production experience with GCP and/or Azure is strongly preferred.
- Experience designing distributed, fault-tolerant processing systems for large document and
- Strong experience processing PDFs, scanned documents, high-resolution technical images, drawings, tables, and complex document layouts.
- Strong experience with data validation, cleansing, normalization, deduplication, schema enforcement, lineage, provenance, and dataset versioning.
- Experience designing data-quality benchmarks and automated validation harnesses with measurable acceptance criteria.
- Experience preparing and curating datasets for LLM training, fine-tuning, evaluation, RAG, and agentic AI systems.
- Experience with distributed processing technologies such as Spark, Ray, Databricks, Kafka, Airflow, Dagster, or equivalent systems.
- Strong experience with relational databases, object storage, data warehouses/lakes/lakehouses, and modern data architectures.
- Strong understanding of testing, observability, profiling, performance optimization, CI/CD, Docker, infrastructure-as-code, and production reliability.
- Experience designing systems where data quality, provenance, and confidence are critical
- Demonstrated technical leadership, including mentoring engineers, reviewing architecture, establishing standards, and driving teams toward high-quality production outcomes.
Preferred Qualifications
- Experience building data infrastructure specifically for LLMs, multimodal models, AI agents, fine-tuning, synthetic data, and evaluation systems.
- Experience with large-scale technical-document intelligence and multimodal extraction.
- Experience benchmarking OCR/document understanding models, parsers, vision-language models, embeddings, rerankers, and retrieval architectures.
- Experience with human annotation, expert review, active learning, and data-curation platforms.
- Experience building golden datasets and ground-truth benchmarks for AI/ML evaluation.
- Experience with vector databases and search technologies such as OpenSearch/Elasticsearch, pgvector, Pinecone, Weaviate, Milvus, or similar systems.
- Experience with data catalogs, governance, lineage, and observability platforms.
- Experience with Terraform, Kubernetes, distributed compute, GPU workloads, and cost optimization.
- Experience handling geospatial, CAD, engineering, scientific, or other highly technical datasets.
- Familiarity with GIS, AutoCAD/Civil 3D, DXF/DWG, engineering drawings, maps, and geospatial data is highly preferred.
- Experience working closely with research scientists and Agentic AI/ML engineers to create