Senior AI Data Engineer

Strategic Systems International

Ciudad de México

Presencial

MXN 900.000 - 1.500.000

Jornada completa

Hace 9 días
Generador de candidaturas

Destaca para este puesto: genera un currículum y una carta de presentación adaptados en cuestión de un minuto.

Supera los filtros ATS

Descripción de la vacante

Strategic Systems International is seeking a Senior AI Data Engineer to design and operate the data foundation powering AI systems. You will model data, build scalable pipelines, and own the data contracts and lineage across source, warehouse, and retrieval layers.

The role emphasizes production-grade Python, Kimball-based modeling, and end-to-end governance for reliable, observable data flows in a fast-moving AI environment.

Formación

  • 5–10+ years in software or data engineering, with substantial time in production data platform work.
  • Expert-level SQL with window functions, CTEs, and performance tuning in a columnar warehouse.
  • Production-grade Python with typing, packaging, dependency management, and testing.
  • Strong design skills: SOLID, domain-driven design, bounded contexts.

Responsabilidades

  • Design and evolve Kimball dimensional models and snowflake schemas with proper grain declarations.
  • Build and operate batch and streaming ETL/ELT pipelines with Airflow, Prefect, or Dagster and ensure idempotent retries.
  • Develop AI data pipelines for embedding generation, vector stores, and feature provisioning.
  • Maintain end-to-end data lineage, governance, and cost management in cloud platforms.
  • Collaborate with data scientists to provide reliable data foundations and governance.

Conocimientos

Data Warehousing
SQL
Python
Data Modeling
Orchestration
Cloud Platforms
Kubernetes
dbt
Data Quality
Kafka/Kinesis

Herramientas

Airflow
Prefect
Dagster
Iceberg
Delta Lake
Hudi
Kubernetes
dbt
SQL tooling

Descripción del empleo

Senior AI Data Engineer
Job Summary

We are seeking a Senior AI Data Engineer to design and operate the data foundation that our AI systems depend on. This role owns the movement, modeling, and quality of data from source systems through the warehouse and into the retrieval and feature layers that power LLM pipelines, agentic workflows, and analytical products.

The ideal candidate is a rigorous software engineer first and a data specialist second: someone who models a warehouse deliberately, writes production Python that other engineers can extend, and treats pipelines as versioned, tested, observable software rather than scripts.

This role partners closely with the AI/ML Data Scientist, who owns model behavior and retrieval strategy. The boundary: you own the pipeline, the schema, and the guarantees; they own the algorithm, the prompt, and the evaluation.

Key Responsibilities
Data Warehousing & Dimensional Modeling

Design and evolve dimensional models using Kimball methodology - star schemas, conformed dimensions, an enterprise bus matrix, and explicit fact-table grain declarations.

Implement transaction, periodic snapshot, and accumulating snapshot fact tables as the business process warrants, and defend the choice of grain.

Manage slowly changing dimensions (Type 1 / 2 / 3, and hybrid variants) with correct effective-dating, surrogate key strategy, and late-arriving dimension handling.

Model semi-structured and unstructured sources - documents, transcripts, event streams -into queryable structures without discarding provenance.

Maintain a governed semantic layer so that both human analysts and agentic consumers resolve the same metric to the same number.

Pipeline & ETL/ELT Engineering

Build and operate batch and streaming pipelines with orchestration frameworks such as Airflow, Prefect, or Dagster, including backfill, replay, and idempotent-retry semantics.

Implement ELT transformation layers with tested, documented, version-controlled SQL.

Own data contracts between producing and consuming systems: schema evolution, compatibility rules, and breaking-change procedure.

Instrument pipelines for observability - freshness, volume, distribution, and schema-drift checks - with alerting that distinguishes a real incident from ordinary variance.

Maintain end-to-end lineage from source record to warehouse fact to retrieved chunk.

Data Engineering for AI

Build embedding generation and refresh pipelines: chunk materialization, embedding jobs, incremental re-embedding on source change, and index lifecycle management.

Operate vector stores (Pinecone, Weaviate, Chroma, Milvus, pgvector) as production data systems - capacity, index build strategy, upsert/delete correctness, and staleness SLAs.

Build extraction pipelines over unstructured sources using OCR, document parsers, and vision-language model outputs, treating extraction confidence as a first-class column.

Implement the preprocessing and feature pipelines that back model training and inference, with train/serve consistency as a design requirement.

Expose data to AI services and applications through well-specified FastAPI or gRPC interfaces.

Engineering Craft

Apply SOLID principles and domain-driven design: bounded contexts that mirror the business domains, ubiquitous language shared with stakeholders, aggregates and repositories that keep domain logic out of transport and persistence layers.

Maintain meaningful test coverage (unit, contract, and data-quality assertions) and treat an untested pipeline as an unfinished one.

Own CI/CD, containerization, and environment promotion for data services.

Contribute to code review, architectural decision records, and internal standards.

Governance & Cost

Implement access control, PII handling, retention, and audit requirements at the data layer.

Manage warehouse and pipeline cost: partitioning, clustering, materialization strategy, and storage tiering.

Collaboration

Translate business problems into data models with product and client stakeholders.

Mentor junior and mid-level data engineers.

Required Qualifications

5-10+ years in software engineering or data engineering, with substantial time in production data platform work.

Data warehousing:

demonstrable command of Kimball dimensional modeling - not just familiarity with the vocabulary, but the judgment to choose a grain, resolve a many-to-many relationship, and know when to denormalize. Working knowledge of alternative approaches (Data Vault, One Big Table, Inman) and the tradeoffs against Kimball.

SQL:

expert-level - window functions, CTEs, query plan reading, and performance tuning on a columnar warehouse.

Python:

expert-level, production-grade - typing, packaging, dependency management, testing.

Design:

SOLID and domain-driven design applied in real systems, with examples you can walk through.

Orchestration:

Airflow, Prefect, Dagster, or equivalent, in production.

Cloud:

expert-level on AWS, Azure, or GCP - storage, compute, IAM, networking, and cost management.

Platform:

containerization, Kubernetes (EKS/AKS/GKE), and CI/CD.

Experience with lakehouse table formats (Iceberg, Delta Lake, Hudi) and their maintenance characteristics: compaction, snapshot expiry, schema and partition evolution.

Preferred Qualifications

Experience building the data layer beneath production RAG systems, including hybrid search infrastructure and index freshness guarantees.

Streaming systems: Kafka, Kinesis, Flink, or Spark Structured Streaming.

dbt or an equivalent transformation and testing framework.

Data quality tooling (Great Expectations, Soda, or similar) and catalog/lineage platforms.

Familiarity with the model-facing side of the stack - MLflow, Weights & Biases, feature stores - sufficient to collaborate credibly with data scientists.

Working knowledge of a second language: TypeScript, Java, Go, Scala, or Rust.

Experience with AI security, governance, and compliance frameworks.

Open-source contributions to data or AI infrastructure projects.

Consigue la evaluación confidencial y gratuita de tu currículum.
o arrastra y suelta tu archivo aquí
Similar jobs

Puestos de trabajo similares que vale la pena comparar

Data & Integration Analyst
Data & Integration Analyst

SCALIS • Región Centro

Presencial
MXN 600.000 - 900.000
Lead Data Engineer ID71008
Lead Data Engineer ID71008

AgileEngine, LLC. • Santiago de Querétaro

Presencial
MXN 2.042.000 - 2.893.000
Growth without limits
Competitive compensation
Flexibility: 100% remote
+3
Lead Data Engineer ID71008
Lead Data Engineer ID71008

AgileEngine, LLC. • Región Centro

Presencial
MXN 900.000 - 1.200.000
Growth opportunities
Competitive pay
Remote work 100%
+3
Senior Data Engineer
Senior Data Engineer

The Home Depot Global Technology Center In Mexico • Estado de México

Presencial
PHP 5.029.000 - 6.805.000
Data Engineer ID52278
Data Engineer ID52278

AgileEngine • León

Híbrido
MXN 1.032.000 - 1.377.000
Mentorship and personalized growth roadmaps
USD-based competitive compensation with budgets
Exciting projects with top companies
+1
Data Engineer ID52278
Data Engineer ID52278

AgileEngine • Puebla de Zaragoza

Híbrido
MXN 1.204.000 - 1.549.000
Professional growth through mentorship
Competitive USD-based compensation
Exciting projects with top companies
+1
Data Engineer ID52278
Data Engineer ID52278

AgileEngine • Santiago de Querétaro

Híbrido
MXN 1.206.000 - 1.723.000
Mentorship
TechTalks
Education budget
+3
Senior Data Engineer
Senior Data Engineer

Nimble Gravity • Región Centro

Presencial
MXN 600.000 - 800.000
Senior AI Data Engineer: Build Scalable AI Data Pipelines
Senior AI Data Engineer: Build Scalable AI Data Pipelines

Strategic Systems International • Ciudad de México

Presencial
MXN 900.000 - 1.500.000
Ingeniero de Datos
Ingeniero de Datos

PROGRAMMING.COM • Azcapotzalco

Presencial
MXN 520.000 - 760.000