Principal Data Engineer

Capmation

Xico

Presencial

MXN 900.000 - 1.500.000

Jornada completa

Hace 4 días
Sé de los primeros/as/es en solicitar esta vacante
Generador de candidaturas

Transforma esta oferta en una entrevista — un currículum y una carta de presentación creados pensando en lo que quiere el empleador.

Supera los filtros ATS

Descripción de la vacante

Capmation is seeking a Principal Data Engineer to lead a real-time data platform initiative. You will architect a lakehouse, land data from diverse sources into Bronze/Silver/Gold layers, and design streaming ingestion patterns that align with business needs.

You will shape workspace governance for Data Agents and partner with Application Development to expose curated data via APIs. You will mentor engineers, drive engineering quality, and balance hands-on work with strategic platform decisions

Formación

  • 5+ years building production data platforms and pipelines.
  • Hands-on production lakehouse experience with Fabric or Databricks.
  • Strong SQL, PySpark, and Power BI modeling skills.
  • Experience with real-time ingestion patterns and data governance.

Responsabilidades

  • Architect and implement lakehouse using Medallion Architecture (Bronze → Gold).
  • Land data from multiple source systems into Bronze/Silver layers.
  • Design real-time ingestion using Fabric or Databricks stacks.
  • Define workspace, governance, and security for Data Agents.
  • Collaborate with app development to expose curated datasets via APIs.

Conocimientos

PySpark
SQL
Power BI
Python
Databricks
Microsoft Fabric
Azure DevOps
Data modeling
Data APIs
DAX

Herramientas

SQL Server
Databricks Platform
Fabric API
OneLake
Power BI Desktop

Descripción del empleo

We are seeking a Principal Data Engineer to join our Engineering Team and help build a real-time data platform that powers self-service analytics, data agents, and business-built applications.

The ideal candidate is a senior technical leader who can architect a lakehouse from the ground up, land data from a large and diverse set of source systems into a medallion (Bronze → Gold) architecture, and design real-time ingestion patterns that meet evolving business needs. This role also shapes workspace and governance strategy in a way that supports Data Agent enablement, and partners with the Application Development team to expose curated data.

This position requires strong collaboration and leadership skills. The Principal Data Engineer must work effectively with engineers, analysts, app developers, AI/agent teams, product stakeholders, and clients, while helping shape ways of working, mentoring data engineers, and driving high engineering quality. The ideal candidate should be proactive, pragmatic, and able to balance hands‑on implementation with strategic decisions about platform design, modeling, governance, and cost.

Mission of the Role
  • Real‑time data: stand up streaming and near‑real‑time ingestion patterns where they meaningfully change the business outcome.
  • Self‑service: enable Data Agents, business‑led report and dashboard creation, and API access for business‑built applications.
  • Trustworthy data products: deliver curated, governed Bronze, Silver and Gold layers from identified systems.
Key Responsibilities
  • SQL Server Builds: architect and build Power BI dashboards using SQL Server and server replication
  • Fabric Data Lake Build: architect and implement a Microsoft Fabric Data Lake using Lakehouse, OneLake, and medallion architecture, with a clear path from raw to curated layers.
  • Source‑to‑Silver Ingestion: land data from source systems (databases, SaaS, files, events, APIs) into Bronze and Silver layers using Fabric Data Pipelines, Dataflows Gen2, notebooks, and shortcuts as appropriate to each source.
  • Real‑Time Patterns: design and implement streaming and near‑real‑time ingestion using Fabric Real‑Time Intelligence (Eventstream, Eventhouse/KQL) and/or Azure Event Hubs and CDC, balancing real‑time goals with cost, complexity, and actual business need.
  • Workspace Strategy for Data Agents: define a Fabric workspace, capacity, and governance strategy that enables Data Agents (data‑aware AI agents) to operate safely against curated data, including domain organization, item ownership, security, and lineage.
  • API Store with App Dev: partner with the Application Development team to design and operate an API store using Fabric API, exposing curated datasets to business‑built applications with appropriate authentication, throttling, and contract design.
  • Modeling & Semantic Layer: lead conceptual, logical, and physical modeling for analytical and operational use cases, and shape Power BI semantic models that make analytics consistent and trustworthy.
  • Quality, Governance & Cost: establish data quality, lineage, and governance standards (Purview, OneLake item‑level controls); treat capacity, storage, and movement cost as first‑class engineering metrics.
  • DataOps & CI/CD: bring software engineering rigor to data — version control, testing, environment promotion, and pipeline observability — through Azure DevOps.
  • Technical Leadership: set engineering standards across the data team, drive strategic platform decisions, and represent Capmation engineering in client whiteboards, architecture reviews, and roadmaps.
  • Team Development: mentor data engineers at all levels, lead training sessions and code reviews, provide constructive feedback, and align technical standards across teams.
Soft Skills
  • Business Acumen: connect data architecture and modeling decisions to business outcomes — real‑time value, self‑service enablement, and time‑to‑insight.
  • Accountability: own the technical and delivery success of the lakehouse platform and pipelines, addressing data quality, cost, and reliability proactively rather than reactively.
  • Communication: communicate complex data and architectural concepts clearly to both technical and non‑technical audiences, aligning stakeholders and enabling confident decisions.
  • Judgement: balance real‑time ambition against cost, complexity, and need; lift platform‑level risks early and recommend pragmatic alternatives.
  • Collaboration: drive alignment across data engineering, analytics, AI/agent, app dev, and business teams, acting as a unifying technical leader who resolves cross‑team friction.
  • Curiosity: stay current on Fabric, Databricks, real‑time platforms, data agents, and emerging lakehouse capabilities; use that understanding to anticipate challenges and guide innovation.
Required Qualifications

Experience: 5+ years of experience in Data Engineering, including significant time spent designing and operating production data platforms and pipelines.

Platform (must have one):

  • Microsoft Fabric — Lakehouse, Warehouse, Data Pipelines, Dataflows Gen2, OneLake, shortcuts, capacities
  • Databricks — Workspaces, Unity Catalog, Delta Lake, Workflows, Auto Loader, DLT, SQL Warehouses, Model Serving

Real‑Time / Streaming:

  • Fabric Real‑Time Intelligence (Eventstream, Eventhouse / KQL), and/or
  • Databricks Structured Streaming, and/or
  • Azure Event Hubs, Kafka, CDC (Debezium and similar)

Shared:

  • Power BI (semantic models, DAX; Direct Lake on Fabric or Power BI on Databricks SQL)
  • PySpark
  • SQL — T‑SQL (Fabric track) and/or Spark SQL / ANSI SQL (Databricks track)
  • Python
  • Data API tooling: Fabric API, Azure API Management, or Databricks SQL endpoints / Model Serving
Must have:
  • Lakehouse Platform Required: deep, hands‑on production experience in Microsoft Fabric OR Databricks — must have built a production lakehouse on one of these, not just experimented, with a strong understanding of the trade‑offs between the two.
  • SQL Server Experience: hands‑on experience with SQL Server and Server Replication with outputting data to Power Bi dashboards.
  • Medallion Architecture at Scale: demonstrated experience landing data from many heterogeneous source systems (databases, SaaS, files, events, APIs) into Bronze and Silver layers with strong reliability, idempotency, and lineage.
  • Real‑Time / Streaming Ingestion: production experience with at least one streaming stack on the candidate's primary platform (Fabric Real‑Time Intelligence / Eventstream / KQL, Databricks Structured Streaming, or Azure Event Hubs / Kafka with CDC), with the judgment to choose batch vs. micro‑batch vs. streaming per use case.
  • Workspace & Capacity Strategy: experience defining workspace, domain, capacity, and security strategy on Fabric or Databricks that scales across multiple teams and data products.
  • API Surface for Data: experience exposing curated data via APIs Fabric API, Azure API Management, Databricks SQL endpoints / Model Serving, or equivalent API gateway / data API patterns) with attention to auth, throttling, contracts, and consumer experience.
  • Strong PySpark and SQL skills (T‑SQL and/or Spark SQL / ANSI SQL) for building and optimizing transformation and modeling workloads in production.
  • Power BI for Engineering: designing semantic models and DAX, integrating with Fabric Direct Lake or Databricks SQL, and enabling self‑service report and dashboard creation.
  • Azure DevOps CI/CD for Data: pipeline‑as‑code, environment promotion, automated tests for data, and release management.
  • Modeling Expertise: deep proficiency in dimensional (Kimball) modeling and at least one of Data Vault 2.0, normalized 3NF, or domain‑driven (Data Mesh) modeling.
  • Strong analytical, problem‑solving, and communication skills.
Language & Location
  • Advanced English level (C1 or higher), spoken and written.
  • Open to nearshore (Latin America) candidates aligned with U.S. business hours.
Consigue la evaluación confidencial y gratuita de tu currículum.

o arrastra y suelta tu archivo aquí

Similar jobs

Puestos de trabajo similares que vale la pena comparar

Senior Data Platform Engineer
Senior Data Platform Engineer

Goods & Services • México

A distancia
MXN 900.000 - 1.300.000
Principal Data Engineer - Real-Time Lakehouse Leader
Principal Data Engineer - Real-Time Lakehouse Leader

Capmation • Xico

Presencial
MXN 900.000 - 1.500.000
Data Architect ID52062
Data Architect ID52062

AgileEngine • Monterrey

Híbrido
MXN 1.748.000 - 2.273.000
Professional growth: Mentorship, TechTalks
Competitive compensation: USD-based pay
Exciting projects: Modern solutions with Fortune 500
+1
25444 - Sr Data Engineer
25444 - Sr Data Engineer

Pediatric Associates • Monterrey

Presencial
MXN 900.000 - 1.300.000
Data Architect + AI
Data Architect + AI

Turtle Trax S.A. • Región Centro

Presencial
MXN 900.000 - 1.500.000
Data & Integration Analyst
Data & Integration Analyst

Bold Business • Región Centro

Presencial
MXN 900.000 - 1.300.000
Senior Data Engineer
Senior Data Engineer

Athenaworks • Región Centro

Presencial
MXN 1.204.000 - 2.063.000
Payment in USD
Flexible work schedule
Learning Budget
+1
Cloud Engineer Azure Databricks
Cloud Engineer Azure Databricks

Turtle Trax S.A. • Región Centro

Presencial
MXN 1.000.000 - 1.500.000
Senior Lakehouse Engineer - Data Quality & Fabric
Senior Lakehouse Engineer - Data Quality & Fabric

Goods & Services • México

A distancia
MXN 900.000 - 1.300.000
Data Analyst – Reporting & Visualization 1861
Data Analyst – Reporting & Visualization 1861

Softgic • Chihuahua

Presencial
MXN 312.000 - 456.000