Principal Data Engineer

Nodi

Minnesota

On-site

USD 150,000 - 190,000

Full time

4 hours ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Capmation is seeking a Principal Data Engineer to design a real-time data platform powering self-service analytics, data agents, and business-built apps. You will lead a lakehouse from Bronze to Gold, land data from diverse sources, and collaborate with product and engineering teams to expose curated data.

The role combines hands-on implementation with strategic platform decisions, governance, and cost optimization, requiring strong leadership and partnership across teams.

Qualifications

  • Experience building production data platforms and pipelines.

Responsibilities

  • Architect and lead real-time data ingestion patterns and lakehouse design.

Skills

Python
PySpark
SQL
Data modeling
Data warehousing
Real-time data
Leadership
Analytics

Education

Bachelor's degree in Computer Science or related field

Tools

Microsoft Fabric
Databricks
Power BI
Azure Event Hubs
Fabric API
Azure DevOps
OneLake

Job description

We are seeking a Principal Data Engineer to join our Engineering Team and help build a real-time data platform that powers self-service analytics, data agents, and business-built applications.

The ideal candidate is a senior technical leader who can architect a lakehouse from the ground up, land data from a large and diverse set of source systems into a medallion (Bronze → Gold) architecture, and design real-time ingestion patterns that meet evolving business needs. This role also shapes workspace and governance strategy in a way that supports Data Agent enablement, and partners with the Application Development team to expose curated data.

This position requires strong collaboration and leadership skills. The Principal Data Engineer must work effectively with engineers, analysts, app developers, AI/agent teams, product stakeholders, and clients, while helping shape ways of working, mentoring data engineers, and driving high engineering quality. The ideal candidate should be proactive, pragmatic, and able to balance hands-on implementation with strategic decisions about platform design, modeling, governance, and cost.

Mission of the Role
  • Real-time data: stand up streaming and near-real-time ingestion patterns where they meaningfully change the business outcome.
  • Self-service: enable Data Agents, business-led report and dashboard creation, and API access for business-built applications.
  • Trustworthy data products: deliver curated, governed Bronze, Silver and Gold layers from identified systems.
Key Responsibilities
  • SQL Server Builds: architect and build Power BI dashboards using SQL Server and server replication
  • Fabric Data Lake Build: architect and implement a Microsoft Fabric Data Lake using Lakehouse, OneLake, and medallion architecture, with a clear path from raw to curated layers.
  • Source-to-Silver Ingestion: land data from source systems (databases, SaaS, files, events, APIs) into Bronze and Silver layers using Fabric Data Pipelines, Dataflows Gen2, notebooks, and shortcuts as appropriate to each source.
  • Real-Time Patterns: design and implement streaming and near-real-time ingestion using Fabric Real-Time Intelligence (Eventstream, Eventhouse/KQL) and/or Azure Event Hubs and CDC, balancing real-time goals with cost, complexity, and actual business need.
  • Workspace Strategy for Data Agents: define a Fabric workspace, capacity, and governance strategy that enables Data Agents (data-aware AI agents) to operate safely against curated data, including domain organization, item ownership, security, and lineage.
  • API Store with App Dev: partner with the Application Development team to design and operate an API store using Fabric API, exposing curated datasets to business-built applications with appropriate authentication, throttling, and contract design.
  • Modeling & Semantic Layer: lead conceptual, logical, and physical modeling for analytical and operational use cases, and shape Power BI semantic models that make analytics consistent and trustworthy.
  • Quality, Governance & Cost: establish data quality, lineage, and governance standards (Purview, OneLake item-level controls); treat capacity, storage, and movement cost as first-class engineering metrics.
  • DataOps & CI/CD: bring software engineering rigor to data — version control, testing, environment promotion, and pipeline observability — through Azure DevOps.
  • Technical Leadership: set engineering standards across the data team, drive strategic platform decisions, and represent Capmation engineering in client whiteboards, architecture reviews, and roadmaps.
  • Team Development: mentor data engineers at all levels, lead training sessions and code reviews, provide constructive feedback, and align technical standards across teams.
Soft Skills
  • Business Acumen: connect data architecture and modeling decisions to business outcomes — real-time value, self-service enablement, and time-to-insight.
  • Accountability: own the technical and delivery success of the lakehouse platform and pipelines, addressing data quality, cost, and reliability proactively rather than reactively.
  • Communication: communicate complex data and architectural concepts clearly to both technical and non-technical audiences, aligning stakeholders and enabling confident decisions.
  • Judgement: balance real-time ambition against cost, complexity, and need; escalate platform-level risks early and recommend pragmatic alternatives.
  • Collaboration: drive alignment across data engineering, analytics, AI/agent, app dev, and business teams, acting as a unifying technical leader who resolves cross-team friction.
  • Curiosity: stay current on Fabric, Databricks, real-time platforms, data agents, and emerging lakehouse capabilities; use that understanding to anticipate challenges and guide innovation.
Required Qualifications

Experience: 5+ years of experience in Data Engineering, including significant time spent designing and operating production data platforms and pipelines.

Platform (must have one):

  • Microsoft Fabric — Lakehouse, Warehouse, Data Pipelines, Dataflows Gen2, OneLake, shortcuts, capacities
  • Databricks — Workspaces, Unity Catalog, Delta Lake, Workflows, Auto Loader, DLT, SQL Warehouses, Model Serving

Real-Time / Streaming:

  • Fabric Real-Time Intelligence (Eventstream, Eventhouse / KQL), and/or
  • Databricks Structured Streaming, and/or
  • Azure Event Hubs, Kafka, CDC (Debezium and similar)

Shared:

  • Power BI (semantic models, DAX; Direct Lake on Fabric or Power BI on Databricks SQL)
  • PySpark
  • SQL — T-SQL (Fabric track) and/or Spark SQL / ANSI SQL (Databricks track)
  • Python
  • Data API tooling: Fabric API, Azure API Management, or Databricks SQL endpoints / Model Serving
Must have:
  • Lakehouse Platform Required: deep, hands-on production experience in Microsoft Fabric OR Databricks — must have built a production lakehouse on one of these, not just experimented, with a strong understanding of the trade-offs between the two.
  • SQL Server Experience: hands-on experience with SQL Server and Server Replication with outputting data to Power Bi dashboards.
  • Medallion Architecture at Scale: demonstrated experience landing data from many heterogeneous source systems (databases, SaaS, files, events, APIs) into Bronze and Silver layers with strong reliability, idempotency, and lineage.
  • Real-Time / Streaming Ingestion: production experience with at least one streaming stack on the candidate's primary platform (Fabric Real-Time Intelligence / Eventstream / KQL, Databricks Structured Streaming, or Azure Event Hubs / Kafka with CDC), with the judgment to choose batch vs. micro-batch vs. streaming per use case.
  • Workspace & Capacity Strategy: experience defining workspace, domain, capacity, and security strategy on Fabric or Databricks that scales across multiple teams and data products.
  • API Surface for Data: experience exposing curated data via APIs Fabric API, Azure API Management, Databricks SQL endpoints / Model Serving, or equivalent API gateway / data API patterns) with attention to auth, throttling, contracts, and consumer experience.
  • Strong PySpark and SQL skills (T-SQL and/or Spark SQL / ANSI SQL) for building and optimizing transformation and modeling workloads in production.
  • Power BI for Engineering: designing semantic models and DAX, integrating with Fabric Direct Lake or Databricks SQL, and enabling self-service report and dashboard creation.
  • Azure DevOps CI/CD for Data: pipeline-as-code, environment promotion, automated tests for data, and release management.
  • Modeling Expertise: deep proficiency in dimensional (Kimball) modeling and at least one of Data Vault 2.0, normalized 3NF, or domain-driven (Data Mesh) modeling.
  • Strong analytical, problem-solving, and communication skills.
Language & Location
  • Advanced English level (C1 or higher), spoken and written.
  • Open to nearshore (Latin America) candidates aligned with U.S. business hours.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Data Platform Engineer
Data Platform Engineer

Jobtailor • Knoxville (TN)

On-site
USD 140,000 - 180,000
Senior Data Engineer, Fabric Platform
Senior Data Engineer, Fabric Platform

IES Holdings • Houston (TX)

On-site
USD 120,000 - 160,000
Manager, Data and Analytics
Manager, Data and Analytics

Forward Air Corp. • Dallas (TX)

On-site
USD 130,000 - 160,000
Lead Data Engineer
Lead Data Engineer

K2 Intelligence, LLC • Northern (KY)

Hybrid
USD 120,000 - 180,000
Principal Data Engineer
Principal Data Engineer

Vomela Commercial Group • Headquarters (KY)

On-site
USD 180,000 - 200,000
Health Care Plan
401k Plan
Life Insurance
+4
Senior Data Engineer
Senior Data Engineer

Fracht Group - North America • Houston (TX)

On-site
USD 120,000 - 150,000
Principal Data Engineer
Principal Data Engineer

Vomela Company • Northern (KY)

Hybrid
USD 180,000 - 200,000
Health care plan
401k retirement plan
Life insurance
+4
Senior Data Engineer
Senior Data Engineer

Fracht Sweden AB • Houston (TX)

On-site
USD 120,000 - 150,000
Senior Data Engineer
Senior Data Engineer

TechRuiter • Northern (KY)

Hybrid
USD 120,000 - 180,000
Manager, Data and Analytics
Manager, Data and Analytics

Forward Air Corporation • Coppell (TX)

On-site
USD 120,000 - 150,000