Senior Data Engineer

Leon Capital Group

Dallas (TX)

On-site

USD 140,000 - 190,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Leon Capital Group, Dallas-based, is seeking a hands-on Data Engineer to own the data foundation across its healthcare, real estate, and financial services portfolio. You will consolidate fragmented data, clean, resolve, and produce a single trusted record for AI workloads.

Reporting to the Chief Innovation Officer, you will build canonical models and curated marts, enable consistent provenance, and scale the platform across the portfolio while enforcing data ownership and regulatory safeguards.

Qualifications

  • 7+ years building data systems end to end, ideally as an early or founding data hire.
  • Deep master data management expertise: canonical data models, matching, survivorship, and golden-record programs.
  • Strong data engineering fundamentals: schema design, ETL/ELT, data quality and de-duplication discipline.
  • Deep SQL and warehouse experience; Azure‑anchored stack (Azure SQL, dbt) with multi‑cloud exposure (Azure/AWS).
  • Proven experience integrating with messy enterprise/vendor systems and with flat-file or SFTP feeds without clean API.
  • Comfort owning data infrastructure in a major cloud environment without a dedicated platform team.
  • Experience handling regulated data (PHI, HIPAA, financial) securely.
  • Excellent judgment under ambiguity and ability to prioritize across initiatives.
  • Bias toward data quality: trustworthy data as the precondition for downstream steps.

Responsibilities

  • Consolidate and harden existing cloud data platforms: fix broken syncs, reduce single points of failure, and meet security standards.
  • Design and own the canonical data model and curated marts for long-term portability.
  • Own master data management end to end: define canonical entities, matching, survivorship, and golden-record rules.
  • Run cleansing and identity resolution as a continuous discipline; standardize and validate source data.
  • Build ingestion pipelines from heterogeneous sources with provenance: APIs, flat-files, SFTP feeds.
  • Stand up the curated data layer that AI depends on: feed data to lead scoring and predictive workloads.
  • Build validation, monitoring, and lineage to catch data-quality issues before models.
  • Treat the platform as a reusable pattern, scaling across portfolio initiatives.

Skills

Data modeling
Master data management
SQL / data warehouse
ETL / ELT pipelines
Data quality & deduplication
Cloud platforms (Azure/AWS)
Vendor systems integration
Regulatory data handling
Decision making under ambiguity

Job description

Leon Capital Group is a multi-billion-dollar holding company that owns and operates businesses across healthcare, real estate, and financial services. We founded, acquire, and scale companies for the long term, backing them with shared capital, talent, and infrastructure while giving each the room to operate and grow. Our portfolio spans provider groups and patient-care platforms, real estate development and investment, and a growing set of financial services businesses. This role reports to the Chief Innovation Officer and supports every business in the portfolio.

About the Role

This is a hands‑on data engineering role where the engineer owns the data foundation that AI is built on. Across our portfolio, the same problem keeps surfacing: the data exists but it is fragmented across vendor systems, duplicated, and untrustworthy, and no model or decision tool is better than the data beneath it. This role consolidates the data we already own, cleans and resolves it into one trustworthy record, and stands up the curated, provenance‑tracked layer that predictive and AI workloads depend on.

Consolidating that platform is the gating prerequisite for everything that follows, so this is where you prove the patterns. As that foundation matures, you will continue to be leveraged across different initiatives, each one modernizing or building up its data foundation for AI, reusing the same patterns and discipline rather than reinventing the approach each time.

We want someone energized by both the consolidation grind and the foundation it unlocks: the engineer who treats dirty, duplicated data as the most important problem in the building because every segment, attribution number, and model downstream depends on getting it right.

Key Responsibilities
  • Consolidate and harden existing cloud data platforms: re‑enable broken syncs, close single‑points‑of‑failure, and bring infrastructure up to architecture and security standards.
  • Design and own the canonical data model and curated marts, built to remain ours regardless of which vendor or CRM sits on top.
  • Own master data management end to end: define the canonical entities, set the matching, survivorship, and merge rules that resolve duplicate and conflicting records into one golden record, and govern that record as the single source of truth across systems.
  • Run cleansing and identity resolution as a continuous discipline, not a one‑time pass: standardize and validate source data on ingestion, match records across systems by shared keys, and keep the golden record clean so attribution and downstream models stay trustworthy.
  • Build ingestion pipelines that pull from heterogeneous, often hostile sources into our schema with full provenance, including vendor servicing output, API feeds, and flat‑file or SFTP partner feeds with no clean API.
  • Stand up the curated data layer that AI depends on: clean, well‑modeled marts that feed lead scoring, attribution, next‑best‑action, and other predictive workloads.
  • Build validation, monitoring, and lineage so data‑quality issues are caught before they reach models, reports, or decisions.
  • Treat the platform as a reusable pattern, standing up each new initiative's own data layer rather than a bespoke build each time, so the foundation scales across the portfolio.
  • Enforce the data‑ownership bar in every buy decision: vendor output must land in our canonical structure, in our schema, and remain portable on exit.
  • Partner with shared IT and security on regulated‑data handling, secrets management and compliance prerequisites, and document runbooks so the platform can be operated and handed off as it matures.
Qualifications
Required
  • 7+ years building data systems end to end, ideally as an early or founding data hire where you owned the whole data function rather than one stage of a large team.
  • Deep master data management expertise: you have designed canonical data models, built matching and survivorship logic, and run golden‑record and identity‑resolution programs that held up at scale. This is core to the role, not a nice‑to‑have.
  • Strong data engineering fundamentals: schema design, ETL and ELT pipeline architecture, and data‑quality and de‑duplication discipline.
  • Deep SQL and warehouse experience. Our current platforms are Azure‑anchored (Azure SQL, dbt for curated marts), and you are comfortable working in and consolidating this stack; you are equally comfortable extending into a multi‑cloud environment (we use a mix across Azure and AWS) as the work broadens across the portfolio.
  • Proven experience integrating against messy enterprise and vendor systems and against flat‑file or SFTP feeds with no clean API.
  • Comfort owning data infrastructure in a major cloud environment without a dedicated platform team; you can stand up and run the warehouse and pipelines yourself.
  • Comfort building under regulatory and fiduciary constraints, and handling sensitive, regulated data (PHI, HIPAA, financial) with the secure practices it requires.
  • Excellent judgment under ambiguity and the ability to prioritize across multiple initiatives without close oversight.
  • A bias toward data quality: you treat trustworthy, well‑resolved data as the precondition for everything downstream, not an afterthought.
Preferred
  • Regulated‑domain experience across healthcare, real estate, or financial services, where you have built to compliance constraints.
  • Hands‑on experience laying the groundwork for predictive or AI workloads: feature engineering, training‑set construction, and an understanding of what models need from the data layer beneath them.
  • Working familiarity with LLM‑assisted pipelines, retrieval‑augmented generation, and document extraction and structuring.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Data Platform & Machine Learning Engineer
Data Platform & Machine Learning Engineer

Leon Capital Group • Dallas (TX)

On-site
USD 120,000 - 150,000
Senior Data Engineer (AI-Native) — Data Layer
Senior Data Engineer (AI-Native) — Data Layer

Proton.ai • Boston (MA)

Hybrid
USD 140,000 - 210,000
Data Platform Architect for AI Foundation & MDM
Data Platform Architect for AI Foundation & MDM

Leon Capital Group • Dallas (TX)

On-site
USD 140,000 - 190,000
Sr. Manager, Data & Analytics
Sr. Manager, Data & Analytics

Specialized Bicycle Components, Inc. • Morgan Hill (CA)

On-site
USD 140,000 - 180,000
Senior Data Engineer
Senior Data Engineer

Intellivo • Memphis (TN)

On-site
USD 90,000 - 120,000
Principal Data Engineer
Principal Data Engineer

Caraa • Indianapolis (IN)

On-site
USD 150,000 - 190,000
Senior Data Engineer - Data Platform
Senior Data Engineer - Data Platform

SunStrong Management, LLC • United States

On-site
USD 130,000 - 170,000
Data Architect
Data Architect

Jobtailor • Cedar Falls (IA)

On-site
USD 140,000 - 210,000
Senior Data Engineer - Data Platform
Senior Data Engineer - Data Platform

SunStrong Management • United States

On-site
USD 130,000 - 170,000
Staff Data Engineer
Staff Data Engineer

Newmark Group • Dallas (TX)

Hybrid
USD 190,000 - 250,000