Data Platform & Machine Learning Engineer

Leon Capital Group

Dallas (TX)

On-site

USD 120,000 - 150,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Leon Capital Group, based in Dallas, Texas, is looking for a Data Platform & Machine Learning Engineer to design, build, and maintain the data platform for AI products. This role involves both data engineering and model building, creating a strong foundation for the company's AI initiatives.

The ideal candidate will have over 5 years of experience in data systems, strong data engineering fundamentals, and a hands-on approach to document extraction. Knowledge of cloud environments is essential. Join a team that values innovation and practical solutions.

Qualifications

  • 5+ years building data systems end-to-end.
  • Proven experience integrating against messy enterprise systems.
  • Comfort owning data infrastructure in a major cloud environment.
  • Applied machine learning depth and judgment regarding model risk.

Responsibilities

  • Design and own the canonical data model and warehouse.
  • Build reusable data platform and models for portfolio companies.
  • Build ingestion pipelines for various data sources.
  • Lay groundwork for risk and recommendation models.
  • Partner with engineers for model and data integration.

Skills

Data systems building
Data engineering fundamentals
PostgreSQL
Hands-on document extraction
Cloud environments (Azure, AWS)
Applied machine learning

Job description

About the Role

Leon Capital Group is a multi-billion-dollar holding company with operating businesses across healthcare, real estate, and financial services. Within the Financial Services Group (FSG), we are building an AI-native operating model from the ground up: a dedicated, full‑lifecycle innovation function that identifies, designs, builds, delivers, and implements AI products across our portfolio companies.

This is two halves of one job: The Data Platform & Machine Learning Engineer will own the data platform every FSG AI product depends on, and you build the first machine learning models on top of it. The two are inseparable, because the models are only as good as the proprietary, provenance‑tracked data beneath them, and that data is yours to design, capture, and own. Not a pipeline specialist, not a modeling specialist, but the engineer who does both, on a foundation they stood up themselves.

The role has a deliberate arc. The first six months lean heavily toward data engineering: the canonical model and warehouse, ingestion that forces messy external data into our schema, and judgment capture that turns every human decision into labeled training data. As that corpus matures, the center of gravity shifts toward building the first risk and recommendation models on it. You work across our portfolio companies, standing up each one’s own data layer and models, and reusing the same patterns and discipline rather than reinventing the approach each time. We want someone energized by both halves, not someone tolerating the foundation to reach the models, or the reverse.

Key Responsibilities
  • Design and own the canonical data model and warehouse, built to remain ours regardless of which vendors we buy.
  • Build the data platform and models as reusable patterns, standing up each portfolio company's own instance rather than a bespoke build each time.
  • Build ingestion pipelines that pull from hostile, heterogeneous sources into our schema with full provenance, including vendor servicing output, flat‑file, and SFTP partner feeds with no API, and public registry data.
  • Stand up judgment capture from the first transaction: structure every decision and correction as labeled training data, the raw material that human‑in‑the‑loop workflows and future models both depend on.
  • Lay the groundwork for risk and recommendation models: feature engineering, training‑set construction, model evaluation, and calibration.
  • Build validation, monitoring, and lineage so data‑quality issues are caught before they reach models or decisions.
  • Enforce the data‑ownership bar in every buy decision: vendor output must land in our canonical structure, in our schema, and remain portable on exit.
  • Partner closely with the Forward‑Deployed Engineer, providing the models, data, and serving interfaces their decision surfaces are built on.
  • Partner with shared IT and security on regulated‑data handling, secrets management (e.g., Doppler), and compliance prerequisites.
Qualifications
Required
  • 5+ years building data systems end‑to‑end, ideally as an early or founding data hire at a startup where you owned the whole data function rather than one stage of a large team.
  • Strong data engineering fundamentals: schema design, ETL and ELT pipeline architecture, and production experience with PostgreSQL.
  • Proven experience integrating against messy enterprise systems and against flat‑file or SFTP feeds with no clean API.
  • Hands‑on document extraction and structuring experience.
  • Comfort owning data infrastructure in a major cloud environment (we use a mix across Azure and AWS); you can stand up and run the warehouse and pipelines without a dedicated platform team.
  • Comfort building under regulatory and fiduciary constraints, and handling sensitive, regulated data (PHI, financial) with the secure practices it requires.
  • Applied machine learning and statistical modeling depth: feature work, model evaluation, calibration, and the judgment to reason about model risk. You orchestrate and apply models; you do not need to publish research.
  • Working fluency with LLM orchestration (e.g., LangGraph), retrieval‑augmented generation, and disciplined prompt and evaluation practices.
  • Excellent judgment under ambiguity and the ability to prioritize across multiple portfolio companies without close oversight.
Preferred
  • Regulated‑domain experience – insurance, reinsurance, healthcare, or financial services – where you have built to compliance constraints.
  • You have been the person the whole company's data depended on, at a place small enough that there was no one else to do it.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Data Engineer
Senior Data Engineer

Leon Capital Group • Dallas (TX)

On-site
USD 140,000 - 190,000
Data Platform Engineering Lead
Data Platform Engineering Lead

Hedge Fund • New York (NY)

On-site
USD 150,000 - 200,000
Senior Data Engineer (AI-Native) — Data Layer
Senior Data Engineer (AI-Native) — Data Layer

Proton.ai • Boston (MA)

Hybrid
USD 140,000 - 210,000
Lead Data Engineer, Data Platform
Lead Data Engineer, Data Platform

crewAI, Inc. • San Francisco (CA)

On-site
USD 140,000 - 230,000
Sr. Manager, Data & Analytics
Sr. Manager, Data & Analytics

Specialized Bicycle Components, Inc. • Morgan Hill (CA)

On-site
USD 140,000 - 180,000
Senior Data Engineer
Senior Data Engineer

Assembl • New York (NY)

On-site
USD 140,000 - 190,000
Data Platform & ML Engineer: Build AI-Ready Data Foundation
Data Platform & ML Engineer: Build AI-Ready Data Foundation

Leon Capital Group • Dallas (TX)

On-site
USD 120,000 - 150,000
Senior Consultant, AI/ML Engineer
Senior Consultant, AI/ML Engineer

Hollstadt Consulting • Minnesota

On-site
USD 140,000 - 200,000
Data Engineer
Data Engineer

FountAI, Inc. • New York (NY)

On-site
USD 100,000 - 140,000
Full Lifecycle Data Engineer
Full Lifecycle Data Engineer

Lockton • Kansas City (MO)

On-site
USD 110,000 - 160,000