Lead Data Engineer – AI/Machine Learning

Core Specialty

Cincinnati (OH)

Hybrid

USD 150,000 - 210,000

Full time

3 days ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

Medical, dental, vision, and life ins.
Disability insurance
401(k) company-match
Employee Assistance Program
Health Savings Account
Flexible Spending Account
Wellness program
Professional development

Job summary

Core Specialty is seeking a Lead AI/ML Data Engineering professional in Cincinnati, OH to drive AI/ML enablement across the organization. You will design and optimize data pipelines, own end-to-end implementations, and define scalable AI infrastructure patterns while collaborating with Data Governance and Enterprise Architecture.

You will lead the definition of AI/ML frameworks, evaluate tools such as LangChain and vector databases, and build prototypes with production-grade practices.

Qualifications

  • Data pipeline design, optimization, and distributed processing (Spark, dbt, Airflow, Kafka, etc.)
  • Hands-on with Snowflake, Databricks, and/or Azure Synapse Analytics; architect/optimize workloads
  • Strong knowledge of AWS, Azure, or GCP and modern data warehouse architectures
  • Strong Python skills; production-grade software engineering practices
  • Experience building ML/AI systems via orchestration, tool-calling, retrieval systems
  • Practical experience with LLM APIs (OpenAI) and open-weight models
  • Prompt engineering and evaluation as established discipline
  • Understanding of context windows, tokenization, embeddings, latency and cost tradeoffs
  • Experience with vector databases and embedding models
  • Chunking strategies, hybrid search, and reranking
  • Agent frameworks (LangChain, LangGraph, LlamaIndex)
  • Function-calling design and multi-step reasoning
  • Model optimization: fine-tune, prompt, or RAG
  • Familiarity with LoRA and parameter-efficient methods
  • MLOps/LLMOps practices and governance awareness
  • Model evaluation, A/B testing, observability, tracing/logging of model calls
  • Latency/cost optimization in deployment; caching and streaming responses
  • Versioning prompts and models; governance of safety and PIIs

Responsibilities

  • Design, build, and optimize data pipelines and platform components for analytics and AI/ML use cases
  • Own complex initiatives from design through rollout with minimal oversight
  • Resolve performance, scalability, reliability issues across data platform
  • Proactively propose improvements and fill gaps in data engineering capabilities
  • Produce clean, well-tested code and infrastructure-as-code with engineering hygiene
  • Define AI/ML frameworks; evaluate tools, platforms, and standards
  • Build working AI/ML prototypes delivering immediate value
  • Shape MLOps strategy including deployment, monitoring, versioning, lifecycle management
  • Collaborate with Data Governance to ensure alignment with governance and compliance
  • Advocate scalable AI infrastructure patterns (feature stores, governed datasets, streaming for training/inference)
  • Partner with Data Science/Engineering/business to assess readiness gaps and roadmap
  • Serve as SME and thought partner on emerging AI/ML technologies and industry trends
  • Document AI/ML standards and decisions for consistent adoption
  • Provide architectural guidance and best practices for AI readiness and ML Ops
  • Collaborate with Enterprise Architecture on architectural blueprints for AI readiness
  • Other duties as assigned

Skills

Data engineering
Data platforms
Cloud
Python
API integration
LLM experience
Prompt engineering
Model fundamentals
Vector search
Search techniques
Agent orchestration
Tool use/agents
Model adaptation
LoRA
MLOps/LLMOps
Evaluation/observability
Deployment patterns
Versioning
Safety/governance

Education

Bachelor's degree

Tools

Snowflake
Databricks
Azure Synapse Analytics
AWS
Azure
GCP
Pinecone
Weaviate
pgvector
LangChain
LangGraph
LlamaIndex
LoRA
MLOps/LLMOps
Data Vault 2.0
Ensemble data modeling

Job description

Lead AI/Machine Learning Data Engineering role partnering with the VP, Head of Data to drive AI/ML enablement and readiness across the organization.

Responsibilities
  • Design, build, and optimize data pipelines, ingestion frameworks, and platform components for analytics, reporting, and AI/ML use cases.
  • Operate with autonomous ownership of complex initiatives, owning technical design through implementation and rollout with minimal oversight.
  • Identify and resolve performance, scalability, and reliability issues across the existing data platform.
  • Propose improvements proactively by recognizing gaps in data engineering and platform capabilities.
  • Produce clean, well-tested, well-documented code and infrastructure-as-code while maintaining strong engineering hygiene.
  • Help define AI/ML frameworks by evaluating and recommending tools, platforms, and standards for building and deploying AI/ML solutions.
  • Build working AI/ML prototypes that deliver immediate value to engineering teams.
  • Shape and support implementation of MLOps strategy, including model deployment, monitoring, versioning, and lifecycle management.
  • Collaborate with Data Governance to ensure AI/ML frameworks align with data governance, security, and compliance requirements.
  • Design and advocate for scalable AI/ML infrastructure patterns such as feature stores, curated/governed datasets, and streaming access for training and inference.
  • Partner with Data Science, Data Engineering, and business stakeholders to assess AI/ML readiness gaps and build a roadmap to close them.
  • Serve as a subject-matter expert and thought partner to the VP, Head of Data on emerging AI/ML technologies, practices, and industry trends.
  • Document AI/ML standards, frameworks, and decisions to support consistent adoption as practices mature.
  • Act as a senior technical resource for architecture guidance, design patterns, and best practices for AI readiness and ML Ops frameworks.
  • Partner with Enterprise Architecture on establishing architectural blueprints for AI readiness.
  • Other Duties as Assigned.
Requirements
  • Data engineering fundamentals: deep expertise in data pipeline design, optimization, and distributed data processing (Spark, dbt, Airflow, Kafka, or equivalent).
  • Data platforms: hands‑on experience with Snowflake, Databricks, and/or Azure Synapse Analytics, with the ability to architect and optimize workloads.
  • Cloud: strong knowledge of AWS, Azure, or GCP, plus modern data warehouse/lakehouse architectures.
  • Python: strong Python skills and production‑grade software engineering practices (testing, version control, code review) for shipping systems beyond notebooks.
  • API and integration: experience building systems around models via orchestration, tool‑calling, and retrieval systems.
  • LLM experience: practical experience with LLM APIs (e.g., OpenAI) and open‑weight models.
  • Prompt engineering: prompt engineering and evaluation as an established discipline.
  • Model fundamentals: understanding of context windows, tokenization, embeddings, and limitations such as hallucination, latency, and cost tradeoffs.
  • Vector search: experience with vector databases (Pinecone, Weaviate, pgvector, etc.) and embedding models.
  • Search techniques: chunking strategies, hybrid search, and reranking.
  • Agent and orchestration: frameworks such as LangChain, LangGraph, LlamaIndex, or custom orchestration.
  • Tool use / agents: function‑calling design, multi‑step reasoning chains, and agent memory/state management.
  • Model adaptation: practical understanding of when to fine‑tune vs. prompt vs. RAG.
  • Parameter‑efficient methods: familiarity with LoRA and related approaches.
  • MLOps / LLMOps: experience with MLOps/LLMOps practices.
  • Evaluation and observability: model evaluation frameworks, A/B testing for model outputs, and observability (tracing, logging model calls).
  • Deployment patterns: latency/cost optimization, caching, streaming responses, and fallback handling.
  • Versioning: versioning prompts and models, not only code.
  • Safety and governance: safety, evaluation, governance awareness; bias/safety evaluation and appropriate handling of PII.
Technologies
  • Spark, dbt, Airflow, Kafka
  • Snowflake, Databricks, Azure Synapse Analytics
  • AWS, Azure, GCP
  • Python, OpenAI
  • Pinecone, Weaviate, pgvector
  • LangChain, LangGraph, LlamaIndex
  • LoRA
  • MLOps, LLMOps
  • Data Vault 2.0, Ensemble data modeling techniques
Experience
  • Minimum 7+ years in data engineering with experience on large‑scale, mature data platforms.
  • 3+ years developing ML or AI deliverables, including deployment to production.
  • Bachelor's degree in a related field or demonstrated equivalent experience in a related field required.
  • Working knowledge of agentic workflows for engineering and architecture.
  • Demonstrated autonomous technical ownership of complex projects from design through delivery with minimal oversight.
  • Experience shaping AI/ML enablement (defining frameworks, evaluating MLOps tooling, or building infrastructure for training and deployment).
  • Experience partnering with Data Governance, Data Science, or Compliance to align technical practices with governance and regulatory requirements.
  • Track record of proposing and driving innovative technical solutions rather than only executing predefined plans.
  • Experience designing or implementing agentic workflows for data engineering preferred.
  • Experience working with Property & Casualty insurance carriers preferred.
  • Experience with Data Vault 2.0 or Ensemble data modeling techniques preferred.
Location
  • Cincinnati, OH (hybrid)
Benefits
  • Medical, dental, vision, and life insurances
  • Short and long‑term disability
  • Company‑match of 100% of a 6% contribution 401(k) plan
  • Employee Assistance Plan
  • Health Savings Account
  • Flexible Spending Account
  • Health Reimbursement Account
  • Wellness program
  • Opportunities for professional development and advancement
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Lead Data Engineer – AI/Machine Learning
Lead Data Engineer – AI/Machine Learning

Kalepa • Cincinnati (OH), Northern (KY)

Hybrid
USD 140,000 - 210,000
Medical insurance
Dental insurance
Vision insurance
+4
Lead Engineer - Data Engg & AI
Lead Engineer - Data Engg & AI

Anblicks Inc. • Dallas (TX), Northern (KY)

On-site
USD 150,000 - 230,000
Lead Data Scientist Applied AI - USA
Lead Data Scientist Applied AI - USA

Socket.dev • Santa Clara (CA)

On-site
USD 150,000 - 170,000
Unlimited PTO
Generous parental leave
Annual bonus program
+3
Lead Engineer - Data Engg & AI
Lead Engineer - Data Engg & AI

Anblicks • Dallas (TX)

On-site
USD 150,000 - 190,000
AI/ML Data Engineer (Fulltime)
AI/ML Data Engineer (Fulltime)

Aptonet • Tampa (FL)

Hybrid
USD 100,000 - 130,000
Employer-matched 401(k)
Company-paid medical insurance
Company-paid vision insurance
+7
Data Engineer (AI) - Cincinnati
Data Engineer (AI) - Cincinnati

Medpace • Cincinnati (OH)

On-site
USD 75,000 - 90,000
Flexible work environment
Competitive compensation and benefits package
Competitive PTO packages
+3
Principal Engineer
Principal Engineer

your Jared • Northern (KY), San Diego (CA)

Hybrid
USD 180,000 - 240,000
Fully remote, work from home
Employee Share Option Plan
Flexible working hours
+4
AI Data Engineer (US)
AI Data Engineer (US)

AVP VIGILANT TECHNOLOGY PVT LTD • San Francisco (CA)

On-site
USD 115,000 - 195,000
Health, dental, and vision benefits
401(k) retirement benefits
Paid time off and holidays
+2
Senior Data Engineer - AI Infrastructure Integration, High Performance Compute
Senior Data Engineer - AI Infrastructure Integration, High Performance Compute

Bank of America • New York (NY)

On-site
USD 170,000 - 210,000
Engineer II - Machine Learning
Engineer II - Machine Learning

PODS • Clearwater (FL)

On-site
USD 120,000 - 170,000