Data Engineer, AI Support

Insurance Institute for Business & Home Safety

Richburg (SC)

On-site

USD 110,000 - 190,000

Full time

13 days ago
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

Health Plan
401k/IRA
Life Insurance
Paid Time Off
Family Leave
Disability Insurance
Training & Development

Job summary

The Insurance Institute for Business & Home Safety (IBHS) is seeking a Data Engineer, AI Support to translate enterprise needs into scalable data solutions for AI/ML workloads. You will collaborate with AI researchers and data engineers to prepare datasets, build pipelines, and ensure data quality and reproducibility across research projects.

Join a team focused on responsible AI adoption, data governance, and experimental workflows while improving decision-making and operational capabilities

Qualifications

  • Bachelor’s degree in data science, statistics, computer science, mathematics, engineering, or a related quantitative field.
  • Experience building ETL/data pipelines to clean, transform, integrate, and prepare structured, semi-structured, and unstructured data for AI/ML workflows.
  • Strong Python data manipulation skills, including efficient use of vectorized libraries for large-scale data processing, exploration, and visualization.
  • Working knowledge of SQL, NoSQL, data modeling, columnar formats, and modern data storage technologies.
  • Familiarity with preparing and versioning LLM/VLM training and evaluation datasets, including QA, preference/RL, multimodal, and human-annotated data.

Responsibilities

  • Work closely with AI researchers and Data Engineering to prepare, organize, and improve the data used across research and machine learning projects.
  • Build reliable data workflows that move research data from raw sources into usable datasets for analysis, experimentation, training, and evaluation.
  • Explore and understand new datasets, identify quality issues or gaps, and help determine the best way to structure and use the data.
  • Support the preparation of datasets for a range of AI applications, including language, vision, multimodal, and retrieval-based systems.
  • Help ensure research datasets are consistent, traceable, reproducible, and well documented as they evolve over time.
  • Automate recurring data preparation and processing tasks to make research workflows more efficient and repeatable.
  • Develop clear summaries and visualizations that help the team understand datasets, patterns, and potential issues.
  • Support the integration of data across research tools, internal systems, and AI platforms.
  • Help protect sensitive information and follow appropriate data handling practices throughout the data lifecycle.
  • Contribute to an experimental research environment where datasets, methods, and requirements may change as projects develop.
  • Stay current on relevant developments in artificial intelligence, machine learning, data science, and emerging analytical technologies.

Skills

Python data manipulation
SQL
NoSQL
Data modeling
Linux Bash
CI/CD
Airflow
Experiment tracking
Langfuse/Weights & Biases

Education

Bachelor’s degree in data science, statistics, CS, math, engineering
Master’s degree (preferred) in data science or AI fields

Tools

Airflow
Dagster
Spark
Terraform
GitHub Actions

Job description

About the Role

The Data Engineer, AI Support is a key technical partner in IBHS’s responsible adoption and application of artificial intelligence, machine learning, and advanced analytics. This position translates enterprise-wide needs across IBHS into practical, scalable data solutions that improve decision‑making, operational effectiveness, and employee capabilities.

The role combines Data Engineering expertise, solution development, technical consultation, and employee support. It works across business and technical teams to evaluate opportunities, develop and implement solutions, assess performance and risk, and help employees use approved AI‑enabled tools effectively.

Why This Role Matters

This role strengthens IBHS’s enterprise-wide ability to use data and artificial intelligence thoughtfully, responsibly, and effectively. By connecting organizational priorities with technical capabilities, the Data Engineer, AI Support helps IBHS identify valuable use cases, improve access to actionable information, and build confidence in AI‑enabled solutions.

The position also helps establish consistent practices for solution quality, documentation, data handling, human review, and responsible AI use.

What You’ll Do
  • Work closely with AI researchers and Data Engineering to prepare, organize, and improve the data used across research and machine learning projects.
  • Build reliable data workflows that move research data from raw sources into usable datasets for analysis, experimentation, training, and evaluation.
  • Explore and understand new datasets, identify quality issues or gaps, and help determine the best way to structure and use the data.
  • Support the preparation of datasets for a range of AI applications, including language, vision, multimodal, and retrieval-based systems.
  • Help ensure research datasets are consistent, traceable, reproducible, and well documented as they evolve over time.
  • Automate recurring data preparation and processing tasks to make research workflows more efficient and repeatable.
  • Develop clear summaries and visualizations that help the team understand datasets, patterns, and potential issues.
  • Support the integration of data across research tools, internal systems, and AI platforms.
  • Help protect sensitive information and follow appropriate data handling practices throughout the data lifecycle.
  • Contribute to an experimental research environment where datasets, methods, and requirements may change as projects develop.
  • Stay current on relevant developments in artificial intelligence, machine learning, data science, and emerging analytical technologies.
What We’re Looking For
  • Bachelor’s degree in data science, statistics, computer science, mathematics, engineering, or a related quantitative field.
  • Experience building ETL/data pipelines to clean, transform, integrate, and prepare structured, semi-structured, and unstructured data for AI/ML workflows.
  • Strong Python data manipulation skills, including efficient use of vectorized libraries for large-scale data processing, exploration, and visualization.
  • Working knowledge of SQL, NoSQL, data modeling, columnar formats, and modern data storage technologies.
  • Familiarity with preparing and versioning LLM/VLM training and evaluation datasets, including QA, preference/RL, multimodal, and human-annotated data.
  • Familiarity with embedding pipelines, vector databases, semantic search, RAG, and metadata-aware retrieval workflows.
  • Exposure to graph databases, knowledge graphs, and graph-based data modeling for AI applications.
  • Understanding of data quality, schema validation, dataset versioning, metadata, lineage, and reproducible train/validation/test splits with leakage prevention.
  • Familiarity with distributed data processing and workflow orchestration concepts such as DAGs, task dependencies, scheduling, and pipeline monitoring.
  • Comfortable working in Linux environments with Bash/shell scripting and basic automation.
  • Familiarity with experiment tracking and LLM observability tools such as Weights & Biases and Langfuse.
  • Basic understanding of PII handling, masking, hashing, tokenization, and de-identification within data pipelines.
  • Familiarity with CI/CD and infrastructure automation tools such as GitHub Actions, GitLab CI, and Terraform.
  • Comfortable working with research datasets that may be incomplete, inconsistent, or evolving, and able to investigate the data before implementing a solution.
  • Strong written communication, presentation, and technical-documentation skills.
  • Ability to build effective working relationships across business and technical functions.
  • Ability to exercise sound judgment, manage multiple priorities, and work independently while contributing to cross-functional initiatives.
Preferred Qualifications
  • Master’s degree in data science, statistics, computer science, artificial intelligence, machine learning, or a related field.
  • Experience supporting AI/ML research, scientific computing, or other data-intensive research environments.
  • Hands‑on experience with cloud data platforms or services in Azure, AWS, or Google Cloud.
  • Experience with data orchestration and distributed processing tools such as Airflow, Prefect, Dagster, Spark, or similar technologies.
  • Familiarity with data annotation and human-in-the-loop platforms such as Label Studio or Prodigy, particularly for machine learning or multimodal datasets.
  • Health Care Plan (Medical, Dental & Vision)
  • Retirement Plan (401k, IRA)
  • Life Insurance (Basic, Voluntary & AD&D)
  • Paid Time Off (Vacation, Sick & Public Holidays)
  • Family Leave (Maternity, Paternity)
  • Short Term & Long Term Disability
  • Training & Development
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Data Engineer, AI Support
Data Engineer, AI Support

Ibhs • Northern (KY)

Hybrid
USD 120,000 - 180,000
Health Insurance
401k/IRA
Life Insurance
+4
AI Analyst
AI Analyst

Maquoketa Area Chamber of Commerce • Maquoketa (IA)

On-site
USD 90,000 - 140,000
Health Insurance
Dental Insurance
401k Retirement Plan
+8
Senior Staff Data Scientist
Senior Staff Data Scientist

Ritchie Bros. • Westchester (IL)

On-site
USD 150,000 - 180,000
Medical insurance
Dental insurance
Vision insurance
+2
AI Data Engineer
AI Data Engineer

C the Signs • Boston (MA)

Hybrid
USD 120,000 - 160,000
Flexible work options
Competitive salary
Healthcare benefits
+1
AI & Machine Learning Engineer
AI & Machine Learning Engineer

OneBlood • Saint Petersburg (FL)

On-site
USD 110,000 - 170,000
AI Engineer
AI Engineer

Goal Solutions • San Diego (CA)

Hybrid
USD 120,000 - 180,000
Competitive salary
401(k) with company match
Medical, dental, vision, and HSA
+2
Senior Data Scientist III
Senior Data Scientist III

LexisNexis • United States

Hybrid
USD 120,000 - 150,000
401(k) with match
Wellness platform
Short-and-Long Term Disability
+1
AI Data Engineer for Research Data Pipelines
AI Data Engineer for Research Data Pipelines

Ibhs • Northern (KY)

Hybrid
USD 120,000 - 180,000
Health Insurance
401k/IRA
Life Insurance
+4
AI Engineer
AI Engineer

InterImage, Inc. • Annapolis (MD)

On-site
USD 100,000 - 130,000
401K with up to 3% discretionary profit sharing
20 days PTO
Free healthcare for single participant
+1
AI Data Architect
AI Data Architect

3Pillar • United States

On-site
USD 180,000 - 250,000
Medical Insurance
Dental Insurance
Vision Insurance
+8