Data Engineer (AI/ML)

Blue Cross Blue Shield Association

Chicago (IL)

On-site

USD 100,800 - 138,600

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Paid time off
Medical/dental/vision insurance
Generous 401(k) matching
Lifestyle spending account

Job summary

The Blue Cross Blue Shield Association is looking for a Data Engineer to design and optimize data pipelines for machine learning and AI workloads. This role involves collaborating with architects and data scientists to ensure pipeline accuracy and compliance with healthcare regulations.

Ideal candidates will have 5+ years of experience in data engineering and expertise in tools like AWS and PySpark. The role offers a competitive salary and a range of employee benefits.

Qualifications

  • 5+ years of experience in data engineering, building and managing pipelines.
  • Hands-on experience with AWS AI/ML and data services.
  • Experience working with healthcare datasets preferred.

Responsibilities

  • Design, build, and maintain reliable data pipelines.
  • Collaborate with teams to translate business needs into solutions.
  • Support compliance with SOC 2, HIPAA, and GDPR.

Skills

Data pipeline management
Machine Learning (ML)
Generative AI (GenAI)
Python
SQL
AWS
Collaboration
Problem-solving

Education

Bachelor’s or Master’s degree in Computer Science, Engineering, or related field

Tools

AWS Glue
Databricks
Airflow
Kubernetes
Snowflake
PySpark

Job description

Job Description Summary

The Data Engineer will design, build, and optimize scalable, secure data pipelines that power analytics and product platforms. For this role specifically, the focus will be on Machine Learning (ML) and Generative Artificial Intelligence (GenAI) workloads, while contributing to innovation and ensuring compliance with healthcare industry standards. This role is expected to provide strong hands‑on technical expertise, collaborate across teams, and contribute to architecture decisions that align engineering practices with organizational goals.

Job Description
  • Design, build, and maintain reliable, high-performance data pipelines for large‑scale structured and unstructured healthcare data.
  • Use PySpark and modern cloud‑based tools (Databricks, AWS Glue, EMR, Snowflake) to transform and process data efficiently.
  • Support ingestion, transformation, and validation processes that ensure data consistency, integrity, and availability.
  • Partner with Data Architects, Data Scientists, and Analysts to translate business needs into scalable engineering solutions.
  • Collaborate with platform and DevOps teams to deploy, scale, and monitor data pipelines using Airflow and Kubernetes.
  • Participate in code reviews, documentation, and continuous improvement efforts across the engineering team.
  • Implement and maintain data validation frameworks to ensure pipeline accuracy and completeness.
  • Contribute to best practices in version control, metadata management, and reproducibility.
  • Stay current with emerging technologies in data engineering and cloud computing, recommending improvements to existing infrastructure.
  • Participate in performance tuning, cost optimization, and scaling strategies for cloud‑based data systems.
  • Identify automation opportunities to streamline ETL/ELT processes and reduce operational overhead.
  • Share knowledge and mentor junior team members on tools, techniques, and best practices.
  • Promote a culture of collaboration, innovation, and continuous learning within the engineering organization.
  • Support compliance with SOC 2, HIPAA, and GDPR by adhering to established data privacy and security practices.
Posting Range For This Position Is

100,800.00 - 138,600.00

Required Education, Certifications and Experience

Bachelor’s or Master’s degree in Computer Science, Engineering, or related field.

Required Experience

5+ years of experience in data engineering, including building and managing pipelines in cloud‑based environments.

Knowledge, Skills, and Abilities
  • Experience with building and operationalizing the data foundations that support machine learning and generative AI use cases, including feature pipelines, training/inference data preparation, and retrieval‑ready datasets (e.g., embeddings and vector stores).
  • Familiarity with GenAI skills and adjacent tooling (foundation models, prompt engineering, RAG, embeddings/vector databases, and GenAI orchestration frameworks).
  • Hands‑on experience with AWS AI/ML and data services, including Amazon Bedrock, Bedrock Agent Core, SageMaker, Glue, and EMR.
  • Experience designing and optimizing data architectures, including data foundations that support ML and GenAI workloads.
  • Hands‑on experience with workflow orchestration (Airflow) and containerization (Kubernetes).
  • Hands‑on technical expertise, cross‑team collaboration, and contributing to architecture decisions.
  • Proficiency in Python, SQL, and distributed data frameworks (PySpark, Databricks, AWS Glue, EMR).
  • Working knowledge of cloud platforms (AWS or Azure) and data warehouses (Snowflake).
  • Familiarity with NoSQL and relational databases, as well as data modeling best practices.
  • Strong analytical, problem‑solving, and communication skills.
  • Understanding of compliance frameworks (SOC 2, HIPAA) and secure data management principles.
  • Experience working with healthcare datasets or knowledge of healthcare standards (HIPAA, HL7, FHIR) preferred.

The posted salary range is the lowest to highest salary we, in good faith, believe we would pay for this role at the time of this posting. We may ultimately pay more or less than the hiring range and this hiring range may also be modified in the future. A candidate’s position within the hiring range may be based on several factors including, but not limited to, specific competencies, relevant education, qualifications, certifications, relevant experience, skills, seniority, performance, shift, travel requirements, and business or organizational needs. This job is also eligible for annual bonus incentive pay.

We offer a comprehensive package of benefits including paid time off, 11 holidays, medical/dental/vision insurance, generous 401(k) matching, lifestyle spending account and many other benefits to eligible employees.

Note: No amount of pay is considered to be wages or compensation until such amount is earned, vested, and determinable. The amount and availability of any bonus, commission, or any other form of compensation that are allocable to a particular employee remains in the Company's sole discretion unless and until paid and may be modified at the Company’s sole discretion, consistent with the law.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Data Engineer (AI/ML)
Data Engineer (AI/ML)

Blue Cross Blue Shield • Chicago (IL)

Hybrid
USD 100,000 - 139,000
Paid time off
Medical/dental/vision insurance
Generous 401(k) matching
+1
Data Engineer (AI/ML)
Data Engineer (AI/ML)

001_BCBSA Blue Cross and Blue Shield Association • Chicago (IL)

On-site
USD 100,000 - 139,000
Paid time off
Medical/dental/vision insurance
Generous 401(k) matching
+1
Data Engineer (4631)
Data Engineer (4631)

Hireclout • Los Angeles (CA)

On-site
USD 150,000 - 220,000
100% employer-paid medical, dental, and vision coverage
401(k) with employer match
Generous PTO policy
+1
Data Engineer
Data Engineer

Appsierra Group • United States

On-site
USD 140,000 - 180,000
Equity
Performance bonuses
Health insurance reimbursement
+3
Data Engineer
Data Engineer

candidhealth • San Francisco (CA)

On-site
USD 165,000 - 205,000
Potential equity in compensation package
Sales incentives and employee benefits
Senior Software Data Architect
Senior Software Data Architect

Complexcare Solutions, Inc. • Northern (KY)

Hybrid
USD 172,000 - 200,000
Senior AI Data Scientist I
Senior AI Data Scientist I

Exelixis • Alameda (CA)

On-site
USD 143,000 - 203,000
401(k) with company contributions
Health, dental, vision
Life and disability insurance
+2
Data Analytics Engineer
Data Analytics Engineer

EXL • San Francisco (CA)

On-site
USD 140,000 - 160,000
Senior Data Engineer
Senior Data Engineer

Cacheflow • New York (NY)

On-site
USD 150,000 - 180,000
18 vacation days
9 company holidays
5 sick days
+4
Data Engineer (Multiple Levels)
Data Engineer (Multiple Levels)

Astrana Health, Inc. • Alhambra (CA)

On-site
USD 105,000 - 115,000