Senior AI Engineer with Spark, AWS Services

EPAM Systems Inc

United States

Remote

USD 150,000 - 210,000

Full time

3 days ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

EPAM Systems is seeking a Senior AI Engineer to design and maintain data pipelines powering AI/GenAI applications for Risk-Based Quality Management in clinical trials.

Focus areas include RAG document ingestion, vector indexing, and AWS-based data infrastructure. The role requires strong Python, Spark SQL, and OpenSearch skills and collaboration with a cross-functional team in a fast-paced environment.

Qualifications

  • 5+ years of hands-on data engineering experience at scale.
  • Expertise in RAG document ingestion pipelines (chunking, embedding, vector indexing).
  • Proficiency in AWS OpenSearch as a vector database for RAG workflows.
  • Advanced proficiency in Python, including SQL and Spark SQL.
  • Skills in unstructured data transformation (PDF, DOCX) for RAG/LLM applications.
  • Familiarity with AWS Services: S3, Lambda, Glue, Athena, Bedrock, Step Functions, API Gateway, CloudWatch, DynamoDB.
  • Knowledge of containerization with Docker.
  • Proficiency in English at a B2+ level.
  • Nice to have Background in pharmaceutical or life sciences domain.

Responsibilities

  • Design and build RAG document ingestion pipelines (chunking, embedding, vector indexing) for clinical trial quality data.
  • Build and manage vector databases (AWS OpenSearch) for RAG-powered AI workflows.
  • Develop batch and streaming ETL/ELT pipelines from scratch for unstructured clinical data (PDF, DOCX, clinical reports).
  • Build and expose data APIs for AI application consumption.
  • Optimize chunking strategies, embedding generation, and retrieval performance for RAG architectures.
  • Manage data quality, lineage, and governance for AI/ML data pipelines.
  • Deploy and maintain AWS data infrastructure (S3, Lambda, Glue, Athena, Step Functions, DynamoDB).
  • Collaborate with Data Scientists and Backend Developers in an integrated pod team.

Skills

RAG pipelines
AWS OpenSearch
Python
SQL
Spark SQL
Unstructured data
Docker
CI/CD
CDK
Terraform
CDISC standards
SageMaker
Snowflake
Pinecone
APIs
English (B2+)

Tools

Snowflake
Pinecone
SageMaker
OpenSearch
Docker
Terraform
CDK

Job description

Introduction

We are seeking a Senior AI Engineer with Spark and AWS Services expertise to join the RBQM Production Pod within the program. You will build and maintain data pipelines that power AI/GenAI applications for Risk-Based Quality Management in clinical trials. This role focuses on RAG document ingestion, vector indexing, and building data APIs for AI applications.

Responsibilities
  • Design and build RAG document ingestion pipelines (chunking, embedding, vector indexing) for clinical trial quality data
  • Build and manage vector databases (AWS OpenSearch) for RAG-powered AI workflows
  • Develop batch and streaming ETL/ELT pipelines from scratch for unstructured clinical data (PDF, DOCX, clinical reports)
  • Build and expose data APIs for AI application consumption
  • Optimize chunking strategies, embedding generation, and retrieval performance for RAG architectures
  • Manage data quality, lineage, and governance for AI/ML data pipelines
  • Deploy and maintain AWS data infrastructure (S3, Lambda, Glue, Athena, Step Functions, DynamoDB)
  • Collaborate with Data Scientists and Backend Developers in an integrated pod team
Requirements
  • 5+ years of hands-on data engineering experience at scale
  • Expertise in RAG document ingestion pipelines (chunking, embedding, vector indexing)
  • Proficiency in AWS OpenSearch as a vector database for RAG workflows
  • Advanced proficiency in Python, including SQL and Spark SQL
  • Skills in unstructured data transformation (PDF, DOCX) for RAG/LLM applications
  • Familiarity with AWS Services: S3, Lambda, Glue, Athena, Bedrock, Step Functions, API Gateway, CloudWatch, DynamoDB
  • Knowledge of containerization with Docker
  • Capability to build custom pipelines from scratch, beyond configuring out-of-the-box services
  • Proficiency in English at a B2+ level
  • Nice to have Background in pharmaceutical or life sciences domain
  • Familiarity with Snowflake, Pinecone (vector DB alternative)
  • Knowledge of SageMaker processing jobs
  • Skills in CI/CD tools (Jenkins, Git/Bitbucket) and infrastructure tools (CDK or Terraform)
  • Understanding of clinical data standards (CDISC, ADaM, SDTM)
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior AI Engineer: AWS Data Pipelines & RAG (Spark)
Senior AI Engineer: AWS Data Pipelines & RAG (Spark)

EPAM Systems Inc • United States

Remote
USD 150,000 - 210,000
ai engineer for clinical trials
ai engineer for clinical trials

HireHi • United States

Remote
USD 120,000 - 180,000
Senior AI Data Engineer for Clinical Trials & RAG Pipelines
Senior AI Data Engineer for Clinical Trials & RAG Pipelines

HireHi • United States

Remote
USD 120,000 - 180,000
AI Engineer – AWS Bedrock, Claude & Databricks
AI Engineer – AWS Bedrock, Claude & Databricks

VeeAR Projects Inc. • Raleigh (NC)

On-site
USD 120,000 - 150,000
Senior AI, Data Engineer
Senior AI, Data Engineer

Sud Recruiting • New York (NY)

On-site
USD 150,000 - 210,000
Senior Platform Data Engineer
Senior Platform Data Engineer

Geisinger • Danville (PA)

On-site
USD 100,000 - 130,000
Senior Consultant, AI/ML Engineer
Senior Consultant, AI/ML Engineer

Hollstadt Consulting • Minnesota

On-site
USD 150,000 - 210,000
Senior AI/ML Engineer, Applications & Automation
Senior AI/ML Engineer, Applications & Automation

Imo Online • United States

On-site
USD 180,000 - 240,000
Senior Data & AI Engineer
Senior Data & AI Engineer

RADcube • Carmel (IN)

On-site
USD 130,000 - 180,000
Software Engineer - AI/RAG
Software Engineer - AI/RAG

Photon • Dallas (TX)

On-site
USD 120,000 - 180,000