AWS PySpark Tech Lead - Data Pipelines & GenAI

Argyllinfotech

South San Francisco (CA)

Hybrid

USD 150,000 - 190,000

Full time

14 days+
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Benefits offered by this job

Health insurance
401(k) plan

Job summary

Saama is seeking a Technical Lead for AWS PySpark to head data engineering initiatives for a Genentech project. You will design and optimize scalable data pipelines on AWS, lead a team of engineers, and drive data governance, quality, and lineage across enterprise workflows.

The role requires deep PySpark, SQL, and AWS ETL experience, with a strong architectural mindset and hands-on leadership in Agile environments. Hybrid onsite in South San Francisco / Remote work options are available.

Qualifications

  • 6+ years in Data Engineering, Big Data, or Cloud Architecture
  • Experience processing unstructured data and metadata extraction
  • Proven experience with LLMs to parse unstructured data
  • Vector databases indexing and storage experience
  • Active AWS data engineer/solutions architect background preferred
  • Strong Python and Spark optimization skills
  • Proficient in SQL tuning for distributed DBs
  • Experience with Agile methodologies and Jira reporting

Responsibilities

  • Design and implement large-scale data pipelines on AWS (EMR, PySpark)
  • Architect unstructured data and GenAI pipelines including metadata extraction
  • Manage embeddings across vector databases for AI workloads
  • Lead data engineers with Agile delivery and Jira dashboards
  • Provide daily/weekly/monthly reporting to stakeholders
  • Ensure data governance, quality, and lineage across workflows
  • Optimize PySpark performance and SQL workloads
  • Deploy serverless data workflows with Lambda and Redshift for analytics

Skills

Data Engineering
AWS
PySpark
SQL
GenAI/LLM
Agile/JIRA
Data Governance
Cloud Architecture

Education

Bachelor's or Master's in IT/CS

Tools

AWS EMR
AWS Glue
AWS DataBrew
AWS Lambda
AWS Redshift
S3/IAM/Step Functions
Vector DB (Pinecone/Milvus/OpenSearch)

Job description

Saama is seeking a Technical Lead for AWS PySpark to head data engineering initiatives for a Genentech project. You will design and optimize scalable data pipelines on AWS, lead a team of engineers, and drive data governance, quality, and lineage across enterprise workflows.

The role requires deep PySpark, SQL, and AWS ETL experience, with a strong architectural mindset and hands-on leadership in Agile environments. Hybrid onsite in South San Francisco / Remote work options are available.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Technical Lead: AWS PySpark & GenAI Data Pipelines
Technical Lead: AWS PySpark & GenAI Data Pipelines

Apptad Inc • California (MO)

Hybrid
USD 150,000 - 210,000
Technical Lead AWS PySpark
Technical Lead AWS PySpark

Argyllinfotech • South San Francisco (CA)

Hybrid
USD 150,000 - 190,000
Health insurance
401(k) plan
Data Engineering Lead: GenAI‑Driven Pipelines
Data Engineering Lead: GenAI‑Driven Pipelines

American International Group • Northern (KY)

Hybrid
USD 125,000 - 135,000
Technical Lead – AWS PySpark
Technical Lead – AWS PySpark

Apptad Inc • California (MO)

Hybrid
USD 150,000 - 210,000
Lead Data Engineer — GenAI-Driven Pipelines on AWS
Lead Data Engineer — GenAI-Driven Pipelines on AWS

AIG • Parsippany-Troy Hills (NJ)

Hybrid
USD 125,000 - 135,000
Hybrid work arrangement
Total Rewards Program
Remote AI Data Engineer for Scalable Pipelines (AWS, Spark)
Remote AI Data Engineer for Scalable Pipelines (AWS, Spark)

S27a • Northern (KY)

Hybrid
USD 120,000 - 180,000
Equity
Bonuses
Health insurance
+2
Remote Senior Data Engineer: Gen AI Powered Pipelines
Remote Senior Data Engineer: Gen AI Powered Pipelines

Resonate • Northern (KY)

Hybrid
USD 120,000 - 180,000
401(k) match
Open PTO
Remote-first environment
Remote Senior Data Platform Engineer — Spark & AI
Remote Senior Data Platform Engineer — Spark & AI

Samsara • United States

On-site
USD 120,000 - 150,000
Senior Data Engineer — AWS Cloud Pipelines (Remote)
Senior Data Engineer — AWS Cloud Pipelines (Remote)

Attain Talent • United States

Hybrid
USD 110,000 - 140,000
Remote Work (Hybrid)
Medical, Dental, Vision
401(k) with matching
+2
AI-Driven Cloud Engineer: AWS Data Pipelines & GenAI
AI-Driven Cloud Engineer: AWS Data Pipelines & GenAI

Innovative Solutions • United States

Hybrid
USD 100,000 - 160,000
Hybrid work model
Competitive salary