Technical Lead AWS PySpark

Argyllinfotech

South San Francisco (CA)

Hybrid

USD 150,000 - 190,000

Full time

13 days ago
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Benefits offered by this job

Health insurance
401(k) plan

Job summary

Saama is seeking a Technical Lead for AWS PySpark to head data engineering initiatives for a Genentech project. You will design and optimize scalable data pipelines on AWS, lead a team of engineers, and drive data governance, quality, and lineage across enterprise workflows.

The role requires deep PySpark, SQL, and AWS ETL experience, with a strong architectural mindset and hands-on leadership in Agile environments. Hybrid onsite in South San Francisco / Remote work options are available.

Qualifications

  • 6+ years in Data Engineering, Big Data, or Cloud Architecture
  • Experience processing unstructured data and metadata extraction
  • Proven experience with LLMs to parse unstructured data
  • Vector databases indexing and storage experience
  • Active AWS data engineer/solutions architect background preferred
  • Strong Python and Spark optimization skills
  • Proficient in SQL tuning for distributed DBs
  • Experience with Agile methodologies and Jira reporting

Responsibilities

  • Design and implement large-scale data pipelines on AWS (EMR, PySpark)
  • Architect unstructured data and GenAI pipelines including metadata extraction
  • Manage embeddings across vector databases for AI workloads
  • Lead data engineers with Agile delivery and Jira dashboards
  • Provide daily/weekly/monthly reporting to stakeholders
  • Ensure data governance, quality, and lineage across workflows
  • Optimize PySpark performance and SQL workloads
  • Deploy serverless data workflows with Lambda and Redshift for analytics

Skills

Data Engineering
AWS
PySpark
SQL
GenAI/LLM
Agile/JIRA
Data Governance
Cloud Architecture

Education

Bachelor's or Master's in IT/CS

Tools

AWS EMR
AWS Glue
AWS DataBrew
AWS Lambda
AWS Redshift
S3/IAM/Step Functions
Vector DB (Pinecone/Milvus/OpenSearch)

Job description

Role Technical Lead AWS PySpark

Location Onsite South SFO / Remote

Client Genentech

Mandatory Skills: Managing Data Ingestion for Unstructured data

Position Overview

We are seeking a highly skilled and technical AWS PySpark Lead to spearhead our data engineering initiatives. In this role, you will lead a team of data engineers to design, build, and optimize robust, scalable data pipelines on the AWS cloud. The ideal candidate brings a strong architectural mindset, deep expertise in distributed data processing, and a proven track record of optimizing PySpark, SQL workloads, and native AWS ETL tools. If you have a background as an AWS Data Engineer or AWS Solution Architect, excel in Agile environments, and are passionate about maintaining high data quality standards, we want you on our team.

Key Responsibilities
  • Pipeline Architecture & ETL: Design and implement large-scale, high-performance data pipelines using AWS EMR, PySpark, AWS Glue, and AWS DataBrew.
  • Unstructured Data & GenAI Pipelines: Architect end-to-end processing pipelines for unstructured data, including optimal ingestion processes, metadata extraction, and integration with Large Language Models (LLMs) to extract key metadata and content.
  • Vector DB Integration: Manage, optimize, and effectively store embeddings and data across enterprise Vector Databases to support advanced search and AI workloads.
  • Technical Leadership & Agile Delivery: Mentor and manage a team of data engineers. Drive Agile Sprints using JIRA, planning and actively managing the backlog and leveraging JIRA reporting/dashboards to track team velocity, bottlenecks, and project health.
  • Reporting: Daily, Weekly, and Monthly reporting to ensure all stakeholders are fully informed and engaged.
  • Data Governance & Quality: Architect solutions that guarantee high Data Quality and implement comprehensive Data Lineage tracking across all enterprise data workflows.
  • Performance Optimization: Take ownership of system performance by applying advanced PySpark optimization techniques (handling data skewness, memory management, partitioning, and broadcasting) and rigorous SQL optimization.
  • Serverless Data Processing: Architect and deploy event-driven data workflows and microservices utilizing AWS Lambda.
  • Data Warehousing: Model, manage, and optimize data storage and querying within AWS Redshift for analytical reporting and BI consumption.
  • Accelerated Delivery: Utilize AI coding assistants (such as GitHub Copilot and Claude Code) to streamline development, improve code quality, and accelerate project delivery timelines.
  • Cloud Architecture: Apply AWS Solution Architect principles to ensure data infrastructure is secure, highly available, cost-efficient, and scalable.
Required Qualifications & Skills (Must-Have)
  • Experience: Minimum of 6 years of hands-on experience in Data Engineering, Big Data, or Cloud Architecture.
  • Unstructured Data Processing: Hands-on experience in processing unstructured data, including designing optimal ingestion processes and metadata extraction pipelines.
  • LLM & GenAI Data Engineering: Proven experience leveraging LLMs to parse unstructured data and perform automated extraction of both metadata and core content.
  • Vector Database Expertise: Deep expertise in Vector Databases (e.g., Pinecone, Milvus, Qdrant, OpenSearch Vector Engine, Pgvector) with demonstrated experience in effectively indexing, managing, and storing vector data.
  • AWS Expertise: Proven experience operating as an AWS Data Engineer or AWS Solutions Architect. (Active AWS certifications are highly preferred).
  • PySpark Mastery: Exceptional proficiency in Python and Apache Spark. Must have a deep understanding of Spark's internal workings and hands-on experience optimizing heavy PySpark workloads.
  • AWS ETL Ecosystem: Strong, demonstrated experience with core AWS data services including AWS EMR, AWS Glue, AWS DataBrew, AWS Lambda, AWS Redshift, AWS S3, IAM, and AWS Step Functions.
  • Data Governance: Deep understanding of and practical experience with implementing automated Data Quality checks and establishing Data Lineage from source to destination.
  • SQL & Database Skills: Advanced SQL proficiency with a track record of tuning complex queries for performance across distributed databases.
  • Agile & JIRA: Demonstrated experience driving Agile methodologies. Must be well-versed in JIRA, specifically in configuring workflows and generating detailed JIRA reports for stakeholders.
  • DevOps & CI/CD: Strong proficiency in version control and DevOps practices using Git and GitHub Actions for automated CI/CD pipelines.
  • AI-Assisted Development: Demonstrated experience successfully integrating AI tools (Copilot, Claude Code) into daily engineering workflows to boost productivity.
Nice to Have
  • Experience with workflow orchestration tools like Apache Airflow.
  • Knowledge of Infrastructure as Code (IaC) tools such as Terraform or AWS CloudFormation.
Education

Bachelors or Masters in Information Technology, Computer Science or relevant field.

Work Environment

This job operates in a professional office environment. This role routinely uses standard office equipment, including but not limited to, computers, phones, and photocopiers.

Physical Demands

This position requires the frequent and repetitive use of a computer, keyboard, and mouse. Hand and finger dexterity is required.

Other Duties

Please note this job description is not designed to cover or contain a comprehensive listing of activities, duties or responsibilities that are required of the employee for this job. Duties, responsibilities, and activities may change at any time with or without notice.

EEO

Saama provides equal employment opportunities to all employees and applicants for employment and prohibits discrimination and harassment of any type without regard to race, color, religion, age, sex, national origin, disability status, genetics, protected veteran status, sexual orientation, gender identity or expression, or any other characteristic protected by federal, state or local laws.

This policy applies to all terms and conditions of employment, including recruiting, hiring, placement, promotion, termination, layoff, recall, transfer, leaves of absence, compensation, and training.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Technical Lead – AWS PySpark
Technical Lead – AWS PySpark

Apptad Inc • California (MO)

Hybrid
USD 150,000 - 210,000
Sr AWS Python Developer
Sr AWS Python Developer

MACHINE LEARNING TECHNOLOGIES LLC • Reston (VA)

On-site
USD 90,000 - 120,000
Senior Data Engineer
Senior Data Engineer

Samsara • United States

On-site
USD 120,000 - 150,000
PySpark Data Engineer
PySpark Data Engineer

Tata Consultancy Services • Irving (TX)

On-site
USD 90,000 - 120,000
Discretionary Annual Incentive
Medical Coverage (Health, Dental & Vis
401K Plan
+1
Apache Spark Developer
Apache Spark Developer

Bright Vision Technologies • Austin (TX)

Remote
USD 125,000 - 185,000
AWS Python Developer with Pyspark
AWS Python Developer with Pyspark

Polarits • Newark (NJ)

Hybrid
USD 140,000 - 180,000
Senior Data Engineer
Senior Data Engineer

EXL • New York (NY)

On-site
USD 140,000 - 190,000
Data Engineer IV- #26-23312
Data Engineer IV- #26-23312

US Tech Solutions • Charlotte (NC)

On-site
USD 138,000 - 152,000
Python/PySpark Developer
Python/PySpark Developer

Software Guidance & Assistance, Inc. (SGA, Inc.) • Rutherford (NJ)

On-site
USD 120,000 - 180,000
Senior Data Analytics Engineer
Senior Data Analytics Engineer

Revel IT • Columbus (OH)

On-site
USD 100,000 - 130,000