Technical Lead – AWS PySpark

Apptad Inc

California (MO)

Hybrid

USD 150,000 - 210,000

Full time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Apptad Inc. is seeking a highly skilled Technical Lead AWS PySpark to drive data engineering initiatives on AWS and lead a team of data engineers in building scalable data pipelines.

You will optimize PySpark workloads, govern data quality, and enable GenAI-enabled ingestion and metadata extraction. The role requires strong AWS expertise, vector database experience, and a proven track record of delivering in Agile environments with CI/CD, while mentoring engineers and driving data governance

Qualifications

  • 6+ years in Data Engineering, Big Data, or Cloud Architecture.
  • Hands-on unstructured data processing and metadata extraction.
  • Experience with LLMs for parsing unstructured data.
  • Expertise in Vector Databases (e.g., Pinecone, Milvus, Qdrant).
  • AWS Data Engineer or AWS Solutions Architect with active certs preferred.
  • Proficient in PySpark and high-performance SQL tuning.
  • Experience with AWS EMR, Glue, DataBrew, Lambda, Redshift, S3, IAM, Step Functions.
  • Strong Agile experience; skilled in JIRA reporting.
  • CI/CD using Git/GitHub Actions; AI-assisted development tools like Copilot.

Responsibilities

  • Design and implement large-scale, high-performance data pipelines on AWS.
  • Architect end-to-end ETL for unstructured data and LLM integration.
  • Manage vector data storage for advanced search and AI workloads.
  • Mentor a team of data engineers; drive Agile sprints and backlog tracking.
  • Provide regular reporting to stakeholders; ensure data governance and quality.
  • Optimize PySpark workloads and SQL performance; implement serverless data processing.

Skills

PySpark
AWS Cloud
Data Engineering
Unstructured Data
LLMs / GenAI
Vector DBs
SQL Tuning
Agile / JIRA
CI/CD
AI in Dev

Education

Bachelor's/Master's in IT/CS

Tools

AWS EMR
AWS Glue
AWS DataBrew
AWS Lambda
Redshift
S3
GitHub Actions
Git

Job description

Role Technical Lead AWS PySpark
Location Onsite South SFO / Remote
Mandatory Skills: Managing Data Ingestion for Unstructured data
Position Overview

We are seeking a highly skilled and technical AWS PySpark Lead to spearhead our data engineering initiatives. In this role, you will lead a team of data engineers to design, build, and optimize robust, scalable data pipelines on the AWS cloud. The ideal candidate brings a strong architectural mindset, deep expertise in distributed data processing, and a proven track record of optimizing PySpark, SQL workloads, and native AWS ETL tools. If you have a background as an AWS Data Engineer or AWS Solution Architect, excel in Agile environments, and are passionate about maintaining high data quality standards, we want you on our team.

Key Responsibilities
  • Pipeline Architecture & ETL: Design and implement large-scale, high-performance data pipelines using AWS EMR, PySpark, AWS Glue, and AWS DataBrew.
  • Unstructured Data & GenAI Pipelines: Architect end-to-end processing pipelines for unstructured data, including optimal ingestion processes, metadata extraction, and integration with Large Language Models (LLMs) to extract key metadata and content.
  • Vector DB Integration: Manage, optimize, and effectively store embeddings and data across enterprise Vector Databases to support advanced search and AI workloads.
  • Technical Leadership & Agile Delivery: Mentor and manage a team of data engineers. Drive Agile Sprints using JIRA, planning and actively managing the backlog and leveraging JIRA reporting/dashboards to track team velocity, bottlenecks, and project health.
  • Reporting: Daily, Weekly, and Monthly reporting to ensure all stakeholders are fully informed and engaged.
  • Data Governance & Quality: Architect solutions that guarantee high Data Quality and implement comprehensive Data Lineage tracking across all enterprise data workflows.
  • Performance Optimization: Take ownership of system performance by applying advanced PySpark optimization techniques (handling data skewness, memory management, partitioning, and broadcasting) and rigorous SQL optimization.
  • Serverless Data Processing: Architect and deploy event-driven data workflows and microservices utilizing AWS Lambda.
  • Data Warehousing: Model, manage, and optimize data storage and querying within AWS Redshift for analytical reporting and BI consumption.
  • Accelerated Delivery: Utilize AI coding assistants (such as GitHub Copilot and Claude Code) to streamline development, improve code quality, and accelerate project delivery timelines.
  • Cloud Architecture: Apply AWS Solution Architect principles to ensure data infrastructure is secure, highly available, cost-efficient, and scalable.
Required Qualifications & Skills (Must-Have)
  • Experience: Minimum of 6 years of hands-on experience in Data Engineering, Big Data, or Cloud Architecture.
  • Unstructured Data Processing: Hands‑on experience in processing unstructured data, including designing optimal ingestion processes and metadata extraction pipelines.
  • LLM & GenAI Data Engineering: Proven experience leveraging LLMs to parse unstructured data and perform automated extraction of both metadata and core content.
  • Vector Database Expertise: Deep expertise in Vector Databases (e.g., Pinecone, Milvus, Qdrant, OpenSearch Vector Engine, Pgvector) with demonstrated experience in effectively indexing, managing, and storing vector data.
  • AWS Expertise: Proven experience operating as an AWS Data Engineer or AWS Solutions Architect. (Active AWS certifications are highly preferred).
  • PySpark Mastery: Exceptional proficiency in Python and Apache Spark. Must have a deep understanding of Spark's internal workings and hands‑on experience optimizing heavy PySpark workloads.
  • AWS ETL Ecosystem: Strong, demonstrated experience with core AWS data services including AWS EMR, AWS Glue, AWS DataBrew, AWS Lambda, AWS Redshift, AWS S3, IAM, and AWS Step Functions.
  • Data Governance: Deep understanding of and practical experience with implementing automated Data Quality checks and establishing Data Lineage from source to destination.
  • SQL & Database Skills: Advanced SQL proficiency with a track record of tuning complex queries for performance across distributed databases.
  • Agile & JIRA: Demonstrated experience driving Agile methodologies. Must be well‑versed in JIRA, specifically in configuring workflows and generating detailed JIRA reports for stakeholders.
  • DevOps & CI/CD: Strong proficiency in version control and DevOps practices using Git and GitHub Actions for automated CI/CD pipelines.
  • AI-Assisted Development: Demonstrated experience successfully integrating AI tools (Copilot, Claude Code) into daily engineering workflows to boost productivity.
Nice to Have
  • Experience with workflow orchestration tools like Apache Airflow.
  • Knowledge of Infrastructure as Code (IaC) tools such as Terraform or AWS CloudFormation.
Education
  • Bachelors or Masters in Information Technology, Computer Science or relevant field.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Technical Lead AWS PySpark
Technical Lead AWS PySpark

Argyllinfotech • South San Francisco (CA)

Hybrid
USD 150,000 - 190,000
Health insurance
401(k) plan
Senior Data Engineer
Senior Data Engineer

EXL • New York (NY)

On-site
USD 140,000 - 190,000
Data Engineer - Python, SQL, AWS
Data Engineer - Python, SQL, AWS

Compunnel, Inc. • Durham (NC)

On-site
USD 95,000 - 120,000
Lead AWS Data Engineer
Lead AWS Data Engineer

Jobtailor • Town of Florida (NY)

On-site
USD 120,000 - 170,000
Data Engineer
Data Engineer

The Value Maximizer • South Carolina

On-site
USD 90,000 - 120,000
PySpark Data Engineer
PySpark Data Engineer

Tata Consultancy Services • Irving (TX)

On-site
USD 90,000 - 120,000
Discretionary Annual Incentive
Medical Coverage (Health, Dental & Vis
401K Plan
+1
Senior Data Analytics Engineer
Senior Data Analytics Engineer

Revel IT • Columbus (OH)

On-site
USD 100,000 - 130,000
Lead Data Engineer
Lead Data Engineer

Strategic Staffing Solutions • Charlotte (NC)

On-site
USD 124,000 - 138,000
Data Engineer
Data Engineer

The Value Maximizer • United States

On-site
USD 90,000 - 120,000
Senior Data Engineer
Senior Data Engineer

Zuven technologies Inc • Malvern

On-site
USD 120,000 - 180,000