Sr. Data Engineer - Services Special Project

Socket.dev

Cupertino (CA)

On-site

USD 180,000 - 240,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Apple is seeking a senior data engineer to design and maintain a massive real-time data platform that converts multimodal data into a searchable foundation. You will build end-to-end pipelines, craft analytics schemas, and enable ML/LLM inference in production, enabling services for billions of customers.

The role requires strong Spark, Kafka, Airflow, AWS, and Python/PySpark skills, plus experience with large-scale data lakes and modern data tooling.

Qualifications

  • Masters degree required and 10+ years of data engineering experience.
  • Experience building and maintaining large-scale ETL/ELT data pipelines.
  • Data modeling and dimensional modeling for analytics and reporting.
  • Experience with SQL/NoSQL databases (Postgres / Cassandra / Redis).
  • Strong experience with Apache Spark.
  • Experience with BigTable/Hadoop for distributed processing.
  • Strong software engineering fundamentals in Scala and Java.
  • Hands-on with Apache Kafka, Iceberg, and Flink.
  • Experience with Airflow and Beam for workflows.
  • Experience with AWS services: S3, EMR, Lambda, Glue, Redshift, Kinesis.
  • Analytics with Trino (Presto), BigQuery, Snowflake.
  • Big data lake architectures experience.
  • Containerization with Docker and Kubernetes/EKS; Jenkins for CI/CD.
  • Python and PySpark proficiency.
  • Familiarity with graph databases like TigerGraph.
  • Pipelines processing multimodal data and ML/LLMs in production.
  • LLMs/ML inference in production using ONNX/TensorRT/TorchServe.

Responsibilities

  • Design and maintain a massive real-time data platform.
  • Develop and optimize end-to-end data pipelines for analytics.
  • Collaborate on data models and analytics schemas for reporting.
  • Build scalable data pipelines processing multimodal data and embeddings.
  • Deploy and optimize ML/LLM inference in production environments.

Skills

ETL/ELT pipelines
Dimensional modeling
SQL/NoSQL databases
Apache Spark
BigTable/Hadoop
Scala
Java
Apache Kafka
Airflow
Beam
AWS
Trino/Presto/BigQuery/Snowflake
Big data lake architectures
Docker
Kubernetes/EKS
Jenkins
Python/PySpark
TigerGraph

Education

Masters Degree

Tools

Postgres
Cassandra
Redis
Snowflake
Redshift
Spark UI
Airflow
Beam

Job description

DESCRIPTION

At Apple, great ideas have a way of becoming phenomenal products, services, and customer experiences very quickly. Our team is building a massive, real-time platform that transforms continuous streams of multimodal data (including structured, image, and log data) into an intelligent, searchable foundation. By enriching this data with language and embedding models, we power critical experiences for billions of Apple customers across multiple downstream applications.


MINIMUM QUALIFICATIONS


  • Masters Degree

  • 10+ years of experience in data engineering, including building and maintaining large-scale ETL/ELT data pipelines

  • Proficiency in data modeling, especially dimensional modeling, and designing schemas optimized for analytics and reporting

  • Experience with leveraging databases including SQL/NoSQL Databases (including Postgres / Cassandra / Redis)

  • Strong experience with distributed data processing frameworks including Apache Spark

  • Strong experience with Parallel processing frameworks: BigTable/Hadoop

  • Strong software engineering fundamentals and proven experience with Scala, Java

  • Hands-on experience with Apache Kafka, Iceberg, and Flink.

  • Experience with workflow orchestration tools including Apache Airflow and Beam

  • Experience with AWS: e.g., S3, EMR, Lambda, Glue, Redshift, BigQuery, Kinesis, or similar services

  • Experience with Analytics frameworks including Trino (Presto, BigQuery, Snowflake)

  • Hands-on experience with big data lake architectures

  • Experience with containerization and orchestration (Docker, Kubernetes/EKS) and CI/CD tooling including Jenkins

  • Experience in Python and PySpark

  • Familiarity with graph databases such as TigerGraph

  • Experience building pipelines that process multimodal data (structured and image) and integrate ML model inference - including LLMs and embedding models - for data enrichment and transformation

  • Hands-on experience deploying, serving, and optimizing LLMs or ML models directly in the production, inference runtimes/compilers (ONNX Runtime, TensorRT/TensorRT-LLM), and serving frameworks (Triton, vLLM, TorchServe or similar).

  • Experience tuning batching, KV-cache, and GPU utilization for low-latency, high-throughput real-time inference in a data pipeline

  • Knowledge of data governance principles, data security best practices, and data privacy regulations


PREFERRED QUALIFICATIONS


  • Experience with data versioning tools and frameworks (e.g., DVC, Delta Lake)

  • Excellent communication skills and a collaborative mindset

  • Experience storing/serving embeddings (e.g., pgvector, Milvus, FAISS)

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Principal Data Architect and Manager - Service Special Projects
Principal Data Architect and Manager - Service Special Projects

Socket.dev • Cupertino (CA)

On-site
USD 180,000 - 260,000
Senior Distributed Systems Engineer - Services Special Projects
Senior Distributed Systems Engineer - Services Special Projects

Socket.dev • Cupertino (CA)

On-site
USD 170,000 - 210,000
Senior Software Engineer, Apple Data Platform
Senior Software Engineer, Apple Data Platform

Socket.dev • Cupertino (CA)

On-site
USD 150,000 - 190,000
Senior AI Engineer - Services Special Projects
Senior AI Engineer - Services Special Projects

Socket.dev • Cupertino (CA)

On-site
USD 190,000 - 270,000
Sr. / Staff Software Engineer, Data Lakehouse, Apple Data Platform
Sr. / Staff Software Engineer, Data Lakehouse, Apple Data Platform

Socket.dev • Seattle (WA)

On-site
USD 160,000 - 260,000
Software Engineer - Data Solutions, AI & Data Platform (AiDP)
Software Engineer - Data Solutions, AI & Data Platform (AiDP)

Socket.dev • Sunnyvale (CA)

On-site
USD 140,000 - 210,000
Staff Data Engineer
Staff Data Engineer

Socket.dev • Austin (TX)

On-site
USD 180,000 - 240,000
Principal Data Engineering Lead - Services Special Project
Principal Data Engineering Lead - Services Special Project

Apple Inc. • Cupertino (CA)

On-site
USD 263,000 - 394,000
Employee stock plan
Medical and dental coverage
Retirement benefits
+3
Machine Learning/ Search Engineer - Services Special Projects
Machine Learning/ Search Engineer - Services Special Projects

Socket.dev • Cupertino (CA)

On-site
USD 180,000 - 240,000
Data Engineer
Data Engineer

Apple Inc. • Austin (TX)

On-site
USD 90,000 - 130,000