APPIT Software Solutions is hiring a Senior Data Engineer (Apache Spark) in San Francisco, USA. Lead the design of large-scale distributed data processing systems using Apache Spark and cloud platforms at APPIT Software in San Francisco.
Responsibilities
- Architect and optimize large-scale Spark jobs processing terabytes of data daily
- Design data lake and lakehouse architectures on AWS S3 or Azure Data Lake Storage
- Mentor junior data engineers on distributed computing best practices and Spark internals
- Build and maintain real-time and batch data pipelines with robust fault-tolerance
- Partner with ML engineers to deliver feature stores and training data sets at scale
- Drive performance tuning including partitioning, caching, and shuffle optimization strategies
Requirements
- 6+ years of data engineering experience with at least 3 years focused on Apache Spark
- Deep understanding of distributed computing, MapReduce paradigms, and cluster resource management
- Expert-level Python and/or Scala programming for Spark application development
- Experience with cloud data services (AWS EMR, Glue, Redshift, or Azure Synapse)
- Strong knowledge of data lake architectures, Delta Lake, or Apache Iceberg table formats
- Proven ability to optimize Spark jobs for cost efficiency and processing speed
Nice to Have
- Experience with Databricks Unified Analytics Platform
- Knowledge of streaming with Spark Structured Streaming or Flink
- Contributions to open-source data projects
About APPIT Software Solutions
APPIT Software Solutions is a leading technology company building AI-powered enterprise products. Join 150+ professionals working on cutting-edge solutions in AI/ML, cloud computing, and digital transformation. We offer competitive salaries, flexible work arrangements, and opportunities for career growth across our global offices.