We are looking for a skilled Scala Developer with a strong Data Engineering background to join our core team The Data Engineering team is the foundational pillar of OpenLMs data strategy, responsible for the entire lifecycle from collection to actionable insight, You will not just be writing code; you will be instrumental in our 2026 strategic mission to decompose monolithic legacy components into independent Spark microservices You will work on a modern stack involving Spark 4, Iceberg, and Kubernetes, ensuring our platform remains agile, resilient, and scalable
Key Responsibilities
- Core ETL & Microservices Development
- Scala/Spark Development: Design, write, test, and deploy robust Scala code for our core Spark ETL processes
- Microservices Transformation: Lead the transition of legacy monolithic engines into independently managed Spark 3 5+ microservices
- Streaming Pipelines: Manage resilient Spark Streaming jobs consuming data from Kafka, handling schema evolution and offset management
- Orchestration: Use Apache Airflow for consistent scheduling and monitoring of batch Spark ETLs
- Spark on Kubernetes: Operationalize and maintain Spark-on-K8s execution to optimize scaling and resource efficiency
- Lakehouse Implementation: Modernize S3 reporting ETLs by integrating the Apache Iceberg table format and Athena query integration
- Infrastructure as Code: Manage the automated provisioning of cloud resources (IAM, S3, Glue) for tenant onboarding via the Seeding Service
- Performance Optimization & Quality Cost & Speed Tuning: Optimize multi-tenant cloud pipelines using multi-threaded execution to reduce processing time (e g,, reducing runtimes from 20 minutes to 5 minutes)
- Storage Optimization: Implement vacuum strategies and snapshot expiration for Iceberg tables to maintain query performance
- Reliability: Maintain Dead-Letter Queues (DLQ) for capturing and analyzing failed records without stopping the pipeline
Why Join This Team
- Modern Tech: We are aggressive about modernizationdeploying Spark 4, moving to Iceberg, and running Spark on Kubernetes
- High Impact: Your work directly enables business intelligence and product innovation Without our pipelines, the data remains inert
- Innovation Culture: We maintain internal sandboxes (e g,, Raspberry Pi clusters) for experimentation and pilot evaluations
- Clear Career Path: We have a defined structure and clear SLAs for operational excellence
Requirements Technical Stack
- Primary Languages: Proficient in Scala and SQL
- Big Data Frameworks: Deep expertise in Apache Spark (Core, SQL, Streaming) Experience with Spark 3 5 is essential; Spark 4 exposure is a plus
- Cloud Platforms: Strong experience with AWS (S3, Glue, Athena, IAM, Secrets Manager)
- Containerization: Experience with Docker (building secure images) and Kubernetes (Spark-on-K8s)
- Streaming & Messaging: Experience with Kafka (integrating with ServiceNow/Data streams)
- Operational Skills Experience with Apache Airflow for workflow orchestration
- Familiarity with table formats like Apache Iceberg
- Ability to troubleshoot complex data issues and performance bottlenecks (L2/L3 support)
- Understanding of CI/CD concepts and unit testing for data pipelines
Benefits
- Competitive salary commensurate with experience
- Strong results are well rewarded
- Ongoing professional development training
- Visible, exciting work with cutting-edge solutions
Perks That You Will Enjoy
- Mandatory vacations
- Flexible work timings
- Working with some of the global business leaders and tech pioneers
- Challenging projects to lead and boost your career growth
- Work with global teams and clients
- No overtime policy