Software Engineer – Data Platform

Alegeus

Bengaluru

On-site

INR 4,200,000 - 6,600,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Alegeus is seeking an Expert Software Engineer to design, build, and scale a next-generation Data Platform and Data-Driven APIs. This role combines distributed data processing with backend platform engineering to enable reliable, scalable, and real-time data access.

The role requires ownership of data pipelines, API design using Java, and implementing data quality, observability, and lineage frameworks across cloud-native architectures (Azure/AWS/GCP).

Qualifications

  • 4+ years of software engineering experience with strong focus on data platforms and/or distributed systems.
  • Hands-on expertise in Apache Spark or Scala or PySpark
  • Strong programming skills in Java (preferred) / Scala / Python
  • Experience developing backend services or APIs (REST/microservices)
  • Deep understanding of: Distributed systems, Data modeling and schema evolution
  • Experience with cloud platforms (Azure/AWS/GCP)
  • Familiarity with workflow orchestration tools (Airflow, Dagster, etc.)
  • Strong system design and performance optimization skills

Responsibilities

  • Design and develop scalable data pipelines using Apache Spark (batch and streaming)
  • Build and maintain data platform layers: ingestion, transformation, and serving
  • Optimize Spark jobs for performance, cost, and reliability (partitioning, skew handling, memory tuning)
  • Implement data quality, observability, and lineage frameworks
  • Contribute to data architecture decisions (Lakehouse, data mesh, storage formats, partition strategies)
  • Define and enforce data contracts and schema evolution practices
  • Design and build data-driven platform APIs using Java (preferred)
  • Develop microservices that expose curated datasets for product and partner consumption
  • Implement RESTful APIs and event-driven services for real-time and near real-time data access
  • Ensure low-latency, high-availability data serving layers
  • Integrate with upstream/downstream systems, including legacy APIs where required
  • Build and deploy solutions on Azure (preferred) / AWS / GCP
  • Leverage cloud-native services for data storage, compute, and messaging
  • Work with event streaming systems (Kafka/Event Hubs) for real-time pipelines
  • Support containerized deployments and orchestration (Kubernetes) where applicable

Skills

Apache Spark
Scala
Python
Java
REST APIs
Distributed systems
Data modeling
Cloud platforms

Tools

Airflow
Dagster
Kubernetes

Job description

We are looking for an Expert Software Engineer to design, build, and scale our next-generation Data Platform and Data-Driven APIs. This role combines distributed data processing (Apache Spark) with platform and microservices engineering (Java) to enable reliable, scalable, and real-time data access.

You will operate at the intersection of data engineering and backend platform engineering-building systems that not only process large volumes of data but also expose that data through robust, well-designed APIs and services.

This role goes beyond implementing requirements. We expect engineers to understand business context, challenge assumptions, and take end-to-end ownership of delivering meaningful outcomes.

Key responsibilities
  • Design and develop scalable data pipelines using Apache Spark (batch and streaming)
  • Build and maintain data platform layers: ingestion, transformation, and serving
  • Optimize Spark jobs for performance, cost, and reliability (partitioning, skew handling, memory tuning)
  • Implement data quality, observability, and lineage frameworks
  • Contribute to data architecture decisions (Lakehouse, data mesh, storage formats, partition strategies)
  • Define and enforce data contracts and schema evolution practices
  • Design and build data-driven platform APIs using Java (preferred)
  • Develop microservices that expose curated datasets for product and partner consumption
  • Implement RESTful APIs and event-driven services for real-time and near real-time data access
  • Ensure low-latency, high-availability data serving layers
  • Integrate with upstream/downstream systems, including legacy APIs where required
  • Build and deploy solutions on Azure (preferred) / AWS / GCP
  • Leverage cloud-native services for data storage, compute, and messaging
  • Work with event streaming systems (Kafka/Event Hubs) for real-time pipelines
  • Support containerized deployments and orchestration (Kubernetes) where applicable
Quality, Observability & Engineering Excellence
  • Champion unit tests across both data and service layers
  • Build automated validation frameworks for data pipelines
  • Implement end-to-end observability (metrics, logging, tracing) across pipelines and APIs
  • Drive CI/CD practices for both data and application code
  • Conduct code reviews and enforce engineering best practices
Product Mindset & Ownership
  • Engage deeply with product and business stakeholders to understand why, not just what
  • Translate business problems into scalable data and platform solutions
  • Take end-to-end ownership from design through production and support
  • Proactively identify performance bottlenecks, data issues, and system gaps
Required qualifications (Hard requirements)
  • 4+ years of software engineering experience with strong focus on data platforms and/or distributed systems
  • Hands-on expertise in Apache Spark or Scala or PySpark
  • Strong programming skills in Java (preferred) / Scala / Python
  • Experience developing backend services or APIs (REST/microservices)
  • Deep understanding of:
  • Distributed systems (partitioning, shuffle, fault tolerance)
  • Data modeling and schema evolution
  • Experience with cloud platforms (Azure/AWS/GCP)
  • Familiarity with workflow orchestration tools (Airflow, Dagster, etc.)
  • Strong system design and performance optimization skills
Preferred qualifications
  • Experience with Spark Structured Streaming
  • Exposure to Lakehouse architectures (Delta Lake, Iceberg, Hudi)
  • Experience with event-driven architectures (Kafka, Event Hubs)
  • Knowledge of data governance, catalog, and lineage tools
  • Experience with CI/CD for data and microservices
  • Familiarity with Kubernetes and containerized workloads
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Software Engineer
Software Engineer

Alegeus • Bengaluru

On-site
INR 3,500,000 - 5,500,000
Senior Data Engineer
Senior Data Engineer

SourcingXPress • Hyderabad

On-site
INR 1,500,000 - 2,500,000
Big Data Developer
Big Data Developer

Metaplore Solutions Pvt Ltd • Bengaluru

On-site
INR 4,500,000 - 7,500,000
Staff Software Engineer - Data Engineer
Staff Software Engineer - Data Engineer

Tekion Corp • Bengaluru

On-site
INR 1,500,000 - 2,500,000
Data Engineer_Spark/Scala
Data Engineer_Spark/Scala

Zorba AI • Maharashtra

On-site
INR 800,000 - 1,200,000
Data Engineer_Spark/Scala
Data Engineer_Spark/Scala

Zorba AI • Kolkata District

On-site
INR 1,000,000 - 1,500,000
Data Engineer_Spark/Scala
Data Engineer_Spark/Scala

Zorba AI • Mumbai

On-site
INR 1,000,000 - 1,500,000
Senior Software Engineer (Data)
Senior Software Engineer (Data)

KSOLVES • Indore District

On-site
Senior Data Engineer
Senior Data Engineer

KSB • Pune District

On-site
INR 800,000 - 1,200,000
Data Platform Engineer
Data Platform Engineer

Salt • Pune District

On-site
INR 1,200,000 - 1,800,000