A technology firm in Bengaluru is seeking a skilled individual proficient in big data technologies, especially Scala and PySpark. This role focuses on designing and building scalable data pipelines while utilizing Apache Spark and distributed data systems. The ideal candidate holds strong problem-solving skills and has experience in Agile and DevOps environments. Opportunities include working with cloud technologies like Azure and various analytics platforms.
Qualifications
Strong experience working with big data technologies and distributed data processing.
Hands-on experience writing production-quality code in Scala (preferred), PySpark, or Java.
Good understanding of Apache Spark architecture and distributed computing.
Comfort with Agile/DevOps and CI/CD practices.
Ability to collaborate with data scientists, analysts, and engineers.
Responsibilities
Design and build scalable data pipelines and data processing frameworks.
Develop high-performance big data applications using Scala / PySpark.
Collaborate with cross-functional teams to deliver data-driven solutions.
Participate in code reviews, CI/CD pipelines, and DevOps practices.
Skills
big data technologies
Scala
PySpark
Java
Apache Spark
data processing frameworks
SQL
NoSQL databases
Agile
DevOps
data pipelines
problem-solving
communication skills
Azure Cloud
Databricks
Snowflake
Tableau
Tools
Apache Spark
Hadoop
Kafka
Impala
MapR filesystem
Maven
SBT
Azure Cloud technologies
Snowflake Data Cloud
Tableau
Job description
Must Have | Good to Have
Must Have
Good to Have
We are looking for someone who:
Has strong experience working with big data technologies and distributed data processing
Is proficient in Scala (preferred), PySpark, or Java for production-level development
Possesses solid understanding of Apache Spark architecture and components
Has hands‑on experience working with large‑scale data pipelines
Has experience working with SQL and NoSQL databases
Is comfortable working in Agile / DevOps environments
Can collaborate effectively with data scientists, analysts, and engineering teams
Demonstrates strong problem‑solving and communication skills
Designs and builds scalable data pipelines and data processing frameworks
Develops high‑performance big data applications using Scala / PySpark
Works with Apache Spark, Hadoop ecosystem, and distributed data systems
Integrates data from multiple sources including relational and NoSQL databases
Supports analytics and reporting platforms for business insights
Collaborates with cross‑functional teams to deliver data‑driven solutions
Participates in code reviews, CI/CD pipelines, and DevOps practices
Required Technical Skills
Hands‑on experience writing production‑quality code in Scala (preferred), PySpark, or Java
Strong experience with Big Data technologies such as Apache Spark, Hadoop, Kafka, and Impala
Good understanding of Apache Spark architecture and distributed computing
Experience with MapR filesystem or similar distributed file systems
Hands‑on experience with scripting languages such as Python and Bash
Strong knowledge of SQL and database technologies
Experience with NoSQL databases such as MongoDB (preferred)
Familiarity with build tools such as Maven and SBT
Experience working with DevOps pipelines and CI/CD environments
Additional / Preferred Skills
Experience with Azure Cloud technologies
Exposure to Azure Databricks and Databricks SQL
Familiarity with Snowflake Data Cloud
Experience with analytics platforms such as Tableau