Data Engineer - Scala/PySpark

Forward Eye Technologies

Pune District

On-site

INR 700,000 - 1,100,000

Full time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Forward Eye Technologies is seeking a data engineer for a 3-month contract to develop and maintain scalable data pipelines using Spark, Hadoop, and SQL. The role involves streaming data, ETL tooling, and cloud services such as AWS.

You will design data models, ensure data quality, and collaborate with DevOps and analytics teams to deliver reliable data infrastructure while adhering to security and compliance standards.

Qualifications

  • Experience building scalable data pipelines with Spark, Hadoop and SQL.
  • Experience with streaming technologies such as Kafka.
  • Proficiency in advanced SQL including window functions.
  • Familiarity with AWS IAM, AWS EMR and Snowflake.
  • Hands-on experience with Airflow, S3 and ETL tooling.

Responsibilities

  • Design, implement, and maintain scalable data pipelines to collect, process, and store large data volumes.
  • Develop and maintain data models, schemas, and metadata to support data initiatives.
  • Integrate data from databases, data warehouses, APIs, and streaming platforms.
  • Optimize data pipelines for performance, scalability, and reliability; resolve bottlenecks.
  • Implement data quality checks, validation rules, and monitoring.
  • Collaborate with DevOps/IT to provision and manage infrastructure (databases, clusters, cloud services).
  • Ensure data security and regulatory compliance (GDPR, HIPAA).
  • Document pipelines and architecture; communicate with cross-functional teams.
  • Stay updated on new technologies and approaches to improve data infrastructure.

Skills

PySpark
Scala with Spark
Spark Architecture
Hadoop
SQL

Tools

Airflow
S3
StreamSets
ETL tools
AWS IAM
AWS EMR
Snowflake

Job description

Contract Duration : 3 months

Skills :
  • - PySpark or Scala with Spark, Spark Architecture, Hadoop, SQL
  • - Streaming Technologies like Kafka etc.
  • - Proficiency in Advanced SQL (Window functions)
  • - Airflow, S3, and Stream Sets or similar ETL tools.
  • - Basic Knowledge on AWS IAM, AWS EMR and Snowflake.
Responsibilities :
  • - Data Pipeline Development: Design, implement, and maintain scalable and efficient data pipelines to collect, process, and store large volumes of structured and unstructured data.
  • - Data Modeling: Develop and maintain data models, schemas, and metadata to support the organization's data initiatives. Ensure data integrity and optimize data storage and retrieval processes.
  • - Data Integration: Integrate data from various sources, including databases, data warehouses, APIs, and streaming platforms, ensuring compatibility, consistency, and quality.
  • - Performance Optimization: Optimize data pipelines and processing systems for performance, scalability, and reliability. Identify and resolve bottlenecks and inefficiencies in data workflows.
  • - Data Quality Assurance: Implement data quality checks, validation rules, and monitoring mechanisms to ensure the accuracy, completeness, and consistency of data across different systems.
  • - Infrastructure Management: Collaborate with DevOps and IT teams to provision, configure, and manage infrastructure components such as databases, clusters, and cloud services required for data processing and storage.
  • - Security and Compliance: Implement data security best practices and compliance standards to protect sensitive information and ensure regulatory compliance (e.g., GDPR, HIPAA). Monitor and audit data access and usage to prevent unauthorized activities.
  • - Documentation and Communication: Document data pipelines, systems architecture, and technical processes. Communicate effectively with cross-functional teams to gather requirements, provide updates, and address issues.
  • - Continuous Learning: Stay updated on emerging technologies, tools, and techniques in data engineering and related fields. Evaluate and recommend new technologies and approaches to improve data infrastructure and workflows.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Data Engineer - Scala/PySpark
Data Engineer - Scala/PySpark

Forward Eye Technologies • Dadri

On-site
INR 700,000 - 1,200,000
Data Engineer - ETL/PySpark
Data Engineer - ETL/PySpark

Forward Eye Technologies • Dadri

On-site
INR 1,800,000 - 2,400,000
Data Engineer-Spark,Scala
Data Engineer-Spark,Scala

Zorba AI • Hyderabad

On-site
INR 1,200,000 - 2,400,000
Data Engineer-Spark,Scala
Data Engineer-Spark,Scala

Zorba AI • Bengaluru

On-site
INR 1,200,000 - 2,400,000
Data Engineer-Spark,Scala
Data Engineer-Spark,Scala

Zorba AI • Chennai District

On-site
INR 900,000 - 1,500,000
Data Engineer - SQL/PySpark
Data Engineer - SQL/PySpark

Forward Eye Technologies • Mumbai, Bengaluru, New Delhi

On-site
INR 1,800,000 - 3,200,000
Data Engineer_Spark/Scala
Data Engineer_Spark/Scala

Zorba AI • Kolkata District

On-site
INR 1,000,000 - 1,500,000
Data Engineer_Spark/Scala
Data Engineer_Spark/Scala

Zorba AI • Mumbai

On-site
INR 1,000,000 - 1,500,000
Data Engineer_Spark/Scala
Data Engineer_Spark/Scala

Zorba AI • Maharashtra

On-site
INR 800,000 - 1,200,000
Data Engineer - ETL/PySpark
Data Engineer - ETL/PySpark

Forward Eye Technologies • Pune District

On-site
INR 1,200,000 - 2,400,000