Data Engineer I

Clear Demand Inc

Chennai District

On-site

INR 1,000,000 - 1,500,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

A data engineering firm in Chennai, India is seeking a DE-I to drive the design and optimization of large-scale data pipelines. You will collaborate with cross-functional teams, mentor junior engineers, and implement robust data processing solutions. The ideal candidate has 3+ years of experience, strong expertise in Apache Airflow, AWS, and data processing frameworks such as Spark. This role offers a dynamic working environment focused on innovation and efficiency.

Qualifications

  • 3+ years of experience with Apache Airflow and orchestration tools.
  • Experience in AWS technologies, including DynamoDB and Glue.
  • Proficiency in web crawling frameworks like Puppeteer.

Responsibilities

  • Lead design and optimization of large-scale data pipelines.
  • Mentor junior engineers and foster collaboration.
  • Implement data partitioning strategies for databases.

Skills

Apache Airflow
AWS
Apache Spark
Kafka
SQL
NoSQL databases
Node.js
Docker
Kubernetes

Tools

Grafana
Prometheus
Terraform

Job description

Job Summary:
Building on the foundation of the SDE-I role, the DE- I position takes on a greater level of responsibility and leadership.
You’ll play a crucial role in driving the evolution and efficiency of our data collection and analytics platform, capable of handling terabyte-scale data and billions of data points.
Key Responsibilities:
  • Lead the design, development, and optimization of large-scale data pipelines and infrastructures using technologies like Apache Airflow, Spark, Kafka, and more.
  • Architect and implement distributed data processing solutions to handle terabyte-scale datasets and billions of records efficiently across multi-region cloud infrastructure (AWS, GCP, DO).
  • Develop and maintain real-time data processing solutions for high-volume data collection operations using technologies like Spark Streaming and Kafka.
  • Optimize data storage strategies using technologies such as Amazon S3, HDFS, and Parquet/Avro file formats for efficient querying and cost management.
  • Build and maintain high-quality ETL pipelines, ensuring robust data collection and transformation processes with a focus on scalability and fault tolerance.
  • Collaborate with data analysts, researchers, and cross-functional teams to define and maintain data quality metrics, implement robust data validation, and enforce security best practices.
  • Mentor junior engineers (SDE-I) and foster a collaborative, growth-oriented environment.
  • Participate in technical discussions, contributing to architectural decisions, and proactively identifying improvements for scalability, performance, and cost-efficiency.
  • Ensure application performance monitoring (APM) is in place, utilizing tools like Datadog, New Relic, or similar to proactively monitor and optimize system performance, detect bottlenecks, and ensure system health.
  • Implement effective data partitioning strategies and indexing for performance optimization in distributed databases such as DynamoDB, Cassandra, or HBase.
  • Stay current with advancements in data engineering, orchestration tools, and emerging cloud technologies, continually enhancing the platform’s capabilities
Qualifications & Experience:
  • 3+ years of hands-on experience with Apache Airflow and other orchestration tools for managing large-scale workflows and data pipelines.
  • Expertise in AWS technologies, Athena, AWS Glue, DynamoDB, Apache Spark, PySpark, SQL, and NoSQL databases.
  • Experience in designing and managing distributed data processing systems that scale to terabyte and billion-scale datasets using cloud platforms like AWS, GCP, or Digital Ocean.
  • Proficiency in web crawling frameworks, including Node.js, HTTP protocols, Puppeteer, Playwright, and Chromium for large-scale data extraction.
  • Experience with monitoring and observability tools such as Grafana, Prometheus, Elasticsearch, and familiarity with monitoring and optimizing resource utilization in distributed systems.
  • Strong understanding of infrastructure as code using Terraform, automated CI/CD pipelines with Jenkins, and event-driven architecture with Kafka.
  • Experience with data lake architectures and optimizing storage using formats such as Parquet, Avro, or ORC.
  • Strong background in optimizing query performance and data processing frameworks (Spark, Flink, or Hadoop) for efficient data processing at scale.
  • Knowledge of containerization (Docker, Kubernetes) and orchestration for distributed system deployments.
  • Deep experience in designing resilient data systems with a focus on fault tolerance, data replication, and disaster recovery strategies in distributed environments.
  • Strong data engineering skills, including ETL pipeline development, stream processing, and distributed systems.
  • Excellent problem-solving abilities, with a collaborative mindset and strong communication skills.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Data Engineer I
Data Engineer I

CLEAR DEMAND Inc. • Chennai District

On-site
INR 1,200,000 - 1,600,000
Data Engineer II
Data Engineer II

Clear Demand Inc • Chennai District

On-site
INR 1,200,000 - 1,500,000
Data Engineer II
Data Engineer II

CLEAR DEMAND Inc. • Chennai District

On-site
INR 1,200,000 - 1,800,000
Data Engineering Pipeline Engineer – Role Description
Data Engineering Pipeline Engineer – Role Description

Innoventes Technologies • Bengaluru

On-site
INR 1,000,000 - 1,500,000
Data Engineer
Data Engineer

Advance Career Solutions • Pune District, Chennai District, Bengaluru

Hybrid
INR 1,200,000 - 2,800,000
Data Engineer
Data Engineer

ConveGenius • Chennai District

On-site
INR 800,000 - 1,200,000
Data Engineering Manager
Data Engineering Manager

Good co India • India

On-site
INR 2,400,000 - 5,400,000
Senior Data Engineer
Senior Data Engineer

GlobalNodes • Gurgaon

On-site
INR 1,500,000 - 2,100,000
Data Engineer
Data Engineer

Saturam • Bengaluru

On-site
INR 800,000 - 1,500,000
Data Engineer
Data Engineer

Pull Skil • Hyderabad

On-site
INR 800,000 - 1,200,000