Senior Data Software Engineer, Spark, Java, Scala

EPAM Systems

Deutschland

Vor Ort

EUR 90.000 - 130.000

Vollzeit

Vor 13 Tagen

Erhalte mehr Antworten von Arbeitgebern

Versende in nur wenigen Minuten einen passgenauen Lebenslauf.

Zusammenfassung

EPAM Systems is seeking a Senior Software Data Engineer to join our team in a software engineering capacity. This role focuses on building data pipelines, processing jobs, and producing datasets for internal and external partners.

You will work on data quality, scalability, and maintainable code, collaborate with product managers and other data engineers, and help evolve platforms using Spark, Hadoop, and related technologies.

Qualifikationen

  • 3+ years of hands-on experience in the big-data field, including Hadoop (HDFS, YARN or Mesos) and Spark.
  • Excellent knowledge and hands-on experience of SQL in context of Big Data: Spark SQL, HiveQL
  • Excellent knowledge of Spark, including ability to understand and optimize Spark execution plans via Spark UI, with upcoming migration to Spark 3
  • Excellent knowledge of Scala or Java
  • Understanding of batch processing and ETL principles in Data Warehouses
  • Familiarity with data completeness signals and orchestration
  • Knowledge of approaches for historical reprocessing and data correction
  • Skills in handling bad data and late data in inputs and outputs
  • Understanding of schema migrations and datasets evolution
  • Strong speaking English, with ability to rely on information heard verbally in meetings and to explain own ideas clearly to native speakers
  • Capability to learn fast new set of tools and technology used internally at the company: platform services, telemetry providers, Spark-as-a-Service, build system, and more

Aufgaben

  • Write new data pipelines and jobs to produce new outputs (datasets) in scope of new features development
  • Adopt existing data pipelines to integrate with new org-wide platforms, tools, services, and languages
  • Fix bugs in code and correct data caused by incorrect logic or implementation
  • Perform ad-hoc data exploration, validation, and investigation to help select the right tech design and support Product Management team decisions
  • Monitor and troubleshoot production issues with pipelines owned by the team
  • Develop and adopt data quality checks to monitor data issues in the systems
  • Scope and plan new development, including assessing level of effort and providing timelines
  • Maintain tickets hygiene in Radar (ticketing system)
  • Evolve jobs, apps, and systems to a better state across all aspects: code quality, complexity, maintainability, and documentation
  • Communicate with other data engineers in the team, peer teams (QA, UAT, Platform, etc), project managers, and engineering managers on status, blockers, estimates, and timelines

Kenntnisse

English proficiency

Tools

Hadoop
Spark
Spark SQL
HiveQL
Scala
Java
Kafka
Spark Streaming
Snowflake
Trino (Presto)
Iceberg
Druid
Cassandra
Amazon S3 / Blob storage

Jobbeschreibung

We are seeking a Senior Software Data Engineer to join our team in a software engineering capacity. This is not a data-science or analytics position centered on ad-hoc data exploration; instead, the role focuses on building software, data processing jobs, and data pipelines consumed by internal and external partners. Data is our main product and first-class citizen, and we value correct, high-quality data as much as clean and maintainable code.

Responsibilities
  • Write new data pipelines and jobs to produce new outputs (datasets) in scope of new features development
  • Adopt existing data pipelines to integrate with new org-wide platforms, tools, services, and languages
  • Fix bugs in code and correct data caused by incorrect logic or implementation
  • Perform ad-hoc data exploration, validation, and investigation to help select the right tech design and support Product Management team decisions
  • Monitor and troubleshoot production issues with pipelines owned by the team
  • Develop and adopt data quality checks to monitor data issues in the systems
  • Scope and plan new development, including assessing level of effort and providing timelines
  • Maintain tickets hygiene in Radar (ticketing system)
  • Evolve jobs, apps, and systems to a better state across all aspects: code quality, complexity, maintainability, and documentation
  • Communicate with other data engineers in the team, peer teams (QA, UAT, Platform, etc), project managers, and engineering managers on status, blockers, estimates, and timelines
Requirements
  • 3+ years of hands-on experience in the big-data field, including Hadoop (HDFS, YARN or Mesos) and Spark
  • Excellent knowledge and hands-on experience of SQL in context of Big Data: Spark SQL, HiveQL
  • Excellent knowledge of Spark, including ability to understand and optimize Spark execution plans via Spark UI, with upcoming migration to Spark 3
  • Excellent knowledge of Scala or Java
  • Understanding of batch processing and ETL principles in Data Warehouses
  • Familiarity with data completeness signals and orchestration
  • Knowledge of approaches for historical reprocessing and data correction
  • Skills in handling bad data and late data in inputs and outputs
  • Understanding of schema migrations and datasets evolution
  • Strong speaking English, with ability to rely on information heard verbally in meetings and to explain own ideas clearly to native speakers
  • Capability to learn fast new set of tools and technology used internally at the company: platform services, telemetry providers, Spark-as-a-Service, build system, and more
  • Nice to have Understanding of functional programming ideas and principles
  • Nice to have Experience in building and using web services
  • Nice to have Familiarity with any of Teradata, Vertica, Oracle, Tableau
  • Nice to have Skills in Spark Streaming and Kafka
  • Nice to have Knowledge of Apache Iceberg, Trino (Presto), Druid, Cassandra, or Blob storage like AWS
  • Nice to have Experience with Splunk
  • Nice to have Experience with Snowflake
Hol dir deinen kostenlosen, vertraulichen Lebenslauf-Check.
oder ziehe deine Datei hierhin.
Similar jobs

Ähnliche Jobs, die dir auch gefallen könnten

Lead Data Software Engineer with Databricks, Apache Kafka, Apache Spark, Kubernetes
Lead Data Software Engineer with Databricks, Apache Kafka, Apache Spark, Kubernetes

EPAM Systems • Deutschland

Remote
EUR 90.000 - 120.000
Senior Data Software Engineer
Senior Data Software Engineer

TrioTech Recruitment • Berlin

Vor Ort
EUR 75.000 - 95.000
Free Breakfast & Lunch
Senior Data Software Engineer with Databricks and Azure
Senior Data Software Engineer with Databricks and Azure

EPAM Systems • Deutschland

Remote
EUR 70.000 - 90.000
Senior Data Software Engineer with AWS and Terraform
Senior Data Software Engineer with AWS and Terraform

EPAM Systems • Deutschland

Remote
EUR 90.000 - 130.000
Big Data Engineer
Big Data Engineer

Embedded Shishya • Deutschland

Hybrid
EUR 70.000 - 110.000
Lead Data Software Engineer with Databricks with Apache Kafka, Apache Spark, Kubernetes
Lead Data Software Engineer with Databricks with Apache Kafka, Apache Spark, Kubernetes

EPAM Systems • Deutschland

Remote
EUR 90.000 - 130.000
Senior Data Engineer
Senior Data Engineer

Onapsis • Heidelberg

Vor Ort
EUR 70.000 - 90.000
Competitive compensation
Supportive and humble colleagues
Career growth opportunities
Senior Software Engineer, Data Transformation Snowflake DE-Berlin-Trion Building
Senior Software Engineer, Data Transformation Snowflake DE-Berlin-Trion Building

Neura Market • Berlin

Vor Ort
EUR 130.000 - 182.000
Data Engineer (Digital Marketing sphere)
Data Engineer (Digital Marketing sphere)

Coherent Solutions • Deutschland

Hybrid
EUR 70.000 - 100.000
Health insurance
Technical training and conferences
Mentorship from experienced engineers
+3
Senior Data Engineer
Senior Data Engineer

COGNIZANT • Deutschland

Remote
EUR 90.000 - 130.000
sharing the costs of sports activities
private medical care
life insurance
+2