Employment Type: Full-time, On-site
Open Positions: 1
Datazza is a specialized data engineering and advanced analytics company that helps organizations build modern, scalable, and secure data platforms.
Our team consists of experienced data professionals with more than 15 years of hands-on expertise in enterprise data warehouse, big data, analytics, and data platform projects. We design and implement on‑premise and cloud‑based data platforms, lakehouse architectures, real‑time and batch data pipelines, analytical data products, machine learning platforms, and business intelligence solutions.
We primarily work with open‑source and cloud‑native technologies to develop flexible, vendor‑independent, and production‑grade data platforms. As an agile and engineering‑focused company, we work closely with our clients and take an active role in architectural design, implementation, optimization, and operationalization.
Role Overview
We are looking for a Senior Data Engineer to join our team and work on a large‑scale data platform project in the telecommunications industry.
In this role, you will design, develop, and optimize high‑volume data pipelines and analytical data structures operating within modern big data and lakehouse environments. You will work with large‑scale batch and streaming datasets generated by telecommunications systems and contribute to the development of scalable, reliable, and high‑performance data platforms.
The ideal candidate has strong hands‑on data engineering experience, understands distributed data processing concepts, and is comfortable working with open‑source technologies in enterprise production environments.
Key Responsibilities
- Design, develop, and maintain scalable batch and real‑time data pipelines.
- Build and optimize data processing workloads for large‑volume telecommunications datasets.
- Develop ELT and ETL processes across data lake, lakehouse, and enterprise data warehouse environments.
- Design analytical data models, data marts, and reusable data layers.
- Work with distributed processing and streaming technologies such as Apache Spark and Apache Kafka.
- Implement lakehouse data structures using technologies such as Apache Iceberg, Delta Lake, or Apache Hudi.
- Develop and optimize SQL workloads running on distributed query engines such as Trino, Spark SQL, or similar platforms.
- Integrate relational databases, NoSQL systems, event streams, object storage, APIs, and enterprise data sources.
- Perform query optimization, data partitioning, file sizing, compaction, and performance tuning.
- Implement data quality controls, monitoring, logging, lineage, and observability practices.
- Contribute to architectural decisions, technology evaluations, and technical standards.
- Produce technical documentation, data mappings, and operational runbooks.
- Collaborate with architects, analysts, data scientists, DevOps teams, and client stakeholders.
- Review code, promote engineering best practices, and mentor less‑experienced team members.
Required Qualifications
- At least 5 years of professional experience in data engineering, big data, data warehouse, or related roles.
- Strong hands‑on experience in building and operating production‑grade data pipelines.
- Advanced SQL skills, including complex transformations, query optimization, and performance tuning.
- Strong experience with Python, Scala, or Java for data processing and automation.
- Hands‑on experience with Apache Spark or a comparable distributed data processing framework.
- Experience with Apache Kafka or another event‑streaming platform.
- Strong understanding of data warehouse, data lake, and lakehouse architecture principles.
- Experience with dimensional modeling, normalized data models, data marts, and analytical data structures.
- Familiarity with workflow orchestration technologies such as Apache Airflow.
- Experience integrating data from relational databases, APIs, files, object storage, and event‑based systems.
- Good understanding of distributed systems, data partitioning, parallel processing, and fault tolerance.
- Experience with Linux‑based environments, Git, CI/CD processes, and software development practices.
- Strong analytical thinking, troubleshooting, communication, and problem‑solving skills.
- Ability to work collaboratively with technical teams and client stakeholders.
- Bachelor's or master's degree in Computer Science, Computer Engineering, Information Systems, or a related field, or equivalent practical experience.
Preferred Qualifications
- Previous experience in the telecommunications industry or with high‑volume customer, network, usage, billing, CDR, XDR, or event data.
- Hands‑on experience with Apache Iceberg, Apache Hudi, or Delta Lake.
- Experience with Spark, Trino, Presto, StarRocks, ClickHouse, Hive, or similar analytical technologies.
- Experience with open‑source data platform components such as MinIO, OpenMetadata, Apache Superset, dbt, or MLflow.
- Experience deploying or operating data workloads on Kubernetes.
- Familiarity with Docker, Helm, GitOps, and cloud‑native platform practices.
- Experience with AWS, Microsoft Azure, Google Cloud Platform, or private cloud environments.
- Knowledge of data governance, metadata management, data lineage, and data security concepts.
- Experience with change data capture technologies such as Debezium or similar tools.
- Familiarity with telecom data domains and large‑scale streaming architectures.
What We Value?
- An engineering mindset focused on maintainability, scalability, and operational reliability.
- Genuine interest in open‑source technologies and modern data architectures.
- Ability to take ownership and independently drive technical tasks.
- Willingness to challenge existing approaches and propose better solutions.
- Clear communication and effective collaboration with both technical and business stakeholders.
- A pragmatic approach that balances architectural quality with project delivery.
- Gain hands‑on experience with modern big data and lakehouse technologies.
- Contribute to architectural and technology decisions rather than only implementing predefined tasks.
- Work with experienced data engineers, architects, and analytics professionals.
- Build solutions using open‑source, cloud‑native, and vendor‑independent technologies.
- Join an agile environment where technical ownership and measurable impact are highly valued.