Overview
Data Engineer role at SourcingXPress. This position focuses on designing and implementing robust data solutions across cloud platforms and big data environments.
Role Summary
We are seeking a highly skilled and experienced Data Engineer with strong expertise in Python, SQL, and PySpark, and proven experience with Databricks and cloud platforms such as Azure, AWS, or GCP. A solid understanding of ETL tools and CI/CD practices is advantageous. This is a fast-paced role focused on delivering scalable data solutions.
Key Responsibilities
- Data Pipeline Development: Design, build, and optimize ETL/ELT workflows using Databricks, SQL, Python/PySpark, and Alteryx (Good to have).
- Develop and maintain robust, scalable data pipelines for large datasets, from source to emerging data.
- Cloud Data Engineering: Build and manage data lakes, data warehouses, and scalable data architectures on cloud platforms (Azure, AWS).
- Utilize cloud services like Azure Data Factory, AWS Glue, for data processing and orchestration.
- Databricks and Big Data Solutions: Use Databricks for big data processing, analytics, and real-time processing; leverage Apache Spark for distributed computing.
- Data Management: Create and manage SQL-based data solutions with high availability, scalability, and performance; develop data quality checks and validation mechanisms.
- Collaboration and Stakeholder Engagement: Work with data scientists, analysts, and business stakeholders to deliver impactful data solutions and translate business requirements into technical solutions.
- DevOps and CI/CD: Leverage CI/CD pipelines and tools (Git, Jenkins, or Azure DevOps) for version control and automation.
- Documentation and Optimization: Maintain clear documentation for data workflows and optimize for performance, scalability, and cost-efficiency.
Education and Experience
- Bachelor’s or Master’s degree in Computer Science, Information Technology, Engineering, or a related field.
- 3–6 years of experience in Data Engineering or related roles.
- Hands-on experience with big data processing frameworks, data lakes, and cloud-native services.
Skills
- Core Skills: Proficiency in Python, SQL, and PySpark for data processing; Databricks and Apache Spark; cloud platforms (Azure, AWS); ETL tools like Alteryx (Good to have).
- Data Engineering Expertise: Data lakes, data warehouses, data pipelines; strong understanding of distributed systems and big data technologies.
- DevOps and CI/CD: Basic understanding of DevOps; experience with Git, Jenkins, or Azure DevOps.
- Additional Skills: Familiarity with Power BI or Tableau; knowledge of streaming technologies (Kafka, Event Hubs) is desirable; strong problem-solving and communication skills.
Seniority level
Employment type
Job function
- Engineering and Information Technology
Industries
- Technology, Information and Internet
Referrals increase your chances of interviewing at SourcingXPress by 2x
We’re unlocking community knowledge in a new way. Experts add insights directly into each article, started with the help of AI.