3 days ago Be among the first 25 applicants
Required Skillset: Streamsets, Python, SQL
Experience Range: 4-10 Years
Job Description
Must Have:
- Hands-on experience with Streamsets (especially with Data Collectors and Control Hub)
- Strong knowledge of SQL, ETL/ELT pipelines, and data integration patterns
- Experience with real-time data processing and stream processing frameworks
- Experience with cloud environments and Data Lake/Warehouse
- Familiarity with source control tools (e.g., Git) and CI/CD pipelines
Good To Have:
- Knowledge of Python, Scala, or Java for scripting and automation
- Understanding of data privacy, compliance, and governance frameworks
Responsibilities / Expectations:
- Design, develop, and deploy end-to-end data pipelines using StreamSets Data Collector and StreamSets Control Hub
- Integrate data from diverse sources including relational databases, cloud storage (AWS/Azure/GCP), APIs, and streaming platforms like Kafka
- Collaborate with data architects and business analysts to understand data requirements and implement solutions accordingly
- Monitor data flows and troubleshoot issues related to data quality, latency, and pipeline failures
- Optimize pipeline performance, reliability, and scalability