A leading tech company in Kentucky is seeking a Data Engineer to lead technical milestones and develop scalable data architectures. Responsibilities include creating POCs, managing data ingestion services, and designing event-driven architectures using AWS. The ideal candidate should have strong skills in Python, AWS services, and NoSQL databases. This position requires collaboration with stakeholders on data architecture and optimization opportunities.
Responsibilities
Lead the team technically to complete milestones on time.
Understand complete requirement, create Architecture and update all stakeholders.
Create POCs.
Manage delivery/release to customer.
Develop Services for data ingestion and synchronization.
Ingest data from multiple sources using Python.
Design event-driven architecture using AWS EventBridge or SNS/SQS.
Implement scalable data pipelines integrating on-prem and AWS environments.
Develop efficient Python scripts to handle large datasets.
Work with various NoSQL databases.
Develop applications in a cloud-native architecture.
Monitor data workflows, troubleshoot issues, and optimize performance.
Transition existing pipeline to MSSQL server.
Collaborate on existing data architecture.
Design and document target data architecture and pipelines.
Identify opportunities for optimization.
Collaborate on business logic decomposition.
Skills
AWS
Glue
SNS/SQS
Python
Py-Spark
Data Lake
Cloud Watch
Cloud Trail
DB Design
SQL
Job description
Skills
AWS
Glue
SNS/SQS
Python
Py-Spark
Data Lake
Cloud Watch
Cloud Trail
DB Design
SQL
Responsibilities
Lead the team technically to complete milestones on time.
Understand complete requirement, create Architecture and update all stakeholders.
Create POCs.
Manage delivery/release to customer.
Develop Services to enable data ingestion from and synchronization with system which exposes required data access mechanisms ensuring near-real-time updates.
Ingest data from multiple sources using Python and any other ETL tools.
Design and implement an event-driven architecture using AWS EventBridge, Kafka, or SNS/SQS for real-time data streaming.
Design, implement, and maintain scalable data pipelines that integrate both on-prem and AWS cloud environments.
Develop efficient Python scripts and applications using libraries like pandas, NumPy, etc., to handle and process large datasets.
Work with various NoSQL databases (e.g., MongoDB, Cassandra, DynamoDB) to support high-performance data storage and retrieval.
Develop and deploy applications in a cloud-native architecture, leveraging modern cloud technologies for scalability and resilience.
Continuously monitor data workflows and systems, troubleshoot issues, and optimize performance for reliability and scalability.
Transition existing pipeline to MSSQL server.
Collaborate with the business application owner on the existing data architecture, including data ingestion, data pipelines, business logic, data consumption patterns, and analytics requirements.
Design and document the target data architecture, pipelines, processing and analytics architecture.
Identify opportunities for optimization and consolidation.
Collaborate with data team on decomposition of business logic and data transformation patterns.