Incedo Inc. is a high-growth Digital, Data, and AI Transformation Specialist firm headquartered in New Jersey. We are a long-term strategy execution partner for Fortune 500 enterprises, operating at the intersection of business and technology. We serve clients across 5 key industries: Banking & Payments, Wealth Management, Telecom, Hi-Tech, and Life Sciences.
Incedo delivers ROI from AI @ Scale by leveraging the 'Power of 3':
- Deep domain expertise: Deep understanding of industries with an ability to identify and solve specific business problems.
- AI & Data: Integration of Data and AI capabilities with business priorities to deliver impact and ROI
- Engineering & Operations: Excellence in building digital platforms, driving modernization, integration, and tech operations.
Core principles that define the DNA of Incedo are:-
- Business outcome focus: Drive measurable business outcomes by developing a client-first view prioritized against quantifiable, business KPI-driven programs.
- Business and technology intersection: Bring together deep domain, AI, data, and engineering capabilities.
- Build for the long-term: Enable sustained growth by building long-term client relationships with Fortune 500 companies and senior client stakeholders.
- Integrated platform and services: Provide a comprehensive offering of the payables IncedoPay platform and Incedo Lighthouse™ solving specific business problems.
- Cross-functional 'one' team: Align multidisciplinary teams including domain, design, data, engineering, and business operations to work towards a collective goal.
- With a team of 4,500+ professionals, we have offices across the U.S., Mexico, Canada, and India.
Job Summary
We are seeking a highly skilled AWS Data Engineer to design, develop, and optimize large-scale data pipelines and ETL workflows on AWS. The ideal candidate will have strong expertise in AWS cloud-native data services, data modeling, and pipeline orchestration, with hands-on experience building robust and scalable data solutions for enterprise environments.
Key Responsibilities
- Design and implement incremental and CDC (Change Data Capture) data pipelines using AWS Glue, DMS, and Iceberg to support near real-time analytics.
- Develop and maintain metadata-driven ETL frameworks to improve reusability, scalability, and operational efficiency across data platforms.
- Create and manage partitioning, compaction, and optimization strategies for Iceberg datasets to reduce query latency and storage costs.
- Build and orchestrate complex workflows using AWS Step Functions, EventBridge, Lambda, and Glue Workflows for automated data processing.
- Perform performance tuning and cost optimization of AWS Glue jobs by optimizing Spark configurations, worker types, partitioning, and job bookmarks.
- Implement CI/CD pipelines for data engineering solutions using AWS CodePipeline, CodeBuild, GitHub, Jenkins, or Terraform.
- Develop and maintain data lake architecture following AWS best practices, ensuring scalability, reliability, and governance.
- Automate data validation and reconciliation processes to ensure data accuracy, completeness, and consistency across multiple systems.
- Create and maintain Athena external tables, Iceberg catalogs, and Glue Data Catalog metadata for efficient data discovery and querying.
- Design and implement role-based access controls (RBAC), data masking, encryption, and audit mechanisms using Lake Formation and IAM policies.
- Support real-time and batch processing architectures integrating Kafka, Kinesis, PostgreSQL, S3, and Redshift.
- Monitor data pipelines using CloudWatch, SNS, AWS Glue Monitoring, and custom alerting mechanisms to ensure SLA compliance.
- Work closely with enterprise architecture and governance teams to establish data standards, retention policies, and compliance frameworks.
- Perform root cause analysis and resolve complex production issues involving Spark, Glue, Iceberg metadata, PostgreSQL connectivity, and permission models.
- Enable self-service analytics by creating curated gold, silver, and bronze data layers within enterprise data lakes.
- Manage schema evolution and version control for Iceberg datasets while maintaining backward compatibility for downstream consumers.
- Develop reusable PySpark utilities, frameworks, and common libraries to standardize data ingestion and transformation patterns.
- Participate in architecture reviews and recommend best practices for data lake modernization, cloud migration, and platform optimization initiatives.
- Implement data lineage, cataloging, and observability solutions to improve data trust, discoverability, and governance.
- Collaborate with DevOps and Infrastructure teams to provision and manage AWS resources using Terraform, CloudFormation, or Infrastructure as Code (IaC) methodologies
Required Qualifications
- 12+ years of experience in Data Engineering, with at least 5+ years designing and implementing AWS cloud-based data platforms and enterprise-scale data lakes.
- Strong hands-on expertise in AWS Glue, AWS DMS, Amazon S3, Amazon Redshift, Athena, Lambda, Step Functions, EventBridge, CloudWatch, SNS, and Glue Data Catalog.
- Advanced proficiency in PySpark, Spark, Python, and SQL, with experience building reusable frameworks, ETL/ELT pipelines, and high-volume data transformation solutions.
- Hands-on experience designing and implementing incremental and Change Data Capture (CDC) pipelines using AWS DMS, Glue, and related AWS services.
- Strong experience with Apache Iceberg, including partitioning strategies, compaction, schema evolution, metadata management, performance optimization, and Iceberg catalog implementation.
- Extensive experience building and supporting enterprise Data Lake architectures using Bronze, Silver, and Gold data layers with strong focus on scalability, reliability, governance, and self-service analytics.
- Experience implementing and orchestrating complex workflows using AWS Step Functions, Glue Workflows, Lambda, and EventBridge.
- Strong knowledge of data modeling, including dimensional, star, snowflake, and lakehouse modeling techniques.
- Experience optimizing AWS Glue and Spark workloads, including partitioning, job bookmarks, worker sizing, Spark configuration tuning, and cost optimization.
- Hands-on experience with real-time and batch data processing architectures integrating Kafka, Kinesis, PostgreSQL, S3, and Redshift.
- Experience implementing data quality, reconciliation, observability, lineage, monitoring, and alerting frameworks across enterprise data platforms.
- Strong understanding of data security and governance, including Lake Formation, IAM, RBAC, encryption, masking, auditing, retention policies, and regulatory compliance requirements.
- Experience with CI/CD and Infrastructure as Code (IaC) using Terraform, CloudFormation, AWS CodePipeline, CodeBuild, GitHub, Jenkins, or similar technologies.
- Strong analytical, troubleshooting, and root cause analysis skills for resolving complex production issues across Spark, Glue, Iceberg, data pipelines, and cloud infrastructure.
- Experience working in Agile/Scrum environments and collaborating with architecture, governance, DevOps, and business stakeholders.
Preferred Qualifications
- Experience in Financial Services, Wealth Management, Brokerage, Capital Markets, or BFSI domains, preferably supporting regulatory and governed data environments.
- AWS Certified Data Engineer - Associate certification required/preferred; additional AWS certifications in Analytics, Data Engineering, or Solutions Architecture are highly desirable.
- Hands-on experience with AWS Glue, PySpark, Apache Iceberg, Lake Formation, Kafka, Kinesis, and Amazon EMR/Spark in large-scale enterprise implementations.
- Experience building metadata-driven ingestion and ETL frameworks and platform accelerators for reusable data engineering patterns.
- Experience implementing data cataloging, lineage, observability, and governance solutions using enterprise data management tools.
- Familiarity with modern Lakehouse architectures, cloud migration initiatives, and data platform modernization programs.
- Experience with Terraform, CloudFormation, GitHub Actions, Jenkins, CodePipeline, and DevOps practices for enterprise data platforms.
- Master's degree in Computer Science, Information Systems, Engineering, or related discipline preferred.
Education
- Bachelor's degree in Computer Science, Information Technology, Engineering, or related field.