Data Engineer

CACI Digital Experience (formerly Cyber-Duck)

Hyderabad

On-site

INR 1,500,000 - 3,000,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

CACI India is seeking a Senior Data Engineer to join our Data Engineering team. You will design and maintain cloud data pipelines, migrate workflows to a modern AWS-based data platform, and collaborate with cross-functional teams to deliver high-integrity data products.

You will work with AWS services (Glue, EMR, S3, Redshift, Sagemaker) and Python/PySpark, develop data pipelines as code, and contribute to data governance within a data mesh approach.

Qualifications

  • Strong experience and knowledge of data architectures implemented in AWS using native AWS services such as S3, DataZone, Glue, EMR, Sagemaker, Aurora and Redshift using Python, PySpark, and SQL
  • Experience developing and administrating databases, data platforms and solutions
  • Good coding discipline in terms of style, structure, versioning, documentation and unit tests
  • A well developed understanding of a data mesh organisation, as well as Master and Reference Data Management
  • Experience migrating legacy to modern, as well as modern to modern.
  • Knowledge and experience of relational databases such as Postgres, Redshift
  • Experience using Git for code versioning, and lifecycle management
  • Experience operating to Agile principles and ceremonies
  • Hands-on experience with CI/CD tools such as GitLab
  • Strong problem-solving skills and ability to work independently or in a team environment.
  • Excellent communication and collaboration skills.
  • A keen eye for detail, and a passion for accuracy and correctness in numbers

Responsibilities

  • Collaborating across CACI departments to develop and maintain data products and the data platform
  • Designing and implementing data processing environments and integrations using AWS PaaS such as Glue, S3, Lambda, Fargate, EMR, Sagemaker, Redshift, Aurora and Snowflake
  • Data architecture and data modelling across full data lifecycles, as well as more detailed modelling requirements of databases and data products.
  • Building data processing and analytics pipelines as code, using python, SQL, PySpark, Spark, CloudFormation, Lambda, Step Functions, Apache Airflow
  • Designing and applying security and access control architectures to secure sensitive data
  • Enabling business units by working with them to deliver complete and manageable solutions; while providing support and expert advice.

Skills

Data architectures
PySpark
Python
SQL
Data mesh
Postgres/Redshift
Git
Agile
CI/CD
Communication
Attention to detail
Teamwork

Tools

AWS Glue
S3
Lambda
Fargate
EMR
Sagemaker
Redshift
Aurora
Snowflake
CloudFormation
Terraform
PostgreSQL

Job description

CACI International Inc is an American multinational professional services and information technology company headquartered in Northern Virginia.

CACI provides expertise and technology to enterprise and mission customers in support of national security missions and government transformation for defense, intelligence, and civilian customers.

CACI has approximately 23,000 employees worldwide.

Headquartered in London, CACI Ltd is a wholly owned subsidiary of CACI International Inc., a publicly listed company on the NYSE with annual revenue in excess of US $6.2bn.

Founded in 2022, CACI India is an exciting, growing and progressive business unit of CACI Ltd. CACI Ltd currently has over 2000 intelligent professionals and are now adding many more from our Hyderabad and Pune offices. Through a rigorous emphasis on quality, the CACI India has grown considerably to become one of the UKs most well-respected Technology centres.

About CACI Data Engineering

CACI has implemented a Data Platform that supports and enables a Data Mesh organisation. It uses AWS technology to provide a unified experience across AWS services to deliver an open federated Lakehouse, and a unified User Experience.

The Data Platform focusses on enabling decentralised management, processing, analysis and delivery of data, while enforcing corporate wide federated governance on data, and project environments across business domains.

The goal is to empower multiple teams to create and manage high integrity data and data products that are analytics and AI ready and consumed internally and externally.

Data Engineers will be part of a community of data engineers and are often embedded in different business units to develop data pipelines and products, maintaining these for use in consulting engagements and service delivery to clients.

Data Engineers will work with the platform, corporate standards and processes. In a business that values continuous improvement, data engineers are expected to contribute to improving the wider data engineering capability in the business.

What does a Data Engineer do?

A Data Engineer partners closely with business units to maintain existing cloud data architectures, as well as assess, design and execute the migration of their existing cloud and on premise/desktop data products and workflows onto a modern cloud data platform. This involves understanding current data architectures, dependencies and transformation logic, and translating that into cloud native solutions in harmony with the Company data platform, governance and strategy. The Senior Data Engineer will need skills in developing data pipelines that operate across a medallion architecture, with an obsession on data quality and integrity of the products and processes and always considering the cost benefit of different approaches. This will need attention to technical rigour and pragmatic delivery. You will leverage skills in AWS Services such as Glue, EMR, S3, MWAA (Apache Airflow), Step Functions, Redshift, Sagemaker Unified Studio, as well as an understanding of traditional RDBMS in Postgres, Oracle and SQL Server. SQL, Python and PySpark will be essential in this role. Experience with IaC such as cloud formation will be an important skill to have, or to develop. You will be able to design architectures and create re-useable solutions to reflect the business needs. It is important that we create “Patterns” that we can reuse for multiple products and services, so that we are note designing new solutions for every need. Some products will have linear data ingest, transformation and publishing requirements, others will require extensive analytical work in the pipelines, including machine learning, deep learning and the potential to integrate more unstructured data as new products are built and matured.

Responsibilities Will Include
  • Collaborating across CACI departments to develop and maintain data products and the data platform
  • Designing and implementing data processing environments and integrations using AWS PaaS such as Glue, S3, Lambda, Fargate, EMR, Sagemaker, Redshift, Aurora and Snowflake
  • Data architecture and data modelling across full data lifecycles, as well as more detailed modelling requirements of databases and data products.
  • Building data processing and analytics pipelines as code, using python, SQL, PySpark, Spark, CloudFormation, Lambda, Step Functions, Apache Airflow
  • Designing and applying security and access control architectures to secure sensitive data
  • Enabling business units by working with them to deliver complete and manageable solutions; while providing support and expert advice.
You Will Have
  • Strong experience and knowledge of data architectures implemented in AWS using native AWS services such as S3, DataZone, Glue, EMR, Sagemaker, Aurora and Redshift using Python, PySpark, and SQL
  • Experience developing and administrating databases, data platforms and solutions
  • Good coding discipline in terms of style, structure, versioning, documentation and unit tests
  • A well developed understanding of a data mesh organisation, as well as Master and Reference Data Management
  • Experience migrating legacy to modern, as well as modern to modern.
  • Knowledge and experience of relational databases such as Postgres, Redshift
  • Experience using Git for code versioning, and lifecycle management
  • Experience operating to Agile principles and ceremonies
  • Hands-on experience with CI/CD tools such as GitLab
  • Strong problem-solving skills and ability to work independently or in a team environment.
  • Excellent communication and collaboration skills.
  • A keen eye for detail, and a passion for accuracy and correctness in numbers
Whilst not essential, the following skills would also be useful:
  • Experience using Jira, or other agile project management and issue tracking software
  • Experience with Spatial Data Processing, technology and approaches.
  • Experience with Machine Learning and AI workflows
  • Experience with IaC such as cloudformation or terraform
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Data Engineer
Data Engineer

CACI Limited • Hyderabad

On-site
INR 1,800,000 - 3,000,000
Principal Data Engineer
Principal Data Engineer

ICIMS • Hyderabad

On-site
INR 1,500,000 - 2,300,000
Principal Data Engineer
Principal Data Engineer

iCIMS, Inc. • Hyderabad

On-site
INR 3,000,000 - 6,000,000
AWS Data Engineer
AWS Data Engineer

Cigres Technologies Private Limited • Bengaluru

On-site
INR 1,200,000 - 1,600,000
Lead AWS Data Engineer
Lead AWS Data Engineer

Weekday (YC W21) • Bengaluru

On-site
INR 1,000,000 - 1,500,000
Data Engineer
Data Engineer

Calfus Inc. • Pune District

On-site
INR 900,000 - 1,500,000
Medical insurance
Parental insurance
Provident fund
+1
Data Engineer
Data Engineer

Pull Skil • Hyderabad

On-site
INR 800,000 - 1,200,000
Lead Data Engineer
Lead Data Engineer

Cynosure Corporate Solutions • Chennai District

On-site
INR 400,000 - 600,000
Data Engineering Senior Analyst
Data Engineering Senior Analyst

The Cigna Group • Bengaluru

On-site
INR 1,800,000 - 2,400,000
Cloud Data Engineer - Data Pipeline and Quality
Cloud Data Engineer - Data Pipeline and Quality

Thakral One • Bengaluru

On-site
INR 1,200,000 - 1,800,000