Data Engineer – Baltimore City, MD

Creative Information Technology India

Falls Church (VA)

On-site

USD 120,000 - 160,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Creative Information Technology Inc is seeking a hands-on Data Engineer to design, develop, and optimize large-scale data pipelines for the Enterprise Data Warehouse and Data Lake solutions. You will implement ingestion, transformation, and integration workflows using PySpark on AWS, ensuring data quality and analytics readiness.

Responsibilities include building data architectures, coordinating with teams, and enforcing CI/CD and IaC practices to improve scalability and cost efficiency.

Qualifications

  • Bachelor’s or master’s degree in CS/Stats/Math/Economics or related field.
  • 3+ years of experience building data pipelines on AWS or equivalent cloud platforms.
  • Proficient in Python and SQL; Spark (PySpark) experience.
  • Experience with ETL processes and data architecture.

Responsibilities

  • Design, build, and maintain data pipelines and infrastructure to support data-driven decisions.
  • Develop data lakehouse architectures (data lake, data warehouse, data marts) for analytics.
  • Implement data reliability, quality checks, and performance improvements.
  • Collaborate with multiple teams and ensure security and compliance of pipelines.

Skills

Python
SQL
Apache Spark (PySpark)
ETL
Airflow
CI/CD
Terraform
AWS
AWS Glue
Delta Lake
Iceberg
Data Pipelines

Education

Bachelor’s or Master’s in CS/Stats/Math/Economics

Tools

AWS Glue
S3
Redshift
Athena
EMR
Iceberg
Delta Lake
Airflow

Job description

If you are unable to complete this application due to a disability, contact this employer to ask for an accommodation or an alternative application process.

Data Engineer – Baltimore City, MD

Contract Falls Church, VA, US

3 days ago Requisition ID: 1853

Data Engineer – Baltimore City, MD

About us

Creative Information Technology Inc (CITI) is an esteemed IT enterprise renowned for its exceptional customer service and innovation. We serve both government and commercial sectors, offering a range of solutions such as Healthcare IT, Human Services, Identity Credentialing, Cloud Computing, and Big Data Analytics. With clients in the US and abroad, we hold key contract vehicles including GSA IT Schedule 70, NIH CIO-SP3, GSA Alliant, and DHS-Eagle II.

Join us in driving growth and seizing new business opportunities.

Background

Client is seeking a hands-on Data Engineer to design, develop, and optimize large-scale data pipelines in support of our Enterprise Data Warehouse (EDW) and Data Lake solutions. This role requires deep technical expertise in coding, pipeline orchestration, and cloud-native data engineering on AWS. The Data Engineer will be directly responsible for implementing ingestion, transformation, and integration workflows — ensuring data is high-quality, compliant, and analytics-ready. This role may support other projects or teams within MDH as needed.

Role and Responsibilities

Responsible for designing, building, and maintaining data pipelines and infrastructure to support data-driven decisions and analytics. The individual is responsible for the following tasks:

  • Design, develop and maintain data pipelines, and extract, transform, load (ETL) processes to collect, process and store structured and unstructured data
  • Build data architecture and storage solutions, including data lakehouses, data lakes, data warehouse, and data marts to support analytics and reporting
  • Develop data reliability, efficiency, and qualify checks and processes
  • Monitor and optimize data architecture and data processing systems
  • Collaboration with multiple teams to understand requirements and objectives
  • Administer testing and troubleshooting related to performance, reliability, and scalability
  • Create and update documentation
  • Design, code, and deploy ETL/ELT pipelines across bronze, silver, and gold layers of the Data Lakehouse.
  • Build ingestion pipelines for structured (SQL), semi-structured (JSON, XML), and unstructured data using PySpark/Python programming language using AWS Glue or EMR.
  • Implement incremental loads, deduplication, error handling, and data validation.
  • Actively troubleshoot, debug, and optimize pipelines for scalability and cost efficiency.

EDW & Data Lake Implementation

  • Develop dimensional data models (Star Schema, Snowflake Schema) for analytics and reporting.
  • Build and maintain tables in Iceberg, Delta Lake, or equivalent OTF formats.
  • Optimize partitioning, indexing, and metadata for fast query performance.
  • Build ingestion and transformation pipelines for EDI X12 transactions (837, 835, 278, etc.
  • Implement mapping and transformation of EDI data with FHIR and HL7 frameworks.
  • Work hands-on with AWS Health Lake (or equivalent) to store and query healthcare data.

Data Quality, Security & Compliance

  • Develop automated validation scripts to enforce data quality and integrity.
  • Implement IAM roles, encryption, and auditing to meet HIPAA and CMS compliance standards.
  • Maintain lineage and governance documentation for all pipelines.
  • Work closely with the Lead Data Engineer, analysts, and data scientists to deliver pipelines that support enterprise-wide analytics.
  • Actively contribute to CI/CD pipelines, Infrastructure-as-Code (IaC), and automation.
  • Continuously improve pipelines and adopt new technologies where appropriate.

Minimum Qualifications

  • The candidate should have experience as data engineer or similar role with a strong understanding of data architecture and ETL processes. The candidate should be proficient in programming languages for data processing and knowledgeable of distributed computing and parallel processing.
  • This position requires a bachelor’s or master’s degree from an accredited college or university with a major in computer science, statistics, mathematics, economics, or a related field. Three (3) years of equivalent experience in a related field may be substituted for the Bachelor’s degree.
  • 3+ years hands-on experience in building, deploying, and maintaining data pipelines on AWS or equivalent cloud platforms.
  • Strong coding skills in Python and SQL (Scala or Java a plus).
  • Proven experience with Apache Spark (PySpark) for large-scale processing.
  • Hands-on experience with AWS Glue, S3, Redshift, Athena, EMR, Lake Formation.
  • Strong debugging and performance optimization skills in distributed systems.
  • Hands-on experience with Iceberg, Delta Lake, or other OTF table formats.
  • Experience with Airflow or other pipeline orchestration frameworks.
  • Practical experience in CI/CD and Infrastructure-as-Code (Terraform, CloudFormation).
  • Practical experience with EDI X12, HL7, or FHIR data formats.
  • Strong understanding of Medallion Architecture for data lake houses.
  • Hands-on experience building dimensional models and data warehouses.
  • Working knowledge of HIPAA and CMS interoperability requirements.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Data Engineer
Data Engineer

Magpie Health Analytics, Inc. • Baltimore (MD)

On-site
USD 110,000 - 165,000
Senior Data Engineer
Senior Data Engineer

Smart Synergies • McLean (VA)

On-site
USD 120,000 - 180,000
Data Architect
Data Architect

Capital Technology Alliance • Tallahassee (FL)

On-site
USD 130,000 - 160,000
Health insurance
401(k) retirement plan
Flexible working hours
Data Engineer - Python, SQL, AWS
Data Engineer - Python, SQL, AWS

Compunnel, Inc. • Durham (NC)

On-site
USD 95,000 - 120,000
Senior Data Engineer
Senior Data Engineer

Compunnel, Inc. • Eden Prairie (MN)

On-site
USD 120,000 - 150,000
Data Engineer
Data Engineer

VTG Defense • McLean (VA)

On-site
USD 100,000 - 130,000
Data Engineer
Data Engineer

Compunnel, Inc. • Columbus (OH)

On-site
USD 85,000 - 115,000
Senior Data Engineer
Senior Data Engineer

Compunnel, Inc. • Charlotte (NC)

On-site
USD 120,000 - 150,000
Data Engineer - Healthcare
Data Engineer - Healthcare

Compunnel, Inc. • Atlanta (GA)

On-site
USD 120,000 - 180,000
Data Engineer
Data Engineer

The Value Maximizer • South Carolina

On-site
USD 90,000 - 120,000