Sr. Data Engineer

Pyramid Systems

Merrifield (VA)

On-site

USD 106,131 - 159,196

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Pyramid Systems is seeking a Senior Data Engineer to design and maintain enterprise data architectures, pipelines, and data lake solutions in a fast-paced environment.

You will optimize data processing with Spark, Databricks, and Delta Lake, while collaborating with AI/ML teams and data science stakeholders to deliver high-quality data services.

Qualifications

  • 8+ years of IT experience focusing on enterprise data architecture and management.

Responsibilities

  • Plan, create, and maintain data architectures aligned with business requirements.

Skills

Databricks
Structured Streaming
Delta Lake
Delta Live Tables
ETL/ELT tools
SQL
Python
Spark
AWS
Kafka

Education

Bachelor's degree in Computer Science or related discipline

Tools

Docker
Jenkins
CloudWatch
Kinesis
Confluent/Kafka
S3
DynamoDB
JSON Schemas
Schema Registry

Job description

Overview

Pyramid Systems is looking for a Data Engineer (Senior) who is passionate about bringing creative architect solutions to end customers.

Key Skills
  • 8+ years of IT experience focusing on enterprise data architecture and management
  • Experience with Databricks, Structured Streaming, Delta Lake concepts, and Delta Live Tables required
  • Experience with ETL and ELT tools such as SSIS, Pentaho, and/or Data Migration Services
  • Advanced level SQL experience (Joins, Aggregation, Windowing functions, Common Table Expressions, RDBMS schema design, Postgres performance optimization)
Responsibilities
  • Plan, create, and maintain data architectures, ensuring alignment with business requirements
  • Obtain data, formulate dataset processes, and store optimized data
  • Identify problems and inefficiencies and apply solutions
  • Determine tasks where manual participation can be eliminated with automation.
  • Identify and optimize data bottlenecks, leveraging automation where possible
  • Create and manage data lifecycle policies (retention, backups/restore, etc)
  • In-depth knowledge for creating, maintaining, and managing ETL/ELT pipelines
  • Create, maintain, and manage data transformations
  • Maintain/update documentation
  • Create, maintain, and manage data pipeline schedules
  • Monitor data pipelines
  • Create, maintain, and manage data quality gates (Great Expectations) to ensure high data quality
  • Support AI/ML teams with optimizing feature engineering code
  • Expertise in Spark/Python/Databricks, Data Lake and SQL
  • Create, maintain, and manage Spark Structured Streaming jobs, including using the newer Delta Live Tables and/or DBT
  • Research existing data in the data lake to determine best sources for data
  • Create, manage, and maintain ksqlDB and Kafka Streams queries/code
  • Data driven testing for data quality
  • Maintain and update Python-based data processing scripts executed on AWS Lambdas
  • Unit tests for all the Spark, Python data processing and Lambda codes
  • Maintain PCIS Reporting Database data lake with optimizations and maintenance (performance tuning, etc)
  • Streamlining data processing experience including formalizing concepts of how to handle lake data, defining windows, and how window definitions impact data freshness.
Qualifications
  • 8+ years of IT experience focusing on enterprise data architecture and management
  • Must be able to obtain a Public Trust security clearance
  • US Citizen
  • Bachelor degree required
  • Experience in Conceptual/Logical/Physical Data Modeling & expertise in Relational and Dimensional Data Modeling
  • Experience with Databricks, Structured Streaming, Delta Lake concepts, and Delta Live Tables required
  • Additional experience with Spark, Spark SQL, Spark DataFrames and DataSets, and PySpark
  • Data Lake concepts such as time travel and schema evolution and optimization
  • Structured Streaming and Delta Live Tables with Databricks a bonus
  • Experience leading and architecting enterprise-wide initiatives specifically system integration, data migration, transformation, data warehouse build, data mart build, and data lakes implementation / support
  • Advanced level understanding of streaming data pipelines and how they differ from batch systems
  • Formalize concepts of how to handle late data, defining windows, and data freshness
  • Advanced understanding of ETL and ELT and ETL/ELT tools such as SSIS, Pentaho, Data Migration Service etc
  • Understanding of concepts and implementation strategies for different incremental data loads such as tumbling window, sliding window, high watermark, etc.
  • Familiarity and/or expertise with Great Expectations or other data quality/data validation frameworks a bonus
  • Understanding of streaming data pipelines and batch systems
  • Familiarity with concepts such as late data, defining windows, and how window definitions impact data freshness
  • Advanced level SQL experience (Joins, Aggregation, Windowing functions, Common Table Expressions, RDBMS schema design, Postgres performance optimization)
  • Indexing and partitioning strategy experience
  • Debug, troubleshoot, design and implement solutions to complex technical issues
  • Experience with large-scale, high-performance enterprise big data application deployment and solution
  • Understanding how to create DAGs to define workflows
  • Familiarity with CI/CD pipelines, containerization, and pipeline orchestration tools such as Airflow, Prefect, etc a bonus but not required
  • Architecture experience in AWS environment a bonus
  • Familiarity working with Kinesis and/or Lambda specifically with how to push and pull data, how to use AWS tools to view data in Kinesis streams, and for processing massive data at scale a bonus
  • Experience with Docker, Jenkins, and CloudWatch
  • Ability to write and maintain Jenkinsfiles for supporting CI/CD pipelines
  • Experience working with AWS Lambdas for configuration and optimization
  • Experience working with DynamoDB to query and write data
  • Experience with S3
  • Knowledge of Python (Python 3 desired) for CI/CD pipelines a bonus
  • Familiarity with Pytest and Unittest a bonus
  • Experience working with JSON and defining JSON Schemas a bonus
  • Experience setting up and management Confluent/Kafka topics and ensuring performance using Kafka a bonus
  • Familiarity with Schema Registry, message formats such as Avro, ORC, etc.
  • Understanding how to manage ksqlDB SQL files and migrations and Kafka Streams
  • Ability to thrive in a team-based environment
  • Experience briefing the benefits and constraints of technology solutions to technology partners, stakeholders, team members, and senior level of management
Education
  • Bachelor's degree in Computer Science or related discipline
Pay Range

The below listed pay range for this position is not a guarantee of compensation or salary. The final offered salary will be influenced by a host of factors including, but not limited to, geographic location, Federal Government contract labor categories and contract wage rates, relevant prior work experience, specific skills and competencies, education, and certifications. Our employees value the flexibility at Pyramid Systems that allows them to balance quality work and their personal lives. We offer competitive compensation, benefits, to include our Employee Stock Ownership Program, FlexPTO, and learning and development opportunities.

Pyramid Min

USD $106,131.00/Yr.

Pyramid Max

USD $159,196.00/Yr.

EEO Statement

Pyramid Systems, Inc. is an Equal Opportunity Employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, or protected veteran status and will not be discriminated against on the basis of disability.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Sr. Data Engineer
Sr. Data Engineer

Pyramid Systems, Inc. • United States

On-site
USD 106,000 - 160,000
Employee Stock Ownership Program
FlexPTO
Learning and development opportunities
Senior Cloud Engineer
Senior Cloud Engineer

Pyramid Systems • Merrifield (VA)

On-site
USD 125,000 - 189,000
Employee Stock Ownership Program
FlexPTO
Learning and development opportunities
Senior Acquisition Support Specialist
Senior Acquisition Support Specialist

Pyramid Systems • Merrifield (VA)

On-site
USD 150,000 - 200,000
SENIOR ACQUISITIONS SUPPORT SPECIALIST
SENIOR ACQUISITIONS SUPPORT SPECIALIST

Pyramid Systems • Merrifield (VA)

On-site
USD 150,000 - 200,000
ESOP
FlexPTO
Learning and development opportunities
Cyber Risk Lead- Security Control Assessor - Senior
Cyber Risk Lead- Security Control Assessor - Senior

Pyramid Systems, Inc. • United States

On-site
USD 104,000 - 107,000
Employee Stock Ownership Plan (ESOP)
Top Workplace recognition
Senior Software Engineer/Developer - AI
Senior Software Engineer/Developer - AI

Pyramid Systems, Inc. • United States

On-site
USD 166,000 - 210,000
Employee Stock Ownership Program
Flexible Paid Time Off
Learning and development opportunities
Data Engineer (GovCon; Public Trust) United States - Remote
Data Engineer (GovCon; Public Trust) United States - Remote

Attain Talent • United States

Hybrid
USD 110,000 - 140,000
Remote Work (Hybrid)
Medical, Dental, Vision
401(k) with matching
+2
SENIOR ACQUISITIONS SUPPORT SPECIALIST
SENIOR ACQUISITIONS SUPPORT SPECIALIST

Pyramid Systems, Inc. • United States

On-site
USD 150,000 - 200,000
ESOP
Senior Acquisition Support Specialist
Senior Acquisition Support Specialist

Pyramid Systems, Inc. • United States

On-site
USD 150,000 - 200,000
Employee Stock Ownership Plan (ESOP)
FlexPTO
Learning and development opportunities
Data Engineer
Data Engineer

Easy Dynamics • McLean (VA)

On-site
USD 150,000 - 180,000