The Data Engineer (Streaming/ETL) will support the design, development, and deployment of data ingestion and processing pipelines for advanced analytics platforms. This role is critical to enabling real-time and batch data integration across multiple sources, including operational systems, telemetry streams, and structured/unstructured datasets.
The position focuses on building scalable data pipelines, supporting instrumentation data ingestion, and enabling downstream analytics, machine learning, and decision support capabilities within a cloud-native environment. The Data Engineer will work closely with data scientists, software engineers, and solution architects to ensure data is accessible, reliable, and optimized for analytics and AI/ML workloads.
Responsibilities
- Design and implement ETL/ELT pipelines for structured, semi-structured, and unstructured data sources
- Develop and maintain scalable data ingestion frameworks for batch and real-time data processing
- Build data connectors and APIs for integrating external systems and operational data sources
- Support ingestion of high-volume telemetry and instrumentation data streams
- Normalize and transform incoming data into standardized formats for analytics use
Streaming & Real-Time Data Processing
- Develop streaming data pipelines for real-time or near-real-time processing (e.g., event-driven architectures)
- Implement data buffering and processing strategies for edge or low-connectivity environments
- Enable event-based data processing and integration into analytics workflows
- Optimize pipelines for high-throughput and low-latency performance
- Collaborate with architects to design scalable data architectures supporting multi-source integration
- Support integration across relational, graph, and data lake environments
- Assist in building unified data ingestion frameworks across multiple programs
- Ensure data is structured and accessible for downstream analytics and machine learning
Data Quality & Governance
- Implement data validation, integrity checks, and monitoring within ingestion pipelines
- Support development of data quality frameworks and verification processes
- Assist in tracking data lineage and ensuring traceability across pipelines
- Enforce data handling and security practices in accordance with system requirements
Collaboration & Stakeholder Support
- Work with data scientists to prepare datasets for machine learning and analytics
- Collaborate with engineers and analysts to support dashboarding and visualization efforts
- Participate in requirements gathering and technical discussions with stakeholders
- Support cross-functional teams in integrating data pipelines into operational workflows
- Evaluate new tools and technologies for data ingestion, streaming, and processing
- Contribute to improving data pipeline performance, scalability, and reliability
- Assist in prototyping new data ingestion approaches for evolving mission needs
- Document data architecture, pipelines, and integration processes
Required Knowledge, Skills, and Abilities
- Strong understanding of data ingestion patterns and pipeline design
- Experience working with structured and unstructured data sources
- Streaming & Data Processing
- Familiarity with streaming or event-driven architectures (Kafka, Kinesis, or similar)
- Experience handling high-volume data ingestion and processing workflows
- Understanding of real-time vs batch processing tradeoffs
- Technical Skills
- Strong proficiency in Python for data engineering workflows
- Experience with SQL and relational databases (PostgreSQL, MySQL, etc.)
- Familiarity with data lake architectures (S3 or similar object storage)
- Experience building APIs or working with REST-based integrations
- Experience working in cloud environments (AWS preferred: EC2, S3, ECS, Lambda, etc.)
- Familiarity with containerized environments (Docker, Kubernetes is a plus)
- Understanding of scalable and distributed data systems
- Ability to work with cross-functional teams including data scientists and software engineers
- Strong written and verbal communication skills
- Ability to translate technical concepts into understandable solutions
- Problem Solving
- Ability to design efficient data pipelines and troubleshoot performance issues
- Experience identifying data integration challenges and proposing solutions
- Comfortable working in fast-paced, evolving environments
Experience With Specific Technologies
- Python (pandas, PySpark, or similar data processing tools)
- SQL / Relational Databases
- Data Streaming Tools (Kafka, Kinesis, or equivalent)
- ETL Tools or Frameworks (custom or commercial)
Additional Information
Education
Bachelor’s degree in Computer Science, Data Engineering, Information Systems, or a related technical field
Physical Requirements
Ability to work independently with minimal guidance and lift up to 25 pounds
Horizon Defense Solutions (HDS) offers a comprehensive and generous benefits package. The HDS benefits package includes medical, dental, and vision insurance for the employee and/or families. HDS also includes basic life insurance plus short- and long-term disability for the employee. Employees may elect to enroll in our company's 401k plan. Employees will also accrue paid time off and holidays.
About the Company
HDS is committed to non-discrimination and equal employment opportunity. All qualified applicants will receive consideration for employment without discrimination based on disability, protected veteran status or any other characteristics protected by law.