Overview
We are seeking a Data & Software Engineer to work with a small team building complex data flows for a custom application. Successful candidates will have advanced Python programming skills, familiarity with Java, an understanding of data security, privacy, governance and compliance principles, and a demonstrated history of building production data pipelines and ETL workflows at scale.
What will you do?
- Build end-to-end data pipelines leveraging Python.
- Use orchestration tools to deploy data pipelines, including configuring and updating Spark jobs.
- Containerize and deploy applications in cloud environments such as AWS.
- Work with MySQL and PostgreSQL, including performance tuning, schema design, and query optimization for complex analytical workloads.
- Use industry-standard tools for code control (Git, IaC control, etc.).
- Work with data catalogs, track data lineage, and handle various data formats, including geospatial.
- Use Bash scripting for automation and data-processing tasks.
- Integrate AI/ML services and models.
- Work with stakeholders to understand data requirements, assess feasibility, and design appropriate solutions with minimal oversight.
- Leverage strong problem-solving and debugging skills for data quality issues, pipeline failures, and performance bottlenecks.
- Leverage a background in large-scale data migration or platform modernization efforts.
- Contribute to data engineering documentation, best practices, and design patterns.
Do you have what it takes?
- Active TS/SCI with polygraph required.
- Bachelor’s degree in Computer Science, Engineering, Finance, or a related technical field, or equivalent practical experience.
- Minimum of 5 years’ experience with Apache Spark & PySpark, advanced Python (Pandas & NumPy), Docker/Podman, AWS S3, Lambda & Step functions, Apache Iceberg, Airflow, SQL (Trino), NoSQL (DynamoDB), Unity Catalog OSS, Apache Polaris, Apache Superset, Terraform or CloudFormation, OpenLineage, H3, PostGIS.