Our client is seeking a Data Engineer II to participate in building our next generation of data pipelines and ETL processing complex, large volumes of structured and unstructured data. This role sits at the center of a major infrastructure transformation—migrating our core data platform to AWS and Snowflake—and will be critical to ensuring that work is delivered with engineering rigor, operational resilience, and clear knowledge transfer back to the internal team.
The ideal candidate is a technical, hands-on engineer who is comfortable in complex cloud environments. They bring production experience with AWS data services, Python, and Snowflake—and they know how to build pipelines and transformation layers that are not just functional, but maintainable, observable, and built to scale. They can work alongside and collaborate with internal and external partners delivering consistent productivity and take pride in owning the quality of what they ship.
You’ll join a high-performing Data & Analytics team operating at the intersection of real estate, operations, and AI—building the data foundation that powers pricing models, operational intelligence tools, and strategic decisions across 40+ divisions of one of the nation’s largest homebuilders.
Responsibilities:- Design, build, and maintain efficient and scalable data pipelines and ETL processing large volumes of structured and unstructured data.
- Optimize data storage and retrieval methods to ensure performance, scalability, and cost-efficiency.
- Collaborate with data scientists, analysts, and business stakeholders to understand data requirements and translate them into technical solutions.
- Implement data governance, data security, and compliance measures to protect sensitive data and ensure regulatory compliance.
- Develop and manage data integrations between internal systems and third-party data providers.
- Conduct performance tuning of SQL queries and data processing workflows to improve system efficiency.
- Create and maintain comprehensive documentation for data pipelines, data models, and related processes.
- Provide technical support and mentorship to junior data engineers, guiding them in best practices and problem-solving.
- Write clean, modular Python for data ingestion, transformation logic, and orchestration—applying software engineering best practices including testing, versioning, and code review.
- Build and maintain data quality frameworks, pipeline observability, and alerting systems that give the team confidence in production data across critical domains including pricing, supply chain, and sales operations.
- Contribute to infrastructure-as-code practices using Terraform to provision and manage cloud resources in a repeatable, auditable way (experience a plus; willingness to learn required).
Education and Experience Requirements:- Bachelor’s degree in Computer Science, Engineering, Information Technology or a related field.
- 2-4 years of experience in Data Engineering or a related role, with demonstrated success in delivering data solutions.
- Proficient in SQL, with the ability to write complex queries, perform query optimization, and conduct performance tuning.
- Experience with NoSQL databases, such as MongoDB, Cassandra, or DynamoDB, and an understanding of their appropriate use cases.
- Strong programming skills in Python, Java, or Scala, with experience in data processing frameworks (e.g. Glue, Lambda, S3, dbt, Qlik, Apache Spark, Hadoop).
- Experience with cloud platforms (AWS, Snowflake, Azure), and data services, such as Redshift, Data Sync, Appflow, Databricks.
- Knowledge of big data technologies, including Hadoop, Spark, Kafka, and HBase, with experience in distributed data processing.
- Familiarity with data orchestration tools, such as Apache Airflow or Luigi, for scheduling and managing data workflows.
- Experience with data versioning and testing tools, such as DVC (Data Version Control) and dbt (data build tool).
- Understanding of data security practices, including encryption, access controls, and data masking.
- Proven ability to collaborate across technical and non-technical teams—communicating clearly, translating business requirements into engineering decisions, and building trust with stakeholders at multiple levels.