JOB SUMMARY
We are seeking a highly motivated Engineer to join our Datastore Migration Factory team. This critical and high-visibility project involves the end-to-end migration of our on-premise Data Lake to an AWS-hosted Lakehouse. The Engineer will play a crucial role in migrating pipelines and consumption patterns, ensuring data integrity and quality, and acting as a technical liaison with stakeholders. The successful candidate will possess strong data engineering skills, experience with data migration, and excellent communication and collaboration abilities. This is a 100% onsite position based in Dallas, TX; local candidates only will be considered.
Key Responsibilities
- Refactoring and migrating extraction logic and job scheduling from legacy frameworks to the new Lakehouse environment.
- Executing the physical migration of underlying datasets while ensuring data integrity.
- Translating and optimizing legacy SQL and Spark-based consumption patterns (raw and modeled) for compatibility with Snowflake and Iceberg.
- Performing usage analysis to understand usage patterns and deliver required data products.
- Working with reconciliation frameworks to validate migrated data and ensure functional equivalence.
- Acting as a technical liaison to internal clients, facilitating "handoff and sign-off" conversations with data owners.
- Collaborating effectively across multiple teams and functions.
- Communicating concise written updates, structured verbal briefings, and proactively managing stakeholders.
Required Qualifications
- Bachelor’s or Master’s degree in Computer Science, Applied Mathematics, Engineering, or a related quantitative field.
- Minimum of 3-5 years of professional "hands‑on‑keyboard" coding experience in a collaborative, team-based environment.
- Ability to troubleshoot (SQL) and basic scripting experience.
- Professional proficiency in Python or Java.
- Deep familiarity with the full Software Development Life Cycle (SDLC).
- Experience with CI/CD best practices & K8s deployment.
- Sophisticated understanding of temporal data modeling (e.g., SCD Type 2).
- Expertise in Schema Evolution (Ref: Iceberg Apache) and enforcement strategies.
- Advanced knowledge of data partitioning and clustering.
- Understanding of balancing Normalization vs. Denormalization and the strategic use of Natural vs. Surrogate Keys.
- Experience with data extraction and logic tools such as Kafka, ANSI SQL, FTP, and Apache Spark.
- Experience with data formats such as JSON, Avro, and Parquet.
- Experience with platforms such as Hadoop (HDFS/Hive), Snowflake, Apache Iceberg, and Sybase IQ.
- Demonstrated integrity and ethical decision‑making.
- Delivery‑focused with a strong sense of ownership.
- Strong intellectual curiosity and problem‑solving skills.
- Ability to work effectively with global teams across time zones and cultures.