Location : Chicago, IL (Work from Office)
- Candidates must work from the Chicago office.
- Hybrid candidates must be available for regular onsite collaboration at least 3 days per week.
- Preference will be given to candidates located in or near Chicago.
Job Summary
We are seeking a highly skilled Data Architect with strong expertise in Azure Databricks, Delta Lake, Spark/PySpark, and Modern Data Lakehouse Architectures. The ideal candidate will be responsible for designing scalable enterprise data platforms, defining data architecture standards, implementing Medallion Architecture, and leading large-scale data modernization initiatives on Azure.
Key Responsibilities
- Design and implement enterprise-scale Data Lakehouse solutions using Azure Databricks and Delta Lake.
- Define and manage Bronze, Silver, and Gold layers following Medallion Architecture best practices.
- Design Delta Lake tables, partitioning strategies, data models, and storage optimization techniques.
- Architect scalable ETL/ELT pipelines using Spark and PySpark.
- Establish data ingestion frameworks for batch and near real-time data processing.
- Implement schema evolution, schema enforcement, and data quality validation mechanisms.
- Define data governance, security, lineage, metadata, and access control standards.
- Collaborate with business stakeholders, data engineers, analysts, and product teams to translate requirements into scalable solutions.
- Perform performance tuning using partitioning, caching, optimization, and workload management techniques.
- Lead architecture reviews, design discussions, and technical governance processes.
- Mentor engineering teams on Databricks and cloud data engineering best practices.
Required Skills
- 8+ years of experience in Data Engineering/Data Architecture.
- Strong hands-on experience with:
- Spark / PySpark
- SQL
- Expertise in designing Delta Tables and Lakehouse architecture.
- Experience implementing Medallion Architecture.
- Strong understanding of schema evolution, data quality frameworks, and metadata management.
- Experience with distributed data processing and large-scale datasets.
- Strong knowledge of CI/CD, Git, DevOps, and deployment automation.
- Excellent communication and stakeholder management skills.
Preferred Skills
- Unity Catalog
- Structured Streaming
- Kafka/Event Hub
- MLflow
- Terraform or Infrastructure as Code
Additional Hiring Criteria
- Strong communication and client-facing skills.
- Ability to work independently in a fast-paced environment.
- Willingness to work onsite in Chicago at least 3 days per week.
- Candidates seeking fully remote opportunities will not be considered.
- Local Chicago candidates are highly preferred. Non-local candidates must demonstrate commitment to the onsite work requirement.