Description:
Hybrid At least 2 days per week in office in Chicago, IL
Our client seeks a Senior Data Engineer to design, build, and operate large-scale data processing for attribution, measurement, forecasting, and privacy-preserving analytics. You will develop Scala and Spark solutions on cloud platforms, implement governance and privacy controls, and partner with cross-functional teams to deliver secure, reliable data products.
We can facilitate w2 and corp-to-corp consultants. For our w2 consultants, we offer a great benefits package that includes Medical, Dental, and Vision benefits, 401k with company matching, and life insurance.
Rate: $70.00 to $80.00/hr. w2
Responsibilities:
- Develop and optimize large-scale data processing solutions using Scala, Spark, and SQL on modern data platforms.
- Build and operate trusted data pipelines across secure cloud environments such as AWS, GCP, or Azure.
- Partner with Product, Data Science, Security, Privacy, and Platform Engineering to deliver privacy-preserving features.
- Build, schedule, and maintain scalable batch and streaming data workflows with orchestration frameworks.
- Implement data classification, access controls, and privacy-preserving techniques aligned to compliance requirements.
- Contribute to clean-room and trusted data-sharing environments with approved aggregated outputs.
- Create observability, monitoring, and operational tooling for reliability and compliance.
- Troubleshoot complex performance and pipeline issues across distributed systems.
- Contribute to technical design, best practices, and operational excellence.
- Mentor junior engineers and perform thorough code reviews.
- Continuously improve attribution, measurement, forecasting, and privacy-preserving analytics capabilities.
- Operate pipelines within trusted environments, clean rooms, or secure data-sharing platforms across cloud and on-premises.
- Apply access controls, data classification, lineage, and governance for PII, PCI, and confidential signals.
- Follow data handling standards with Security, Privacy, and Compliance teams to keep sensitive data within trust boundaries.
- Enforce aggregation, anonymization, tokenization, and approved outputs for data leaving trusted environments.
- Build monitoring and alerting to detect anomalous data movement and policy violations.
- Apply privacy-preserving computation when outputs cross trust boundaries, including aggregation-before-export, pseudonymization, tokenization, differential privacy concepts, and privacy-aware reporting.
- Implement encryption, key management, and secure handling with cloud-native security services.
- Document trust boundaries, data contracts, lineage, and permitted data movement.
- Support audits, compliance requirements, governance reviews, and secure data-sharing initiatives.
- Participate in architecture and design reviews to embed governance, privacy, lineage, and trust-boundary requirements.
- Contribute to engineering standards for secure data processing and trusted platform operations.
Experience Requirements:
- 5+ years of data engineering with strong Scala and Apache Spark on AWS and/or GCP.
- Strong Python for pipelines, tooling, automation, and infrastructure modules.
- Advanced SQL across RDBMS, cloud data warehouses, and lakehouse platforms with TB-scale datasets.
- Designing and maintaining batch and streaming data pipelines.
- Data warehousing, dimensional modeling, data quality, partitioning, and performance optimization.
- Distributed processing and modern lakehouse architectures such as Databricks, Delta Lake, or Apache Spark.
- Operating distributed data platforms at scale.
- Workflow orchestration with Airflow, Databricks Workflows, AWS Step Functions, or equivalent.
- Source control with Git and test automation frameworks.
- Cloud-native development on AWS and/or GCP.
- Software engineering practices including CI/CD, code reviews, observability, and production support.
- Ownership of features and pipelines with cross-team collaboration and mentoring.
- Trusted environment execution with clean rooms or secure data-sharing platforms handling PII and regulated data.
- Fine‑grained access controls, governance policies, and policy‑based enforcement for sensitive datasets.
- Privacy‑preserving techniques such as tokenization, pseudonymization, aggregation‑before‑export, and differential privacy concepts.
- Experience with clean‑room, measurement, attribution, audience analytics, or privacy‑preserving reporting solutions.
- Understanding of trust boundaries, secure data‑sharing patterns, and zero‑trust principles.
- Encryption, key management, and secure handling of sensitive data with cloud‑native services.
- Observability and alerting to detect anomalous data movement and potential leakage events.
- Strong understanding of cloud‑native security and governance.
- Good to have: Databricks, AWS Clean Rooms, advertising measurement platforms, collaboration with Security/Privacy/Risk/Compliance, ELK/Grafana/OpenTelemetry, Docker and Kubernetes, lineage and governance tooling, and documenting data contracts and flows.
- Strong written and verbal English communication skills.
- Experience with Agile or SCRUM in cross‑functional product teams.