Transforma esta função numa entrevista — um currículo e uma carta de apresentação criados à volta do que este empregador procura.
Valtech is seeking a Senior Data Engineer to design, build, and optimize modern cloud-based data platforms powering analytics, AI, and data products. You’ll work across batch, streaming, and near-real-time pipelines with a strong emphasis on governance, security, and observability.
You’ll collaborate with data scientists and ML engineers to deliver production-grade data assets and enable GenAI pipelines on AWS Databricks or Azure Fabric tooling, with Snowflake/GCP as pluses.
We're the experience innovation company - a trusted partner to the world's most recognized brands. To our people we offer growth opportunities, a values -driven culture, international careers and the chance to shape the future of experience.
At Valtech, you'll find an environment designed for continuous learning, meaningful impact, and professional growth. Whether you're pioneering new digital solutions, challenging conventional thinking or building the next generation of customer experiences, your work will help transform industries.
We are looking for an experienced Senior Data Engineer to design, build, and optimize modern, cloud-based data platforms that power analytics, AI, and data products across the organization. Beyond technical delivery, we're looking for someone genuinely curious about the business problems behind the data - someone who can apply common sense thinking, detailed analysis, and experience-based recommendations to help derive and shape business requirements, not just implement them as given.
You will work on scalable batch, streaming, and near-real-time pipelines, enabling high-quality, curated datasets while ensuring robust data governance, security, and observability across the data ecosystem. You will also play a key role in supporting AI and GenAI systems, enabling pipelines for machine learning, causal modeling, and LLM-powered applications such as RAG and agent-based systems.
This role can be delivered on either of our two core cloud ecosystems - AWS (with Databricks) or Azure/Fabric (with Databricks or native Fabric tooling) - and you'll be staffed on projects that match your strongest platform. You don't need experience in both. Additional experience across Snowflake or GCP is considered a strong plus.
You will collaborate closely with data scientists, ML engineers, and platform teams to ensure the data foundation supports production-grade, decision-oriented AI systems.
Design and implement scalable data platforms and pipelines on whichever of our two core ecosystems matches your strengths - AWS (with Databricks) or Azure/Fabric (with Databricks or native Fabric tooling) - with exposure to other environments (Snowflake, GCP) considered a plus. This includes developing reliable batch, streaming, and near-real-time pipelines using technologies such as Spark and Delta Lake, and building ingestion, transformation, and curation workflows for both structured and unstructured data.
You will implement modern data architectures including lakehouse patterns and medallion layering (bronze, silver, gold) within Databricks or Fabric, ensuring systems are reusable, scalable, and aligned with enterprise needs.
Deliver high-quality datasets that support analytics, machine learning, causal modeling, and optimization systems. You will enable data pipelines for GenAI use cases (including LLMs, RAG pipelines, and vector-based data flows), as well as agent-based architectures and intelligent workflows, ensuring that data is model-ready and production-grade - leveraging Databricks' MLflow and Unity Catalog, or the equivalent Azure/Fabric and Azure ML tooling, depending on the platform you work in.
Design scalable logical and physical data models for analytical and operational use cases, ensuring consistency across domains. Orchestrate workflows using tools such as Airflow, dbt, Databricks Workflows, Azure Data Factory, or equivalents, with strong focus on automation, reliability, and maintainability of end-to-end pipelines.
Apply modern architecture patterns including event-driven and streaming architectures, and ensure adherence to best practices in data governance, lineage, quality, and access control (RBAC/ABAC), using tools such as Unity Catalog and AWS Lake Formation, or Microsoft Purview, depending on the platform in use.
Establish strong data observability, including monitoring of data freshness, pipeline reliability, and SLA adherence, ensuring systems remain trustworthy and production-ready.
Enable data serving layers (APIs, feature inputs, analytical endpoints) to support downstream systems, including ML and AI platforms. Continuously monitor and optimize pipelines and infrastructure for performance, scalability, and cost efficiency across our core cloud ecosystems.
Bring genuine curiosity to every engagement - ask the right questions to understand not just what stakeholders are asking for, but why. Apply common sense thinking, detailed an