Context and Mission of the Position
We are expanding our data engineering structure and are seeking a professional with a technical profile, focused on advanced data engineering, solid foundations in DevOps, and knowledge of MLOps. This individual will be responsible for designing, optimizing, and maintaining the infrastructure and pipelines that support local data projects, data science projects, and the analytical operations of the organization, serving as a critical bridge between model development and large‑scale production.
Main Responsibilities and Activities
- Develop, monitor, and automate robust data ingestion pipelines (batch and streaming) that handle complex scenarios such as advanced integration of third‑party APIs and web scraping of application APIs.
- Design and organize data storage in Google BigQuery using best modeling practices, ensuring correct transition and logical segregation of environments from raw data reception (data swamp/data lake) to optimized availability layers (publish layers).
- Support the Data Science stack by creating the technical bases necessary for autonomy, local testing, product packaging, deployment, and continuous monitoring of predictive models in production.
- Ensure code lifecycle and implementations are automated through CI/CD.
- Contribute to highly scalable, secure, and well‑documented solutions; promote and structure modern repositories in a monorepo environment, facilitating collaboration and independence of data science teams.
Technical Profile and Candidate Requirements
Mandatory Requirements (Hard Skills)
- Consolidated experience in data engineering or software engineering focused on large‑scale data systems.
- Practical experience in the cloud ecosystem GCP (Google Cloud Platform), with a specialized focus on BigQuery.
- Experience with containerization and orchestration tools: Docker and Kubernetes.
- Proficiency in designing automated CI/CD pipelines, preferably using GitHub Actions or equivalent tools.
- Advanced knowledge in organizing and managing code in monorepo architectures.
- Strong understanding of modern data architectures (lakehouse, data lake) and their data layer conventions.
- Proficiency in English at minimum intermediate–advanced level (B2 or C1).
Valued Requirements (Differentials)
- Direct practical experience in MLOps concepts and tools (model lifecycle management, model serving APIs, feature stores).
- Experience deploying solutions in orchestrated clusters (Kubernetes) when architecture requires.
- Experience with observability using logging tools such as Datadog or Grafana.
Key Technologies / Concepts
Cloud & Storage
Google Cloud Platform (GCP), Google BigQuery
Infrastructure & DevOps
Kubernetes, Docker, Virtual Machines
CI/CD & Repositories
GitHub Actions, Monorepo
ML Methodologies
MLOps (operationalization and deployment of data science models)
Data Architecture
Data swamp, data lake, publish layers, data warehouse conventions
Languages
English (B2 or C1)