Thinkproject builds construction intelligence software for firms delivering Europe’s largest infrastructure, energy, and real‑estate projects. The platform manages the information flow across the full lifecycle of a built asset – from design and construction through operation and eventual decommissioning – and empowers customers to digitise, connect, and control their construction workflows.
About The Role
We are looking for a Data Integration Engineer to own the data pipelines and integration layer that powers our AI Search Platform. The role involves designing, building, and maintaining secure, reliable data workflows that move data from source systems into Google Cloud Platform services including Vertex AI, and providing well‑designed APIs for internal and external consumption. This hands‑on engineering position requires writing production code, owning reliability, and collaborating closely with DevOps, Network, and Platform Engineering teams.
Key Responsibilities
Data Integration & Pipeline Development
- Design, implement, and optimise scalable data integration workflows supporting inference and data synchronisation across GCP services (Cloud Run, Pub/Sub, Cloud Storage, Cloud Spanner, Vertex AI)
- Build and maintain event‑driven pipelines and ETL/ELT workflows that deliver clean, reliable data to the AI Search Platform
- Automate deployment, testing, and pipeline orchestration using Cloud Run, Pub/Sub triggers, and Terraform
API Development for AI Integration
- Build and maintain APIs that expose data integration and AI inference capabilities to internal and external systems
- Ensure secure, reliable, and performant access to the AI Search Platform – correct authentication, rate limiting, and error handling by default
Permissions & Compliance Layer
- Integrate and enforce API and IAM policies for compliant access control across all AI Search Platform components
- Own and evolve the permissions API layer to meet growing scalability and security requirements
Data Quality & Reliability
- Ensure data integrity through monitoring, validation, and alerting across all integrated systems and services
- Continuously monitor workflows for latency, reliability, and cost efficiency – implement improvements proactively
Documentation & Standards
- Maintain architecture documentation and runbooks
- Contribute to best practices for data integration, reproducibility, scalability, and security
What You Need to Fulfil the Role
Required Skills & Qualifications
- 8+ years of professional experience in data engineering, cloud integrations, or backend development
- Strong proficiency in Python and SQL
- Production experience with Google Cloud Platform services: Cloud Run, Pub/Sub, Cloud Storage, Cloud Spanner, and Vertex AI
- Experience with event‑driven architectures and cloud‑based ETL/ELT workflows
- Experience with relational databases (PostgreSQL, Cloud Spanner) and exposure to NoSQL
- Proficient with Git and familiar with CI/CD workflows and containerisation (Docker)
- Experience with Terraform or equivalent Infrastructure‑as‑Code tooling
- Working knowledge of IAM, data governance, and access management principles
Nice‑to‑Have (Bonus Skills)
- Azure DevOps or cross‑cloud integration experience
- API design experience (REST or gRPC)
- Experience with AI/ML inference pipelines or Vertex AI in production
- Prior work in construction, engineering, or real‑estate software domains
Soft Skills
- Engineering rigour – care about pipeline reliability and data correctness, not just throughput
- Ownership mindset – monitor what you build and rectify issues promptly
- Clear written communication – document integration contracts and architecture decisions for non‑specialists
- Collaborative – work smoothly with DevOps, Network, and Platform teams
- Comfortable with ambiguity – scope and deliver integration work from incomplete upstream specifications
What Success Looks Like
- Month 3: Core data pipelines understood and contributing to production; first reliability or latency improvement shipped
- Month 6: Own at least one integration area end‑to‑end; permissions API layer extended with evidence‑backed design decisions
- Month 12: Data integration reliability measurably improved; pipeline documentation and monitoring coverage complete; at least one material cost or latency inefficiency identified and closed
You're Probably NOT a Fit If
- Your data engineering experience is primarily batch ETL without event‑driven or streaming context
- You are not comfortable working across cloud‑native GCP services in production
- You treat IAM and access control as someone else’s concern
- You need fully defined requirements before designing an integration
Compensation (Pune, Mid‑Senior)
- Competitive fixed salary – disclosed on request
- Variable performance bonus: 5% of fixed salary
- Continuous learning & certification budget; learning programmes; career growth; international exposure