Job Description
PRIMARY OBJECTIVES:
- Build and maintain scalable data pipelines and datasets that support analytics, reporting, and downstream business systems.
- Develop data solutions on Databricks using established engineering patterns, reusable frameworks, and enterprise standards.
- Ensure reliable, high-quality, and performant data delivery across batch and, where relevant, streaming use cases.
- Support Takeda’s data transformation journey through strong engineering practices, collaboration, and scalable platform-aligned development.
RESPONSIBILITIES:
- Design, develop, test, and maintain scalable data pipelines and integrations usingDatabricks, PySpark, and SQL.
- Build datasets optimized for analytics, BI, and downstream consumption while ensuring data quality, reconciliation, and production reliability.
- Work within establisheddata frameworks, design patterns, and reusable componentscreated by other engineering teams.
- Read, understand, troubleshoot, and extend existing codebases and pipeline logic in line with engineering standards.
- Collaborate with analytics, product, and business teams to support data models and data products for enterprise use cases.
- Contribute to unit, integration, and performance testing, documentation, and engineering best practices.
- Partner with platform, architecture, security, and DevOps teams to deploy and support pipeline solutions in cloud environments.
- Troubleshoot data and pipeline issues and drive continuous improvement in performance, scalability, and maintainability.
SCOPE OF SUPERVISION:
NUMBER SUPERVISED WORKERS
Direct
Indirect
Employees
0-3
0-3
Non-Employees
0-3
0-3
EDUCATION AND EXPERIENCE:
- Bachelor’s or Master’s degree in Computer Science, Engineering, Information Systems, or related field.
- 5+ years of experiencein data engineering, data warehousing, or large-scale data platform development.
- Strong hands-on experience withDatabricksand distributed data processing.
- Strong hands-on experience withPySparkfor pipeline development and transformation of large datasets.
- Strong hands-on experience withSQL, including joins, aggregations, optimization, and analytical data processing.
- Experience building and maintaining data pipelines for batch processing; exposure to streaming is a plus.
- Experience working with existing enterprise frameworks, shared libraries, and engineering standards.
- Experience reading, understanding, debugging, and enhancing existing code developed by other teams.
- Experience with cloud data platforms such asAWS or Azure.
- Experience working in agile, cross-functional engineering environments.
KEY SKILLS AND COMPETENCIES:
- Strong proficiency inPySpark and SQL;Python alone is not sufficientfor this role.
- Strong understanding of distributed data processing, performance optimization, and scalable pipeline design.
- Ability to work effectively within predefinedpatterns, frameworks, and architectural guardrails.
- Strong code reading and code comprehension skills across shared enterprise codebases.
- Good understanding of data modeling, schema design, and data quality controls.
- Strong engineering discipline in testing, version control, documentation, and maintainable development.
- Strong problem-solving skills and ability to troubleshoot production data issues.
- Effective communication and collaboration with technical and non-technical stakeholders.
NICE TO HAVE:
Experience with streaming technologies such asSpark Structured StreamingorKafka.
Experience with orchestration and workflow tools in enterprise data environments.
Experience with Infrastructure as Code, preferablyTerraform.
Experience designing and developing API-based integrations.
LICENSES/CERTIFICATIONS:
- Preferred - Databricks Certified Data Engineer Associate / Professional
- Preferred - AWS or Azure Data Engineering certification
PHYSICAL DEMANDS:
N/A
TRAVEL REQUIREMENTS:
Access to transportation to attend meetings.
Ability to fly to meetings regionally and globally.
Locations
IND - Bengaluru
Worker Type
Employee
Worker Sub-Type
Regular
Time Type
Full time