Associate Data Engineer

Biopharma Careers

Hyderabad

On-site

INR 1,200,000 - 2,200,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Amgen, a leading biotechnology company, invites applications for a data engineering role in Hyderabad. You will design, build, test, and optimize data pipelines using Databricks, PySpark, and Python to support analytics and ML initiatives.

You will ingest and transform data from multiple sources, validate quality, and collaborate with engineers, analysts, and data scientists in an Agile setting. Strong communication and problem-solving are essential.

Qualifications

  • Must have a Bachelor’s degree in a relevant field or equivalent practical experience.
  • Hands-on Python for data processing, scripting, and automation.
  • Strong PySpark and distributed data processing knowledge.
  • Proven Databricks experience including notebooks, clusters, jobs, and performance tuning.
  • Ability to build and troubleshoot scalable ETL/ELT pipelines in Databricks.
  • Experience with Delta Lake and lakehouse concepts.
  • SQL proficiency for data querying and validation.
  • Familiarity with CSV/JSON/Parquet/Delta data formats.
  • Understanding of ETL/ELT, data pipelines, data lakes, and data warehouses.
  • Basic AI/ML concepts and feature engineering familiarity.
  • Exposure to cloud data platforms (AWS/Azure/GCP).
  • Git version control and collaborative development.
  • Strong problem solving and communication skills.

Responsibilities

  • Develop, test, and maintain data pipelines using Databricks, PySpark, and Python.
  • Ingest, transform, and process structured and semi-structured data from multiple sources.
  • Support scalable ETL/ELT workflows for analytics, reporting, and ML use cases.
  • Collaborate with data engineers, analysts, and data scientists to define data needs.
  • Perform data cleansing, validation, and quality checks for accuracy.
  • Optimize Spark jobs and Databricks notebooks for performance and cost.
  • Create and maintain documentation for pipelines and data definitions.
  • Troubleshoot pipeline failures, data issues, and performance bottlenecks.
  • Follow code quality, testing, and deployment best practices.
  • Support basic AI/ML data prep, feature engineering, and model inputs.
  • Monitor scheduled jobs to ensure timely data delivery.
  • Work in an Agile or iterative development environment.

Skills

Python
PySpark
Databricks
SQL
Git
Cloud platforms
ETL/ELT pipelines
Agile
Communication
Troubleshooting

Education

Bachelor’s degree in Computer Science, Data Engineering, Information Systems, Engineering, Mathematics, or equivalent

Tools

Databricks notebooks
Delta Lake
MLflow
Tableau/Power BI

Job description

Career Category

Engineering

Job Description
ABOUT AMGEN

Amgen harnesses the best of biology and technology to fight the world’s toughest diseases, and make people’s lives easier, fuller and longer. We discover, develop, manufacture and deliver innovative medicines to help millions of patients. Amgen helped establish the biotechnology industry more than 40 years ago and remains on the cutting-edge of innovation, using technology and human genetic data to push beyond what’s known today.

Roles & Responsibilities
  • Develop, test, and maintain data pipelines using Databricks, PySpark, and Python.
  • Ingest, transform, and process structured and semi-structured data from multiple sources.
  • Support the development of scalable ETL/ELT workflows for analytics, reporting, and machine learning use cases.
  • Work with data engineers, analysts, and data scientists to understand data requirements and deliver reliable datasets.
  • Perform data cleansing, validation, and quality checks to ensure accuracy and consistency.
  • Optimize Spark jobs and Databricks notebooks for performance, reliability, and cost efficiency.
  • Create and maintain documentation for data pipelines, workflows, data definitions, and processes.
  • Assist in troubleshooting pipeline failures, data issues, and performance bottlenecks.
  • Follow best practices for version control, code quality, testing, and deployment.
  • Support basic AI/ML data preparation activities, including feature engineering, dataset creation, and model input preparation.
  • Monitor scheduled jobs and workflows to ensure timely and successful data delivery.
  • Collaborate with cross-functional teams in an Agile or iterative development environment.
Basic Qualifications and Experience

2-6 years of experience with Bachelor’s degree in Computer Science, Data Engineering, Information Systems, Engineering, Mathematics, or a related field, or equivalent practical experience

Must-Have Qualifications
  • Bachelor’s degree in Computer Science, Data Engineering, Information Systems, Engineering, Mathematics, or a related field, or equivalent practical experience.
  • Hands-on experience with Python for data processing, scripting, and automation.
  • Strong working knowledge of PySpark and distributed data processing concepts.
  • Proven hands-on experience using Databricks for data engineering, including notebooks, clusters, jobs, workflows, Delta tables, and performance optimization.
  • Ability to build, maintain, and troubleshoot scalable ETL/ELT pipelines in Databricks.
  • Experience working with Delta Lake and lakehouse architecture concepts.
  • Working knowledge of SQL for querying, transforming, and validating data.
  • Ability to work with structured and semi-structured data formats such as CSV, JSON, Parquet, and Delta.
  • Understanding of data engineering concepts such as ETL/ELT, data pipelines, data lakes, data warehouses, batch processing, and data quality.
  • Basic understanding of AI and machine learning concepts, including features, training datasets, model inputs/outputs, and model evaluation basics.
  • Experience supporting data preparation or feature engineering for AI/ML use cases.
  • Familiarity with cloud-based data platforms, preferably AWS, Azure, or GCP.
  • Understanding of Git or other version control tools.
  • Strong analytical, problem-solving, and troubleshooting skills.
  • Good communication skills and ability to work collaboratively with technical and non-technical stakeholders.
  • Willingness to learn new tools, technologies, and data engineering best practices.
Preferred Qualifications
  • Exposure to Delta Lake, Unity Catalog, or Lakehouse architecture.
  • Experience with workflow orchestration tools or Databricks Jobs.
  • Familiarity with CI/CD practices for data engineering projects.
  • Exposure to machine learning workflows using MLflow, scikit-learn, or similar tools.
  • Experience with Tableau, Power BI, or similar data visualization tools to create dashboards, support reporting needs, validate datasets, and perform exploratory analysis.
  • Understanding of data governance, security, and access control concepts.
  • Experience working in an Agile/Scrum environment.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Associate Data Engineer
Associate Data Engineer

Amgen SA • Hyderabad

On-site
INR 1,200,000 - 1,800,000
Associate Data Engineer
Associate Data Engineer

Amgen Technology Private Limited • Hyderabad

On-site
INR 900,000 - 1,400,000
Data Engineer
Data Engineer

Amgen Technology Private Limited • Hyderabad

On-site
INR 1,800,000 - 2,400,000
Associcate Data Engineer
Associcate Data Engineer

Amgen • Hyderabad

On-site
INR 1,400,000 - 2,100,000
Sr Data Engineer
Sr Data Engineer

Biopharma Careers • Hyderabad

On-site
INR 1,800,000 - 2,600,000
Associcate Data Engineer
Associcate Data Engineer

Biopharma Careers • Hyderabad

On-site
INR 3,000,000 - 5,500,000
Data Engineer
Data Engineer

Amgen • Hyderabad

On-site
INR 800,000 - 1,200,000
Sr Associate IS Engineer, Commercialization Technology
Sr Associate IS Engineer, Commercialization Technology

Biopharma Careers • Hyderabad

On-site
INR 1,200,000 - 1,800,000
Senior Associate Data Engineer
Senior Associate Data Engineer

Amgen Technology Private Limited • Hyderabad

On-site
INR 1,500,000 - 2,500,000
Senior Data Engineer
Senior Data Engineer

GlobalNodes • Gurgaon

On-site
INR 1,500,000 - 2,100,000