Una candidatura hecha para este puesto de trabajo — un currículum y una carta de presentación adaptados que responden directamente a la oferta.
Insud Pharma invites applications for a Data Engineer MLE to join AI Labs in Madrid. You will design and deploy scalable data pipelines powering analytics and machine learning, collaborating with data scientists to productionize models.
The role emphasizes end-to-end ML workflow implementation, data quality, and scalable infrastructure. Fluency in English and Spanish is required, with a focus on robust, maintainable code.
Want to know more? INSUD PHARMA operates across the entire pharmaceutical value chain, providing specialized knowledge and experience in scientific research, development, manufacturing, sales, and marketing of a wide range of active pharmaceutical ingredients (API), finished dosage forms (FDF), and branded pharmaceutical products, adding value to human and animal health. The activities of INSUD PHARMA are organized into three synergistic business areas: Industrial (Chemo), Branded (Exeltis), and Biotech (mAbxience), with over 9,000 professionals in more than 50 countries, 20 state-of-the-art facilities, 15 specialized R&D centers, 12 commercial offices, and more than 35 pharmaceutical subsidiaries, serving 1,150 customers in 96 countries worldwide. INSUD PHARMA believes in innovation and sustainable development. Ready to be a #Challenger? What are we looking for? We're AI Labs — the applied AI team at Insud Pharma. 30 people. AI Engineers, Data Scientists, DevOps Engineers, Product Managers building the systems that power how trials get designed, how patients get recruited, and how everything gets monitored once the trial is live. Clinical trials run on data. Bad pipelines, slow models, and infrastructure that breaks under pressure can cost months — or worse, the trial itself.
AI Labs operates with a startup mindset within Insud Pharma. The department is young, and the culture reflects that: flat, collaborative, and fast-moving . You will work alongside Data Scientists, AI Engineers, DevOps Engineers, and Product Managers who are equally committed to delivering high-quality work. We hold regular demo days where teams present their work, as well as whiteboard sessions where we tackle problems together. The cross-disciplinary dynamic is genuinely strong. The office is located in central Madrid (Chamberí, near Eloy Gonzalo), well connected and situated in a vibrant part of the city.
Design, build, and maintain scalable data pipelines for data ingestion, transformation, and serving, supporting both analytics and machine learning use cases. Develop and productionize machine learning pipelines , covering training, validation, deployment, and monitoring. Collaborate closely with Data Scientists to translate notebooks and prototypes into robust, production-ready ML systems . Implement model deployment patterns (batch, real-time, or hybrid) using APIs, scheduled jobs, or event-driven architectures. Build and maintain feature pipelines and data abstractions that enable reproducible and reliable model behavior. Ensure data quality, versioning, and traceability across datasets and models. Optimize pipelines and ML workloads for performance, scalability, and cost efficiency. Work with DevOps and Platform teams to deploy solutions using containerization and CI/CD best practices. Contribute to defining data engineering and MLOps standards across AI Labs. Participate in code reviews, documentation, and mentoring to foster a culture of engineering excellence.
Proficient in Spanish and English , written and verbal communication. Strong proficiency in Python , including clean code practices, packaging, and modular design. Solid understanding of software engineering principles (OOP, SOLID, testing, version control). Hands-on experience building data pipelines (ETL / ELT) using Python-based frameworks or custom solutions. Experience working with machine learning workflows , including model training, evaluation, and deployment. Familiarity with REST APIs and service-based architectures (FastAPI, Flask, or similar). Strong experience with Git and collaborative development workflows. Experience with containerization (Docker) and cloud environments (AWS or Azure). Experience with MLOps practices (model versioning, monitoring, drift detection, retraining strategies). Familiarity with orchestration tools (e.G., Airflow, Prefect, Dagster). Experience with data storage systems (SQL / NoSQL databases, data lakes, object storage). Exposure to streaming or event-driven architectures. Experience deploying or operating ML systems in regulated or high-reliability environments. Familiarity with ML frameworks and scientific libraries (NumPy, Pandas, Scikit-learn, PyTorch, TensorFlow). Interest in applied AI topics such as NLP, LLM-based systems, or scientific computing.