DataZymes Analytics Pvt. Ltd. | Full time
We empower the Pharma industry with our innovative products.
The idea of DataZymes germinated with the realization that Pharma Commercial teams had few alternatives to the antiquated and inefficient solutions offered by traditional consulting and technology companies. Already a laggard in analytical maturity, the Pharma industry had been facing challenges to adapt to a Big Data world.
We saw that the products offered by technology companies were too rigid and generic to handle novel problems. The custom solutions offered by consulting organizations took too long to deploy and required many services to maintain and improve.
We felt the need for a different approach to finding solutions and we knew it would take a different kind of company to build it. That's why DataZymes.
We're focused on creating the world's best user experience for working with data, one that empowers people to ask and answer complex questions without requiring them to master querying languages, statistical modeling, or the command line. To achieve this, we are building platforms for integrating, managing, and securing data on top of which we layer applications for fully interactive machine-driven, human-assisted analysis.
Job Description
The role will be responsible for setting up the data warehousesnecessary to handle large volumes of data, create meaningful analyses, anddeliver recommendations to leadership.
CoreResponsibilities
- Create andmaintain optimal data pipeline architecture ETL/ ELT into structured data
- Assemblelarge, complex data sets that meet business requirements and create andmaintain multi-dimensional modelling like Star Schema and SnowflakeSchema, normalisation, de-normalization, joining of datasets.
- Expert levexperience in creating a scalable data warehouse including Fact tables,Dimensional tables and ingest datasets into cloud based tools.
- Identify,design, and implement internal process improvements including automatingmanual processes, optimising data delivery and re-designing infrastructurefor greater scalability.
- Collaboratewith stakeholders to ensure seamless integration of data with internaldata marts, enhancing advanced reporting
- Setup andmaintain data ingestion, streaming, scheduling, and job monitoringautomation using AWS services. Setup Lambda, code pipeline (CI/CD), Glue,S3, Redshift needs to be maintained for uninterruptedautomation.
- Buildanalytics tools that utilize the data pipeline to provide actionableinsight into customer acquisition, operational efficiency, and other keybusiness performance metrics
- Work withstakeholders to assist with data-related technical issues and supporttheir data infrastructure needs.
- Utilize GitHubfor version control, code collaboration, and repository management.Implement best practices for code reviews, branching strategies, andcontinuous integration.
- Create datatools for analytics and data scientist team members that assist them inbuilding and optimising our product into an innovative industry leader
- Ensure dataprivacy and compliance with relevant regulations (e.g., GDPR) whenhandling customer data.
- Maintain dataquality and consistency within the application, addressing data-relatedissues as they arise.
Requirements
Required
- 4-8 years ofrelevant experience
- Advancedworking SQL knowledge and experience working with relational databases,query authoring (SQL) as well as working familiarity with a variety ofdatabases and Cloud Data warehouse like AWS Redshift
- Experience increating scalable, efficient schema designs to support diverse businessneeds.
- Experiencewith database normalization, schema evolution, and maintaining dataintegrity
- Proactivelyshare best practices, contributing to team knowledge and improving schemadesign transitions.
- Develop datamodels, create dimensions and facts, and establish views and procedures toenable automation programmability.
- Collaborateeffectively with cross-functional teams to gather requirements,incorporate feedback, and align analytical work with business objectives
- Datacompression into PARQUET to improve processing and finetuning SQLprogramming skills.
- Experiencebuilding and optimizing “big data” data pipelines, architectures and datasets.
- Experienceperforming root cause analysis on internal and external data and processesto answer specific business questions and identify opportunities forimprovement.
- Experiencewith manipulating, processing and extracting value from large disconnectedunrelated datasets
- Stronganalytic skills related to working with structured and unstructureddatasets.
- Workingknowledge of message queuing, stream processing, and highly scalable “bigdata” stores.
- Experiencesupporting and working with cross-functional teams and Global IT.
- Familiarity ofworking in an agile based working models.
PreferredQualifications/Expertise
- Experiencewith relational SQL and NoSQL databases, especially AWS Redshift.
- Experiencewith AWS cloud services Preferable: S3, EC2, Lambda, Glue, EMR, Codepipeline highly preferred. Experience with similar services on anotherplatform would also be considered.
Education:
- Bachelor’s ormaster’s degree on Technology and Computer Science background