Job Summary
Owns the data science lifecycle end-to-end across multi-industry engagements - presales and opportunity shaping, greenfield ML/AI platform architecture, model development and productionization, and program governance through steady-state delivery, with a roadmap toward autonomous, self-optimizing AI/ML operations.
Key Responsibilities
- Lead presales engagements - RFP/RFI response, AI/ML solution architecture, effort estimation and commercial shaping - and present win themes to executive level stakeholders.
- Own the end-to-end data science lifecycle - problem framing, exploratory data analysis, feature engineering, model development, validation and deployment.
- Architect greenfield data and ML platforms (cloud-native data lakes/lakehouses, feature stores, MLOps pipelines), including vendor and tooling selection.
- Define the roadmap toward autonomous, self-optimizing ML operations - automated retraining, drift detection and closed-loop model monitoring.
- Lead large-scale brownfield data and analytics platform modernization and legacy-to-target migrations with minimal business disruption.
- Partner with data engineering to ensure pipeline quality, governance and lineage, while personally owning model design, validation and business impact.
- Own delivery governance for multi-year, multi-workstream AI/ML transformation programs - scope, schedule, risk, quality and financials.
- Act as single technical point of accountability across design, build, deploy and operate phases, coordinating data engineering, MLOps and business teams.
- Establish and track model and program KPIs (accuracy, drift, business impact), steering committee reporting and executive dashboards.
- Mentor data science/delivery leads and build reusable accelerators, ML frameworks and playbooks across engagements.
Technical Skills Tools
- Languages ML/DL: Python, R, SQL, Scikit-learn, TensorFlow, PyTorch, XGBoost
- MLOps Deployment: MLflow, Kubeflow, Amazon SageMaker, Azure ML, Vertex AI
- Data Platforms: Databricks, Snowflake, Apache Spark, Hadoop
- Cloud: AWS, Azure, GCP
- Visualization: Power BI, Tableau
- Generative AI: LangChain, Azure OpenAI/OpenAI, RAG frameworks
- Frameworks Program Delivery: CRISP-DM, Agile/SAFe, TOGAF, MS Project, Jira, Confluence, PMP/Prince2
Primary Skills
- Data Science ML Experience in designing and deploying advanced analytics and AI solutions using traditional Machine Learning techniques including Classification, Regression, Clustering, Recommendation Systems, Anomaly Detection, Time Series Forecasting, and Reinforcement Learning.
- Statistics - Deep understanding of statistical concepts such as Probability Theory, Hypothesis Testing, Confidence Intervals, Bayesian Statistics, A/B Testing, Experimental Design, Correlation Analysis, Multivariate Statistics, Sampling Techniques, and Predictive.
- Data Modelling - Modeling. Ability to evaluate data quality, identify bias and fairness concerns, perform causal inference, and develop explainable AI solutions using industry-standard methodologies.
- Big Data - Experienced in working with large-scale datasets and Big Data technologies including Spark, Hadoop, Databricks, Kafka, and distributed computing frameworks. Proficient in Python, SQL, and modern data science ecosystems.
- Deep Learning GenAI - Deep Learning, Generative AI, Large Language Models (LLMs), Retrieval-Augmented Generation (RAG), Prompt Engineering, Vector Databases, Agentic AI frameworks, and MLOps practices for enterprise-scale AI deployments. Demonstrated ability to translate complex business problems into data-driven solutions while ensuring Responsible AI, model governance, transparency, and measurable business outcomes.