In this role, you will collaborate closely with one of our esteemed clients—a globally recognized leader in their industry, distinguished by their commitment to excellence, innovation, and delivering exceptional value. As a trusted IT consulting partner, Dautom is supporting their strategic initiatives by connecting them with exceptional talent to drive business growth and transformation.
About the Role :
Job Type: 1year + extendable long term contract
Remote from India
We are looking for a Data Scientist and ML Engineer who can build machine learning models and run them reliably in production. You will own the full model lifecycle: problem framing, feature engineering, model development, validation, deployment, and monitoring. You will also build the data pipelines that feed your models, so strong data engineering fundamentals are essential.
Key Responsibilities :
- Frame business problems as ML problems, define success metrics, and agree on evaluation criteria with stakeholders
- Develop, tune, and validate models for forecasting, classification, regression, segmentation, anomaly detection, and recommendation
- Engineer robust features from large structured and time-series datasets, and publish them as reusable, governed feature tables
- Deploy models to production as real-time endpoints and scheduled batch inference jobs
- Build end-to-end MLOps workflows: experiment tracking, model registry, versioning, CI/CD, and champion-challenger promotion
- Monitor production models for data drift, prediction drift, and performance decay, and own retraining strategies
- Build and maintain scalable data pipelines using medallion architecture principles, with data quality checks at every layer
- Explain model behavior using interpretability techniques, and communicate results clearly to non-technical audiences
- Apply statistical rigor through hypothesis testing, experiment design, and uncertainty quantification
- Contribute to team standards for ML code quality, reproducibility, documentation, and model governance
Technical Skills :
- Algorithms: Strong command of supervised and unsupervised learning, gradient boosting, time-series foecasting, clustering, and ensemble methods
- Frameworks: scikit-learn, XGBoost or LightGBM, statsmodels, and PyTorch or TensorFlow
- Model quality: Cross-validation strategies (including time-based splits), hyperparameter optimization (Optuna or similar), and explainability (SHAP)
- Statistics: Solid foundation in probability, inference, experimental design, and causal reasoning
- Model Deployment and MLOps (Core)
- Lifecycle: MLflow for experiment tracking, model registry, and model packaging
- Serving: Real-time model serving endpoints and distributed batch inference at scale
- Operations: Model and data monitoring, drift detection, automated retraining, and alerting
- Engineering: Production-quality Python, unit and integration testing, Git, Docker, and CI/CD pipelines
- Processing: PySpark and advanced SQL on large datasets
- Lakehouse: Delta Lake, Lakeflow Declarative Pipelines (DLT) or equivalent, and orchestration through Databricks Workflows or Azure Data Factory
- Governance: Unity Catalog or equivalent for access control, lineage, and data discovery
Preferred
- Databricks ML stack: Feature Engineering in Unity Catalog, Model Serving, and Lakehouse Monitoring
- Mathematical optimization with OR-Tools, PuLP, or Gurobi
- Marketing analytics: attribution, media mix modeling, customer lifetime value, or churn modeling
- Familiarity with GenAI and LLM integration in ML workflows
- Experience with ERP or CRM source data (for example, SAP)
Qualifications :
- 4 to 10 years of experience in data science or ML engineering, with at least 2 years deploying and maintaining models in production.
- A demonstrable track record of models that delivered measurable business impact, not only offline accuracy gains.
- Bachelor's or Master's degree in Computer Science, Statistics, Mathematics, Engineering, or a related quantitative field.
- Ability to own problems end to end, from raw data to a monitored production model.
- Clear communication of technical findings to business stakeholders.
- Relevant certifications (Databricks Machine Learning Professional, Azure Data Scientist Associate) are a plus