Job Overview
Job Title: Data Scientist & Machine Learning Engineer
Experience: 4 to 15 Years
Location: Madurai, Tamil Nadu, India
Shift: 2:00 PM – 11:30 PM IST
Working Mode: Remote/Hybrid
Job Type: Full-Time
Key Skills: Machine Learning Engineer, deep learning, insurance, Scikit learn, PyTorch, TensorFlow, Python, MLOps, CI/CD, XGBoost, Retrieval Augmented Generation, Spark, Airflow, Azure, AWS, FastAPI, Docker, version control, claims, LangChain, LlamaIndex, Semantic Kernel, enterprise AI solution development
Key Responsibilities
- Design, develop, and deploy end-to-end machine learning and deep learning models for real-world business problems.
- Build scalable solutions for large-volume data processing (structured & unstructured).
- Develop and optimize Generative AI applications using LLMs (e.g., RAG pipelines, copilots, summarization, Q&A systems).
- Implement predictive analytics models such as classification, regression, clustering, and anomaly detection.
- Work on insurance-focused use cases, including:
- Claims anomaly/fraud detection
- Risk scoring and underwriting support
- Document processing (OCR + NLP pipelines)
- Build and maintain data pipelines and feature engineering workflows
- Fine-tune and evaluate LLMs and embedding models for domain-specific use cases
- Ensure model performance, scalability, and monitoring in production environments
- Collaborate with cross-functional teams (product, data engineering, business stakeholders)
- Maintain best practices in MLOps, model versioning, and CI/CD pipelines
Required Skills & Qualifications
- Strong foundation in machine learning algorithms (supervised & unsupervised).
- Experience with anomaly detection, time-series, and predictive modeling.
- Proficiency in Python and ML libraries (Scikit-learn, XGBoost, PyTorch/TensorFlow).
- Experience with data preprocessing, feature engineering, and model evaluation.
- Hands‑on experience with LLMs (OpenAI, Claude, Hugging Face, etc.).
- Strong understanding of RAG (Retrieval Augmented Generation), prompt engineering & evaluation, embeddings & vector databases (e.g., FAISS, Milvus).
- Experience building GenAI applications (chatbots, document, summarization systems).
- Experience working with large datasets (batch + streaming) and data pipelines and tools (Spark, Airflow, or similar).
- Familiarity with cloud platforms (Azure, AWS, or GCP).
- Experience deploying models using APIs (FastAPI/Flask).
- Understanding of Docker, CI/CD pipelines, and model monitoring.
- Knowledge of version control and experiment tracking tools.
Preferred Qualifications
- Experience in the Insurance domain (claims processing, fraud detection, underwriting analytics).
- Familiarity with document AI / OCR / NLP pipelines for insurance workflows.
- Experience with graph-based or network-based anomaly detection.
- Exposure to multi-agent systems or AI orchestration frameworks.
- Understanding of regulatory and compliance considerations in insurance AI.
Key Use Cases You Will Work On
- AI-powered claims anomaly & fraud detection systems.
- Intelligent document processing and insights extraction.
- Generative AI-based assistants for claims and underwriting teams.
- Predictive models for risk assessment and customer insights.
Soft Skills
- Strong problem-solving and analytical thinking.
- Ability to work in a fast-paced, ambiguous environment.
- Effective communication with both technical and business stakeholders.
- Ownership mindset with a focus on delivering production-ready solutions.
Nice to Have
- Experience with LangChain / LlamaIndex / Semantic Kernel.
- Knowledge of knowledge graphs and hybrid search systems.
- Prior experience in enterprise AI solution development.