Technical Lead, AI/ML

Rakuten Kobo Inc.

Bengaluru

Hybrid

INR 5,500,000 - 7,500,000

Full time

4 days ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Rakuten Kobo Inc. in Bengaluru invites a Technical Lead to shape AI/ML strategies and production-ready software for telecom OSS. You will lead ML modeling and AI harness development, collaborating with data pipelines to deploy reliable, observable models at scale.

The role requires 6+ years in ML/AI engineering, strong Python and SQL skills, and experience with PyTorch, Kafka, ClickHouse, and Kubernetes. You will own end-to-end delivery and on-call reliability in a hybrid environment.

Qualifications

  • 6+ years of experience in ML/AI engineering, including production deployment of models on live, high-volume data.
  • Hands-on experience developing time-series ML models (anomaly detection, forecasting or classification) - deep learning and classical - with real operational impact.
  • Experience building ML infrastructure: training pipelines, evaluation frameworks, feature handling, model serving and monitoring (frameworks such as PyTorch, scikit-learn, XGBoost/LightGBM; MLOps tooling such as MLflow, Kubeflow or equivalent).
  • Strong Python engineering; comfortable with large datasets via SQL (ClickHouse/YugabyteDB/PostgreSQL experience a plus).
  • Solid software engineering fundamentals - version control, testing, CI/CD, containerization, Kubernetes.

Responsibilities

  • ML modeling: develop anomaly detection models over network KPIs and counters to surface degraded cells and network elements before outages.
  • Predictive fault & incident analysis: build forecasting models and correlate alarms for root-cause analysis.
  • AI harness: train/evaluate/inference/monitoring infrastructure to productionize models.
  • Serving/inference pipelines: deploy low-latency inference services with canary/rollback; batch and streaming inference.
  • Monitoring & lifecycle: model observability, drift detection, retraining loops.
  • LLM & agentic AI: guard-railed LLM-based capabilities for RAG, intent interfaces and root-cause analysis.
  • Collaboration with data pipelines: feature/training-data contracts; governed interfaces.
  • On-call for AI services; reliability of harness in production.

Skills

ML/AI engineering
Time-series modeling
Python programming
SQL proficiency
MLOps
Model deployment
Kubernetes
CI/CD
Containerization
PyTorch
scikit-learn

Education

Bachelor's or Master's in CS/Engineering/Statistics or related field

Tools

PyTorch
scikit-learn
XGBoost
LightGBM
MLflow
Kubeflow
Kafka
Flink
Spark
ClickHouse
YugabyteDB
PostgreSQL
Kubernetes

Job description

**Job Title: Technical Lead, AI/ML \\_RIO(OSS)****Location: Bangalore (Hybrid)****Why should you choose us?**Rakuten Symphony is reimagining telecom, changing supply chain norms and disrupting outmoded thinking that threatens the industry’s pursuit of rapid innovation and growth. Based on proven modern infrastructure practices, its open interface platforms make it possible to launch and operate advanced mobile services in a fraction of the time and cost of conventional approaches, with no compromise to network quality or security.Rakuten Symphony has operations in Japan, the United States, Singapore, India, South Korea, Europe, and the Middle East Africa region. For more information, visit: https://symphony.rakuten.com.Building on the technology Rakuten used to launch Japan’s newest mobile network, we are taking our mobile offering global.To support our ambitions to provide an innovative cloud-native telco platform for our customers, Rakuten Symphony is looking to recruit and develop top talent from around the globe. We are looking for individuals to join our team across all functional areas of our business – from sales to engineering, support functions to product development. Let’s build the future of mobile telecommunications together! About Rakuten Group, Inc. (TSE: 4755) is a global leader in internet services that empower individuals, communities, businesses and society. Founded in Tokyo in 1997 as an online marketplace, Rakuten has expanded to offer services in e-commerce, fintech, digital content and communications to 2 billion members around the world. The Rakuten Group has over 30,000 employees, and operations in 30 countries and regions. For more information visit https://global.rakuten.com/corp/.**About the RIO Team**RIO (Rakuten Intelligent Operations) is Rakuten Symphony's AI-first operational intelligence platform and the OSS engineering team behind it. RIO replaces fragmented OSS tooling with a single, intelligent control layer for complex, multi-vendor telecom networks: unified real-time observability across RAN, Core, Transport and Cloud; AI-driven service assurance with anomaly detection, predictive analytics and proactive fault resolution; and closed-loop, intent-based automation.**What Do We Expect From You**RIO's promise is AI-native network operations. As a Senior AI Engineer on the AI Model & Harness team, you will span two equally important halves: developing the ML models that detect anomalies, predict faults and localize root causes across telecom networks; and building the AI harness - the training, evaluation, inference and monitoring infrastructure - that takes those models from notebook to production with confidence. You will work directly on top of RIO's ClickHouse and YugabyteDB (YBSQL) data platforms, in tight partnership with the Data Management Pipelines team.**Key Responsibilities****ML Modeling (approximately 50%)*** **Anomaly detection:** Develop and productionize multivariate time-series anomaly detection models over network KPIs and counters to surface degraded cells and network elements before outages.* **Predictive fault & incident analysis:** Build forecasting and classification models that predict faults and incidents ahead of failure, and correlate alarms across domains for root-cause analysis and noise reduction.* **Model portfolio:** Own models across the operations lifecycle - KPI forecasting, event correlation, root cause localization, capacity prediction and optimization - from framing through production monitoring.* **AI-defined KPIs:** Move beyond vendor KPI formulas by learning leading/lagging indicators directly from raw counters, surfacing signals humans and rule-based systems miss.**AI Harness & Inference Infrastructure (approximately 50%)*** **Training & experimentation:** Build reproducible training, experiment-tracking and hyperparameter-tuning workflows over large-scale telemetry datasets drawn from ClickHouse/YBSQL.* **Evaluation frameworks:** Design rigorous offline and online evaluation (backtesting, precision/recall on incident ground truth, A/B and shadow deployments) so every model's operational impact is measurable.* **Serving & inference pipelines:** Deploy low-latency, high-availability inference services with canary/rollback strategies; implement batch and streaming inference over near-real-time pipeline data.* **Monitoring & lifecycle:** Own model observability - drift detection, data-quality gates, alerting on degradation - and automated retraining/evaluation loops.* **LLM & agentic AI:** Contribute guard-railed LLM-based capabilities such as RAG over fault manuals and specifications, natural-language intent interfaces and LLM-assisted root-cause analysis.* **Partnership with data pipelines:** Define feature and training-data contracts with the Data Management Pipelines team; ensure models consume and produce data through governed, well-modeled interfaces.* **Operational excellence:** Participate in on-call for AI services; own the reliability of everything the harness runs in production.**Required Qualifications**:* Bachelor's or Master's degree in Computer Science, Engineering, Statistics or a related field (or equivalent experience).* 6+ years of experience in ML/AI engineering, including production deployment of models on live, high-volume data.* Hands-on experience developing time-series ML models (anomaly detection, forecasting or classification) - deep learning and classical - with real operational impact.* Experience building ML infrastructure: training pipelines, evaluation frameworks, feature handling, model serving and monitoring (frameworks such as PyTorch, scikit-learn, XGBoost/LightGBM; MLOps tooling such as MLflow, Kubeflow or equivalent).* Strong Python engineering; comfortable with large datasets via SQL (ClickHouse/YugabyteDB/PostgreSQL experience a plus).* Solid software engineering fundamentals - version control, testing, CI/CD, containerization, Kubernetes.* Ability to own problems end-to-end: from ambiguous operational question to productionized, monitored model.**Preferred Qualifications:*** Domain experience in telecom OSS, network operations, or AIOps - e.g., KPI anomaly detection, alarm correlation, root cause analysis on network telemetry.* Experience with LLM applications in production: RAG pipelines, prompt/eval harnesses, guard-railed agent orchestration.* Familiarity with streaming/data platforms: Kafka, Flink/Spark, ClickHouse, or YugabyteDB.* Experience with online evaluation, shadow deployments and human-in-the-loop feedback loops.* Public contributions to ML/AIOps research or open source (papers, repos, talks).**Rakuten Shugi Principles****:**Our worldwide practices describe specific behaviours that make Rakuten unique and united across the world. We expect Rakuten employees to model these 5 Shugi Principles of Success.* **Always improve, always advance.** Only be satisfied with complete success - Kaizen.* **Be passionately professional.** Take an uncompromising approach to your work and be determined to be the best.* **Hypothesize - Practice - Validate - Shikumika.** Use the Rakuten Cycle to success in unknown territory.* **Maximize Customer Satisfaction.** The greatest satisfaction for workers in a service industry is to see their customers smile.* **Speed!! Speed!! Speed!!** Always be conscious of time. Take charge, set clear goals, and engage your team.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Technical Lead AI/ML
Technical Lead AI/ML

Rakuten Kobo Inc. • Bengaluru

On-site
INR 1,200,000 - 1,800,000
Software Engineer 1, Data Science
Software Engineer 1, Data Science

Rakuten Kobo Inc. • Bengaluru

On-site
INR 1,500,000 - 2,800,000
Specialist - Core Network AI & Automation
Specialist - Core Network AI & Automation

Rakuten Symphony • Bengaluru

On-site
INR 2,500,000 - 4,000,000
Senior Consultant, Data Monetization
Senior Consultant, Data Monetization

Rakuten Symphony • Bengaluru

On-site
INR 2,500,000 - 4,500,000
Specialist, Core Network AI & Automation
Specialist, Core Network AI & Automation

Rakuten Kobo Inc. • India

On-site
INR 3,500,000 - 6,000,000
Software Engineer 1, Data Science
Software Engineer 1, Data Science

Rakuten Symphony • Bengaluru

On-site
INR 1,800,000 - 3,200,000
Cutting-edge AI/ML
Ownership of projects
Collaborative culture
Senior Software Engineer
Senior Software Engineer

Rakuten Symphony • Indore District

On-site
INR 1,200,000 - 2,600,000
Senior Specialist, Data Monetization
Senior Specialist, Data Monetization

Rakuten Symphony • Bengaluru

On-site
INR 4,000,000 - 7,000,000
Technical Lead - Platform
Technical Lead - Platform

Rakuten Kobo Inc. • Bengaluru

On-site
INR 2,800,000 - 4,800,000
None
Staff Engineer, Packet Core
Staff Engineer, Packet Core

Rakuten Kobo Inc. • Indore District

On-site
INR 6,000,000 - 9,000,000