Senior Machine Learning Engineer

Persistent

Pune District

Hybrid

INR 3,000,000 - 6,000,000

Full time

4 days ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

Culture of growth and education
Hybrid work arrangements
Opportunity to work with cutting-edge技

Job summary

Persistent in Pune seeks a Senior Machine Learning Engineer to own the product-sizing capability for a cloud-based analytics platform. You will design sizing models, optimize deployments, and work with product, engineering, and customer teams to deliver scalable, data-driven infrastructure solutions.

The role requires 5–10 years of ML/ sizing experience, strong Python and ML tooling, and hands-on Kubernetes/Linux expertise.

Qualifications

  • 5–10 years of experience in ML, performance engineering, or sizing.
  • Strong Python development and data-analysis skills.
  • Experience building models using telemetry and production data.
  • Experience with Kubernetes and Linux administration.
  • Ability to design and validate sizing models.

Responsibilities

  • Design, implement, and maintain sizing models for platform deployments.
  • Model infra requirements (CPU, memory, storage, IOPS, network) from workloads.
  • Develop data-driven sizing methods using statistics and ML.
  • Collaborate with Product, Eng, Support, and Customer Ops.
  • Modernize telemetry collection and ensure safe production diagnostics.

Skills

Python
Pandas
NumPy
SciPy
Scikit-learn
TensorFlow/PyTorch
Linux
Kubernetes
Data analysis
Cross-functional collaboration

Tools

Jupyter Notebooks

Job description

About Position:

We are seeking a highly skilled Senior Machine Learning Engineer to own and enhance the product-sizing capability of a cloud-based, cloud-agnostic telecom analytics platform. The platform is designed to support consistent deployments across public, private, and hybrid cloud environments. This role blends machine learning, performance engineering, capacity planning, software development, and telemetry analytics to build accurate sizing models and optimize platform deployments. The ideal candidate will work closely with product, engineering, support, and customer teams to deliver scalable, data-driven solutions that improve infrastructure planning, resource utilization, and customer experience.


  • Role: Senior ML Engineer
  • Location: Pune
  • Experience: 5 to 10 Years
  • Job Type: Full Time Employment

What You'll Do:

  • Design, implement, validate, and maintain sizing models for product components and end-to-end platform deployments.
  • Model infrastructure requirements including CPU, memory, storage capacity, disk IOPS, throughput, network bandwidth, pod counts, worker threads, Redis shards, and operational headroom based on customer workloads.
  • Develop and improve data-driven sizing methodologies using statistical analysis and machine learning techniques.
  • Support telecom workload dimensions such as PPS, EPS, OPS, protocol mix, call rates, active subscribers, concurrent sessions, recording volumes, retention periods, replication requirements, and high-availability topologies.
  • Enhance existing regression-based sizing approaches using multivariate regression, clustering, time-series forecasting, nonlinear models, and other machine learning techniques where appropriate.
  • Evaluate model performance using holdout datasets, cross-validation strategies, confidence intervals, prediction intervals, sensitivity analysis, and error metrics.
  • Detect and investigate outliers, missing data, configuration drift, throttling, system saturation, and non-representative workloads while preserving critical operational evidence.
  • Define model applicability boundaries, minimum data-quality thresholds, and conservative fallback mechanisms to prevent unsafe extrapolation.
  • Build and maintain version-controlled Python-based telemetry collectors for Linux and Kubernetes environments.
  • Modernize legacy collection frameworks by replacing Bash, Perl, and Python 2 based implementations where feasible.
  • Safely execute approved telemetry collection tools in customer environments, both remotely and on-site, while collaborating with support and operations teams.
  • Collect workload, infrastructure, network, storage, application, and Redis telemetry with minimal impact on production systems.
  • Engineer telemetry collectors that are read-only by default and designed with resource limits, timeout controls, retry handling, dry-run capabilities, graceful shutdown, progress indicators, and auditable logging.
  • Collaborate with cross-functional teams including Product Management, Engineering, Support Services, Professional Services, and Customer Operations.
  • Contribute to performance analysis, capacity planning recommendations, deployment guidance, and continuous platform improvements.

Expertise You'll Bring:

  • 5 to 10 years of experience in Machine Learning, Performance Engineering, Capacity Planning, or Solution Sizing.
  • Strong experience modeling CPU, memory, storage, network, and Kubernetes resource requirements from operational and workload data.
  • Hands-on experience analyzing distributed systems performance, scalability, latency, throughput, reliability, and high-availability architectures.
  • Knowledge of telecom workload characteristics including PPS, EPS, OPS, protocol mix, subscriber volumes, call rates, retention requirements, replication, and HA deployments.
  • Experience defining model applicability limits, uncertainty measurements, confidence intervals, sensitivity analysis, and error-based model validation.
  • Strong understanding of outlier detection, data quality validation, anomaly identification, and safe model extrapolation practices.
  • Advanced Python development skills with experience using pandas, NumPy, SciPy, scikit-learn, Jupyter Notebooks, and automated testing frameworks.
  • Proven experience building, validating, and deploying statistical and machine learning models using operational and production data.
  • Strong knowledge of multivariate regression, time-series forecasting, clustering techniques, nonlinear modeling approaches, and predictive analytics.
  • Experience evaluating model quality using cross-validation, holdout testing, performance metrics, and iterative model improvement techniques.
  • Strong Linux administration and troubleshooting experience.
  • Hands-on expertise with Kubernetes, including containers, pods, resource allocation, CPU throttling, memory utilization, storage management, networking, and observability tools.
  • Experience collecting and analyzing production telemetry in cloud-native and distributed environments.
  • Understanding of secure production diagnostics, customer privacy requirements, least-privilege principles, and change-controlled operational processes.
  • Strong analytical, problem-solving, communication, and stakeholder management skills.
  • Ability to work independently and drive technical solutions from concept to deployment.

Benefits:

  • Competitive salary and benefits package
  • Culture focused on talent development with quarterly growth opportunities and company-sponsored higher education and certifications
  • Opportunity to work with cutting-edge technologies
  • Employee engagement initiatives such as project parties, flexible work hours, and Long Service awards
  • Annual health check-ups
  • Insurance coverage: group term life, personal accident, and Mediclaim hospitalization for self, spouse, two children, and parents

Values-Driven, People-Centric & Inclusive Work Environment:

Persistent is dedicated to fostering diversity and inclusion in the workplace. We invite applications from all qualified individuals, including those with disabilities, and regardless of gender or gender preference. We welcome diverse candidates from all backgrounds.


  • We support hybrid work and flexible hours to fit diverse lifestyles.
  • Our office is accessibility-friendly, with ergonomic setups and assistive technologies to support employees with physical disabilities.
  • If you are a person with disabilities and have specific requirements, please inform us during the application process or at any time during your employment

"Persistent is an Equal Opportunity Employer and prohibits discrimination and harassment of any kind."

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Telecom Analytics Domain Expert
Telecom Analytics Domain Expert

Persistent Systems • Pune District

On-site
INR 4,500,000 - 6,500,000
Talent development
Education sponsorship
Flexible hours
+2
Telecom Analytics Domain Expert
Telecom Analytics Domain Expert

Persistent Systems Limited • Pune District

Hybrid
INR 4,000,000 - 7,000,000
Hybrid work
Growth opportunities
Education sponsorship
+2
Sr. ML Engineer
Sr. ML Engineer

Persistent • Pune District

On-site
INR 3,500,000 - 7,000,000
Hybrid work
Competitive salary
Education sponsorship
+3
Machine Learning Operations Site Reliability Engineer
Machine Learning Operations Site Reliability Engineer

Persistent • Bengaluru

On-site
INR 2,800,000 - 4,200,000
Competitive salary and benefits
Culture focused on talent development
Cutting-edge technologies
+3
Telecom Analytics Domain Expert
Telecom Analytics Domain Expert

Persistent • Pune District

On-site
INR 4,000,000 - 7,000,000
Competitive salary
Hybrid work options
Professional development
+2
Machine Learning Specialist
Machine Learning Specialist

Persistent Systems • Pune District

On-site
INR 1,500,000 - 2,500,000
Competitive salary
Quarterly growth opportunities
Company-sponsored higher education
+3
Infrastructure Architect
Infrastructure Architect

TymblHub • Pune District

On-site
INR 4,500,000 - 7,000,000
Competitive salary
Hybrid work options
Healthcare benefits
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

Persistent Systems Limited • Pune District

On-site
INR 1,400,000 - 2,200,000
Hybrid work
Long Service awards
Company-sponsored education
Senior Backend Engineer
Senior Backend Engineer

Persistent Systems Limited • Pune District

Hybrid
INR 1,800,000 - 2,400,000
Hybrid work
Professional development
Insurance coverage
+1
Senior ClickHouse Database Engineer
Senior ClickHouse Database Engineer

Persistent • Pune District

On-site
INR 500,000 - 900,000
Competitive salary
Hybrid work
Health insurance
+4