We are an AI-led, platform-driven Digital Engineering and Enterprise Modernization partner, combining deep technical expertise and industry experience to help our clients anticipate what?s next. Our offerings and proven solutions create a unique competitive advantage for our clients by giving them the power to see beyond and rise above. We work with many industry-leading organizations across the world, including 20 Fortune 50 companies and 4 of the 5 top banks in both the US and India, and numerous innovators across the healthcare ecosystem.
We are looking for an R&D Developer to join the team responsible for theend-to-end machine learning platformthat powers our network analytics anomaly detection capability. The platform spans three interconnected components: atraining pipelinethat ingests time-series data, trains deep learning and clustering models, and exports them for production serving; anorchestration servicethat manages the full ML application lifecycle on Kubernetes; and aweb-based control panelthat gives operations and client-facing teams a unified interface to configure, train, deploy, and monitor models.
- Location: All Persistent Location
- Experience: 6 to 12years
- Job Type: Full Time Employment
What You'll Do:
- The role combines applied machine learning engineering with MLOps platform work.
- You will write training code that directly affects production model quality, maintain the orchestration layer that deploys and manages those models across environments, and develop the UI surface through which non-technical users interact with the ML lifecycle.
- Design, implement, and maintain ML training pipelines: data ingestion from columnar databases, dataset preparation, LSTM-based anomaly detection model training, and KMeans-based classifier training.
- Log, version, and register trained models usingML flow; export models toONNXformat for downstream inference deployment.
- Maintain theFast API-based orchestration service: per-application YAML configuration storage and update, REST API endpoints for workflow lifecycle management, drift metrics exposure, and config artefact generation.
- Integrate withArgo Workflows(via the Hera Python SDK) to trigger, monitor, and manage distributed ML training jobs on Kubernetes.
- ManageHelm release deployment and teardownvia Ansible Runner, including programmatic chart parameterization from application configuration.
- Build and maintain theStreamlit-based ML control panel: screens for application configuration, model training initiation, deployment management, retraining triggers, and drift metrics visualization.
- Implement interactivetime-series drift chartsand tabular workflow status displays within the control panel UI.
- Write and maintain unit and integration tests for training pipelines, orchestration API endpoints, and UI logic.
- Package all components asDocker images; contribute to Kubernetes and Helm deployment configurations.
- Collaborate with inference service teams on the model artefact contract (ONNX format, metadata, versioning).
- Participate in code reviews, architecture discussions, and maintain technical documentation.
Expertise You'll Bring:
- Proficiency inPython 3.11
- Experience structuring batch workflow scripts and packaging them for production execution (Docker, Poetry, pip)
- Familiarity with async Python and framework-level dependency injection (Fast API/Pydantic patterns)
- Strong understanding of type annotations, data validation withPydantic, and clean API design
- Hands-on experience trainingLSTM or other recurrent neural networksfor time-series tasks usingPyTorch
- Understanding ofanomaly detectionapproaches: reconstruction-based, threshold-based, and statistical methods applied to time-series data
- Working knowledge ofclustering algorithms(KMeans and variants) and their use in classification or segmentation tasks
- Ability to evaluate model quality, tune hyperparameters, and interpret results on time-series datasets
- Bachelor?s or master?s degree in computer science, Data Science, Mathematics, or equivalent practical experience in machine learning engineering or MLOps platform development.
- Competitive salary and benefits package
- Culture focused on talent development with quarterly growth opportunities and company-sponsored higher education and certifications
- Opportunity to work with cutting-edge technologies
- Employee engagement initiatives such as project parties, flexible work hours, and Long Service awards
- Insurance coverage: group term life, personal accident, and Mediclaim hospitalization for self, spouse, two children, and parents
Values-Driven, People-Centric & Inclusive Work Environment:
Persistent is dedicated to fostering diversity and inclusion in the workplace. We invite applications from all qualified individuals, including those with disabilities, and regardless of gender or gender preference. We welcome diverse candidates from all backgrounds.
- We support hybrid work and flexible hours to fit diverse lifestyles.
- Our office is accessibility-friendly, with ergonomic setups and assistive technologies to support employees with physical disabilities.
- If you are a person with disabilities and have specific requirements, please inform us during the application process or at any time during your employment
?Persistent is an Equal Opportunity Employer and prohibits discrimination and harassment of any kind.?