We are seeking a Senior AI/ML Engineer (MLOps) to design, deploy, monitor, and continuously improve large-scale AI and machine learning solutions in production. This role focuses on ensuring the reliability, performance, governance, and operational excellence of both Generative AI and traditional ML systems, while collaborating closely with data scientists, engineers, platform teams, and business stakeholders.
Key responsibilities
- Design and implement monitoring, alerting, and observability solutions for ML and AI systems, including model performance, data drift, latency, and AI operational health.
- Build and maintain automated evaluation, testing, and validation pipelines for ML models, GenAI workflows, prompts, and agentic systems.
- Investigate and resolve production issues related to model behavior, AI applications, data quality, integrations, and RAG/retrieval pipelines.
- Manage model, prompt, embedding, and vector index lifecycles, including versioning, rollout strategies, and evaluation gates.
- Collaborate with platform and infrastructure teams to optimize AI deployments for scalability, reliability, and cost efficiency.
- Establish feedback loops using user interactions, business metrics, UAT findings, and expert reviews to drive continuous improvement.
- Ensure compliance with governance, security, privacy, and operational best practices.
- Create operational documentation, incident reports, and performance benchmarks.
- Communicate technical insights, risks, and system health to both technical and non-technical stakeholders.
- Mentor junior team members through coaching, knowledge sharing, and code reviews.
About you
- Bachelor's degree in Computer Science, Data Science, Engineering, Information Technology, or a related field.
- 6+ years of experience in Machine Learning, AI Engineering, MLOps, or related software engineering roles.
- Experience supporting production AI/ML systems in cloud environments and agile delivery teams.
- Strong problem-solving, communication, stakeholder management, and mentoring capabilities; willingness to participate in on-call support when required.
- Strong Python expertise including Pandas, NumPy, and Scikit-learn for model maintenance, automation, and production support.
- Advanced SQL skills for data analysis, operational checks, troubleshooting, and data quality validation.
- Hands-on experience building or supporting production Generative AI solutions, including RAG, AI agents, LlamaIndex, CrewAI, or similar orchestration frameworks.
- Strong understanding of Machine Learning fundamentals, model evaluation techniques, drift detection, bias monitoring, and model performance management.
- Proven experience in MLOps and model monitoring, using tools such as Prometheus, Grafana, or cloud-native monitoring platforms to track model performance, data quality, latency, and operational metrics.
- Experience working with Microsoft Azure AI/ML ecosystem, including services such as Azure Machine Learning, Databricks, or related cloud-based AI deployment platforms.
About us
Our client is a global technology consulting and digital transformation organization known for delivering innovative software, engineering, and business solutions to enterprise clients worldwide. With a strong culture of collaboration, learning, and continuous improvement, the client focuses on building inclusive teams and creating meaningful career experiences for both employees and candidates.