Stand out for this role — generate a tailored resume and cover letter in about a minute.
Tap Growth ai is hiring a Senior AI/ML Engineer (MLOps) to design, deploy, monitor, and improve large-scale AI/ML solutions in production. You will focus on reliability, governance, and operability for Generative AI and traditional ML systems while collaborating with data scientists, platform teams, and business stakeholders.
The role requires 6+ years in ML/AI, cloud production support, and strong mentoring capabilities, with duties spanning monitoring, pipelines, lifecycle management, and
We are seeking a Senior AI/ML Engineer (MLOps) to design, deploy, monitor, and continuously improve large-scale AI and machine learning solutions in production. This role focuses on ensuring the reliability, performance, governance, and operational excellence of both Generative AI and traditional ML systems, while collaborating closely with data scientists, engineers, platform teams, and business stakeholders.
Design and implement monitoring, alerting, and observability solutions for ML and AI systems, including model performance, data drift, latency, and AI operational health.
Build and maintain automated evaluation, testing, and validation pipelines for ML models, GenAI workflows, prompts, and agentic systems.
Investigate and resolve production issues related to model behavior, AI applications, data quality, integrations, and RAG/retrieval pipelines.
Manage model, prompt, embedding, and vector index lifecycles, including versioning, rollout strategies, and evaluation gates.
Collaborate with platform and infrastructure teams to optimize AI deployments for scalability, reliability, and cost efficiency.
Establish feedback loops using user interactions, business metrics, UAT findings, and expert reviews to drive continuous improvement.
Ensure compliance with governance, security, privacy, and operational best practices.
Create operational documentation, incident reports, and performance benchmarks.
Communicate technical insights, risks, and system health to both technical and non-technical stakeholders.
Mentor junior team members through coaching, knowledge sharing, and code reviews.
Bachelor's degree in Computer Science, Data Science, Engineering, Information Technology, or a related field.
6+ years of experience in Machine Learning, AI Engineering, MLOps, or related software engineering roles.
Experience supporting production AI/ML systems in cloud environments and agile delivery teams.
Strong problem-solving, communication, stakeholder management, and mentoring capabilities; willingness to participate in on-call support when required.
Strong Python expertise including Pandas, NumPy, and Scikit-learn for model maintenance, automation, and production support.
Advanced SQL skills for data analysis, operational checks, troubleshooting, and data quality validation.
Hands-on experience building or supporting production Generative AI solutions, including RAG, AI agents, LlamaIndex, CrewAI, or similar orchestration frameworks.
Strong understanding of Machine Learning fundamentals, model evaluation techniques, drift detection, bias monitoring, and model performance management.
Proven experience in MLOps and model monitoring, using tools such as Prometheus, Grafana, or cloud-native monitoring platforms to track model performance, data quality, latency, and operational metrics.
Experience working with Microsoft Azure AI/ML ecosystem, including services such as Azure Machine Learning, Databricks, or related cloud-based AI deployment platforms.
Practical experience working in Agile environments (Scrum/Kanban).
Experience with Apache Spark, Dask, or other distributed data processing frameworks.
Familiarity with model serving and deployment platforms such as Azure ML Inference, Triton, SageMaker Endpoints, and API frameworks like FastAPI or Flask.
Knowledge of Docker and Kubernetes for scalable ML deployments.
Experience with orchestration and ML lifecycle tools such as Airflow, Kubeflow, or MLflow.
Exposure to Infrastructure as Code (IaC) tools such as Terraform.
Experience mentoring engineers and driving engineering best practices across teams.