Job Summary
AI/ML Platform Engineering focuses on building scalable engineering platforms, infrastructure, automation, and operational capabilities that enable AI and Machine Learning teams to develop, deploy, integrate, monitor, and operate AI/ML solutions efficiently and securely. AI/ML platforms provide standardized tools and services for model lifecycle management, MLOps, model serving, Generative AI, LLM applications, agentic AI, experimentation, deployment, observability, and governance. AI/ML Platform Engineering combines cloud infrastructure, software engineering, DevOps, MLOps, Kubernetes, automation, and AI technologies to create reliable and self-service platforms that accelerate enterprise AI adoption.
Data AI/ML (Artificial Intelligence and Machine Learning) Engineering involves the use of algorithms and statistical models to enable systems to analyze data, learn patterns, and make data-driven predictions or decisions without explicit human programming. AI/ML applications leverage vast amounts of data to identify insights, automate processes, and solve complex problems across a wide range of fields, including healthcare, finance, e-commerce, and more. AI/ML processes transform raw data into actionable intelligence, enabling automation, predictive analytics, and intelligent solutions. Data AI/ML combines advanced statistical modeling, computational power, and data engineering to build intelligent systems that can learn, adapt, and automate decisions.
Responsibilities
- Independently design and implement scalable machine learning solutions and data systems, ensuring end to end workflows, large scale analytics and reliability
- Collaborate with stakeholders to translate business needs into data engineering solutions, evaluate user journeys and challenge business requirements to ensure seamless, value driven delivery and integration of solutions
- Implement and refine feature engineering, monitoring, ML pipelines, deploy models in production, and address challenges in data pipelines
- Apply innovative problem-solving techniques, leveraging advanced methodologies to find unique approaches to complex problems and improve outcomes
- Investigate and resolve complex challenges in data models and deployment to ensure reliable solutions that meet performance benchmarks
- Mentor team members through code reviews, pairing sessions, knowledge-sharing sessions, and contribute to Communities of Practice
- Communicate technology, infrastructure, and deployment decisions clearly to both technical and non-technical stakeholders while maintaining detailed documentation to ensure reproducibility, scalability, and understanding
- Ensure readiness for production releases, focusing on testing, monitoring, observability, and maintaining scalability and reusability of models for future projects
- Drive cross-team and cross-discipline initiatives to optimize workflows, remove redundant applications and processes, share best practices, and enhance collaboration between teams
- Demonstrate awareness of shared platform capabilities and actively identify opportunities to leverage them in designing efficient and scalable data engineering solutions
Core Skills
- AI Machine Learning: Creating AI-powered solutions using Generative AI, Agentic AI and machine learning technologies to solve business problems and automate processes. Proficiency Level: Proficient
- Cloud Platform Engineering: Designing, deploying and operating secure, scalable and highly available cloud-native platforms on AWS. Proficiency Level: Advanced
- Programming: Writing production-grade applications, platform services and automation using languages such as Python, Java and SQL. Proficiency Level: Advanced
- MLOps AI Operations: Automating the end-to-end lifecycle of AI models including deployment, monitoring, governance and optimization. Proficiency Level: Advanced
- DevSecOps Automation: Using CI/CD pipelines, Infrastructure as Code and security controls to automate software development and platform operations. Proficiency Level: Advanced
- Distributed Systems: Designing scalable, resilient and fault-tolerant systems capable of supporting enterprise AI workloads. Proficiency Level: Advanced
Specialized Skills
- Generative AI LLMs: Building applications using Large Language Models, Retrieval Augmented Generation (RAG), prompt engineering and AI agents.
- Agentic AI Frameworks: Developing autonomous and multi-agent solutions using modern orchestration frameworks and enterprise AI patterns.
- AI Platform Architecture: Designing reusable AI platform capabilities including model serving, inference orchestration and governance frameworks.
- AWS Cloud Services: Leveraging services such as EKS, Lambda, Bedrock, API Gateway, S3, IAM, Step Functions and CloudWatch.
- Containerization Kubernetes: Deploying, managing and scaling containerized applications using Docker and Kubernetes.
- Infrastructure as Code: Automating infrastructure provisioning and platform operations using Terraform and cloud-native tooling.
- Observability Reliability Engineering: Implementing monitoring, logging, tracing and performance optimization for platform services.
- Security Governance: Applying enterprise security practices including IAM, secrets management, compliance and responsible AI controls.
- API Integration Engineering: Building scalable APIs and integrating enterprise platforms, cloud services and AI solutions.
- Technical Documentation: Creating architecture documents, operational runbooks and technical standards for enterprise platforms.