A different kind of consulting company that delivers offshore and nearshore product development, user experience, app modernization, business intelligence, and business apps services.
Job Description
Job Title: AI DevOps Engineer —Mid/Senior Level
Experience: 4–7 Years
Location: Raidurg Main Road, Hyderabad.
Work Mode: On-site
Work Hours: 2-11 PM
Notice Period: Immediate Joiner (15-30 days)
About theRole
We arelooking for a Mid-Level AI DevOps Engineer with 4-7 years ofexperience in DevOps, cloud infrastructure, automation, and productiondeployment environments.
The idealcandidate should have strong hands-on experience with cloud platforms,CI/CD, Docker, Kubernetes, infrastructure as code, monitoring, and automation,along with a working understanding of AI/ML deployment workflows.
KeyResponsibilities
CloudInfrastructure & DevOps
- Design, deploy, and managecloud-based infrastructure for AI and software applications.
- Work with cloud platforms such as AWS, Azure, or GCP.
- Build and maintain infrastructureusing tools such as Terraform, CloudFormation, Ansible.
- Support scalable, secure, andreliable environments for production workloads.
- Optimize infrastructure forperformance, cost, availability, and operational efficiency.
CI/CD& Automation
- Build and maintain CI/CDpipelines for application and AI service deployments.
- Automate build, testing,deployment, and rollback processes.
- Improve deployment reliabilityand reduce manual operational tasks.
- Work with tools such as AzureDevOps, GitHub Actions, Jenkins.
- Create reusable scripts,templates, and automation workflows for engineering teams.
Containerization& Orchestration
- Deploy and manage containerizedapplications using Docker.
- Work with Kubernetes forapplication deployment, scaling, networking, and troubleshooting.
- Manage Helm charts and Kubernetesmanifests.
- Troubleshoot container, cluster,and infrastructure-related issues.
AI / MLOpsSupport
- Support deployment and monitoringof AI/ML models in production environments.
- Collaborate with data scientists,ML engineers, and backend engineers to streamline model deploymentworkflows.
- Assist with model versioning,model serving, and release automation.
- Work with MLOps tools such as MLflow,Kubeflow, SageMaker, Vertex AI, Azure ML, Airflow, or similar platforms.
- Support infrastructure for AIservices, APIs, and model inference workloads.
- Implement and maintainmonitoring, logging, tracing, and alerting systems.
- Use tools such as Prometheus,Grafana, ELK Stack, Datadog, New Relic, CloudWatch, or Azure Monitor.
- Monitor application andinfrastructure performance.
- Participate in incident response,root cause analysis, and production support.
- Help improve system reliability,uptime, and operational visibility.
Security& Compliance
- Apply DevSecOps practices acrossinfrastructure and deployment pipelines.
- Manage access controls, IAMroles, secrets, and secure configuration.
- Support vulnerability scanning,patching, and security hardening.
- Ensure cloud and deploymentenvironments follow security best practices.
- Work with tools such as HashiCorpVault, AWS Secrets Manager, Azure Key Vault, or GCP Secret Manager.
RequiredQualifications
- 4-7 years of experience in DevOps, Cloud Engineering,Site Reliability Engineering, Platform Engineering, or InfrastructureEngineering.
- Strong hands-on experience withat least one cloud platform: AWS, Azure, or GCP.
- Experience building and managingCI/CD pipelines.
- Strong experience with Docker and containerized deployments.
- Working experience with Kubernetes in production or near-production environments.
- Experience withinfrastructure-as-code tools such as Terraform, Ansible, CloudFormation.
- Strong scripting skills using Python,Bash, or PowerShell.
- Experience with monitoring andlogging tools such as Prometheus, Grafana, ELK, Datadog, New Relic, orCloudWatch.
- Good understanding of networking,Linux systems, security, and cloud architecture.
- Familiarity with AI/ML workflows,model deployment, or MLOps concepts.
- Experience supporting productionapplications and troubleshooting infrastructure issues.
PreferredQualifications
- Experience supporting AI/MLapplications or model deployment pipelines.
- Exposure to LLM applications,vector databases, RAG pipelines, or generative AI infrastructure.
- Experience with GPU-basedworkloads or AI inference infrastructure.
- Familiarity with tools such as MLflow,Kubeflow, SageMaker, Vertex AI, Azure ML, Airflow, or Argo Workflows.
- Experience with Helm, servicemesh, or Kubernetes operators.
- Knowledge of DevSecOps practicesand cloud security controls.
- Cloud, Kubernetes, or DevOpscertifications are a plus.
Required TechnicalSkills
CloudPlatforms: AWS,Azure, GCP
Containers & Orchestration: Docker, Kubernetes, Helm
Infrastructure as Code: Terraform, Ansible, CloudFormation
CI/CD: GitHub Actions, Jenkins, Azure DevOps
Scripting: Python, Bash, PowerShell
Monitoring & Logging: Prometheus, Grafana, ELK Stack, Datadog, New Relic, CloudWatch
MLOps / AI Tools: MLflow, Kubeflow, SageMaker, Vertex AI, Azure ML,Airflow
Security: IAM, secrets management, vulnerability scanning, DevSecOps
Operating Systems: Linux, Unix-based systems