As a Junior SRE/DevOps Engineer, you will support the setup and maintenance of Infra, CI/CD and help sustain different key projects used throughout globally, working under the guidance of senior engineers.
As a member of our geographically distributed development team your communication and analytical skills are essential to the role.
Key Responsibilities
- Assist in designing and maintaining cloud infrastructure that is secure, scalable, and highly available on AWS/Azure
- Work collaboratively with software engineering to define infrastructure and deployment requirements
- Provision, configure and maintain AWS cloud infrastructure defined as code.
- Containerization using Docker and Kubernetes
- Troubleshoot problems across a wide array of services and functional areas
- Build and maintain operational tools for deployment, monitoring, and analysis of AWS infrastructure and systems
- Assist with infrastructure cost analysis and support optimization efforts.
- Support the development of self-healing and automated remediation mechanisms using AI/ML techniques
- Assist in integrating AI/LLM capabilities into DevOps workflows (e.g., log analysis, automated RCA, deployment insights)
- Support monitoring strategy enhancements by leveraging intelligent alerting, noise reduction, and pattern-based anomaly detection across logs, metrics, and traces.
- Assist in building and maintaining MLOps pipelines for model training, deployment, and continuous improvement.
- Collaborate with a global team of engineers in a highly agile DevOps environment, focused on efficient operation of daily activities, developer productivity and continuous improvement of the framework.
- Support the development, implementation, and maintenance of CI/CD frameworks, and contribute to tools development for hybrid environments (Cloud, On premise) with a vision to achieve “CI/CD” objectives for large-scale integration of systems in order to reduce manual build and deploy efforts.
- Work with geographically dispersed teams including multi-vendor into Scrum teams to meet “CI/CD”
Required Knowledge & Skills
- 2-4 years of experience building and maintaining AWS infrastructure (VPC, EC2, Security Groups, IAM, ECS/EKS, CloudFront, S3, RDS, SQS, SNS, Lambda Function, Batch jobs, AWS Glue)
- Working understanding of how to secure AWS environments and meet compliance requirements
- Working knowledge of deploying and managing infrastructure with Terraform.
- Exposure to or working knowledge of LLMs (OpenAI, Azure OpenAI, Claude etc.)
- Basic awareness of LLMOps concepts (prompt management, model evaluation, versioning, fine-tuning lifecycle)
- Familiarity with MLOps tools such as MLflow, SageMaker, Kubeflow or equivalent.
- Familiarity with AIOps platforms/tools for intelligent monitoring and incident management.
- Ability to apply AI techniques to improve deployment speed, reliability, and monitoring effectiveness.
- Working experience on windows & Linux based environments.
- Experience with Docker, GitHub, Jenkins, Azure DevOps, ELK and deploying applications on AWS.
Good command on scripting languages like Python, Bash/Shell, Powershell etc- Knowledge in log analytics tools like Elastic search and Kibana.
- Basic knowledge of Cloud Migration/Disaster Recovery/Blue Green Deployment implementation.
- Good understanding about monitoring the services and alerting using Cloudwatch, Datadog, Prometheus or Azure monitor.
- Good to hire a candidate with certification
Personal Attributes
- Very good communication skills.
- Ability to easily fit into a distributed development team.
- Customer service oriented.
- Enthusiastic/High initiative.
- Ability to manage timelines of multiple initiatives.
- Very good attention to detail and the ability to always follow up.