The Senior DevOps Engineer will be responsible for designing cloud infrastructure, implementing CI/CD pipelines, automating deployments, improving application reliability, and mentoring junior engineers. The role requires deep expertise in cloud platforms, Kubernetes, Infrastructure as Code (IaC), automation, monitoring, security, and production support.
The ideal candidate has strong technical leadership skills, hands-on implementation experience, and a proactive approach to solving complex infrastructure challenges.
Key Responsibilities
- Design and implement enterprise DevOps practices across multiple client engagements.
- Define infrastructure standards, deployment strategies, and automation frameworks.
- Mentor and guide junior DevOps engineers.
- Conduct architecture and code reviews related to infrastructure and automation.
- Collaborate with development, QA, security, and project teams to improve delivery efficiency.
Design, implement, and manage cloud infrastructure on one or more platforms:
AWS
- Hybrid Infrastructure
- EC2
- VPC
- IAM
- S3
- Route 53
- Auto Scaling
- ECS
- EKS
- RDS
- Aurora
- CloudFront
- CloudWatch
- Lambda
OR
- Compute Engine
- GKE
- Cloud Functions
CI/CD Implementation
Design and manage automated CI/CD pipelines using:
Responsibilities include:
- Build automation
- Release automation
- Rollback strategies
- Canary deployments
- Automated testing integration
Lead container platform implementation.
Required experience:
Responsibilities include:
- Cluster management
- RBAC configuration
- Ingress Controllers
- Stateful Applications
- Autoscaling
- Resource optimization
Monitoring & Observability
- Grafana
- ELK Stack
- Loki
- Splunk
- New Relic
- AWS CloudWatch
Responsibilities:
- Alerting
- Performance monitoring
- Root cause analysis
- Capacity planning
Security & DevSecOps
Implement security throughout the software delivery lifecycle.
Responsibilities include:
- IAM management
- Secrets management
- Security scanning
- Container security
- Infrastructure security
- Vulnerability management
- Policy enforcement
- OpenLDAP
Experience with tools such as:
Automation & AI Infrastructure
- Python
- Bash
- Shell Scripting
- PowerShell
- Go
Automate:
- Deployments
- Backups
- Monitoring
- Scaling
Architect, provision, and maintain on-premises GPU/AI servers:
- Manage high-performance, on-premises AI training clusters, optimizing hardware layouts for specialized workloads
Networking
Strong understanding of:
- TCP/IP
- DNS
- VPN
- SSL/TLS
- Reverse Proxy
- CDN
- Firewalls
- Load Balancers
- NAT
- Private Networking
Experience in restricting access to internal tools and administrative dashboards by routing traffic through dedicated corporate VPN endpoints
- Production SupportTroubleshoot production issues.
- Perform root cause analysis (RCA).
- Ensure high availability and uptime.
- Participate in on-call support rotations where required.
- Plan disaster recovery and business continuity strategies.
Required Technical skills
DevOps Tool
Infrastructure as Code
- Terraform
- Ansible
- CloudFormation
Container
Version Control
Monitoring
- Grafana
- ELK
- Splunk
- CloudWatch
Operating System
Databases
Experience Required
- 4-6 years of experience in DevOps, Cloud Engineering, or Site Reliability Engineering.
- Strong hands-on experience with AWS, Azure, or GCP.
- Experience supporting production systems with high availability requirements.
- Experience working with distributed or remote teams.
- Experience in client-facing IT services or consulting environments is preferred.
Soft Skills
- Excellent communication and presentation skills
- Technical leadership and mentoring ability
- Strong analytical and troubleshooting skills
- Ownership and accountability
- Ability to work independently in a remote environment