As a Senior Cloud Engineer you will lead and collaborate with engineering and development teams to evaluate, architect, and identify optimal cloud solutions. You will be responsible for designing, developing, and deploying enterprise-scale modular cloud-based systems, and ensuring the functioning of data storage and processing functions are in accordance with company security policies and best practices in cloud security. You will also mentor junior team members, provide strategic cloud guidance to clients, and make informed recommendations based on complex client needs and business requirements.
You will do the following in this job:
- Lead, mentor, and provide technical direction for engineering teams; define cloud strategy and drive innovation across initiatives.
- Architect and design complex, enterprise-scale cloud infrastructure solutions with high availability, scalability, performance optimization, and security design best practice.
- Build and maintain modern DevOps workflows, CI/CD pipelines, system policies, and enterprise-level automation tooling.
- Manage and administer IT infrastructure across on-premises and hybrid multi-cloud environments to deliver services in high availability.
- Enforce cloud security policies and best practices across data storage, processing functions, and system operations.
- Lead identification, analysis, and resolution of complex infrastructure vulnerabilities and production issues, and implement preventive measures.
- Conduct comprehensive reviews of existing systems and lead technology evaluation initiatives to drive operational improvements.
- Serve as the primary technical point of contact for clients and lead vendor relationships to deliver strategic cloud recommendations.
- Design and deploy cloud infrastructure for AI/ML workloads, including GPU-accelerated compute, model serving environments, and data pipelines.
- Implement and maintain MLOps pipelines covering model training, versioning, deployment, and monitoring in collaboration with AI and data science teams.
We are looking for you if you are:
- Minimum 5+ years of experience in Cloud or Application Development, with at least 2 years in a senior role.
- Advanced experience with container orchestration — Docker and Kubernetes — in production environments.
- Deep knowledge of DevOps practices, CI/CD pipelines (GitLab CI/CD, Jenkins), and enterprise-level automation.
- Expert-level proficiency across AWS, GCP, and Azure, with hands-on experience in multi-cloud strategies and cloud migration projects.
- Advanced experience in Infrastructure as Code using Terraform, including resource lifecycle and state management.
- Extensive experience with SQL and NoSQL databases (e.g. Mongo, Redis, Cassandra, DynamoDB) and designing highly scalable, fault‑tolerant systems.
- Expert knowledge of High Availability, Disaster Recovery, and Business Continuity best practices.
- Advanced experience with observability and monitoring tools (e.g. Prometheus, Grafana, OpenTelemetry, Elastic Stack).
- Foundational to intermediate knowledge of MLOps practices, managed ML platforms (AWS SageMaker, GCP Vertex AI, Azure ML), and AI application infrastructure including LLM API integration patterns.
- Demonstrated leadership, communication, and mentoring skills with a proven ability to lead technical teams and advise clients on cloud strategy.
Would be great if you are:
- Experience in application development with frameworks such as Spring, NodeJS, or similar, with familiarity in microservices and reactive programming patterns.
- Multiple cloud platform certifications (AWS Solutions Architect Professional, GCP Professional Cloud Architect, Azure Solutions Architect Expert, etc).
- Expert-level scripting and coding skills (e.g. Python, Go, Java, Bash, PowerShell).
- Deep expertise in streaming and distributed data technologies such as Kafka and Spark, and cloud-native big data services.
- Advanced Linux system administration including virtualization, network configuration, and systems security.
- Experience with compliance frameworks (SOC2, HIPAA, PCI-DSS, GDPR).
- Experience with FinOps practices and AI/GPU workload cost optimization.
- Hands-on experience with GPU instance management and model serving infrastructure (e.g. Triton Inference Server, vLLM, TorchServe).
- Exposure to LLMOps practices and vector database technologies (e.g. Pinecone, Weaviate, pgvector).
- Knowledge of emerging technologies including serverless computing, edge computing, and containerization at scale.
Final and Most Important Requirement
"Enjoy What You Do!!"
We offer an attractive remuneration package, comprehensive on‑the‑job training on new technologies, and excellent career development opportunities to the right candidate.
Telecommunications & Internet Service Providers