Get more replies from employers
Send a job-specific resume in minutes.
V2 Solutions is seeking an experienced DevOps/SRE lead to design and manage hybrid and on-prem architectures for intensive compute workloads. You will deploy containerized microservices with Kubernetes across on-prem and cloud, automate provisioning with Terraform/Ansible/Helm, and scale data platforms like Kafka and Spark.
You will also implement robust observability, CI/CD pipelines, and SRE practices to ensure high availability, security, and auditability across services and storage layers.
Role & Responsibilities
Hybrid & On-Prem Architecture: Architect, build, and maintain robust hybrid and on-premises environments for intensive compute and storage workloads.
Container Orchestration: Deploy, scale, and manage containerized microservices using Kubernetes across both on-prem and cloud infrastructures.
Infrastructure as Code (IaC): Automate end-to-end infrastructure provisioning and configuration management using Terraform, Ansible, and Helm .
Data Platform Support: Configure, optimize, and scale distributed data streaming and storage platforms (including Kafka clusters, Apache Spark, and Data Lakes ).
Scalability, Security & Governance Performance Engineering: Design frameworks for elastic scaling, high availability, load balancing, and resource isolation for high-throughput data services and APIs.
Hybrid Security: Secure hybrid deployments by implementing firewalls, VPNs, and strict identity/access management (IAM).
Data Protection: Implement granular data access controls, audit trails, and end-to-end encryption across services and storage layers while managing secrets securely.
Observability & CI/CD Pipelines Full-Stack Observability: Set up and maintain distributed monitoring and logging stacks using Prometheus, Grafana, ELK, and OpenTelemetry for real-time system insights.
Continuous Delivery: Build, enhance, and maintain CI/CD pipelines to deploy microservices, configurations, and heavy data jobs with zero-downtime and automated rollback capabilities.
SRE Practices: Drive System Reliability Engineering (SRE) practices, including active incident management, root-cause post-mortems, and comprehensive runbook creation.
Experience: 810+ years of dedicated experience in DevOps and System Reliability Engineering (SRE) roles.
Education: Bachelors or Masters degree in Computer Science, Engineering, or a related technical field.
Core Kubernetes Expertise: Proven, deep experience handling on-premises Kubernetes deployments alongside cloud variants (AWS/Azure).
Data Infrastructure Familiarity: Direct experience supporting infrastructure for real-time data platforms, event-driven architectures, microservices orchestration, and ETL pipelines.
Automation & Scripting: Strong proficiency in Bash, Python, or Golang for systems automation.
Networking Foundations: Solid understanding of core networking, distributed storage systems, firewalls, and security topologies for hybrid cloud models.