Get more replies from employers
Send a job-specific resume in minutes.
Ardent SoftSol Inc. is seeking a Senior DevOps Engineer to own the AWS cloud platform, focusing on infrastructure as code, secure multi-account patterns, reliable delivery, and production-grade EKS operations.
The role requires driving AI adoption for infrastructure, incident workflows, and runbooks, while aligning security, compliance, and change control in regulated environments.
We are hiring a Senior DevOps Engineer to own and evolve our AWS cloud platform, with a strong focus on infrastructure as code, secure multi-account patterns, reliable delivery, and production-grade Amazon EKS operations. This person will shape the DevOps roadmap across standards, tooling, automation, operational excellence, and cost-conscious platform practices.
Amazon EKS is central to how we run workloads. We need someone who has built and owned Kubernetes on AWS end-to-end, including cluster lifecycle, networking, security, observability, capacity, reliability, and incident response—not someone whose experience is limited to deploying applications to a cluster managed by others.
The role also includes leading practical AI adoption for infrastructure and platform work, such as AI-assisted authoring and review for IaC and automation, stronger runbooks, incident workflows, and evaluation of tools that improve speed without weakening security, compliance, or change control.
DevOps strategy: Define and socialize priorities across security, reliability, cost, and delivery velocity. Align teams on AWS Well-Architected practices, tagging, guardrails, and repeatable patterns for networking, identity, secrets, and data.
Infrastructure as code: Design, review, and implement changes using Terraform and Terragrunt, with clear module boundaries, environment-specific configuration, and safe promotion across dev, non-prod, and production.
EKS ownership: Build, operate, and own the Kubernetes platform on AWS, including cluster lifecycle, upgrades, node capacity, networking, security, add-ons, cost tuning, reliability, workload standards, namespaces, safe rollouts, and escalation support for cluster-level incidents.
Broader AWS platform: Operate and improve adjacent services such as RDS/Aurora, DynamoDB, object storage and CDN, KMS, Secrets Manager, SNS, Lambda, EventBridge, CI/CD, IAM, VPC, and multi-tenant or multi-namespace patterns where applicable.
Release engineering: Partner with development teams on release processes, deployment strategies, change management, rollbacks, and post-release verification in regulated or high-stakes environments.
Production support: Participate in on-call or escalation rotation as defined by the team; troubleshoot incidents, drive root-cause analysis, and implement preventive fixes through runbooks, dashboards, alarms, and automation.
Observability and operations: Improve monitoring, logging, tracing, and alerting; tune thresholds; reduce noise; and document operational procedures.
Collaboration and governance: Work with security, architecture, and engineering leads to implement least-privilege access, encryption, backup/DR posture, and audit-friendly operations, including expectations for AI-assisted workflows.
AI adoption for infrastructure: Drive a pragmatic AI strategy for IaC, pipeline changes, documentation, runbooks, incident triage, and operational workflows. Establish review gates, testing expectations, drift detection, and guardrails so AI tooling fits regulated or high-stakes environments.
Experience: 8+ years in DevOps / SRE roles, including 4+ years focused on AWS in production.
Infrastructure as code: Strong command of Terraform and modular, environment-driven layouts. Experience with Terragrunt or similar composition patterns is a plus.
Amazon EKS: Deep, mandatory experience building and owning Kubernetes on AWS, including cluster design and lifecycle, upgrades, patching, networking, identity and security, observability, capacity, performance, and production troubleshooting. Surface-level "kubectl-only" experience is not sufficient.
CI/CD and change control: Solid grasp of artifact promotion, secrets injection, rollback planning, and safe change practices across multi-environment pipelines.
Production operations: Experience with incident triage, communication, root-cause analysis, and durable remediation.
AI for DevOps: Demonstrated interest or experience applying AI to platform work, such as AI-assisted IaC review, internal tooling, operational documentation, or incident workflows, with sound judgment around verification, risk, and production limits.
Influence and standards: Ability to influence without authority through written standards, design reviews, and roadmap proposals that engineering teams can adopt.
Communication: Excellent communication skills and comfort working with distributed teams and stakeholders outside pure engineering.