SDE - III DevOps

AiDASH

Bengaluru

On-site

INR 4,000,000 - 6,500,000

Full time

7 days ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

AiDASH is seeking a senior DevOps engineer to own the infrastructure for a satellite-imagery and ML-inference platform. You will drive end-to-end reliability, security, and operability, collaborating with dev, QA, security, and product teams in a fast-paced, AI-forward environment.

Expect to design scalable CI/CD, manage Kubernetes clusters with GPU workloads, and implement robust observability and incident response practices while advancing an existing DevSecOps program.

Qualifications

  • 6+ years in DevOps, SRE, or infrastructure engineering with ownership of production systems at scale.
  • Deep cloud experience (AWS/GCP/Azure) and understanding of others' differences.
  • Strong IaC and CI/CD: Jenkins, GitHub/GitLab Actions, or similar.
  • Production Kubernetes expertise with pod debugging, autoscaling, and cost awareness.
  • Strong scripting in Python or Bash; good debugging and automation skills.
  • Hands-on with observability stacks and defining meaningful SLOs.

Responsibilities

  • Own scalable, secure infra across major clouds with AI-assisted IaC and cost reviews.
  • Architect and maintain CI/CD pipelines enabling rapid, safe deployments with AI aids.
  • Lead Kubernetes orchestration for satellite-data and ML workloads, incl. GPU scheduling.
  • Own observability standards and ensure platform availability against SLO targets.
  • Automate with Terraform/Ansible and promote AI-assisted pair programming.
  • Define secure secrets management and access-control practices with DevSecOps.
  • Establish disaster-recovery and backups with regular game days.
  • Build/internal tooling using LLMs to speed up log triage and runbook drafting.

Skills

Cloud knowledge
SRE practices
Python scripting
Automation mindset
Influencing without authority

Tools

Terraform
Ansible
Jenkins
GitLab CI
GitHub Actions
Kubernetes
Prometheus
Grafana
ELK
Airflow
Argo
LLM APIs

Job description

We are looking for an SDE-III DevOps who takes end-to-end ownership of the infrastructure that runs our satellite-imagery and ML-inference platform. This is a senior, hands-on, individual-contributor role you will set the bar across the DevOps function, influence technical direction, and lead by example, without managing a team. What sets this role apart at AiDASH is what runs on the infrastructure: a globally deployed platform that ingests [petabytes of satellite imagery], runs [millions of inference requests per day] across CPU and GPU fleets, and serves utilities, transportation, and construction customers on tight SLOs. You will build the systems that underpin that and you will do it with AI as a first-class tool, not an afterthought. You will work closely with developers, QA, security, and product teams to design systems that are reliable, secure, and easy to operate while extending an already mature DevSecOps program.

Responsibilities
  • Own scalable, secure infrastructure across AWS, Azure, or GCP using AI coding assistants to accelerate IaC authoring, policy validation, and cost reviews.
  • Architect and maintain CI/CD pipelines that support rapid, safe deployments, with AI-assisted failure triage, smart test selection, and automated release notes.
  • Lead container orchestration on Kubernetes for production satellite-data and ML-inference workloads, including GPU scheduling, autoscaling, and model-serving infrastructure.
  • Own observability standards (SLIs, SLOs, error budgets, alerting) and be accountable for keeping platform availability at our committed SLO targets.
  • Implement automation using Terraform, Ansible, or equivalent, with AI pair-programming as a normal part of the workflow, not a side experiment.
  • Define and harden security, secrets management, and access-control practices in partnership with the DevSecOps function.
  • Establish disaster-recovery strategies and backups for critical systems and prove them with regular game days.
  • Build or extend internal tooling that uses LLMs to make engineers faster at log triage, runbook drafting, alert summarization, and ChatOps for routine ops tasks.
  • Participate in a follow-the-sun on-call rotation with the global DevOps team.
  • Raise the bar on engineering standards, including how the team adopts and governs AI tooling in infrastructure workflows.
Requirements
  • 6+ years in DevOps, infrastructure engineering, or SRE, with proven ownership of production systems at meaningful scale.
  • Deep experience with at least one major cloud (AWS, Azure, or GCP) and a working grasp of what the others do differently.
  • Strong infrastructure-as-code (Terraform or equivalent) and CI/CD pipeline experience (Jenkins, GitLab CI, GitHub Actions, or similar).
  • Production Kubernetes beyond what I have used with Helm. You can debug a stuck pod, design an autoscaling strategy, and reason about cost.
  • Strong scripting in Python, Bash, or equivalent.
  • Hands-on with at least one observability stack (Prometheus, Grafana, ELK, or equivalent) and able to define meaningful SLOs.
  • Shipped real work using AI coding assistants, IaC, debugging, incident triage, and internal tooling and can speak to where they helped and where they got in your way.
  • Built or extended at least one internal tool using LLM APIs (or are clearly keen to). A hacky prototype counts.
  • Comfortable in a fast-moving, engineering-driven environment and good at influencing without authority.
Nice To Have
  • Production experience with ML infrastructure model serving (Triton, KServe, TorchServe, or similar), GPU workload management, feature stores, or data-pipeline orchestration (Airflow, Argo, or equivalent).
  • Familiarity with compliance frameworks relevant to critical infrastructure (SOC 2 ISO 27001 NERC, or similar).
  • Certifications such as AWS DevOps Engineer, CKA / CKAD, or equivalent.
  • Experience with serverless platforms (Lambda, Cloud Functions, etc. ).
  • Exposure to managing and tuning production databases (PostgreSQL, MySQL, or NoSQL).

This job was posted by Akanksha Negi from AiDASH.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

DevOps Engineer
DevOps Engineer

KnowledgeWorks Global Ltd. • Mumbai

On-site
INR 1,800,000 - 3,000,000
Program Manager, AI Data Ops
Program Manager, AI Data Ops

AiDash • Bengaluru

On-site
INR 1,800,000 - 2,400,000
Manager, AI Data Ops
Manager, AI Data Ops

ProducePay • Bengaluru

On-site
INR 450,000 - 750,000
Devops Engineer
Devops Engineer

Larsen & Toubro • Chennai District

On-site
INR 1,000,000 - 1,500,000
Program Manager, AI Data Ops
Program Manager, AI Data Ops

AiDASH, Inc. • Bengaluru

On-site
INR 2,400,000 - 3,400,000
Lead SDE - DevOps
Lead SDE - DevOps

Flourish Ventures • Chennai District

On-site
INR 2,000,000 - 3,000,000
Inclusive and people-first culture
Health & wellness programs
Comprehensive medical insurance
+2
DevOps Engineer
DevOps Engineer

NAVVYASA CONSULTING PRIVATE LIMITED • Gurugram District

On-site
INR 800,000 - 1,200,000
DevOps Engineer
DevOps Engineer

Larsen & Toubro-Vyoma • Chennai District

On-site
INR 2,200,000 - 3,800,000
Senior AI DevOps Engineer
Senior AI DevOps Engineer

Autonomize AI • Bengaluru

On-site
INR 1,500,000 - 2,500,000
Manager, AI Data Ops
Manager, AI Data Ops

AI Chopping Block • India

On-site
INR 2,500,000 - 4,200,000