AI DevOps Engineer

nCircle Tech Co

Pune District

On-site

INR 1,500,000 - 3,000,000

Full time

4 days ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

nCircle Tech Private Limited in India is seeking an AI DevOps Engineer to own deployment, observability, and reliability of our agentic AI platform. You will build CI/CD pipelines, manage AWS-based environments, and establish security and cost controls to make releases safe and repeatable.

The role emphasizes telemetry, SLOs, on-call readiness, and an opinionated platform template that other teams will follow, enabling scalable production AI across desktop, mobile, and cloud.

Qualifications

  • 6–8 years in DevOps, SRE, platform, or infrastructure engineering with production systems ownership.
  • Deep CI/CD experience with pipelines that gate, test, and safely release production software.
  • Strong cloud operations, preferably AWS, with infrastructure-as-code approaches.
  • Hands-on observability experience: metrics, logging, tracing, dashboards, and alerting.
  • SRE fundamentals: SLOs, incident response, on-call, capacity planning, blameless postmortems.
  • Enough software fluency to read application code and debug deployments.

Responsibilities

  • Own CI/CD for the platform and establish the default safe path for deployments.
  • Manage cloud infrastructure as code (AWS) with Terraform/CDK and cost controls.
  • Run the reliability practice: SLOs, on-call, incident response, and postmortems.
  • Build agent telemetry and observability: traces, cost accounting, dashboards and alerts.
  • Set the operational standard by example; create reusable pipelines and observability patterns.

Skills

CI/CD pipelines
AWS cloud operations
Observability
SRE fundamentals
Software fluency
DevOps experience

Tools

Terraform
CDK
GitHub Actions
GitLab CI
OpenTelemetry
Prometheus/Grafana

Job description

nCircle Tech Private Limited (Incorporated in 2012) empowers passionate innovators to createimpactful 3D visualization software for desktop, mobile and cloud. Our domain expertise in CADand BIM customization is driving automation with the ability to integrate advanced technologieslike AI/ML and AR/VR, which empowers our clients to reduce time to market and meet businessgoals. nCircle has a proven track record of technology consulting and advisory services for AECand Manufacturing industry across the globe. Our team of dedicated engineers, partnerecosystem and industry veterans are on a mission to redefine how you design and visualize.

Job Description
About the role :-

We are hiring an AI DevOps Engineer to own how our AI systems get deployed, stay up, and stay observable. You will own the infrastructure and operational substrate fornCircle Tech's agentic platform — the CI/CD pipelines, the cloud environments, the release machinery, the reliability practice, and the telemetry layer that turns opaque agent behavior into something you can measure, alert on, and debug.

Our engineers build the AI agents; you make shipping them safe, repeatable, and observable. Right now most of our apps have no CI/CD, the infrastructure is fragmented, and the agentic systems being stood up have little more than print statements for observability. Building that operational foundation is the job.

Agentic systems fail differently from ordinary services — nondeterministic output, silent quality drift, runaway tool‑call loops, and cost that spikes without warning. Standard DevOps is necessary but not sufficient. This role exists because someone has to own reliability and telemetry for systems that don’t fail the way the runbooks assume.

What you’ll do :-

Own CI/CD for the platform. Build the pipelines that take AI systems from commit to production — automated testing, evaluation gates, security and dependency checks, controlled and canary releases, and one-command rollback. Most of our apps have no pipeline today; you will establish the pattern and make the safe path the default path.

Manage cloud infrastructure as code. Own the cloud environments (AWS primarily) the platform runs on — provisioning, networking, secrets, environment parity, and cost controls — as versioned, reviewable infrastructure-as-code, not hand-tuned consoles. You are accountable for environments that are reproducible, least-privilege by default, and cheap to stand up and tear down.

Run the reliability practice. Own production reliability: SLOs, on-call and incident response, capacity and cost management, self‑healing loops that detect and recover from failures, and blameless post‑incident review. You will help define what “up” and “healthy” even mean for a nondeterministic system.

Build the agent telemetry and observability layer. Instrument the platform so agent behavior is legible: structured traces of agent runs and tool calls, token and cost accounting, latency and success metrics, output‑quality tracking over time, and the dashboards and alerts that surface a regression before a user does. When an agent misbehaves in production, the telemetry you built is how the team finds out and figures out why.

Set the operational standard by example. On a small, high‑leverage team, your pipelines and dashboards are the template. You establish the deployment patterns others adopt, the observability every new system gets wired into by default, and the operational discipline that lets a lean team run production systems well.

What we’re looking for :-
  • 6–8 years in DevOps, SRE, platform, or infrastructure engineering, with a track record of running production systems you were accountable for
  • Deep CI/CD experience — you have built and owned pipelines (GitHub Actions, GitLab CI, or similar) that gate, test, and safely release real production software
  • Strong cloud operations, ideally AWS — provisioning, networking, secrets, and cost management as infrastructure-as-code (Terraform, CDK, or similar)
  • Hands‑on observability experience — metrics, logging, distributed tracing, dashboards, and alerting (OpenTelemetry, Prometheus/Grafana, Datadog, CloudWatch, or similar) — and the instinct to instrument first
  • SRE fundamentals: SLOs, incident response, on‑call, capacity planning, and blameless postmortems
  • Enough software fluency to read application code, wire telemetry into it, and debug a failing deploy without waiting for someone else
Nice to have :-
  • Experience operating AI or LLM systems in production — token and cost accounting, prompt and evaluation‑score tracking, or LLM observability tooling (LangSmith, Langfuse, Arize, or similar)
  • Familiarity with the failure modes of nondeterministic systems: quality drift, runaway loops, cost spikes, non‑reproducible output
  • Experience with Databricks or a similar lakehouse platform, and with tool‑integration layers such as MCP
  • Container and orchestration experience (Docker, Kubernetes, or serverless equivalents)Experience in a non‑software-company engineering organization — internal tools, corporate IT transformation, or similarExperience standing up an internal platform or golden‑path deployment pattern that other teams adopted
    Why this role :-

    You will build the operational foundation an entire organization’s AI runs on — the pipelines, the environments, and the telemetry — with a clear mandate and a direct line to the Director of APEX. The platform is early, and that is the appeal. You are not tuning someone else’s mature platform; you are building the deployment and observability substrate nCircle Tech will run AI on for the next decade, and defining what running AI in production looks like here.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

AI DevOps Engineer
AI DevOps Engineer

nCircle Tech Co • Pune District

On-site
INR 1,200,000 - 2,000,000
Tech Lead – Agentic AI Platform
Tech Lead – Agentic AI Platform

Multiscale AI • Hyderabad

On-site
INR 3,600,000 - 6,000,000
Senior Software Engineer - AI Platform Engineer
Senior Software Engineer - AI Platform Engineer

CloudBees • Chennai District

On-site
INR 3,500,000 - 5,500,000
DevOps , AWS & AI Agent Engineer
DevOps , AWS & AI Agent Engineer

Theepaan Technologies • Viluppuram

On-site
INR 1,400,000 - 2,500,000
AI Platform Engineer — Agentic SDLC
AI Platform Engineer — Agentic SDLC

Sutherland • Bengaluru

On-site
INR 1,800,000 - 2,400,000
Director of Software Engineering - Java/Python, AI
Director of Software Engineering - Java/Python, AI

JPMorgan Chase & Co. • Bengaluru

On-site
INR 6,000,000 - 9,000,000
Senior DevOps Engineer (CI/CD & AI-Enabled Platform)
Senior DevOps Engineer (CI/CD & AI-Enabled Platform)

Regnology • Pune District

On-site
INR 3,000,000 - 6,000,000
Senior AI Applications Engineer
Senior AI Applications Engineer

GE HealthCare • Bengaluru

On-site
INR 1,500,000 - 2,700,000
AI Engineer
AI Engineer

AlgoLeap Technologies Pvt Ltd. • Hyderabad

On-site
INR 1,400,000 - 2,800,000
Software Engineer III
Software Engineer III

Arcadia Power, Inc. • Chennai District

On-site
INR 2,800,000 - 5,600,000
Employee stock options
Hybrid work model
Medical insurance (self + family)
+2