58603 Platform Engineer

Cephas Consultancy Services Private Limited

Pune District

On-site

INR 600,000 - 800,000

Full time

14 days+
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

The AI Platform Operations Engineer (L1) role is part of a 24x7 operations team supporting Cephas AI Kitchen platform on on‑prem OpenShift. You monitor dashboards, handle first‑level incidents, perform basic recovery steps, and escalate as needed to L2/L3 teams.

We value foundational Linux, networking and Kubernetes/OpenShift knowledge, strong documentation, and clear communication. This entry‑level position offers exposure to GitOps, runbooks and incident drills while building SRE skills in AI

Qualifications

  • Bachelor's degree in Computer Science, Information Technology, Engineering or related discipline.
  • 2 years of IT operations, infrastructure, cloud or application support experience.
  • Basic Linux command-line knowledge.
  • Basic networking knowledge including IP addressing, DNS, ports and connectivity troubleshooting.
  • Basic understanding of containers and Kubernetes concepts.
  • Ability to follow technical procedures accurately.
  • Good written and verbal communication skills.
  • Willingness and ability to work in a 24×7 shift environment.

Responsibilities

  • Monitor platform and application dashboards, alerts and operational mailboxes.
  • Identify availability, infrastructure, application and service degradation events.
  • Acknowledge alerts promptly and determine initial severity based on established procedures.
  • Maintain accurate shift handover and operational records.
  • Perform basic network connectivity checks including ping, DNS resolution and endpoint connectivity.
  • Check OpenShift/Kubernetes resource status including nodes, pods, deployments and services.
  • Review basic application and platform logs to identify common failure conditions.
  • Check infrastructure and service health using approved dashboards and operational tools.
  • Execute documented recovery procedures and operational runbooks.
  • Restart or redeploy affected workloads using approved GitOps processes.
  • Verify service recovery through dashboards, health checks and application endpoints.
  • Escalate when recovery procedures are unsuccessful or when an incident falls outside the approved operating scope.
  • Create and maintain incident tickets with accurate timestamps, symptoms, actions and observations.
  • Engage the appropriate application, platform, infrastructure or network teams based on established escalation procedures.
  • Support incident bridges by providing operational information and executing actions requested by L2/L3 engineers.
  • Ensure effective handover of unresolved incidents between shifts.
  • Follow established SOPs, runbooks and change-management processes.
  • Document newly encountered symptoms and successful troubleshooting steps.
  • Highlight recurring alerts or operational problems to senior platform engineers.
  • Participate in operational drills and recovery exercises.

Skills

Linux basics
Networking basics
Kubernetes/OpenShift
GitOps awareness
SOP adherence
Communication

Education

Bachelors in Computer Science/IT/Engineering

Tools

OpenShift
Kubernetes
Grafana
Prometheus
Git

Job description

About this position

Positions:2 Full Time
Experience
5 - 10 Years


Job Description

AI Platform Operations Engineer (L1)


Position Summary

The AI Platform Operations Engineer is part of the 24×7 operations team supporting the Central AI Kitchen platform and its AI services.


The role is responsible for continuous monitoring, first-level incident detection and troubleshooting, execution of approved recovery procedures, and timely escalation to the appropriate platform, infrastructure and application support teams.


The platform consists of AI services running on Red Hat OpenShift, hosted on on‑prem infrastructure. The engineer will operate primarily using established dashboards, alerts, runbooks and GitOps‑based recovery procedures.


This is an entry‑level operations role designed for engineers who have foundational Linux, networking and Kubernetes/OpenShift knowledge and are interested in developing deeper platform engineering and SRE capabilities.


Key Responsibilities


  • Monitor platform and application dashboards, alerts and operational mailboxes.

  • Identify availability, infrastructure, application and service degradation events.

  • Acknowledge alerts promptly and determine initial severity based on established procedures.

  • Maintain accurate shift handover and operational records.

  • Perform basic network connectivity checks including ping, DNS resolution and endpoint connectivity.

  • Check OpenShift/Kubernetes resource status including nodes, pods, deployments and services.

  • Review basic application and platform logs to identify common failure conditions.

  • Check infrastructure and service health using approved dashboards and operational tools.

  • Execute documented recovery procedures and operational runbooks.

  • Restart or redeploy affected workloads using approved GitOps processes.

  • Verify service recovery through dashboards, health checks and application endpoints.

  • Escalate when recovery procedures are unsuccessful or when an incident falls outside the approved operating scope.

  • Create and maintain incident tickets with accurate timestamps, symptoms, actions and observations.

  • Engage the appropriate application, platform, infrastructure or network teams based on established escalation procedures.

  • Support incident bridges by providing operational information and executing actions requested by L2/L3 engineers.

  • Ensure effective handover of unresolved incidents between shifts.

  • Operational Procedures

    • Follow established Standard Operating Procedures (SOPs), runbooks and change-management processes.

    • Document newly encountered symptoms and successful troubleshooting steps.

    • Highlight recurring alerts or operational problems to senior platform engineers.

    • Participate in operational drills and recovery exercises.




Scope of Authority

Engineer may independently:


  • Acknowledge and investigate alerts.

  • Perform approved diagnostic commands and health checks.

  • Execute documented L1 recovery procedures.

  • Open incidents and engage predefined support teams.

  • Escalate incidents based on severity and runbook criteria.


Engineer must escalation:


  • Changes requiring manual modification of production infrastructure.

  • Unauthorised configuration or source-code changes.

  • Security incidents or suspected compromises.

  • Incidents where documented recovery procedures fail.

  • Major incidents requiring business or management decisions.


Qualifications & Experience


  • Bachelor's degree in Computer Science, Information Technology, Engineering or related discipline.

  • 2 years of IT operations, infrastructure, cloud or application support experience.

  • Basic Linux command-line knowledge.

  • Basic networking knowledge including IP addressing, DNS, ports and connectivity troubleshooting.

  • Basic understanding of containers and Kubernetes concepts.

  • Ability to follow technical procedures accurately.

  • Good written and verbal communication skills.

  • Willingness and ability to work in a 24×7 shift environment.


Good to Have


  • Exposure to Kubernetes or Red Hat OpenShift.

  • Exposure to Git and GitOps concepts.

  • Familiarity with monitoring tools such as Grafana and Prometheus.

  • Basic scripting experience with Bash or Python.

  • Familiarity with incident-management or ITIL processes.

  • Exposure to cloud or data‑centre infrastructure.

  • Interest in AI/ML infrastructure and GPU-based platforms.

  • Systematic troubleshooting

  • Ability to remain structured during incidents

  • Clear communication and escalation

  • Discipline in following operational procedures

  • Willingness to learn

  • Teamwork across geographically distributed support teams


any graduate

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Platform Engineer – Pune
Platform Engineer – Pune

Umanist Staffing LLC • Pune District

On-site
INR 900,000 - 1,400,000
Associate - AI Tooling Ops - Platform Reliability Engineer
Associate - AI Tooling Ops - Platform Reliability Engineer

Jefferies Financial Group Inc. • Pune District

On-site
INR 1,600,000 - 2,800,000
OpenShift Admin - L3 Level (with OpenShift AI & CLI Experience)
OpenShift Admin - L3 Level (with OpenShift AI & CLI Experience)

Paramaah It Services • India

Remote
INR 3,000,000 - 5,000,000
Certification sponsorship
Continuous learning opportunities
Collaborative environment
+1
CloudOps Engineer (L3)
CloudOps Engineer (L3)

Larsen & Toubro-Vyoma • Chennai District

On-site
INR 2,500,000 - 4,500,000
OpenShift Platform Engineer - DevOps
OpenShift Platform Engineer - DevOps

GSS Group • Bengaluru

On-site
INR 4,000,000 - 6,500,000
OpenShift Platform Engineer - DevOps
OpenShift Platform Engineer - DevOps

Global Software Solutions Group • Bengaluru

On-site
INR 3,500,000 - 6,500,000
CloudOps Engineer (L2)
CloudOps Engineer (L2)

Larsen & Toubro • Mumbai

On-site
INR 1,400,000 - 2,100,000
OpenShift & Linux Administrator
OpenShift & Linux Administrator

Zycus Inc. • India

On-site
INR 2,800,000 - 5,000,000
Senior DevOps Engineer - (OpenShift & CI/CD)
Senior DevOps Engineer - (OpenShift & CI/CD)

Global Software Solutions Group • Bengaluru

On-site
INR 1,200,000 - 2,100,000
Technology Services Engineer III
Technology Services Engineer III

Jobtailor • Chennai District

On-site
INR 1,400,000 - 2,000,000