Cognizant Hiring For HPC , Github Pipelines, AWS

Cognizant

Hyderabad, Chennai District, Bengaluru

Hybrid

INR 1,500,000 - 2,800,000

Full time

8 days ago
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Cognizant is hiring for HPC administration and AWS DevOps roles across Hyderabad, Chennai, Bangalore, Coimbatore, Pune, and Kolkata. The position requires 5 to 8 years of experience and focuses on deploying and managing high-performance computing clusters, CI/CD pipelines with GitHub, and scalable AWS-based infrastructure.

You will work in a fast-paced environment, optimize workloads, configure schedulers like Slurm and LSF, ensure InfiniBand connectivity, and collaborate with researchers to

Qualifications

  • Strong Linux administration experience.
  • Proficient in Python scripting.
  • Experience with HPC clusters and CI/CD pipelines.
  • Knowledge of AWS DevOps practices.
  • Experience with workload managers (LSF, SLURM, PBS).

Responsibilities

  • Deploy HPC clusters and automate with CI/CD pipelines.
  • Configure workload managers and resources.
  • Maintain HPC storage and interconnects.
  • Provide user support and performance tuning.
  • Ensure security policies and monitoring.

Skills

Linux
Python
CI/CD
AWS
HPC
Networking

Tools

GitHub
AWS
Slurm
Lustre
GPFS
InfiniBand
Kubernetes

Job description

Cognizant is hiring HPC , Github Pipelines, AWS
Primary & Mandatory Skill: HPC administration & AWS DevOps
Yrs of exp : 5 years to 8 Years
Location : Chennai, Kolkata, Bangalore, Coimbatore, Pune, Hyderabad
Shift timing: 1pm IST to 11pm IST
Detailed JD
Scope of Work
HPC Cluster Deployment
  • Automate the deployment process of HPC clusters using CI/CD pipelines by utilizing GitHub pipeline and AWS Systems Manager.
  • Implement CI/CD pipelines to manage and deploy updates to the HPC cluster efficiently.
  • Set up and configure HPC clusters to meet specific requirements and workloads.
  • Manage and maintain HPC hardware components such as CPUs and GPUs, along with the necessary software.
  • Conduct regression testing to verify the functionality and performance of non-GXP HPC clusters.
Workload Scheduler Management
  • Install and configure workload managers and schedulers like LSF, SLURM, and PBS Pro.
  • Manage the addition and removal of compute nodes and adjust the priority of master and slave nodes.
  • Develop and manage resource policies and rules to optimize cluster performance.
  • Configure and allocate resources such as CPU and memory, and profile applications for optimal performance.
  • Address and resolve issues related to schedulers, daemons, and license servers.
Network and High-Performance Connectivity Management
  • Install and configure HPC interconnect networks.
  • Design and configure the network topology for HPC clusters.
  • Ensure the maintenance and monitoring of InfiniBand connectivity.
  • Resolve connectivity issues related to InfiniBand, RoCE, and Ethernet.
Monitoring and Reports
  • Produce daily health check reports for the HPC cluster.
  • Automate monitoring scripts to streamline the monitoring process.
  • Conduct periodic reviews of reports and audit trails.
OS Administration and Management
  • Install and configure operating systems for HPC clusters.
  • Address OS-related issues such as CPU, memory, and SWAP utilization, and perform application file system cleanup.
  • Ensure application service continuity by performing pre and post checks from both OS and application perspectives during planned and unplanned outages.
Applications and Tools
  • Install HPC libraries and tools such as MPI and compilers.
  • Install and configure HPC applications, both commercial off-the-shelf (COTS) and open source, and manage packages using Spack.
  • Apply patches and upgrades to HPC applications.
  • Resolve issues related to HPC applications.
HPC Storage Management
  • Administer and configure HPC storage systems.
  • Oversee the administration of HPC file systems.
  • Monitor and troubleshoot HPC storage systems.
  • Manage backup and tape library systems.

Below is the key responsibility, essential skills of the resources we will deploy.

Key Responsibilities
  • Cluster Management: Install, configure, and maintain compute nodes, GPUs (NVIDIA), high-speed storage (Lustre, GPFS), and interconnects (InfiniBand, RoCE).
  • Performance Tuning: Optimize scientific applications, kernels, and workflows for maximum throughput, scalability, and minimal queue times.
  • User Support: Act as a technical expert for researchers, debugging jobs, resolving complex issues, and providing training on tools and best practices.
  • Software Management: Manage workload managers (Slurm, LSF), schedulers, software licensing (FlexLM), OpenPBS, containers (Singularity), and compilers.
  • Infrastructure: Administer high-speed interconnects (InfiniBand), storage (Lustre, CEPH), and potentially cloud/hybrid solutions.
  • Implement and manage monitoring (Grafana, Prometheus) and orchestration tools (Slurm, Kubernetes).
  • Automation: Develop scripts (Python, Ansible) for provisioning, monitoring, and automating routine tasks.
  • Security & Policy: Implement and enforce security policies, manage user access, and oversee lifecycle management.
Essential Skills & Qualifications
  • Technical Expertise: Strong Linux, Python, scripting (Ansible, Terraform), HPC schedulers (Slurm), networking (InfiniBand), and GPU computing.
  • HPC Domain Knowledge: Experience with parallel file systems, workload management, and performance analysis tools.
  • Problem Solving: Excellent analytical and debugging skills for complex distributed systems.
  • Communication: Ability to explain complex technical issues to scientists and non-technical stakeholders.
  • Experience: Hands-on experience in data centers, managing large clusters, and supporting diverse scientific/AI workloads

Team will have knowledge of Gilead systems and AWS CICD pipelines.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Cognizant Hiring For Sr. HPC Engineer
Cognizant Hiring For Sr. HPC Engineer

Cognizant • Pune District, Chennai District, Bengaluru

On-site
INR 3,000,000 - 6,000,000
Sr. HPC ENGINEER
Sr. HPC ENGINEER

Cognizant • Hyderabad

On-site
INR 800,000 - 1,200,000
Lead HPC Engineer
Lead HPC Engineer

Clovertex • Hyderabad

On-site
INR 2,000,000 - 3,000,000
HPC Admin
HPC Admin

5 Star Recruitment • Chennai District

On-site
INR 2,000,000 - 4,000,000
HPC ADMINISTRATOR
HPC ADMINISTRATOR

Vrinda International • Bengaluru

On-site
INR 1,260,000 - 1,540,000
HighPerformance Computing ( HPC) Administrator
HighPerformance Computing ( HPC) Administrator

VIT • Chennai District

On-site
INR 900,000 - 1,300,000
HPC Engineer
HPC Engineer

Whiteblue • Chennai

On-site
INR 1,500,000 - 2,500,000
Linux System Administrator
Linux System Administrator

SISL Global • Chennai District

On-site
INR 800,000 - 1,200,000
Senior HPC Engineer SME/Architect
Senior HPC Engineer SME/Architect

Tata Consultancy Services • Hyderabad, Chennai District, Bengaluru

On-site
INR 1,800,000 - 3,000,000
HPC Engineer
HPC Engineer

Clovertex • Hyderabad

On-site
INR 1,200,000 - 2,400,000