HPC Operations Engineer

P2P

Mumbai

On-site

INR 800,000 - 1,200,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

P2P is seeking a dedicated individual for operational support in Linux HPC environments. This role involves providing 24/7 support, managing complex projects, and collaborating across teams to automate tasks and solve problems effectively.

The ideal candidate has 2+ years of experience with Linux systems, a solid foundation in programming languages like Go or Python, and strong communication skills. You will work primarily in the office and participate in an on-call rotation.

Qualifications

  • 2+ years of professional experience with Linux systems.
  • High proficiency with at least one programming/scripting language.
  • Ability to manage complex projects independently and multiple workstreams.

Responsibilities

  • Provide operational support for 24/7 Linux HPC compute and storage.
  • Solve problem reports and questions by members of the research community.
  • Participate in large maintenance operations during evenings and weekends.
  • Collaborate with team members to write code for various infrastructures.

Skills

Linux systems
High performance computing (HPC)
Programming/scripting languages (Go, Python, C)
Collaboration skills
Root cause analysis
Strong verbal and written communication

Job description

Jump Trading Group is committed to world class research. We empower exceptional talents in Mathematics, Physics, and Computer Science to seek scientific boundaries, push through them, and apply cutting edge research to global financial markets. Our culture is unique. Constant innovation requires fearlessness, creativity, intellectual honesty, and a relentless competitive streak. We believe in winning together and unlocking unique individual talent by incenting collaboration and mutual respect. At Jump, research outcomes drive more than superior risk adjusted returns. We design, develop, and deploy technologies that change our world, fund start-ups across industries, and partner with leading global research organizations and universities to solve problems.

We are looking for an adaptable hands‑on individual, passionate about the details and nuances of managing Linux HPC environments at scale, and eager to tackle complex and unpredictable operational work as their primary job function.

What You'll Do:
  • Provide front‑line operational support for 24/7 Linux HPC compute, storage, and interconnects. Technologies involved include RDMA fabrics, parallel filesystems, HPC batch schedulers, FUSE filesystems, internal Jump software, multi‑vendor hardware, cybersecurity requirements, a challenging and unpredictable client workload, and high user expectations.
  • Solve problem reports and questions posed by members of Jump's research community, escalating as needed and managing the entire problem lifecycle.
  • Respond to alerts in a timely fashion.
  • Participate in large, coordinated maintenance operations, including during evenings and weekends.
  • Work on global projects across a wide range of infrastructure.
  • Write code for diagnosing, resolving, and triaging difficult problems and automating frequently performed tasks.
  • Collaborate with team members and across teams to write code and testing infrastructures spanning both new and existing codebases in multiple programming languages.
  • Manage relationships with outside vendors, including traveling both domestically and internationally to meet with current and potential vendors.
  • Implement and support performance monitoring and fault monitoring systems.
  • Develop and improve systems and user documentation.
  • Develop and monitor the tools used to maintain a production computing environment.
  • Provide operational support as primary job function.
  • Adhere to all company cybersecurity and IT policies, including performing all work using only approved hardware and software.
  • Participate in an on‑call rotation.
  • Other tasks as assigned or needed.
  • Work from company office an average of 5 days a week.
  • Must be willing to work a maintenance window of either Friday evening or Saturday morning on a rotationary basis.
Skills You'll Need:
  • A desire for operational work as primary job function.
  • At least 2+ years of professional experience with Linux systems.
  • High performance computing (HPC), including parallel filesystems (e.g., Lustre, GPFS), batch systems (e.g., Slurm, Grid Engine), and high‑performance network interconnects experience is a plus, but not required.
  • High proficiency with at least one programming/scripting language (e.g., Go, Python, C) and ability to learn additional languages quickly.
  • Ability to perform root cause analysis.
  • Strong verbal and written communication skills, including the ability to communicate effectively and efficiently with both coworkers and third‑party vendors.
  • Strong collaboration skills with a willingness to undertake tasks of various technologies and complexities.
  • Ability to independently manage complex projects and multiple workstreams.
  • Strong sense of urgency.
  • Willingness to perform regular operational maintenance work during evenings and weekends and as needed.
  • Ability to work effectively in a busy, open floor plan office environment.
  • Reliable and predictable availability.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

HPC Systems Engineer
HPC Systems Engineer

Jump Trading • Mumbai

On-site
INR 10,163,000 - 12,873,000
Infrastructure Engineer
Infrastructure Engineer

P2P • Mumbai

On-site
INR 1,000,000 - 1,400,000
HPC Engineer
HPC Engineer

PlusWealth Group • Gurugram District

On-site
INR 900,000 - 1,300,000
Medical insurance
Meals at office
Generous paid time off
HPC Admin
HPC Admin

SHI Solutions India Pvt. Ltd. • Maharashtra

On-site
INR 1,000,000 - 1,500,000
HPC Senior System Integrator/System Administrator
HPC Senior System Integrator/System Administrator

GBB • Mumbai

On-site
INR 1,000,000 - 1,500,000
Linux Administrator
Linux Administrator

Pacefin • Gurugram District

On-site
INR 1,800,000 - 3,000,000
Catered breakfast & lunch
Group health insurance
4 weeks of annual leave
+1
Linux - HPC Administrator
Linux - HPC Administrator

Atos SE • Dadri

On-site
INR 900,000 - 1,300,000
Engineer II Systems (HPC)
Engineer II Systems (HPC)

Microchip Technology Inc. • Chennai District

On-site
INR 900,000 - 1,500,000
Lead Software Engineer
Lead Software Engineer

Premium Aerotec • Bengaluru

On-site
INR 2,500,000 - 4,200,000
HPC Operations Engineer
HPC Operations Engineer

NVIDIA • Bengaluru

On-site
INR 1,200,000 - 2,200,000