Senior HPC Scheduler & Grid Reliability Engineer

Qualcomm

Ciudad de México

Presencial

MXN 900.000 - 1.300.000

Jornada completa

Hace 12 días
Generador de candidaturas

Una candidatura completa en un minuto — currículum adaptado y carta de presentación, listos para enviar.

Supera los filtros ATS

Descripción de la vacante

Qualcomm is seeking a Senior HPC Scheduler Operations Engineer to operate, scale, and optimize the IBM Spectrum LSF-based scheduling stack for large-scale EDA workloads. You will work within Engineering IT Hardware Infrastructure — EDA Compute, influencing reliability and performance across sites.

Proven Linux administration, scripting skills, and strong problem-solving are required. Preference for Slurm exposure and container workload experience, with close collaboration to CAD/engineering

Formación

  • Bachelor's degree in Computer Science, Electrical Engineering, or related field, or equivalent practical experience.
  • 5+ years operating and supporting large-scale Linux-based compute infrastructure in HPC or silicon design environments.
  • Strong hands-on experience with IBM Spectrum LSF, including queue configuration, policy tuning, fairshare, job arrays, and multi-cluster operations.
  • Proficiency in SLES administration including OS-level troubleshooting, workload profiling and user environment management.
  • Ability to independently analyze complex system behavior under load; experience with root cause analysis across scheduler, OS, and hardware layers.
  • Proficiency in Python and/or Shell scripting for operational automation, data processing, and tooling development.
  • Demonstrated ability to clearly articulate technical tradeoffs, reliability metrics, and operational status to both engineering and management audiences.

Responsabilidades

  • Manage, scale, and optimize IBM Spectrum LSF job scheduling systems for EDA and HPC compute-intensive workloads across multi-site environments.
  • Administer scheduler configuration including queue policies, resource limits, fairshare, job arrays, and preemption strategies.
  • Analyze scheduler and infrastructure performance data to identify bottlenecks and improve utilization, throughput, and job turnaround time.
  • Perform capacity planning and workload characterization to support SKU-level packing models and tiered job scheduling strategies.
  • Implement and maintain scheduler tuning parameters to minimize dispatch latency and maximize grid efficiency under high-saturation conditions.
  • Troubleshoot and resolve service-impacting issues across scheduler, OS, and workload layers with minimal time-to-resolution.
  • Define and track SLOs for service performance and reliability; partner with customer teams to set and communicate realistic expectations.
  • Implement automation and process improvements to reduce manual toil and prevent recurring incidents.
  • Develop and enforce operational standards, runbooks, and best practices ensuring consistency across all sites.
  • Build and maintain observability systems including metrics pipelines, dashboards, and alerting frameworks for scheduler and compute infrastructure health.
  • Leverage accounting and telemetry data (e.g., LSF stream logs, bjobs, bacct) to quantify workload coverage, grid efficiency, and tier performance.
  • Develop automation tooling (Python, Shell, APIs) to streamline operations, enforce policy guardrails, and surface actionable insights.
  • Produce operational reports and management summaries that combine automated metrics with human operational narratives.
  • Collaborate directly with CAD and engineering teams to clarify workload requirements, translate technical tradeoffs, and drive issues to closure.
  • Communicate scheduler performance metrics, reliability posture, and capacity status clearly to engineering and management audiences.
  • Partner with peer infrastructure teams to coordinate cross-functional changes and align on shared operational standards.

Conocimientos

LSF
Slurm
Linux
Python
Shell scripting
Problem solving
Communication

Educación

Bachelor's degree in CS/EE or related field

Herramientas

IBM Spectrum LSF
Slurm
Docker/Singularity

Descripción del empleo

Qualcomm is seeking a Senior HPC Scheduler Operations Engineer to operate, scale, and optimize the IBM Spectrum LSF-based scheduling stack for large-scale EDA workloads. You will work within Engineering IT Hardware Infrastructure — EDA Compute, influencing reliability and performance across sites.

Proven Linux administration, scripting skills, and strong problem-solving are required. Preference for Slurm exposure and container workload experience, with close collaboration to CAD/engineering

Consigue la evaluación confidencial y gratuita de tu currículum.

o arrastra y suelta tu archivo aquí

Similar jobs

Puestos de trabajo similares que vale la pena comparar

Senior Linux Systems Engineer – Hybrid Cloud HPC
Senior Linux Systems Engineer – Hybrid Cloud HPC

Qualcomm • Ciudad de México

Híbrido
MXN 900.000 - 1.300.000
Remote Linux HPC Administrator & Job Scheduler
Remote Linux HPC Administrator & Job Scheduler

ALTEN Mexico • Santiago de Querétaro

Presencial
MXN 700.000 - 1.000.000
Hybrid scheme
Indeterminate Contract
Professional growth opportunities
Senior Linux & Unix Infrastructure Engineer
Senior Linux & Unix Infrastructure Engineer

Qualcomm • Ciudad de México

Presencial
MXN 650.000 - 900.000
Senior Linux Systems Engineer for HPC & Kubernetes
Senior Linux Systems Engineer for HPC & Kubernetes

Luxoft • Región Centro

Presencial
MXN 300.000 - 450.000
Senior Linux Systems Engineer – Guadalajara
Senior Linux Systems Engineer – Guadalajara

Luxoft • Región Centro

Presencial
MXN 900.000 - 1.200.000
Site Reliability Engineer (SRE) – Regional Multi Project Platform
Site Reliability Engineer (SRE) – Regional Multi Project Platform

Qualcomm • Tijuana

Presencial
MXN 1.214.000 - 1.735.000
Linux Systems Engineer
Linux Systems Engineer

Luxoft • Región Centro

Presencial
MXN 300.000 - 450.000
Senior HPC CAE Analyst | Remote, AWS ParallelCluster Expert
Senior HPC CAE Analyst | Remote, AWS ParallelCluster Expert

ALTEN MÉXICO • Santiago de Querétaro

Presencial
MXN 680.000 - 1.000.000
Seguro de Gastos Médicos Mayores (incl
plan dental y visión)
15 días de aguinaldo
+5
Senior Linux & UNIX Systems Engineer
Senior Linux & UNIX Systems Engineer

Qualcomm • Ciudad de México

Presencial
MXN 700.000 - 900.000
Senior Linux Administrator — HA & Pacemaker Expert
Senior Linux Administrator — HA & Pacemaker Expert

SAP • Monterrey

Presencial
MXN 420.000 - 700.000