We are seeking a HPC/Kubernetes SW Engineer to design, develop, and scale high-performance software systems supporting distributed computing workloads. This role focuses on building and optimizing solutions that run on HPC on-prem clusters and cloud/on-prem hybrid environments, with an emphasis on performance, scalability, and maintainability.
You will be a key member of the R&D team developing next-generation HPC-enabled applications and infrastructure supporting simulation, AI/ML, and data‑intensive workloads.
Responsibilities
- Design, develop, and optimize high-performance software systems using C++ and Python for distributed and compute-intensive workloads.
- Build, deploy, and maintain applications on HPC clusters using workload schedulers (e.g., Slurm, LSF, PBS).
- Architect and implement containerized solutions using Docker and Kubernetes for scalable distributed execution.
- Develop and optimize software on Linux-based systems, including performance tuning, debugging, and system-level troubleshooting.
- Contribute to system design, code reviews, and architectural decisions across distributed systems.
- Profile and improve performance, scalability, and resource utilization of HPC applications.
- Collaborate with cross-functional teams (DevOps, infrastructure, product) to deliver production-grade solutions.
- C++ (performance‑critical systems, memory management, multithreading).
- Python (automation, orchestration, tooling, or data workflows).
- Hands‑on experience with HPC clusters, including:
- Distributed workload execution and scaling
- Experience with containerization and orchestration:
- Docker for building and packaging applications
- Kubernetes for deployment, scaling, and orchestration
- Performance tuning and troubleshooting
- Experience designing or building distributed or cloud-based systems
- Strong problem‑solving skills and ability to work on complex systems
Preferred / Nice‑to‑Have Skills
- Experience with Angular (preferred for UI/dashboard development).
- Familiarity with:
- Cloud environments (AWS, Azure, or GCP)
- Networking or distributed communication (e.g., gRPC, REST APIs)
- GPU computing or AI/ML workloads in HPC environments
What We're Looking For
- How you used C++ and/or Python in production systems (not just coursework).
- Your direct involvement with HPC clusters (setup, usage, scaling, troubleshooting, etc.).
- Concrete examples of using Docker and/or Kubernetes (e.g., deployment, orchestration, CI/CD pipelines).
- Practical experience working in Linux environments (debugging, performance tuning, scripting).
Resumes that only list technologies without describing actual usage will not be considered competitive.
Education
Bachelor’s or Master’s degree in Computer Science, Software Engineering, or related discipline (or equivalent experience).
Equal Opportunity Employer
Keysight is an Equal Opportunity Employer.