Turn this role into an interview — a resume and cover letter built around what this employer wants.
MulticoreWare, Inc. is seeking a Senior Software Engineer to develop and optimize deep learning models for efficient inference across CPU, GPU, and edge devices.
You will implement and tune quantization techniques and model compression to boost latency and throughput. The role requires strong Python and C++ skills, hands-on experience with CNNs, Transformers, and LLMs, and deployment pipelines using PyTorch/ONNX.
MulticoreWare is a global software solutions & products company with its HQ in San Jose, CA, USA. With worldwide offices, it serves its clients and partners in North America, EMEA and APAC regions. Started by a group of researchers, MulticoreWare has grown to serve its clients and partners on HPC & Cloud computing, GPUs, Multicore & Multithread CPUS, DSPs, FPGAs and a variety of AI hardware accelerators.
MulticoreWare was founded by a team of researchers that wanted a better way to program for heterogeneous architectures. With the advent of GPUs and the increasing prevalence of multi-core, multi-architecture platforms, our clients were struggling with the difficulties of using these platforms efficiently.
We started as a boot-strapped services company and have since expanded our portfolio to span products and services related to compilers, machine learning, video codecs, image processing and augmented/virtual reality. Our hardware expertise has also expanded with our team; we now employ experts on HPC and Cloud Computing, GPUs, DSPs, FPGAs, and mobile and embedded platforms. We specialize in accelerating software and algorithms, so if your code targets a multi-core, heterogeneous platform, we can help.
Develop and optimize deep learning models (CNNs, LLMs, MoE) for efficient inference across CPU, GPU, and hardware accelerators / edge devices.