Get more replies from employers
Send a job-specific resume in minutes.
Oho Group is hiring a senior architect to shape the software stack for a programmable accelerator platform. You will connect models, frameworks, compilers, runtimes and hardware from high-level workload behavior to kernel execution and code generation.
You will work across graph capture, MLIR/LLVM lowering, runtime and driver architecture, with a focus on SIMT execution, memory hierarchy and performance profiling, guiding senior engineers through design reviews and pre-silicon validation.
I’m working with a well-funded AI compute company building a programmable accelerator platform and the complete software stack required to run modern AI workloads efficiently.
They are hiring a senior architect to shape the full path from models and frameworks through graph capture, compiler, runtime and driver layers, down to kernels and hardware execution. This is a rare role for someone who can connect high-level workload behaviour with low-level GPU architecture and code generation.
You will work across:
They are looking for:
Experience with custom accelerators, PyTorch compilation, XLA, Triton, quantization, performance modelling or pre-silicon software development would be particularly valuable.
This is an opportunity to define the software architecture for a new compute platform as it moves from a working stack into large-scale production.