Stand out for this role — generate a tailored resume and cover letter in about a minute.
Oho Group in San Francisco is seeking a Director of Engineering to build and lead the team responsible for turning its frontier-model inference platform into production-grade software and services.
You will own the architecture and delivery of the end-to-end inference stack, including model enablement, runtimes, scheduling, and serving infrastructure, while mentoring an exceptional team and guiding performance decisions.
A heavily funded AI hardware company is building a vertically integrated platform for frontier-model inference, spanning custom silicon, rack-scale systems, interconnects, compilers, runtimes and serving software.
Its first production silicon has returned and the team is now validating rack-scale systems with major AI customers. The platform targets demanding workloads including large mixture-of-experts models, long-context inference and agentic applications, with the goal of substantially improving throughput, latency, power efficiency and cost per token.
The company is hiring a Director of Engineering to build and lead the organisation responsible for turning its custom accelerator into a production-grade inference platform.
You will own the technical direction and delivery of the inference stack across model enablement, distributed execution, runtime scheduling, kernel performance and serving infrastructure. This is a hands-on leadership position requiring someone capable of building an exceptional team while remaining closely involved in architecture and performance decisions.
This is an opportunity to define the inference organisation around a new computing platform rather than inherit an established stack. You will influence both the software and silicon roadmap while helping move a technically ambitious architecture into large-scale customer deployment.
A highly competitive compensation and equity package is available.