About Us
We are a tech company specializing in the design and development of cutting-edge, customized server hardware solutions optimized for artificial intelligence and machine learning applications. Our mission is to empower businesses and researchers to accelerate their AI initiatives by providing them with high-performance, scalable, and energy-efficient hardware infrastructure.
As a rapidly growing company at the forefront of AI hardware innovation, we are constantly seeking talented and motivated individuals to join our team. We offer a dynamic and challenging work environment, with opportunities to make a significant impact on the future of AI technology.
You’ll Collaborate With
Compiler engineers, firmware & driver engineers, hardware architects, ML framework engineers and performance & validation teams
What You’ll Own
- The core runtime architecture that bridges the compiler and the hardware
- Runtime execution engine performance and stability
- Device memory management design and efficiency
- Command submission and hardware interaction layer (in collaboration with firmware/driver teams)
- Concurrency and scheduling model for model execution
- Observability tooling (profiling, logging, tracing) within the runtime
- Technical design decisions that ensure long-term scalability and maintainability
Minimum Qualifications
- 5–8+ years of experience in systems software, runtime systems, or performance-critical infrastructure
- Strong proficiency in C/C++and Python
- Solid understanding of fundamentals of operating systems, multi-threading, synchronization, memory management, and cache behavior
- Experience working close to hardware (GPU, accelerator, driver, or firmware environments)
- Familiarity with command queues, execution engines, and DMA-based systems
- Understanding of computational graphs, tensor execution, and memory layouts
- Experience profiling and optimizing latency and throughput in performance-sensitive systems
Preferred Qualifications
- Experience building or architecting a runtime from scratch
- Familiarity with MLIR, LLVM, or compiler-runtime interfaces
- Experience with one or many from CUDA, ROCm, TensorRT, TVM, XLA, ONNX Runtime or similar runtimes
- Experience integrating custom hardware backends into PyTorch, ONNX or OpenXLA
- Knowledge of quantization and mixed-precision inference
- Experience with multi-device or distributed execution
- Background in AI accelerator or semiconductor environments
What Success Looks Like
- You enable reliable end-to-end execution of compiled models on our NPU
- Runtime overhead is minimal and performance targets are consistently met
- Memory management and scheduling are efficient and scalable
- The runtime architecture is robust, maintainable, and extensible
- Compiler, firmware, and hardware teams can build confidently on top of your execution layer
- The system supports real-world AI workloads with stability and observability
Join us in our mission to democratize AI compute — where your firmware expertise becomes the bedrock of tomorrow's AI breakthroughs.