A complete application in a minute — tailored resume and cover letter, ready to send.
Modular is seeking a senior leader to head a high‑impact team driving LLM inference on Modular Cloud. You will partner across GTM, Product, and Engineering to redefine inference performance, shaping the product direction and building scalable systems that meet demand.
You’ll own the technical direction to achieve Pareto‑optimal performance across GPUs and ASICs, translating customer workloads into actionable engineering roadmaps while growing and mentoring a top-tier team.
Sound judgment in evaluating technical tradeoffs and setting priorities, paired with strong communication and technical leadership skillsA track record of shipping durable, reusable software tools and libraries adopted across teams and functions, and of guiding a team to do the sameThe ability to translate ambiguous customer and product needs into focused engineering directionCreativity and curiosity in solving complex problems, a collaborative and team oriented mindset, and alignment with our culture5+ years in distributed systems or performance engineering, including experience leading or managing engineering teamsHands on background in GPU kernel programming, inference engine internals, or distributed inference architecturesExperience with Kubernetes and cloud native ecosystemsFamiliarity with modern LLM architectures and the latest inference optimization techniques