Meta is seeking a Software Engineer to join the MTIA (Meta Training & Inference Accelerator) Software Tooling team, which develops and maintains the tooling ecosystem for Meta's in-house AI accelerator ASICs. The Tooling team provides debugging, profiling, memory analysis, and monitoring capabilities for the whole MTIA Ecosystem, advancing ML accelerator tooling by leveraging Meta's full-stack ownership from silicon specs to fleet observability.
In this role, you will be a senior technical contributor responsible for designing and building developer tools that help engineers debug, profile, measure, and monitor AI workloads running on MTIA hardware at scale. You will work at the intersection of compilers, runtime, hardware, and ML frameworks, collaborating with cross-functional partners to deliver a high-quality developer experience for Meta's custom AI accelerators.
- Bachelor's degree in Computer Science, Computer Engineering, relevant technical field, or equivalent practical experience6+ years of experience in systems software engineering, performance engineering, developer tooling, or a closely related field
- Experience building debugging, profiling, or diagnostic tools for complex software/hardware systems
- Proficiency in C++ and Python, including low-level systems programming and scripting for tool automation
- Experience working across multiple layers of a system stack (compiler, runtime, OS/driver, hardware)
- Experience leading the technical design and delivery of tooling or infrastructure projects from inception through production deployment
- Experience using data-driven methods and experimentation to evaluate and validate tooling effectiveness and systems performance improvements
- Familiarity with ML framework internals (PyTorch graph execution, torch.compile, operator dispatch) and AI compiler stacks (MLIR, LLVM, TVM, Triton)
- Demonstrated ability to integrate AI tools to optimize/redesign workflows and drive measurable impact (e.g., efficiency gains, quality improvements)
- Experience with accelerator ecosystems (GPU/CUDA, TPU, custom ASICs) including performance profiling, memory analysis, and runtime debugging using their toolchains (cuda-gdb, nsight-compute, nsight-systems, cuda-memcheck)
- Demonstrated cross-stack debugging ability, including Linux kernel and driver-level debugging, with capacity to trace issues across application, OS, and hardware boundaries
- Demonstrated ongoing AI skill development (e.g., prompt/context engineering, agent orchestration) and staying current with emerging AI technologies
- Experience with distributed systems debugging or profiling (multi-device, multi-node/multi-rank)
- Experience with Linux debugging and profiling infrastructure (gdb, perf, eBPF, ftrace, coredump analysis, hardware performance counters) and familiarity with binary formats and debugging metadata (ELF/DWARF)
- Experience adhering to and implementing responsible, ethical AI practices (e.g., risk assessment, bias mitigation, quality and accuracy reviews)
7+ years of experience in systems software, developer tooling, or accelerator software development (or equivalent with advanced degree)- Track record of building developer tools adopted by large engineering populations, ideally with contributions to open-source projects (gdb, LLVM sanitizers, Valgrind, Triton, etc.)