Staff Engineer, ML Profiling Tools

Engg

San Jose (CA)

On-site

USD 170,000 - 260,000

Full time

6 days ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Benefits offered by this job

4+ weeks PTO
Fertility care stipend
Medical travel support
Virtual therapy sessions
Charitable giving match

Job summary

Engg is seeking a senior engineer to design and build interactive web analytics tools for AI/ML workloads in a daily onsite role in San Jose, CA. You will render large telemetry datasets into clear visualizations and collaborate with hardware and software teams to translate profiling data into practical workflows.

The role demands extensive experience with data visualization, developer tooling, and performance profiling across CPUs, GPUs, and accelerators.

Qualifications

  • Experience designing rich, interactive visualizations for developers.
  • Ability to translate large telemetry data into intuitive visuals.
  • Strong experience with web-based UI tools and dashboards.

Responsibilities

  • Develop and maintain interactive web analytics tools for AI/ML workloads.
  • Create visualizations like flame graphs and call trees for large datasets.
  • Collaborate with compiler engineers to translate profiling metrics into workflows.
  • Improve developer experience for ML performance profiling and diagnostics.
  • Communicate with stakeholders to deliver systems on time and budget.

Skills

Data visualization
Web UI
Developer tools
Telemetry analysis
Profiling tools
Computing architectures
Machine learning frameworks
Collaboration
Problem solving

Education

Bachelor's degree
Master's degree
PhD

Tools

Perfetto
Chrome Tracing
NVIDIA Nsight
PyTorch Profiler
TensorBoard
eBPF

Job description

Advancing the World’s Technology Together Our technology solutions power the tools you use every day--including smartphones, electric vehicles, hyperscale data centers, IoT devices, and so much more. Here, you’ll have an opportunity to be part of a global leader whose innovative designs are pushing the boundaries of what’s possible and powering the future. We believe innovation and growth are driven by an inclusive culture and a diverse workforce. We’re dedicated to empowering people to be their true selves. Together, we’re building a better tomorrow for our employees, customers, partners, and communities. The AGI (Artificial General Intelligence) Computing Lab is dedicated to solving the complex system-level challenges posed by the growing demands of future AI/ML workloads. Our team is committed to designing and developing scalable platforms that can effectively handle the computational and memory requirements of these workloads while minimizing energy consumption and maximizing performance. To achieve this goal, we collaborate closely with both hardware and software engineers to identify and address the unique challenges posed by AI/ML workloads and to explore new computing abstractions that can provide a better balance between the hardware and software components of our systems. Additionally, we continuously conduct research and development in emerging technologies and trends across memory, computing, interconnect, and AI/ML, ensuring that our platforms are always equipped to handle the most demanding workloads of the future. By working together as a dedicated and passionate team, we aim to revolutionize the way AI/ML applications are deployed and executed, ultimately contributing to the advancement of AGI in an affordable and sustainable manner. Join us in our passion to shape the future of computing! Location: Daily onsite presence at our San Jose, CA office / U.S. headquarters in alignment with our Flexible Work policy.

What You’ll Do
  • Develop and maintain highly interactive, large-scale and complex web-based visual analytics tools on AI/ML workloads.
  • Build interactive data visualization (flame graphs, call trees, timeline charts) capable of fluidly rendering massive, multi-dimensional infrastructure telemetry datasets.
  • Partner closely with compiler engineers and hardware architects to translate highly complex performance profiling metrics into clear, intuitive developer workflows.
  • Analyze and optimize the developer experience for machine learning performance profiling, designing intuitive tooling and workflows to help engineers diagnose and eliminate inference bottlenecks.
  • Communicate effectively with stakeholders, including users, partners, and management, to ensure that the systems are delivered on time and within budget.
  • Complete other responsibilities as assigned.
What You Bring
  • Bachelor's with 10+ years, or Master's with 8+ years, or PhD's with 5+ years of industry experience.
  • Abundant experience designing and implementing rich, interactive data visualizations and user interfaces using modern web technologies.
  • Strong background or interest in creating developer tools, IDE extensions, or complex diagnostic dashboards.
  • Proven ability to translate massive, unstructured, or multidimensional telemetry/profiling data into intuitive, human-readable visual representations.
  • Strong track record of designing workflows specifically for technical users (software engineers, data scientists, or researchers), focusing on minimizing cognitive load and streamlining root-cause diagnosis.
  • Hands-on experience with, or a strong curiosity about, systems profiling and tracing tools (e.g., Perfetto, Chrome Tracing, NVIDIA Nsight Systems/Compute, PyTorch Profiler, TensorBoard, eBPF).
  • Baseline understanding of computing architectures (CPUs, GPUs, TPUs, interconnects/networking) and typical performance bottlenecks (memory bandwidth, compute utilization, synchronization stalls).
  • Bonus: Working knowledge of modern machine learning frameworks (e.g., PyTorch, JAX, TensorFlow) and execution paradigms (LLM inference, pipeline/tensor parallelism, kernel dispatch).
  • Excellent problem-solving skills and ability to think critically and creatively.
  • You’re inclusive, adapting your style to the situation and diverse global norms of our people.
  • An avid learner, you approach challenges with curiosity and resilience, seeking data to help build understanding.
  • You’re collaborative, building relationships, humbly offering support and openly welcoming approaches.
  • Innovative and creative, you proactively explore new ideas and adapt quickly to change.
What We Offer

The pay range below is for all roles at this level across all US locations and functions. Pay within this range varies by work location and may also depend on job-related knowledge, skills, and experience. We also offer incentive opportunities that reward employees based on individual and company performance. This is in addition to our diverse package of benefits centered around the wellbeing of our employees and their loved ones. In addition to the usual Medical/Dental/Vision/401k, our inclusive rewards plan empowers our people to care for their whole selves. An investment in your future is an investment in ours. Give Back With a charitable giving match and frequent opportunities to get involved, we take an active role in supporting the community.

Enjoy Time Away
  • You’ll start with 4+ weeks of paid time off a year, plus holidays and sick leave, to rest and recharge.
Care for Family
  • Whatever family means to you, we want to support you along the way—including a stipend for fertility care or adoption, medical travel support, and virtual vet care for your fur babies.
Prioritize Emotional Wellness
  • With on-demand apps and free confidential therapy sessions, you’ll have support no matter where
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior ML Engineer (Applied AI)
Senior ML Engineer (Applied AI)

Internetwork Expert • Massachusetts

On-site
USD 140,000 - 210,000
Annual paid vacation
Health Insurance
Remote-first culture
+1
Founding Forward Deployed Machine Learning Engineer
Founding Forward Deployed Machine Learning Engineer

adaption • San Francisco (CA)

On-site
USD 100,000 - 140,000
Flexible work
Annual travel stipend
Weekly meal allowance
+2
Principal, AI Engineer, Solution Delivery
Principal, AI Engineer, Solution Delivery

Tata Consultancy Services • New York (NY)

On-site
USD 120,000 - 130,000
Discretionary Annual Incentive
Comprehensive Medical Coverage
401K Plan
+1
Senior Software Engineer
Senior Software Engineer

Runware • Northern (KY)

On-site
USD 140,000 - 210,000
Generous stock options
Remote-first setup
Flexible hours
+1
Senior AI Engineer - USA
Senior AI Engineer - USA

Cogniify, Inc. • San Francisco (CA)

On-site
USD 140,000 - 165,000
Unlimited PTO
Generous parental leave
Entrepreneurial culture
+6
Lead Data Scientist Applied AI - USA Onsite (Santa Clara, CA)
Lead Data Scientist Applied AI - USA Onsite (Santa Clara, CA)

Dover • Santa Clara (CA), Northern (KY)

On-site
USD 150,000 - 170,000
Unlimited PTO
Generous parental leave
Employee stock purchase program
+2
REMOTE Senior AI/ML Engineer - SaaS
REMOTE Senior AI/ML Engineer - SaaS

CyberCoders, Inc. • United States

Remote
USD 140,000 - 240,000
Remote work
Unlimited vacation
Comprehensive benefits
+3
Software Engineer AI/ML Systems - USA Onsite (Santa Clara, CA)
Software Engineer AI/ML Systems - USA Onsite (Santa Clara, CA)

Dover • Santa Clara (CA), Northern (KY)

On-site
USD 150,000 - 170,000
Unlimited PTO
Generous parental leave
Stock Purchase Program (ESPP)
+3
Senior Full Stack Software Engineer
Senior Full Stack Software Engineer

Artificial Analysis • San Francisco (CA)

On-site
USD 140,000 - 180,000
Equity
Staff Engineer, ML Profiling Tools
Staff Engineer, ML Profiling Tools

Samsung-Semiconductor • San Jose (CA)

On-site
USD 163,000 - 253,000
Medical/Dental/Vision/401k
Diversity and inclusion programs
Accommodations for disabilities