Enable job alerts via email!

Principal Data Platform Architect

NVIDIA Corporation

Santa Clara (CA)

On-site

USD 272,000 - 426,000

Full time

5 days ago
Be an early applicant

Boost your interview chances

Create a job specific, tailored resume for higher success rate.

Job summary

NVIDIA seeks a Principal Data Platform Architect to define and implement a vision for distributed data platforms in AI and HPC applications. In this role, you will collaborate with engineering teams, lead the development of observability systems, and ensure the efficient management of data. The ideal candidate will possess extensive experience in large-scale system design and a strong technical background.

Benefits

Equity opportunities
Comprehensive benefits package

Qualifications

  • 15+ years of relevant experience in data platform architecture.
  • Experience designing large-scale, distributed observability systems.
  • Technical lead level experience in Python, JavaScript, and Java programming.

Responsibilities

  • Define vision for AI/HPC observability and architect necessary systems.
  • Lead teams to develop and deploy data collection pipelines.
  • Continuously improve process observability and data management.

Skills

Collaboration
Data Analysis
System Design
Observability
Programming (Python, JS, Java)
Database Management
Problem-Solving

Education

MS in Computer Science or Electrical Engineering
BS in Computer Science or Electrical Engineering

Tools

Apache Spark
Elastic/Open Search
Grafana
Prometheus

Job description

NVIDIA’s Hardware Infrastructure organization is seeking a Principal Data Platform Architect. We serve and collaborate directly with NVIDIA’s rapidly growing AI, HW, and SW engineering and research teams across the company. We are looking for a technical leader to define a vision and roadmap for distributed data platform and observability systems for large-scale AI and HPC clusters and workloads and guide implementation towards this vision. You will architect systems for data collection, aggregation, enrichment, storage, retrieval, and visualization to spectacularly improve efficiency, performance, and productivity of AI and HPC workloads. You will lead technical teams to develop, deploy, and operate observability solutions for multiple compute clusters around the world.

What You’ll Be Doing:

Collaborate with AI, HW, and SW engineering and research teams to define a vision and roadmap for AI/HPC cluster observability.

Architect and lead teams to develop, test, and deploy data collectors, pipelines, visualization and retrieval services.

Define data collection and retention polices to balance network bandwidth, system load, and storage capacity costs with data analysis requirements.

Work in a diverse team to provide operational and strategic data to empower our engineers and researchers to improve performance, productivity, and efficiency.

Continuously improve quality, workloads, and processes through better observability.

What We Need to See:

Experience designing and building large scale, distributed observability systems.

Ability to collaborate with data scientists, researchers, and engineering teams to identify high value data for collection and analysis.

Experience with turning raw data into actionable reports

Experience with observability platforms such as Apache Spark, Elastic/Open Search, Grafana, Prometheus, and other similar open-source tools

Technical lead level Python, JS and Java programming experience.

Thorough understanding of databases (relational and non-relational)

Passion for improving the productivity of others

Excellent planning and interpersonal skills

Flexibility/adaptability working in a dynamic environment with changing requirements

MS (preferred) or BS in Computer Science, Electrical Engineering, or related field or equivalent experience

15+ years of relevant experience.

Ways To Stand Out from The Crowd:

Background in computer science, machine learning, deep learning, open-source software, infrastructure technologies, and GPU technology.

Prior experience in infrastructure software, production application software development, software development, release and support methodology and devops

Experience in the management of datacenters and large-scale distributed computing

Experience in working with AI researchers and/or EDA developers

Consistent track record of driving process improvements and measuring efficiency and a passion for sharing knowledge and experience driving complex projects end-to-end.

NVIDIA’s Hardware Infrastructure organization is seeking a Principal Data Platform Architect. We serve and collaborate directly with NVIDIA’s rapidly growing AI, HW, and SW engineering and research teams across the company. We are looking for a technical leader to define a vision and roadmap for distributed data platform and observability systems for large-scale AI and HPC clusters and workloads and guide implementation towards this vision. You will architect systems for data collection, aggregation, enrichment, storage, retrieval, and visualization to spectacularly improve efficiency, performance, and productivity of AI and HPC workloads. You will lead technical teams to develop, deploy, and operate observability solutions for multiple compute clusters around the world.

What You’ll Be Doing:

  • Collaborate with AI, HW, and SW engineering and research teams to define a vision and roadmap for AI/HPC cluster observability.

  • Architect and lead teams to develop, test, and deploy data collectors, pipelines, visualization and retrieval services.

  • Define data collection and retention polices to balance network bandwidth, system load, and storage capacity costs with data analysis requirements.

  • Work in a diverse team to provide operational and strategic data to empower our engineers and researchers to improve performance, productivity, and efficiency.

  • Continuously improve quality, workloads, and processes through better observability.

What We Need to See:

  • Experience designing and building large scale, distributed observability systems.

  • Ability to collaborate with data scientists, researchers, and engineering teams to identify high value data for collection and analysis.

  • Experience with turning raw data into actionable reports

  • Experience with observability platforms such as Apache Spark, Elastic/Open Search, Grafana, Prometheus, and other similar open-source tools

  • Technical lead level Python, JS and Java programming experience.

  • Thorough understanding of databases (relational and non-relational)

  • Passion for improving the productivity of others

  • Excellent planning and interpersonal skills

  • Flexibility/adaptability working in a dynamic environment with changing requirements

  • MS (preferred) or BS in Computer Science, Electrical Engineering, or related field or equivalent experience

  • 15+ years of relevant experience.

Ways To Stand Out from The Crowd:

  • Background in computer science, machine learning, deep learning, open-source software, infrastructure technologies, and GPU technology.

  • Prior experience in infrastructure software, production application software development, software development, release and support methodology and devops

  • Experience in the management of datacenters and large-scale distributed computing

  • Experience in working with AI researchers and/or EDA developers

  • Consistent track record of driving process improvements and measuring efficiency and a passion for sharing knowledge and experience driving complex projects end-to-end.

The base salary range is 272,000 USD - 425,500 USD. Your base salary will be determined based on your location, experience, and the pay of employees in similar positions.

You will also be eligible for equity and benefits . NVIDIA accepts applications on an ongoing basis.

NVIDIA is committed to fostering a diverse work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.

About the company

9637389 Nvidia Corporation is an American multinational technology company incorporated in Delaware and based in Santa Clara, California.

Notice

Talentify is an Equal Opportunity Employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, disability, or protected veteran status.

Talentify provides reasonable accommodations to qualified applicants with disabilities, including disabled veterans. Request assistance at accessibility@talentify.io or 407-000-0000.

Federal law requires every new hire to complete Form I-9 and present proof of identity and U.S. work eligibility.

An Automated Employment Decision Tool (AEDT) will score your job-related skills and responses. Bias-audit & data-use details: www.talentify.io/bias-audit-report . NYC applicants may request an alternative process or accommodation at aedt@talentify.io or 407-000-0000.

Get your free, confidential resume review.
or drag and drop a PDF, DOC, DOCX, ODT, or PAGES file up to 5MB.

Similar jobs

Principal Data Platform Architect

Nvidia Corporation in

Santa Clara null

On-site

On-site

USD 272,000 - 426,000

Full time

4 days ago
Be an early applicant

Principal Data Platform Engineer

ServiceTitan

null null

Remote

Remote

USD 244,000 - 327,000

Full time

6 days ago
Be an early applicant

Principal Data Platform Engineer

ServiceTitan, Inc.

Snowflake null

Remote

Remote

USD 244,000 - 327,000

Full time

8 days ago

Lead Data Engineer Architect (Exp. with Data Vault 2.0)

Technogen, Inc.

null null

Remote

Remote

USD 97,000 - 720,000

Full time

5 days ago
Be an early applicant

Principal Systems Architect

The Rundown AI, Inc.

San Jose null

On-site

On-site

USD 270,000 - 315,000

Full time

-1 days ago
Be an early applicant

Principal System Architect

AMD

San Jose null

On-site

On-site

USD 180,000 - 280,000

Full time

6 days ago
Be an early applicant

Principal System Cloud Architect

NVIDIA Corporation

New York null

Remote

Remote

USD 272,000 - 426,000

Full time

5 days ago
Be an early applicant

Principal Software Architect - Martech

CVS Health

Olympia null

Remote

Remote

USD 144,000 - 289,000

Full time

3 days ago
Be an early applicant

Principal System Cloud Architect

NVIDIA

null null

Remote

Remote

USD 272,000 - 426,000

Full time

4 days ago
Be an early applicant