Senior Data Platform Engineer: Fleet Telemetry & Insights

NVIDIA

United States

Remote

USD 140,000 - 224,000

Full time

6 days ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Benefits offered by this job

Equity
Benefits

Job summary

NVIDIA’s DGX Cloud organization is seeking a Senior Data Engineer to join the data team and help build a reliable data foundation supporting fleet health, capacity, utilization, cost, reliability, and operational decision‑making across DGX Cloud. This hands‑on, platform‑minded role focuses on ingestion, transformation, quality, and self‑service consumption to create dependable data products.

You will own end‑to‑end systems from requirements through deployment and monitoring, construct batch and

Qualifications

  • BS or MS in Computer Science, Engineering, or a related field, or equivalent experience.
  • 5+ years of experience building and operating production software, data platforms, backend infrastructure, databases, or distributed systems.
  • Strong software‑engineering fundamentals and production proficiency in Python or another backend or systems language, with the ability and willingness to work primarily in Python and SQL.
  • Deep hands‑on experience in at least one of the following areas: Distributed data processing using Spark or a comparable compute framework, Relational, distributed, or analytical database architecture and operation at scale, Production ETL, change‑data‑capture, streaming, or event‑processing systems, Backend or cloud‑platform systems that process, transform, or serve substantial data volumes, Strong SQL and data‑modeling skills, including a practical understanding of query performance, schema evolution, incremental processing, consistency, and analytical consumption patterns.
  • Demonstrated ability to debug unfamiliar systems across multiple layers using logs, metrics, traces, query plans, profiles, and controlled experiments to find root causes.
  • Experience operating services or pipelines in a cloud or similarly complex production environment, including testing, CI/CD, monitoring, alerting, rollback, and incident response.
  • Working knowledge of secure platform development, including identity and access management, least privilege, secret handling, trust boundaries, and safe multi‑environment deployments.
  • Ability to make sound architectural tradeoffs, own work through ambiguity, and communicate effectively with users, partner teams, and engineers from different fields.
  • A track record of learning unfamiliar technologies and domains and turning that learning into maintainable systems and reusable team practices.
  • Experience with AI agents and LLM‑supported workflow automation, particularly as applied to engineering and operational activities.

Responsibilities

  • Own systems end to end. For example, work from ambiguous customer and operational needs through architecture, implementation, deployment, observability, incident response, and ongoing support.
  • Construct data pipelines and products. Such as designing and maintain batch and streaming ingestion, transformation, reconciliation, and serving paths for fleet, capacity, utilization, cost, scheduling, and operational telemetry.
  • Build shared libraries, workflow and DAG or equivalent experience abstractions to evolve the data platform. Develop deployment tooling, data contracts, and paved‑road patterns that improve team speed and safety.
  • Engineer reliable distributed workloads. As well as diagnose correctness and performance issues across applications, SQL engines, Spark jobs, storage systems, networks, and cloud services. Build for retries, idempotency, backfills, schema evolution, and partial failure.
  • Treat security as part of the build. For example, applying least privilege, service identities, secrets management, access controls, environment isolation, auditability, and safe operational practices throughout the system lifecycle.
  • Improve quality and operations: Establish automated tests, data‑quality checks, lineage, freshness and completeness monitoring, actionable alerting, SLOs, and clear ownership.
  • Deliver consumption experiences. Such as making trusted data usable through well‑modeled tables, APIs, automation, dashboards, and focused internal applications—not only through one‑off queries.
  • Raise the engineering bar. Lead build reviews, communicate tradeoffs, mentor other engineers, and improve the team's architecture, testing, debugging, and operational practices.

Skills

Python
SQL
Distributed systems
Data pipelines
Spark
Cloud platforms

Education

BS or MS in Computer Science, Engineering, or related field

Tools

Databricks
Apache Spark
PySpark
Spark SQL
Delta Lake
Unity Catalog
Kafka
OpenSearch
Elasticsearch
AWS
Azure
GCP
Kubernetes

Job description

NVIDIA’s DGX Cloud organization is seeking a Senior Data Engineer to join the data team and help build a reliable data foundation supporting fleet health, capacity, utilization, cost, reliability, and operational decision‑making across DGX Cloud. This hands‑on, platform‑minded role focuses on ingestion, transformation, quality, and self‑service consumption to create dependable data products.

You will own end‑to‑end systems from requirements through deployment and monitoring, construct batch and

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Data Platform Engineer — Fleet Telemetry
Senior Data Platform Engineer — Fleet Telemetry

NVIDIA Corporation • Northern (KY)

On-site
USD 140,000 - 270,000
Equity
Benefits
Senior Data Platform Engineer (Fleet & Observability)
Senior Data Platform Engineer (Fleet & Observability)

NVIDIA Corporation • Santa Clara (CA)

On-site
USD 140,000 - 270,000
Senior Data Platform Engineer: Equity & Trusted Pipelines
Senior Data Platform Engineer: Equity & Trusted Pipelines

NVIDIA • Washington

On-site
USD 140,000 - 270,000
Senior Data Platform Engineer: Ingestion, Spark & Data Quality
Senior Data Platform Engineer: Ingestion, Spark & Data Quality

NVIDIA Corporation • Santa Clara (CA)

On-site
USD 200,000 - 322,000
Equity and benefits
Senior Data Platform & Analytics Product Lead (Cloud)
Senior Data Platform & Analytics Product Lead (Cloud)

NVIDIA Corporation • California (MO)

Hybrid
USD 208,000 - 328,000
DGX Cloud Data Platform & Analytics Lead
DGX Cloud Data Platform & Analytics Lead

NVIDIA Corporation • Santa Clara (CA)

Hybrid
USD 208,000 - 328,000
Equity
Benefits
Hybrid work model
DGX Cloud Data Platform & Analytics Lead
DGX Cloud Data Platform & Analytics Lead

NVIDIA • California (MO)

Hybrid
USD 208,000 - 328,000
Equity
Benefits
Data and Platform Engineer
Data and Platform Engineer

NVIDIA Corporation • Santa Clara (CA)

On-site
USD 200,000 - 322,000
Equity and benefits
Senior Data and Platform Engineer
Senior Data and Platform Engineer

NVIDIA Corporation • Santa Clara (CA)

On-site
USD 140,000 - 270,000
Senior Software Engineer - Capacity & Resource Management
Senior Software Engineer - Capacity & Resource Management

NVIDIA • Santa Clara (CA)

On-site
USD 140,000 - 270,000