Data Infra - Telemetry and Observability (IC)

Matter Intelligence

El Segundo (CA)

On-site

USD 170,000 - 210,000

Full time

7 days ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Benefits offered by this job

100% employer-paid health, dental, and
vision coverage

Job summary

Matter Intelligence is seeking a Telemetry and Observability Engineer to connect flight and mission reality with model behavior through shared telemetry, traces, and reliability mechanisms, reporting to Ignacio Cases Martin. This role focuses on observability across datasets, training runs, models, and environments within a space- and defense-grade AI system.

You will instrument training, inference, and agent workflows, linking mission state with hardware context, time, geometry, and geolocation

Qualifications

  • Experience with ML observability, telemetry, time-series systems, distributed training or inference, cloud infrastructure, or production AI platforms.
  • Strong hands-on engineering skills in observability, automation, APIs, infrastructure, incident tooling, and systems debugging.
  • Understanding of metrics, logs, traces, events, model and data monitoring, service objectives, capacity, release safety, and incident diagnosis.
  • Ability to reason about AI failure modes including bad datasets, training instability, stale checkpoints, drift, version mismatch, environment bugs, and agent-tool failure.
  • Experience designing telemetry schemas and correlation workflows across multiple systems, time domains, and ownership boundaries.

Responsibilities

  • Build a common telemetry model linking mission state, datasets, code, checkpoints, accelerators, model and prompt versions, environment episodes, agent steps, tools, decisions, cost, latency, and feedback.
  • Instrument distributed training, data loading, GPU utilization, checkpointing, experiment health, batch and online inference, and model-serving behavior.
  • Instrument agent workflows across prompts, context, retrieval, memory, planning, tool calls, graph state, human intervention, evidence, outcomes, and safety controls.
  • Connect planned-versus-actual mission state, hardware context, time, geometry, and geolocation to relevant datasets and model results.
  • Establish actionable service objectives, alerts, incident interfaces, rollout and rollback signals, and escalation paths across platform, model, environment, and agent failures.
  • Build capacity and cost visibility across storage, processing, accelerators, training, evaluation, inference, retrieval, and agent execution.

Skills

ML observability
telemetry
time-series systems
distributed training
cloud infrastructure
production AI platforms

Job description

About Matter Intelligence

Welcome to Matter, where we are building the future of vision AI: pairing a world-first sensor that sees molecular chemistry, temperature, and 3D shape with a Large World Model that will be the most powerful intelligence engine for the physical world. This system doesn’t just see what something looks like; it understands everything from a single pixel. We call this Superintelligent Vision.

Our team has delivered technologies to Mars for NASA/JPL, designed advanced sensors for U.S. Defense, and built core infrastructure at OpenAI. We are now building the next generation of space- and airborne-based sensing systems.

About the Role

Matter is hiring a Telemetry and Observability Engineer to make datasets, training runs, models, environments, agents, and missions observable as one AI system. Reporting to Ignacio Cases Martin, this individual contributor will connect flight and mission reality with infrastructure, model, and product behavior through shared telemetry, traces, alerts, reliability mechanisms, and operational context.

Key Responsibilities
  • Build a common telemetry model linking mission state, datasets, code, checkpoints, accelerators, model and prompt versions, environment episodes, agent steps, tools, decisions, cost, latency, and feedback.

  • Instrument distributed training, data loading, GPU utilization, checkpointing, experiment health, batch and online inference, and model-serving behavior.

  • Instrument agent workflows across prompts, context, retrieval, memory, planning, tool calls, graph state, human intervention, evidence, outcomes, and safety controls.

  • Connect planned-versus-actual mission state, hardware context, time, geometry, and geolocation to relevant datasets and model results.

  • Establish actionable service objectives, alerts, incident interfaces, rollout and rollback signals, and escalation paths across platform, model, environment, and agent failures.

  • Build capacity and cost visibility across storage, processing, accelerators, training, evaluation, inference, retrieval, and agent execution.

Qualifications
Required
  • Experience with ML observability, telemetry, time-series systems, distributed training or inference, cloud infrastructure, or production AI platforms.

  • Strong hands-on engineering skills in observability, automation, APIs, infrastructure, incident tooling, and systems debugging.

  • Understanding of metrics, logs, traces, events, model and data monitoring, service objectives, capacity, release safety, and incident diagnosis.

  • Ability to reason about AI failure modes including bad datasets, training instability, stale checkpoints, drift, version mismatch, environment bugs, and agent-tool failure.

  • Experience designing telemetry schemas and correlation workflows across multiple systems, time domains, and ownership boundaries.

Preferred
  • Experience with MLOps, GPU workloads, model serving, agent observability, reinforcement-learning environments, mission telemetry, or scientific instrumentation.

  • Experience operating high-throughput inference, streaming systems, geospatial pipelines, or mixed cloud and edge deployments.

  • Experience with reliability platforms or observability products used by multiple engineering teams.

  • Familiarity with aerospace, defense, regulated operations, or other settings requiring formal change control and incident evidence.

What Success Looks Like
  • Teams can correlate mission, data, infrastructure, model, environment, and agent behavior during normal operation and incidents.

  • Alerts and service objectives identify actionable failures without obscuring scientific or operational context.

  • Capacity, cost, version, and reliability signals support safe deployments and faster diagnosis across the AI stack.

Location

This role is based in El Segundo, CA, and requires onsite work.

ITAR Requirements

To comply with U.S. export regulations, applicants must be one of the following:

  • A U.S. citizen or national

  • A lawful permanent resident (green card holder)

  • Eligible to obtain required authorizations from the U.S. Department of State

Employee Offerings and Benefits

At Matter, we believe in rewarding high performance and providing the support you need to thrive. Our compensation and benefits package includes:

  • Competitive compensation based on experience

  • Early-stage equity package

  • 100% employer-paid health, dental, and vision coverage

  • Opportunity to work on novel sensing, data, and AI systems with real-world deployment paths to the largest industries in the world

Matter Intelligence is an equal opportunity employer. We welcome candidates from all backgrounds who can raise the ambition and performance of the team.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Data Infra - Telemetry and Observability (IC)
Data Infra - Telemetry and Observability (IC)

Matter Intelligence • San Francisco (CA)

On-site
USD 180,000 - 240,000
Equity package
Health insurance
Dental & vision
Full-Stack Software Engineer
Full-Stack Software Engineer

Matter • El Segundo (CA)

Hybrid
USD 120,000 - 170,000
Competitive compensation
Early-stage equity
Health, dental, vision coverage
Applied AI Engineer (Product)
Applied AI Engineer (Product)

Matter Intelligence • San Francisco (CA)

On-site
USD 140,000 - 210,000
Equity
Health coverage
Onsite work allowance
Product Intelligence Engineer
Product Intelligence Engineer

Matter Intelligence • San Francisco (CA)

On-site
USD 150,000 - 230,000
Early-stage equity package
Health, dental, and vision coverage
Competitive compensation
Applied AI Engineer - Internal
Applied AI Engineer - Internal

SupportFinity™ • San Francisco (CA)

On-site
USD 180,000 - 240,000
Employer-paid health, dental, vision
Early-stage equity compensation
Exposure to hardware and AI projects
Head of Product & Deployment
Head of Product & Deployment

Matter Intelligence • San Francisco (CA)

On-site
USD 180,000 - 280,000
Equity
Health, dental & vision
Competitive compensation
Head of Talent & Hiring
Head of Talent & Hiring

Matter Intelligence • San Francisco (CA)

On-site
USD 180,000 - 280,000
Health coverage
Early-stage equity
Competitive compensation
Optical Engineer - Airborne
Optical Engineer - Airborne

Matter Intelligence • El Segundo (CA)

On-site
USD 120,000 - 190,000
Equity package
100% employer-paid health, dental, and
Vision coverage
Scheduler and Analyst (Mid Career)
Scheduler and Analyst (Mid Career)

Matter Intelligence • El Segundo (CA)

On-site
USD 90,000 - 120,000
Early-stage equity
Health, dental, and vision coverage
Competitive compensation
Control Systems Engineer
Control Systems Engineer

Matter Intelligence • El Segundo (CA)

On-site
USD 120,000 - 170,000
Equity package
Employer-paid health, dental, and vis
Opportunity to work on novel sensing/3