Staff Software Engineer, Systems Infrastructure - Agent Evaluation

Linkedin3

Mountain View (CA)

Hybrid

USD 180,000 - 240,000

Full time

5 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

LinkedIn is seeking a Staff Engineer to own the end‑to‑end EOS platform, shaping the data infrastructure for capturing and labeling interactions and building models to evaluate AI agents in production. You will collaborate with AI product teams, ML engineers, and infrastructure partners to embed evaluation across the development lifecycle and ensure high‑quality AI systems across the company.

The role combines distributed systems, data platforms, and ML, demanding an ability to lead complex

Qualifications

  • Bachelor's Degree in Computer Science or a related technical discipline, or equivalent practical experience.
  • 4+ years of experience in the industry with leading/building deep learning systems.
  • 4+ years of experience with Java, C++, Python, Go, Rust, C# and/or functional languages such as Scala or other relevant coding languages.
  • Hands‑on experience developing distributed systems or other large‑scale systems.
  • Hands‑on experience building or evaluating AI Agents in Production.

Responsibilities

  • Own the technical vision, architecture, and execution of the Evaluation Operating System (EOS).
  • Design and build large‑scale evaluation infrastructure enabling teams to measure AI agents’ quality, reliability, safety, and performance.
  • Work on reliable and scalable tracing infrastructure for LinkedIn AI Agents with trace debuggability features.
  • Architect scalable data pipelines and platforms for capturing, processing, labeling, and managing AI interactions and datasets.
  • Build and evolve evaluation systems powered by LLM‑as‑judge, reward models, and other automated evaluators.
  • Develop experimentation and testing frameworks, including adversarial testing and champion/challenger experiments.
  • Establish real‑time observability, monitoring, and feedback loops to detect regressions and model drift.
  • Partner with AI product teams, ML engineers, and infra to integrate evaluation into the lifecycle.
  • Lead cross‑functional initiatives, influencing technical strategy across AI Platforms.
  • Mentor engineers and help shape engineering culture in a growing AI platform.",

Skills

Distributed systems
AI/ML systems
Mentorship
Cross-team collaboration

Education

Bachelor's Degree in Computer Science or related field
Master's Degree (preferred)

Tools

Java
C++
Python
Go
Rust
C#
Scala

Job description

This role will be based in Mountain View, CA.

At LinkedIn, our approach to flexible work is centered on trust and optimized for culture, connection, clarity, and the evolving needs of our business. The work location of this role is hybrid, meaning it will be performed both from home and from a LinkedIn office on select days, as determined by the business needs of the team.

LinkedIn's Core AI is building the Evaluation Operating System (EOS), a foundational Agent Evaluation platform that defines how all AI agents and GenAI products at LinkedIn are measured, evaluated, and continuously improved in production. This is a brand-new, industry-defining problem space with no established playbook, focused on evaluating multi-step, non-deterministic, and personalized AI systems where traditional metrics and testing approaches fall short.

EOS acts as the central intelligence layer for AI quality, combining large-scale data pipelines, evaluator models (e.g., LLM-as-a-judge, reward models), and real-time production monitoring to understand how AI systems behave, where they fail, and how to improve them. The platform includes capabilities like synthetic data generation, adversarial testing, golden dataset management, recursive Self Improving Agents and live "agent arena" experimentation frameworks (champion/challenger testing) to measure performance across multiple dimensions of quality. This platform also is responsible for tracing infrastructure for all LinkedIn AI Agents.

As a Staff Engineer, you will own the end-to-end technical vision, architecture, and execution of this platform. This includes designing the data infrastructure for capturing and labeling interactions, building systems to train and deploy evaluation models, and creating real-time monitoring and feedback loops that detect regressions, model drift, and quality degradation in production. You'll work closely with AI product teams, ML engineers, and infrastructure partners to embed evaluation deeply into the development lifecycle, making it possible for teams across LinkedIn to ship high‑quality AI systems with confidence.

This role sits at the intersection of distributed systems, data platforms, and machine learning, and is ideal for engineers who want to define how AI quality is measured at scale. The impact is company-wide: the systems you build will directly determine the quality ceiling, safety, and trustworthiness of every AI‑powered experience at LinkedIn.

Responsibilities
  • Own the technical vision, architecture, and execution of the Evaluation Operating System (EOS), solving complex, open‑ended challenges at the intersection of distributed systems, data infrastructure, and machine learning.
  • Design and build large‑scale evaluation infrastructure that enables LinkedIn teams to measure, understand, and continuously improve the quality, reliability, safety, and performance of AI agents and GenAI products.
  • Work on reliable and scalable Tracing Infrastructure for LinkedIn AI Agents along with trace debuggability features.
  • Architect scalable data pipelines and platforms for capturing, processing, labeling, and managing large volumes of AI interactions, evaluation data, golden datasets, and synthetic data.
  • Build and evolve evaluation systems powered by LLM-as-judge, reward models, and other automated evaluators to assess AI systems across multiple dimensions of quality and performance.
  • Develop experimentation and testing frameworks, including adversarial testing, champion/challenger experiments, and agent arena capabilities, to identify weaknesses and drive continuous improvement of AI systems.
  • Establish real‑time observability, monitoring, and feedback loops that detect regressions, model drift, quality degradation, and unexpected behavior in production AI systems.
  • Partner closely with AI product teams, ML engineers, and infrastructure organizations to integrate evaluation deeply into the AI development lifecycle and establish consistent evaluation standards across LinkedIn.
  • Lead multiple high‑impact, cross‑functional initiatives, influencing technical strategy and architectural decisions across AI Platforms and the broader engineering organization.
  • Mentor and develop engineers, raise the technical bar, and help shape the engineering culture and practices of a growing AI platform organization.
  • Build and Platformitize Recursive Self Improving Agents
Basic Qualifications
  • Bachelor's Degree in Computer Science or related technical discipline, or equivalent practical experience
  • 4+ years of experience in the industry with leading/ building deep learning systems.
  • 4+ years of experience with Java, C++, Python, Go, Rust, C# and/or Functional languages such as Scala or other relevant coding languages
  • Hands‑on experience developing distributed systems or other large‑scale systems.
  • Hands‑on experience building or evaluating AI Agents in Production.
Preferred Qualifications
  • BS and 8+ years of relevant work experienceMS and 7+ years of relevant work
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Staff Software Engineer, Systems Infrastructure - Agent Evaluation
Staff Software Engineer, Systems Infrastructure - Agent Evaluation

LinkedIn • Mountain View (CA)

Hybrid
USD 175,000 - 287,000
Staff Engineer, AI Evaluation Infrastructure
Staff Engineer, AI Evaluation Infrastructure

LinkedIn • Mountain View (CA)

Hybrid
USD 175,000 - 287,000
Sr. Evaluation Engineer
Sr. Evaluation Engineer

logicmonitor • San Francisco (CA)

On-site
USD 150,000 - 190,000
Sr. Software Engineer, AI Infrastructure
Sr. Software Engineer, AI Infrastructure

LinkedIn • California (MO)

Hybrid
USD 139,000 - 229,000
Staff AI Software Engineer
Staff AI Software Engineer

Harnham • San Francisco (CA)

On-site
USD 150,000 - 200,000
Manager, Software Engineering, Machine Learning
Manager, Software Engineering, Machine Learning

LinkedIn • Mountain View (CA)

Hybrid
USD 170,000 - 277,000
Senior Staff Software Engineer, AI Infrastructure
Senior Staff Software Engineer, AI Infrastructure

LinkedIn • Sunnyvale (CA)

Hybrid
USD 198,000 - 326,000
Staff Software Engineer - AI Platform
Staff Software Engineer - AI Platform

LinkedIn • California (MO)

Hybrid
USD 175,000 - 287,000
Senior Software Engineer - AI Platform
Senior Software Engineer - AI Platform

LinkedIn • California (MO)

Hybrid
USD 144,000 - 236,000
Sr. AI / Machine Learning Platform Engineer - Voice Agents
Sr. AI / Machine Learning Platform Engineer - Voice Agents

Skyrocket Ventures • San Francisco (CA)

On-site
USD 120,000 - 150,000