Forward-Deployed Engineer, AI Fabric

Upscale AI

United States

Hybrid

USD 180,000 - 285,000

Full time

14 days+
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Upscale AI is seeking a Forward-Deployed Engineer to address complex AI fabric issues, reproduce problems in the lab, and provide actionable fixes. You will work across support, engineering, and product to deliver robust field-ready solutions and contribute to agent improvements.

The role involves direct customer interaction, on-call rotations, and travel to on-site engagements to ensure trust and impactful resolution in production environments.

Qualifications

  • 5+ years hands-on experience with high-speed data center switching platforms, AI fabric deployments is a plus.
  • Deep Ethernet troubleshooting experience: L2/L3 forwarding, ECMP, packet-level analysis.
  • BGP operations: route reflectors, convergence, fabric-scale deployments.
  • QoS configuration and troubleshooting: PFC, ECN, DSCP, queue management.
  • Hardware support background: ASICs, FPGAs, chassis-based platforms, pluggable optics.
  • Linux proficiency: command line, administration, log analysis.

Responsibilities

  • Triage and resolve complex AI fabric issues from intake to resolution.
  • Reproduce issues in lab by building test scenarios to isolate problems.
  • Act as escalation buffer for engineering when possible, packaging clear problem statements when needed.
  • Train and improve on-box and off-box AI agents with field-validated signatures and content.
  • Validate AI agent accuracy by reviewing data collection and triage classifications against outcomes.
  • Work directly with customers to understand fabric topology and operational constraints, communicating findings clearly.
  • Contribute to product by identifying supportability gaps and feeding field insights.
  • Maintain and evolve AI agent detection capabilities with updates and retraining.
  • Read and work with source code to understand behavior and verify fixes.

Skills

L2/L3 Troubleshooting
High-speed Networking
Field Engineering
Customer Interaction
On-call Experience

Education

BSc in CS/EE

Tools

NOS/Firmware
Lab Testing

Job description

Why join Upscale AI

Upscale AI is building the high-performance infrastructure powering the next generation of artificial intelligence. Backed by over $300M in funding and rapid global adoption, we are scaling systems designed for the world’s most demanding AI workloads.

We focus on first-principles engineering across silicon, systems, and networking—where performance, scale, and execution are critical.Our team is talent-dense andhigh-performing. We value ownership, technical rigor, and speed, and we offer the opportunity to work on foundational problems with immediate, real-world impact.

If you’re looking to do high-impact work, move fast, and help define the infrastructure behind the future of AI—Upscale AI is where you can produce meaningful work at the frontier—and operate at a high standard.

Upscale is building the engineering team thatdrives product improvementsthoughhands-on field engineering.The Forward-Deployed Engineer (FDE) owns complex fabricand systemissues end-to-end: from first report through root-cause analysis, workaround delivery, and verified fix. You resolve themwith the rest of the team,or you drive them to resolution across whatever boundary stands in the way.

This role is new toUpscaleAIand sits at the intersection of support, engineering, and product management. You will work directly with customers running production AI fabrics, reproduce issues in the lab, develop workarounds under pressure, and contribute fixes and diagnostic content back into the product and our AI-driven monitoring agent.

The environment is often unstructured. Problem definitions are incomplete. Documentation may not exist yet. The right candidate sees that as an opportunity, not an obstacle.

Large AI infrastructure operators increasingly expect theirvendor'sengineering team to function as an extension of their own infrastructure organization. This role is built to meet that expectation.

This is a small team. There will be an on-callcomponent, but we are building a global team to reduce out-of-hourscalls. We work with customers directly, so some travel is involved. Remote meeting tools handle much of thecollaborationbutbuilding strong trust relationships will require time on site.


Key Responsibilities

  • Triage and resolve complex AI fabric issues:including silent packet drops, queue anomalies, NCCL stalls, gray failures, and performance degradation in production environments. You own the problem from intake to resolution.

  • Reproduce issues in the lab:building test scenarios that isolate L2 and L3 problems (packet loss, latency, retransmits, ECMP behavior, PFC/ECN interactions) and deliver reproducible cases to the development team when code fixes are needed.

  • Act as the escalation buffer for engineering:resolving issues without engaging development when possible, and packaging clean, reproducible problem statements when development engagement isrequired.

  • Train and improve the on-boxand off-boxAI agents:contributing field-validated detection signatures, classification logic, and resolution recommendations based on real cases. Your field experience directly shapes what the agent canidentifyand handle autonomously.

  • ValidateAIagent accuracy:reviewing the agent's data collection, anomaly detection, and triage classifications against real-world outcomes. Youdeterminewhen the agent is ready to advance from data collection to active triageto mitigation

  • Work with customers directly:understanding their fabric topology, workload patterns, and operational constraints. Communicate findings clearly to both technical and non-technical stakeholders.

  • Contribute to the product:identifyingsupportability gaps, proposing diagnostic improvements, and feeding field insights into the product management process. You are an active voice in what the product needs tobecome.

  • Maintain and evolve theAIagent's detection capabilities:update detection signatures, retrain classification models, and tune thresholds as customer fabricsscale,new hardware is deployed, and new failure modes are discovered in the field. TheAIagent is anevolvingsystem, not a shipped product.

  • Read and work with source code:engaging with developers at the code level when needed to understand behavior,identifyroot causes, orvalidatefixes. You do not need to be a full-time developer, but you must be comfortable in the codebase.

Requirements

  • 5+ years of hands-on experience with high speed data center switching platforms at scale. Experience with AI fabric deployments is a strong plus but not required if switching and troubleshooting background depth is there.People who take ownership. The FDE role requires someone who treats every problem as theirs until it's resolved, regardless of where the root cause sits. If the issue crosses into ASIC behavior, NOS code, optics firmware, or customer configuration, you follow it there. Pointing to another team is not a resolution.
  • Deep Ethernet troubleshooting experience, including L2/L3forwarding, ECMP, and packet-level analysis

  • BGP operations experience, including route reflectors, convergence behavior, and fabric-scale deployments

  • QoS configuration and troubleshooting: PFC, ECN, DSCP, queue management

  • Hardware support background: ASICs, FPGAs, chassis-based platforms, pluggable optic modules

  • Fiber and optical link troubleshooting, including DOM telemetry interpretation

  • Software support history: working with NOS, firmware, and driver-level issues

  • Linuxproficiency(command line, system administration, log analysis)

  • Telemetry and monitoring: experience with streaming telemetry, counters, and event-driven diagnostics

  • Experience contributing to or training ML/AI systems (classification models, labeled data, feedback loops)

  • Lab skills: ability to design, build, and execute complex test scenarios that isolate specific failure modes

  • Clear written and verbal communication, including customer-facing interaction

Nice to have

  • Experience withSONiCor other open network operating systems

  • Familiarity with AI data center designs: low-latency fabrics, GPU cluster networking, NCCL, RDMA/RoCEv2

  • Experience with NVIDIA Spectrum switch family,BlueFieldSuperNICs, orConnectXadapters

  • Understanding of Ultra Ethernet Consortium specifications and goals

  • PCIe architecture knowledge (relevant to NIC and accelerator integration)

  • Experience with Ixia/Keysight or similar network test equipment and test suites

  • Exposure to high-frequency telemetry (HFT) or WJH (What Just Happened) event data


$180,000 - $285,000 a year

Where you fall within that range depends on your experience, skills, and impact—we benchmark against internal levels to keep things fair and consistent.

Equal Opportunity

Upscale AI is building a team that reflects a wide range of perspectives, backgrounds, and experiences. We’re proud to be an Equal Opportunity Employer and consider all qualified applicants regardless of race, color, religion, national origin, sex, sexual orientation, gender identity, disability, or veteran status.

Accessibility & Accommodations

We’re committed to making our hiring process accessible to everyone. If you need accommodations at any stage, just reach out to us athiring@upscaleai.com—we’re happy to help. Note: This inbox is only for accommodation requests.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Technical Support Principal Engineer – AI Network
Technical Support Principal Engineer – AI Network

Upscale AI • United States

On-site
USD 248,000 - 269,000
Technical Support Senior Staff Engineer
Technical Support Senior Staff Engineer

Upscale AI • United States

On-site
USD 200,000 - 216,000
Senior Staff Engineer, Lab Operations
Senior Staff Engineer, Lab Operations

Upscale AI • Northern (KY)

Hybrid
USD 205,000 - 230,000
Technical Support Senior Staff Engineer
Technical Support Senior Staff Engineer

The Consensus • United States

On-site
USD 200,000 - 216,000
Technical Support Principal Engineer – AI Network
Technical Support Principal Engineer – AI Network

The Consensus • United States

On-site
USD 248,000 - 269,000
Senior Staff DevOps Engineer – Orchestration
Senior Staff DevOps Engineer – Orchestration

Upscale AI • United States

Hybrid
USD 263,000 - 284,000
Director, Orchestration
Director, Orchestration

Upscale AI • United States

On-site
USD 140,000 - 200,000
Senior Staff Engineer - Platform
Senior Staff Engineer - Platform

Upscale AI • Santa Clara (CA), Northern (KY)

Hybrid
USD 237,000 - 260,000
Design Verification Engineer
Design Verification Engineer

Upscale AI • Northern (KY)

Hybrid
USD 212,000 - 246,000
Software Engineer, Compute Infrastructure
Software Engineer, Compute Infrastructure

OpenAI • California (MO)

On-site
USD 180,000 - 260,000