Staff Engineer — Agentic AI

Praxy

San Francisco, Northern (CA, KY)

Hybrid

USD 160,000 - 250,000

Full time

13 days ago
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Praxy is seeking a senior technical leader to own the core agent intelligence layer in an AI-native engineering software environment. You will report to the CTO and lead a small team of AI engineers, a user researcher, and contractors, shaping real-world product value for enterprise customers.

You will drive the full product loop from user story definition to implementation and benchmarking across CAD/CAE/PLM workflows, optimize costs, and ensure reliable multi-step automation in desktop

Qualifications

  • 7+ years in software engineering with real-world agentic LLM systems.
  • Deep experience with LLM architecture: model selection, context/window management, retrieval strategies.
  • Strong evaluation and benchmarking instincts for agentic systems.
  • Proven track record of shipped AI systems with measurable outcomes.
  • Strong Python skills and hands-on familiarity with LLM tooling.
  • Technical leadership experience for small teams (3–6 engineers).
  • Hands-on background with mechanical engineering software (CAD/CAE/PLM).
  • Experience shipping AI tooling on engineering data or desktop apps.

Responsibilities

  • Lead development of the agent intelligence layer across CAD/PLM software.
  • Own the full product loop from user story to implementation and benchmarking.
  • Define evaluation frameworks and cost controls for per-task budgets and workflows.
  • Collaborate with cross-functional teams during POCs to align agent behavior with real-world usage.
  • Write production code, review designs, and raise the engineering bar.
  • Architect tool-calling strategies, context management, and recovery for reliable workflows.

Skills

7+ years software engineering
Agentic LLM systems
Python skills
LLM tooling familiarity
Technical leadership
Desktop engineering tools
Cost optimization

Tools

Function calling
Tool use APIs
Tracing/observability
Evaluation frameworks

Job description

This is a senior technical leadership role at the heart of an AI-native engineering software company, owning the core agent intelligence layer that turns mechanical engineers' intent into reliable, cost-efficient multi-step workflows across complex desktop engineering tools. You'll report directly to the CTO and serve as the technical lead for a small team of AI engineers, a user researcher, and domain expert contractors. The work you do here will define real-world product value for enterprise customers.

WHAT YOU'LL DO
  • Lead development of the agent intelligence layer that executes multi-step workflows across CAD, simulation, and PLM software.
  • Own the full product loop — from user story definition to implementation to benchmarking against real engineering workflows.
  • Drive agent task success rate by defining evaluation frameworks, establishing baselines, and iterating on performance.
  • Set and enforce per-task token budgets and track cost per completed workflow to ensure commercial viability.
  • Design rigorous, reproducible evaluation infrastructure grounded in validated user stories — think SWE-bench-level rigor applied to engineering workflows.
  • Lead user story mapping and validation through direct interviews and collaboration with domain experts.
  • Translate validated user stories into testable evals, closing the loop between user research and benchmarking.
  • Own agent architecture decisions: tool-calling strategies, state management, error recovery, model routing, and context management.
  • Act as a player-coach — write production code, review designs, unblock the team, and raise the engineering bar.
  • Collaborate cross-functionally with integrations, product, and customers during POCs to align agent behavior with real-world usage.
WHAT WE'RE LOOKING FOR
  • 7+ years in software engineering, including at least 2 years building and shipping real-world agentic LLM systems (tool calling, multi-step workflows, failure recovery, cost control).
  • Deep experience with LLM application architecture: model selection, context/window management, retrieval strategies, tool-calling frameworks, and orchestration patterns.
  • Strong evaluation and benchmarking instincts for agentic systems — task completion rates, cost efficiency, failure mode analysis; familiarity with benchmarks such as SWE-bench, GAIA, or τ-bench is a plus.
  • Proven track record of shipped AI systems with measurable outcomes — not just demos or prototypes.
  • Strong Python skills and hands‑on familiarity with the LLM tooling ecosystem (function calling, tool use APIs, tracing/observability tools, evaluation frameworks).
  • Technical leadership experience setting direction for small teams (3–6 engineers) and performing meaningful code review and architecture decisions.
  • Hands‑on background with mechanical engineering software — CAD/CAE/PLM or simulation tooling (e.g. Siemens NX/NXOpen, Teamcenter, CATIA, Creo, SolidWorks, Ansys, Abaqus, or similar) — either as a builder of these tools or as a power user inside an engineering or manufacturing org.
  • Experience shipping AI/LLM tooling on top of proprietary engineering data or desktop engineering software (e.g. an agent or MCP server over CAD/PLM APIs, RAG over engineering repos or schematics).
  • Familiarity with enterprise deployment constraints, including behavior on locked-down corporate workstations.
  • Experience with desktop automation or programmatic control of applications (COM or similar) is a strong plus.
  • Published work, open-source contributions, or benchmark contributions in agentic AI is a plus.
COMPENSATION & BENEFITS

Salary range: $160,000 – $250,000 USD annually. Visa sponsorship is not available for this role.

LOCATION

On-site in San Francisco, California, USA.

Pay

COMPENSATION & BENEFITS Salary range: $160,000 – $250,000 USD

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Staff Engineer - Agentic AI
Staff Engineer - Agentic AI

Clera • San Francisco (CA)

On-site
USD 160,000 - 250,000
Equity
Visa sponsorship not available
Staff Engineer — Agentic AI for Desktop Engineering
Staff Engineer — Agentic AI for Desktop Engineering

Clera • San Francisco (CA)

On-site
Senior Agent Systems Engineer
Senior Agent Systems Engineer

DeepRec.ai • San Francisco (CA)

On-site
USD 160,000 - 185,000
Senior Software Engineer - Agentic Systems
Senior Software Engineer - Agentic Systems

Strativ Group • San Francisco (CA)

On-site
USD 200,000 - 400,000
Staff Software Engineer
Staff Software Engineer

Newmark Group • New York (NY)

On-site
USD 190,000 - 250,000
Principal AI Engineer - AI/ML Engineer Agentic
Principal AI Engineer - AI/ML Engineer Agentic

TekVizor • Pittsburgh

On-site
USD 120,000 - 150,000
Unlimited sick days
Tuition reimbursement
401k with 4% match
AI Engineer
AI Engineer

re-zoo-me • San Francisco (CA), Northern (KY)

On-site
USD 180,000 - 250,000
Founding Agentic Engineer
Founding Agentic Engineer

Praxy • San Francisco (CA), Northern (KY)

Hybrid
USD 150,000 - 300,000
Visa sponsorship
Staff Software Engineer
Staff Software Engineer

Newmark Group • Chicago (IL)

Hybrid
USD 190,000 - 250,000
Competitive compensation
Collaborative Culture
Growth & Learning
+1
AI Engineer
AI Engineer

Goliath-Partners • San Francisco (CA)

Hybrid
USD 220,000 - 275,000
Ownership potential
Hybrid work model
Career growth in AI