Staff Engineer — Agentic AI

Getclera

San Francisco (CA)

On-site

USD 180,000 - 240,000

Full time

14 days+
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Cosmon is seeking a senior technical leader to own the core agent intelligence that turns engineers' intent into reliable, cost‑efficient multi‑step workflows across desktop tools. You will report to the CTO and guide a small AI team, researchers, and contractors to ship real‑world value for enterprise customers.

You will drive architecture decisions, define evaluation baselines, and ensure the product delivers measurable performance and cost efficiency in fast‑paced, early‑stage settings.

Qualifications

  • 7+ years in software engineering, including 2+ years building agentic LLM-based agents.
  • Experience designing LLM application architectures, including model selection, context/window management, retrieval and orchestration patterns.
  • Proven ability to build evaluation and benchmarking frameworks measuring task completion, cost efficiency, and failure modes.
  • Technical leadership experience setting direction for small teams (3–6 engineers) and performing meaningful code review.
  • Strong Python skills and familiarity with LLM tooling (function calling, tool APIs, observability/tracing, evaluation frameworks).
  • Experience with desktop automation or programmatic control of applications (COM or similar).
  • Domain experience in mechanical engineering, CAD/CAE, PLM, or adjacent industries.
  • Understanding of enterprise deployment constraints on locked‑down corporate workstations.
  • Track record contributing to public benchmarks, publications, or open-source agentic AI projects.

Responsibilities

  • Lead development of the core agent intelligence layer that executes multi‑step workflows across complex desktop engineering software.
  • Report to the CTO and serve as technical lead for a small team of AI engineers, a user researcher, and domain expert contractors.
  • Own the full product loop: define agent capabilities from user stories, build implementations, and benchmark against real workflows.
  • Drive agent task success rate by defining evaluation frameworks, establishing baselines, and iterating to improve completion metrics.
  • Set and enforce per‑task token budgets and track cost per completed workflow to ensure commercial viability.
  • Build rigorous, reproducible evaluation infrastructure grounded in validated user stories.
  • Lead user story mapping and validation through interviews and close collaboration with domain experts.
  • Translate validated user stories into testable evals and close the loop between research and benchmarking.
  • Own agent architecture decisions including tool‑calling, state management, error recovery, model routing, and context management.
  • Act as a player‑coach: write production code, review designs, unblock the team, and raise engineering standards.
  • Collaborate cross‑functionally with integrations, product, and customers during POCs to align agent behavior with real‑world usage.
  • Operate in an early‑stage, high‑impact environment (small team, Series A, Fortune 100 customers, direct line to the CTO).

Job description

Cosmon is hiring a senior technical leader to own the core agent intelligence that turns mechanical engineers' intent into reliable, cost‑efficient multi‑step workflows across desktop engineering tools—this role sits at the intersection of applied agentic AI, user research, and product delivery and will determine the product's real-world value to enterprise customers.

What you'll do
  • Lead development of the core agent intelligence layer that executes multi-step workflows across complex desktop engineering software.
  • Report to the CTO and serve as technical lead for a small team of AI engineers, a user researcher, and domain expert contractors.
  • Own the full product loop: define agent capabilities from user stories, build implementations, and benchmark against real workflows.
  • Drive agent task success rate by defining evaluation frameworks, establishing baselines, and iterating to improve completion metrics.
  • Set and enforce per-task token budgets and track cost per completed workflow to ensure commercial viability.
  • Build rigorous, reproducible evaluation infrastructure grounded in validated user stories.
  • Lead user story mapping and validation through interviews and close collaboration with domain experts.
  • Translate validated user stories into testable evals and close the loop between research and benchmarking.
  • Own agent architecture decisions including tool-calling, state management, error recovery, model routing, and context management.
  • Act as a player‑coach: write production code, review designs, unblock the team, and raise engineering standards.
  • Collaborate cross-functionally with integrations, product, and customers during POCs to align agent behavior with real-world usage.
  • Operate in an early-stage, high-impact environment (small team, Series A, Fortune 100 customers, direct line to the CTO).
What Cosmon is looking for
  • 7+ years in software engineering, including at least 2 years building agentic LLM-based agents that act in the real world.
  • Deep experience designing LLM application architectures, including model selection, context/window management, retrieval, and orchestration patterns.
  • Proven ability to build evaluation and benchmarking frameworks measuring task completion, cost efficiency, and failure modes.
  • Technical leadership experience setting direction for small teams (3–6 engineers) and performing meaningful code review.
  • Strong Python skills and familiarity with LLM tooling (function calling, tool APIs, observability/tracing, evaluation frameworks).
  • Experience with desktop automation or programmatic control of applications (COM or similar).
  • Domain experience in mechanical engineering, CAD/CAE, PLM, or adjacent industries.
  • Understanding of enterprise deployment constraints on locked‑down corporate workstations.
  • Track record contributing to public benchmarks, publications, or open-source agentic AI projects.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Staff Engineer: Lead Agentic AI & Workflow Orchestrator
Staff Engineer: Lead Agentic AI & Workflow Orchestrator

Getclera • San Francisco (CA)

On-site
USD 180,000 - 240,000
Staff Engineer — Agentic AI for Desktop Engineering
Staff Engineer — Agentic AI for Desktop Engineering

Clera • San Francisco (CA)

On-site
Staff Engineer — Agentic AI
Staff Engineer — Agentic AI

Praxy • San Francisco (CA), Northern (KY)

Hybrid
USD 160,000 - 250,000
Software Agentic Engineer
Software Agentic Engineer

penlink • United States

On-site
USD 110,000 - 170,000
Staff Engineer - Agentic AI
Staff Engineer - Agentic AI

Clera • San Francisco (CA)

On-site
USD 160,000 - 250,000
Equity
Visa sponsorship not available
AI Engineer, Internal Enablement & Productivity
AI Engineer, Internal Enablement & Productivity

Air • Pittsburgh

On-site
USD 120,000 - 190,000
Principal Software Engineer, Agentic Engineering
Principal Software Engineer, Agentic Engineering

Jobtailor • Missouri

On-site
USD 180,000 - 240,000
Staff Engineer, AI/LLM Platform
Staff Engineer, AI/LLM Platform

Simulations Plus • Northern (KY)

Hybrid
USD 100,000 - 120,000
Fully remote work
Flexible schedules
Generous vacation policy
+2
Artificial Intelligence Engineer
Artificial Intelligence Engineer

iT Resource Solutions.net,inc • Boston (MA)

On-site
USD 150,000 - 230,000
Agentic AI Engineer
Agentic AI Engineer

Compunnel, Inc. • Dallas (TX)

On-site
USD 120,000 - 150,000