Senior/Staff Software Engineer (Platform and Execution Model)

Trase

United States

On-site

USD 180,000 - 240,000

Full time

4 days ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Trase Systems is hiring a Senior or Staff Software Engineer to own critical parts of Trase OS, a platform powering AI deployments in regulated environments. You’ll tackle distributed systems, ensure reliability, and guide architectural decisions across teams.

Ideal candidates have 8+ years of distributed systems experience, with a focus on durable execution, security, observability, and modern languages such as Go, Rust, Java, or TypeScript. Strong governance and auditable designs are essential.

Qualifications

  • 8+ years building distributed/platform systems.
  • 4+ years owning mission‑critical runtimes or workflow/orchestration systems.
  • Durable execution and idempotency expertise.
  • Security & governance in production systems (auth, RBAC, audit, policy).
  • Observability with tracing across async boundaries.
  • Strong system design across storage, queues, schedulers, and evented architectures.

Responsibilities

  • Design, build, and own critical components of the core execution model (state machine, lifecycle, resource model, failure semantics).
  • Build reliable distributed systems that behave predictably across retries, restarts, partial failures, and concurrent execution.
  • Develop platform APIs/SDKs connecting workflows, agents, tools, and product surfaces; drive versioning & compatibility.
  • Guarantee correctness via idempotency, deterministic replays, compensating actions, and data integrity.
  • Engineer reliability at scale: concurrency controls, rate limits, backpressure, sharding/partitioning, and workload isolation.
  • Build security & governance into the core: RBAC/ABAC, policy enforcement, audit & lineage.
  • Deliver observability: distributed tracing, structured logs, metrics, evaluation hooks; explainable trail of agent actions.
  • Own quality: design reviews, test strategy, performance baselines, SLOs, incident response, postmortems.
  • Mentor engineers and establish engineering patterns across the platform.

Skills

Distributed systems
Architecture leadership
Workflow/orchestration
Security & governance
Observability
Systems design
Go/Rust/Java/TypeScript
Regulated environments

Tools

Grafana
Kubernetes
CI/CD

Job description

About Us Co-founded in 2023 by Joe Laws and Grant Verstandig , Trase Systems is AI, Uncomplicated. Trase empowers enterprise leaders to harness the full potential of AI without the associated complexity and risks. We are an end-to‑end solution for deploying, managing, and optimizing AI in the enterprise. Our platform specializes in bridging the “last mile” of AI adoption, unlocking AI's full potential while driving efficiency and significant cost savings. Trase is at the forefront of AI Agent innovation, topping the Hugging Face GAIA Leaderboard for Generalized AI Assistants, ahead of industry giants such as Google, Meta, Microsoft, and OpenAI. We are leveraging our cutting‑edge technologies to develop mission‑critical agentic applications in complex industries such as Healthcare, Oil & Gas, and National Security.

About The Role

As a Senior or Staff Software Engineer, you'll build and own critical parts of Trase OS, the shared platform that powers Trase deployments in regulated environments. You'll work on the distributed systems and platform primitives connecting workflows, agents, tools, and product surfaces, with a particular focus on reliability, scalability, and correctness. You'll take on technically ambiguous problems, design solutions, and drive them through production. This is a hands‑on engineering role for someone who is comfortable going deep on distributed systems while helping other engineers make sound architectural decisions. Clean abstractions and correctness‑under‑failure are critical because we operate long‑lived agents in healthcare/defense environments where auditability and reliability are non‑negotiable. The level will reflect your experience and demonstrated scope. Staff‑level candidates will be expected to lead architecture across teams, establish engineering standards, and mentor other engineers.

Why This Role is Needed

Trase OS is an orchestration‑heavy system coordinating long‑lived workflows, agents, and tools across multiple services and environments. As the platform evolves, the primary risks shift from implementation to system design quality: Poor abstractions create tight coupling across services Workflow execution becomes difficult to reason about under failure Platform capabilities fragment instead of becoming reusable primitives Scaling introduces complexity instead of leverage This role exists to: Design and implement clean, durable abstractions for the platform execution model Ensure correctness and determinism in workflow execution Translate evolving product requirements into coherent platform architecture Enable teams to build on Trase OS without introducing systemic complexity

What Makes This Role Challenging

You are designing systems where failure is the norm, not the exception, and correctness must be preserved across retries, restarts, and partial execution You must balance clean abstractions with real‑world constraints (performance, security, multi-tenant environments) Decisions made here become foundational primitives used across all products and teams The system must remain understandable and auditable, even as complexity and scale increase

Responsibilities
  • Design, build, and own critical components of the core execution model (state machine, lifecycle, resource model, failure semantics)
  • Build reliable distributed systems that behave predictably across retries, restarts, partial failures, and concurrent execution.
  • Develop platform APIs/SDKs connecting workflows, agents, tools, and product surfaces; drive versioning & compatibility
  • Guarantee correctness via idempotency, deterministic replays, compensating actions, and data integrity
  • Engineer reliability at scale: concurrency controls, rate limits, backpressure, sharding/partitioning, and workload isolation
  • Build security & governance into the core: RBAC/ABAC, policy enforcement, fine‑grained audit & lineage
  • Deliver observability: distributed tracing, structured logs, metrics, and evaluation hooks; build an “explainable trail” of agent actions
  • Own quality: design reviews, test strategy (unit, property, chaos), performance baselines, SLOs, incident response, and postmortems
  • Mentor engineers, contribute to design reviews, and help establish strong engineering patterns across the platform.
Requirements
  • 8+ years of experience building distributed/platform systems, including significant experience defining architecture across teams or domains
  • 4+ years owning mission‑critical runtimes or workflow/orchestration systems
  • Deep expertise with durable execution (e.g., state machines, event sourcing, saga/compensation, idempotency, exactly/at‑least‑once semantics)
  • Proven track record with security & governance in production systems (auth, RBAC, audit, policy)
  • Hands‑on with observability (Grafana or equivalent), including trace correlation across async boundaries
  • Strong systems design across storage, queues, schedulers, and evented architectures; performance tuning under load
  • Excellence in a modern language (e.g., Go, Rust, Java, or TypeScript) and cloud‑native stacks (containers, CI/CD, IaC)
  • Comfortable operating in regulated or high‑assurance environments; bias toward correctness, clarity, and documentation
  • Strong technical judgment and the ability to influence design decisions within a team and across closely related engineering areas.
  • Ability to incorporate advance LLM capabilities into system design and platform architecture decisions where appropriate
  • Nice to Have Prior work on workflow engines (Temporal/Cadence/AWS Step Functions, Argo, Airflow) or serverless runtimes
  • Experience with policy engines (OPA), secrets/KMS, or data‑handling controls (PII/PHI)
  • ML/LLM evaluation frameworks, tool/plugin architectures, or embedding model governance into execution
  • Government or healthcare experience (HIPAA, audit readiness) and multi‑tenant isolation

Salary Range: $180,000-240,000.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior/Staff Software Engineer (Platform and Execution Model)
Senior/Staff Software Engineer (Platform and Execution Model)

redcellpartners • United States

On-site
USD 180,000 - 260,000
Senior/Staff Software Engineer (Platform and Execution Model) New Seattle, WA (Preferred) or McLean, VA or Remote (USA)
Senior/Staff Software Engineer (Platform and Execution Model) New Seattle, WA (Preferred) or McLean, VA or Remote (USA)

Trase Systems, Inc. • Seattle (WA), Northern (KY)

Hybrid
USD 144,000 - 240,000
Health care (medical, dental, vision)
Maternity/Paternity leave
Unlimited PTO
+4
Senior/Staff Software Engineer (Platform and Execution Model)
Senior/Staff Software Engineer (Platform and Execution Model)

Red Cell Partners, LLC. • Seattle (WA), McLean (VA)

Hybrid
USD 180,000 - 240,000
Health care benefits
Unlimited PTO
Professional development
+3
Senior/Staff DevOps Engineer, Platform Infrastructure
Senior/Staff DevOps Engineer, Platform Infrastructure

Trase Systems, Inc. • Seattle (WA)

On-site
USD 180,000 - 260,000
Senior/Staff Platform Engineer - AI Orchestration
Senior/Staff Platform Engineer - AI Orchestration

Trase • Seattle (WA)

On-site
USD 180,000 - 240,000
Career growth opportunities
Comprehensive health benefits for you
Professional development and equity
Senior/Staff Software Engineer (Platform and Execution Model)
Senior/Staff Software Engineer (Platform and Execution Model)

Red Cell Partners, LLC. • United States

On-site
USD 180,000 - 240,000
Career track advancement
Employer-paid health care
Paid maternity/paternity leave
+2
Senior/Staff Software Engineer (Platform and Execution Model)
Senior/Staff Software Engineer (Platform and Execution Model)

Trase • Seattle (WA)

On-site
USD 180,000 - 240,000
Career growth opportunities
Comprehensive health benefits for you
Professional development and equity
Senior/Staff Distributed Systems Engineer — Platform Lead
Senior/Staff Distributed Systems Engineer — Platform Lead

redcellpartners • United States

On-site
USD 180,000 - 260,000
Sales Engineer/Solution Architect (Contract)
Sales Engineer/Solution Architect (Contract)

Trase Systems, Inc. • United States

On-site
USD 170,000 - 200,000
Health coverage for you and family
Unlimited PTO
401K and equity incentives
Senior Product Manager
Senior Product Manager

Trase Systems, Inc. • United States

Remote
USD 190,000 - 220,000
Career track growth
100% employer paid health care for you
Paid maternity and paternity leave
+4