Member of Technical Staff

Chakra Labs

New York (NY)

On-site

USD 100,000 - 140,000

Full time

14 days+
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

A tech innovation company in New York is looking for a skilled professional to manage infrastructure for AI agents. The role involves orchestrating container environments, maintaining distributed systems, and ensuring LLM workloads run efficiently. Ideal candidates will have experience with Kubernetes, SQS, and Kafka, alongside a passion for working in a fast-paced, early-stage team environment. You will have the chance to collaborate with AI researchers directly and influence the future of agent technology.

Qualifications

  • Experience with container orchestration tools in production environments.
  • Knowledge of distributed systems and message-driven architectures.
  • Experience in running LLM workloads at scale.

Responsibilities

  • Manage agent orchestration and dispatch layer for AI workloads.
  • Design environments and tasks for agent evaluation.
  • Monitor and maintain model behavior and infrastructure.

Skills

Container orchestration
Distributed systems
LLM infrastructure
Debugging token throughput

Tools

Kubernetes
SQS
Kafka

Job description

What You'd Work On
  • Agent orchestration at scale. Hundreds of agent runs at once, each with its own stateful environment. 100M tokens per minute across the fleet. You own the dispatch layer: SQS, concurrency control, failure handling.

  • Environment and task design. We need environments that feel real and scenarios that actually push agents to their limits. You'd figure out how to build new evaluations and design the tasks that test what matters, not just what's easy to measure.

  • New frontiers. The agent evaluation space is moving fast. You'd stay on that edge, supporting new environment modalities and shipping integrations with external orchestration frameworks.

  • Observability. Prometheus and OpenTelemetry across services, Grafana dashboards, structured logging.

About You
  • Container orchestration. You're comfortable running Kubernetes or similar in production. Auto-scaling, pod lifecycle, persistent storage, networking. You can figure out why something won't schedule and reason about resource contention.

  • Distributed systems. You've built or maintained message-driven architectures. SQS, Kafka, or similar. You know how to keep jobs moving when things back up, retry without duplicating, and fail without losing work.

  • LLM infrastructure. You've run LLM workloads at scale. Token instrumentation, rate limit handling, prompt caching, multi-provider routing. You've built the plumbing between models and external tools, and you know what it takes to keep it all running under load.

  • Experience. No hard rule. Ideally at least 3 years at this level, but less works if the above sounds like you.

What Makes This Different
  • It's infra, but the workload is AI agents. You're monitoring model behavior alongside pod health, debugging token throughput alongside network throughput.

  • Our customers are AI researchers and labs. You'd work directly with the people pushing the frontier of what agents can do, and build the infrastructure they run it on.

  • Early-stage team. You own whole systems, not tickets in a queue. One week you're shipping a new environment type, the next you're scaling the dispatch layer to handle 10x the throughput.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Staff Software Engineer, Agent Eval Platform
Staff Software Engineer, Agent Eval Platform

Servicenow • Santa Clara (CA)

On-site
USD 180,000 - 320,000
Software Engineering Superbuilder, AI-DNA, $200k/year USD
Software Engineering Superbuilder, AI-DNA, $200k/year USD

IgniteTech • United States

On-site
USD 120,000 - 160,000
Member of Technical Staff (Applied AI Engineer, Agent Capabilities)
Member of Technical Staff (Applied AI Engineer, Agent Capabilities)

United States Digital Space LLC • San Francisco (CA)

On-site
USD 190,000 - 230,000
AI Engineer, Agent Builder (Remote).
AI Engineer, Agent Builder (Remote).

Catalyst Wayfare • United States

Remote
USD 120,000 - 170,000
Python AI/ML with Full stack
Python AI/ML with Full stack

Programmers.io • Dallas (TX)

On-site
USD 180,000 - 230,000
Member of Technical Staff, Forward Deployed
Member of Technical Staff, Forward Deployed

Chakra Labs • New York (NY)

On-site
USD 130,000 - 190,000
Staff ML Engineer, Agent Training & Environments
Staff ML Engineer, Agent Training & Environments

EngineersOfAI • San Francisco (CA)

On-site
USD 170,000 - 260,000
Staff AI Engineer - End-to-End Agent Systems
Staff AI Engineer - End-to-End Agent Systems

Vela • San Francisco (CA)

On-site
Member of Technical Staff, AI Research - Post-Training
Member of Technical Staff, AI Research - Post-Training

Chakra Data Warehouse • New York (NY)

On-site
USD 120,000 - 160,000
Software Engineer, Agent Infrastructure
Software Engineer, Agent Infrastructure

Sapiom • San Francisco (CA)

On-site
USD 150,000 - 210,000