Member of Technical Staff, Forward Deployed

Chakra Labs

New York (NY)

On-site

USD 130,000 - 190,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Chakra Labs in New York, NY is seeking an engineer to tackle frontier problems by building high‑fidelity environments, evaluation tasks, and datasets for frontier AI research. You will push the boundaries of what agents can do, shipping practical deliverables.

You will own end-to-end results, collaborate with researchers and customers, and iterate rapidly in a fast-paced setting that values measurable impact and reliability.

Qualifications

  • Full-stack engineering across backend services, data pipelines, and frontend to ship a usable interface.
  • Turn loosely-defined research questions into concrete environments, tasks, evals, or datasets.
  • Communicate clearly with researchers and technical customers, balancing requests with needs.
  • Interest in how agents fail and how to measure it during design and testing.
  • Ideally 3+ years shipping production software, with background aligning to frontier ML work.

Responsibilities

  • Build and deploy high-fidelity environments, evaluation tasks, and datasets for frontier AI research.
  • Extend and generalize custom builds into core products.
  • Collaborate with researchers and customers to ensure outputs are useful and measurable.

Skills

Full-stack development
Ambiguity handling
Customer communication
Evals & datasets
Production software

Job description

About Us

Chakra Labs' mission is to encode human taste into intelligence. We build high-fidelity environments, evals, and datasets for frontier AI research, working with several of the top labs.


Our work sits at the frontier of post-training, agent environments, data quality, and research infrastructure. We care about building systems that make models better in ways that are measurable, useful, and hard to fake.


What You'd Work On


  • The hardest problems at the frontier. A new environment modality, an eval targeting a failure mode nobody's measured, a dataset that doesn't exist yet. You take problems like these from a researcher's hunch to a shipped deliverable, working at the edge of what agents can currently do.


  • Environments, evals, and datasets. One project is a high-fidelity environment, the next is a task distribution with grading logic, the next is a dataset built to a demanding spec. The bar is frontier-lab quality and the pace is relentless - you're writing whatever the deliverable needs: environment code, task specs, scoring harnesses.


  • Pulling the frontier into the platform. The best one-offs don't stay one-offs. You'd recognize when a custom build proves out a capability worth generalizing, and help fold it into the core product - so it compounds instead of sitting on a shelf.



About You


  • Full-stack range. You're a strong generalist engineer - comfortable across backend services, data pipelines, and enough frontend to ship a usable interface. TypeScript/Python or similar. You'd rather own a whole deliverable than a layer of one.


  • Fast in ambiguity. Requirements arrive as a hunch, not a spec. You can turn a loosely-defined research question into a concrete environment, task set, eval, or dataset - and validate that it measures what was meant, not just what was easy to build.


  • Customer instincts. You communicate clearly with researchers and technical customers, push back when they're asking for the wrong thing, and know the difference between what someone requests and what they need. The work moves between deep async stretches and high-bandwidth sessions with customers, in person when it counts.


  • Evals curiosity. You don't need an ML research background, but you're genuinely interested in how agents fail and how to measure it. You'll learn task design, LLM-judged scoring, and reward hacking detection on the job.


  • Experience. No hard rule. Ideally at least 3 years shipping production software, but less works if the above sounds like you.



What Makes This Different


  • You see the frontier first. The problems that land on your desk are the ones frontier researchers haven't solved yet. What you build is often the first working version of something the field will need - and you're the one who proves it's possible.


  • Real customers, real deadlines. Every project has a named customer and a researcher waiting on the output. This isn't building a platform and hoping someone uses it.


  • Ownership, not theater. You own whole deliverables and customer relationships, not tickets in a queue. One week you're shipping a custom eval for a lab, the next you're generalizing it into the core product.


  • The team. Our team is ex-Stripe, Snap, AWS, Microsoft, Airtable - you'll work with a small team who has years of shipping high-impact products over the last decade.


  • Cutting edge. You will get to touch the latest and greatest technologies across the data, AI, and infrastructure stack.


Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Member of Technical Staff, AI Research - Post-Training
Member of Technical Staff, AI Research - Post-Training

Chakra Data Warehouse • New York (NY)

On-site
USD 120,000 - 160,000
Applied Research Scientist
Applied Research Scientist

Fleet AI, Inc. • New York (NY)

On-site
USD 100,000 - 150,000
Software Engineering Superbuilder, AI-DNA, $200k/year USD
Software Engineering Superbuilder, AI-DNA, $200k/year USD

IgniteTech • United States

On-site
USD 120,000 - 160,000
Senior Backend Engineer – Agents (USA Only - 100% Remote)
Senior Backend Engineer – Agents (USA Only - 100% Remote)

close • United States

Remote
USD 140,000 - 190,000
Member of Technical Staff
Member of Technical Staff

Chakra Labs • New York (NY)

On-site
USD 100,000 - 140,000
Software Engineer (New Grad)
Software Engineer (New Grad)

Maximor AI • New York (NY)

On-site
USD 150,000 - 210,000
Agent Harness Engineer (Remote)
Agent Harness Engineer (Remote)

Viktor • New York (NY)

Hybrid
USD 180,000 - 280,000
Remote-first
AI Engineer
AI Engineer

Valsoft Corporation • Northern (KY)

Hybrid
USD 120,000 - 180,000
Applied Research Scientist
Applied Research Scientist

Fleet AI, Inc. • Buffalo (NY)

On-site
USD 150,000 - 210,000
Staff Engineer
Staff Engineer

Mira Mace • San Francisco (CA)

On-site
USD 180,000 - 240,000