Member of Technical Staff, Coding Research

Pro Integrate LLC

New York (NY)

Remote

USD 200,000 - 260,000

Full time

12 hours ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Benefits offered by this job

Equity compensation
Bonuses
Health insurance premium reimbursement
Paid time off
401(k) with company match
Remote-first work environment

Job summary

Pro Integrate LLC is seeking a Member of Technical Staff to advance the evaluation and development of frontier coding agents. You will work at the intersection of AI research, software engineering, and model evaluation to design benchmarks, data systems, and experiments that shape how coding models are measured and improved.

You will collaborate with researchers, engineers, and applied AI teams to design experiments, measure model capabilities, and communicate findings through technical reports

Qualifications

  • Strong software engineering background with Python or C++.
  • 3+ years of experience in software engineering, ML, AI research, evaluation, or related disciplines.
  • Experience designing, reviewing, or validating technical assessments, benchmarks, coding tasks, or evaluation methodologies.
  • Familiarity with LLMs, coding agents, RL, model evaluation, or related AI systems.
  • Proven ability to build tooling, automate workflows, and improve technical processes.
  • Strong analytical skills to investigate model behavior and derive insights.
  • Excellent written and verbal communication to convey findings.
  • Comfortable in fast-moving research environments with ambiguity.

Responsibilities

  • Design and own evaluation frameworks for coding agents, including benchmarks and scoring rubrics.
  • Lead end-to-end research initiatives measuring coding model performance.
  • Develop high-quality datasets, golden examples, and evaluation protocols.
  • Analyze model behavior and identify weaknesses for improvements in training/evaluation.
  • Build tooling and infrastructure for large-scale experimentation and review workflows.
  • Establish best practices for coding-agent assessment, ensuring rigor and reproducibility.
  • Partner with researchers, engineers, and applied AI teams to design experiments.
  • Contribute to technical reports and benchmark studies.

Skills

Python
C++
Software engineering
ML/AI research
Benchmark design
Experimentation
Analytical skills
Communication

Tools

Git
CI/CD

Job description

Job Title: Member of Technical Staff, Coding Research
Job Type: Full-time
Location: Remote

The Role

We are seeking a Member of Technical Staff to help advance the evaluation and development of frontier coding agents. Sitting at the intersection of AI research, software engineering, and model evaluation, you will design the benchmarks, methodologies, and data systems that shape how next-generation coding models are measured and improved.

What You'll Do
  • Design and own evaluation frameworks for coding agents, including benchmark specifications, scoring methodologies, rubrics, and quality standards.
  • Lead end-to-end research initiatives focused on measuring and improving coding model performance across diverse software engineering tasks.
  • Develop high-quality datasets, golden examples, and evaluation protocols that enable reliable assessment of frontier coding systems.
  • Analyze model behavior and failure modes, identifying systematic weaknesses and translating findings into actionable improvements for training and evaluation.
  • Build tooling and infrastructure that support large-scale experimentation, data generation, review workflows, and evaluation pipelines.
  • Establish best practices for coding-agent assessment, ensuring methodological rigor, reproducibility, and measurement quality.
  • Partner closely with researchers, engineers, and applied AI teams to design experiments and evaluate emerging model capabilities.
  • Contribute to technical reports, benchmark studies, and client-facing research initiatives that communicate model performance and insights.
What We're Looking For
  • Strong software engineering background with expertise in Python, C++, or comparable programming languages.
  • 3+ years of experience in software engineering, machine learning, AI research, evaluation, or related technical disciplines.
  • Experience designing, reviewing, or validating technical assessments, benchmarks, coding tasks, or evaluation methodologies.
  • Familiarity with large language models, coding agents, reinforcement learning, model evaluation, or related AI systems.
  • Proven ability to build tooling, automate workflows, and improve technical processes through systematic experimentation.
  • Strong analytical skills with the ability to investigate model behavior and derive insights from complex technical systems.
  • Excellent written and verbal communication skills, including the ability to clearly articulate technical findings to diverse audiences.
  • Comfortable operating in fast-moving research environments with significant ambiguity and evolving priorities.
Preferred
  • Experience working on frontier AI systems, coding agents, or model evaluation research.
  • Deep interest in understanding how data, evaluations, and feedback mechanisms influence model capabilities.
  • Track record of independently driving ambiguous technical or research projects from conception to execution.
  • Experience designing benchmarks or datasets for machine learning systems at scale.
  • Familiarity with agentic workflows, tool use, reinforcement learning, or post-training methodologies.
  • Publications, open-source contributions, or demonstrated technical leadership in AI, machine learning, or software engineering.
Compensation & Benefits Notice

The national pay range for this full-time position is base salary of $200,000 $260,000 USD. All employees are eligible for equity compensation, and employees may also receive performance-based bonuses, dependent on role and subject to company policies. micro1 provides a comprehensive benefits package, including up to 100% reimbursement for health-insurance premiums, paid time off, a 401(K) plan with a company match, and additional benefits designed to support a high-performing, remote-first workforce.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Member of Technical Staff, Coding Research
Member of Technical Staff, Coding Research

Remotebridge • United States

Remote
USD 200,000 - 260,000
Equity compensation
Bonus eligibility
Health-insurance reimbursement
+2
Remote | Member of Technical Staff, Coding Research — $600,000–$1,300,000/year
Remote | Member of Technical Staff, Coding Research — $600,000–$1,300,000/year

24-MAG • United States

Remote
USD 600,000 - 1,300,000
Member of Technical Staff, Frontier AI
Member of Technical Staff, Frontier AI

Pro Integrate LLC • New York (NY)

Remote
USD 240,000 - 350,000
Equity compensation
Health insurance reimbursement
401(k) with company match
+2
Coding Research Evaluation Engineer
Coding Research Evaluation Engineer

Remotebridge • United States

Remote
USD 200,000 - 260,000
Equity compensation
Bonus eligibility
Health-insurance reimbursement
+2
Member of Technical Staff, Enterprise AI
Member of Technical Staff, Enterprise AI

Pro Integrate LLC • New York (NY)

Remote
USD 200,000 - 250,000
Equity compensation
Performance-based bonuses
Health insurance reimbursement
+3
Member of Technical Staff, Finance Research
Member of Technical Staff, Finance Research

Pro Integrate LLC • New York (NY)

Remote
USD 200,000 - 250,000
Equity compensation
Performance-based bonuses
Remote-first workforce
+2
Software Engineering Expert
Software Engineering Expert

Weekday AI • United States

Remote
USD 83,000 - 124,000
Forward Deployed Engineer
Forward Deployed Engineer

Pro Integrate LLC • New York (NY)

Remote
USD 180,000 - 250,000
Equity compensation
Remote-first workforce
Health insurance reimbursement
Remote | Research Engineer - Code Generation & Model Evaluation — $50–$100/hour
Remote | Research Engineer - Code Generation & Model Evaluation — $50–$100/hour

24-MAG • United States

Remote
USD 69,000 - 138,000
Remote work
Flexible hours
Contractor engagement
Machine Learning Engineer - Model Evaluation & Experimentation
Machine Learning Engineer - Model Evaluation & Experimentation

Weekday 1 • United States

Remote
USD 83,000 - 124,000