MCP & Tools Python Developer - Agent Evaluation Infrastructure

Mindrift

Town of Florida (NY)

Remote

USD 90,921 - 129,494

Part time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Competitive pay rates up to $80/hour
Flexible working hours
Remote work opportunity
Valuable experience in AI projects

Job summary

A tech consulting firm in New York is seeking experienced Python engineers for a part-time, remote role. You will develop Model Context Protocol servers and tools while collaborating with infrastructure engineers. The ideal candidate has over 4 years of Python development experience, particularly in backend systems, and is familiar with Docker and APIs. Join us to influence the future of AI while enjoying flexible work arrangements and competitive pay rates of up to $80/hour based on experience.

Qualifications

  • 4+ years of Python development experience, ideally in backend or tools.
  • Solid experience building APIs, testing frameworks, or protocol-based interfaces.
  • Familiarity with how LLM agents are prompted, executed, and evaluated.

Responsibilities

  • Developing and maintaining MCP-compatible evaluation servers.
  • Implementing logic to check agent actions against scenario definitions.
  • Working closely with infrastructure engineers to ensure compatibility.

Skills

Python development
API development
Understanding of Docker
Linux CLI
Clear documentation

Tools

Docker
FastAPI

Job description

5 days ago Be among the first 25 applicants

This opportunity is only for candidates currently residing in the specified country. Your location may affect eligibility and rates. Please submit your resume in English and indicate your level of English proficiency.

At Mindrift, innovation meets opportunity. We believe in using the power of collective human intelligence to ethically shape the future of AI.

What We Do

The Mindrift platform, launched and powered by Toloka, connects domain experts with cutting‑edge AI projects from innovative tech clients. Our mission is to unlock the potential of GenAI by tapping into real‑world expertise from across the globe.

Who We're Looking For

Calling all security researchers, engineers, and penetration testers with a strong foundation in problem‑solving, offensive security, and AI‑related risk assessment. If you thrive on digging into complex systems, uncovering hidden vulnerabilities, and thinking creatively under constraints, join us!

We're looking for someone who can bring a hands‑on approach to technical challenges, whether breaking into systems to expose weaknesses or building secure tools and processes. We value contributors with a passion for continuous learning, experimentation, and adaptability.

About The Project

We're on the hunt for hands‑on Python engineers for a new project focused on developing Model Context Protocol (MCP) servers and internal tools for running and evaluating agent behavior. You'll implement base methods for agent action verification, integrate with internal and client infrastructures, and help fill tooling gaps across the team.

What you'll be doing
  • Developing and maintaining MCP‑compatible evaluation servers
  • Implementing logic to check agent actions against scenario definitions
  • Creating or extending tools that writers and QAs use to test agents
  • Working closely with infrastructure engineers to ensure compatibilityOccasionally helping with test writing or debug sessions when needed

Although we're only looking for experts for this current project, contributors with consistent high‑quality submissions may receive an invitation for ongoing collaboration across future projects.

How to get started

Apply to this post, qualify, and get the chance to contribute to a project aligned with your skills, on your own schedule. Shape the future of AI while building tools that benefit everyone.

Requirements

The ideal contributor will have:

  • 4+ years of Python development experience, ideally in backend or tools
  • Solid experience building APIs, testing frameworks, or protocol‑based interfaces
  • Understanding of Docker, Linux CLI, and HTTP‑based communication
  • Ability to integrate new tools into existing infrastructures
  • Familiarity with how LLM agents are prompted, executed, and evaluated
  • Clear documentation and communication skills – you'll work with QA and writers

We also value applicants who have:

  • Experience with Model Context Protocol (MCP) or similar structured agent‑server interfaces
  • Knowledge of FastAPI or similar async web frameworks
  • Experience working with LLM logs, scoring functions, or sandbox environments
  • Ability to support dev environments (devcontainers, CI configs, linters)
  • JS experience
Benefits
  • Get paid for your expertise, with rates that can go up to $80/hour depending on your skills, experience, and project needs
  • Take part in a flexible, remote, freelance project that fits around your primary professional or academic commitments
  • Participate in an advanced AI project and gain valuable experience to enhance your portfolio
  • Influence how future AI models understand and communicate in your field of expertise
Seniority level

Mid‑Senior level

Employment type

Part‑time

Job function

Other

Industries

IT Services and IT Consulting

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

MCP & Tools Python Developer - Agent Evaluation Infrastructure
MCP & Tools Python Developer - Agent Evaluation Infrastructure

Mindrift • North Carolina

Remote
Competitive pay up to $80/hour
Flexible remote freelance work
Valuable experience in advanced AI
MCP & Tools Python Developer - Agent Evaluation Infrastructure
MCP & Tools Python Developer - Agent Evaluation Infrastructure

Mindrift • Arizona

Remote
Flexible freelance project
Competitive pay rates
Valuable experience on an AI project
MCP & Tools Python Developer - Agent Evaluation Infrastructure
MCP & Tools Python Developer - Agent Evaluation Infrastructure

Mindrift • Missouri

Remote
Flexible hours
Competitive compensation
Advanced AI project experience
MCP & Tools Python Developer - Agent Evaluation Infrastructure
MCP & Tools Python Developer - Agent Evaluation Infrastructure

Mindrift • Dallas (TX)

Remote
USD 64,000 - 96,000
Flexible, remote work
Competitive hourly rates up to $80/hour
Participation in advanced AI projects
MCP & Tools Python Developer - Agent Evaluation Infrastructure
MCP & Tools Python Developer - Agent Evaluation Infrastructure

Mindrift • San Antonio (TX)

Remote
Flexible, remote work
Competitive pay up to $80/hour
Experience in advanced AI projects
+1
MCP & Tools Python Developer - Agent Evaluation Infrastructure
MCP & Tools Python Developer - Agent Evaluation Infrastructure

Mindrift • Michigan

Remote
Flexible remote work
Participation in advanced AI projects
Competitive pay rates up to $80/hour
MCP & Tools Python Developer - Agent Evaluation Infrastructure
MCP & Tools Python Developer - Agent Evaluation Infrastructure

Mindrift • Mississippi

Remote
Flexible working hours
Competitive pay up to $80/hour
Valuable project experience in AI
MCP & Tools Python Developer - Agent Evaluation Infrastructure
MCP & Tools Python Developer - Agent Evaluation Infrastructure

Mindrift • New York (NY)

Remote
Up to $80/hour depending on skills
Flexible remote work
Valuable experience in AI project
MCP & Tools Python Developer - Agent Evaluation Infrastructure
MCP & Tools Python Developer - Agent Evaluation Infrastructure

Mindrift • Rhode Island

Remote
Flexible remote work
Competitive pay
Valuable portfolio experience
MCP & Tools Python Developer - Agent Evaluation Infrastructure
MCP & Tools Python Developer - Agent Evaluation Infrastructure

Mindrift • Houston (TX)

Remote
Competitive pay up to $80/hour
Flexible remote workload
Engagement in advanced AI projects