Python Infrastructure Engineer — LLM Training & Agent Tooling [AS‐L]

OpenTrain AI

Northern (KY)

Hybrid

USD 12,000 - 22,000

Part time

14 days+
Application generator

Get a reply from this recruiter — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

OpenTrain AI is seeking senior-minded Python engineers to design and maintain infrastructure for LLM training workflows and agent tooling. You will architect secure sandboxes and task environments, build reusable repositories and scoring pipelines, and support researchers who depend on these tools.

The listing identifies the experience level as intermediate while requiring at least five years of professional Python experience.

Qualifications

  • 5+ years of professional Python experience, including production code, packaging, async I/O, and legacy-module refactoring.
  • Strong testing mindset with pytest, coverage, and reliability practices.
  • Advanced Linux command-line skills using bash, grep, curl, jq, and systemd.
  • Understanding of basic networking, permissions, and Linux operations.
  • Hands-on Docker expertise, including multi-stage Dockerfiles and docker-compose.
  • Experience designing GitHub Actions or similar CI/CD workflows.
  • Knowledge of CI/CD secrets management and caching.
  • FastAPI or Flask proficiency for modular REST or asynchronous services.
  • Experience with authentication, validation using Pydantic, and service logging.
  • Ability to create devcontainer.json files, Makefiles, .env workflows, and pre-commit hooks.
  • Exposure to LLM or agent infrastructure, such as sandboxes, scoring pipelines, or evaluation frameworks.
  • Security awareness, including least-privilege design, image hardening, and CI scanners such as Trivy or Snyk.
  • Version-control discipline, including semantic commits, PR templates, and code-review practices.

Responsibilities

  • Design and maintain secure sandboxes and task environments.
  • Build reusable repositories, scoring pipelines, and developer environments.
  • Write clean, production-grade, test-driven Python code.
  • Create unit, integration, and functional tests with pytest.
  • Containerize services with Docker and debug multi-stage Dockerfiles.
  • Design CI/CD pipelines that lint, test, build, and deploy.
  • Support researchers who rely on the infrastructure.
  • Explain tools, write concise documentation, and pair-program when needed.
  • Use AI coding assistants safely, including tools such as Cursor, Claude Code, or Copilot.

Skills

Python
Docker
CI/CD
pytest
Linux
FastAPI/Flask
GitHub Actions
Security best practices

Tools

Cursor
Claude Code
Copilot

Job description

About OpenTrain

OpenTrain AI is the hiring and contracting organization for this opportunity and the #1 platform for finding and building careers in AI training and data labeling. OpenTrain helps people discover projects, build their AI-work profile, and apply in minutes. Creating an account is free.

  • Fully remote contract work
  • Part-time schedule of 20+ hours per week
  • Open to candidates in Bangladesh, Georgia, India, Indonesia, Malaysia, Pakistan, the Philippines, Sri Lanka, Thailand, and Vietnam
About AI Training And Agent Evaluation

AI training is the human side of building modern artificial intelligence. Engineers and contributors create the tools, environments, test cases, and evaluation systems that help researchers measure how reliably AI models and agents perform.

This fast-growing field supports cutting-edge systems through work such as coding evaluation, model-output review, scoring pipelines, and developer tooling. Your infrastructure work will help researchers run repeatable, secure experiments.

  • Contribute to infrastructure used in LLM training workflows
  • Support evaluation of AI agent performance
  • Work remotely with flexible project participation
The Role

OpenTrain AI is seeking senior-minded Python engineers to design and maintain infrastructure for LLM training workflows and agent tooling. You will architect secure sandboxes and task environments, build reusable repositories and scoring pipelines, maintain CI/CD systems, and support researchers who depend on these tools.

The listing identifies the experience level as intermediate while requiring at least five years of professional Python experience. The work is fully remote, available to eligible Asia-Low candidates, and conducted in English.

  • Role type: Part-time contractor
  • Time requirement: 20+ hours per week
  • Data type: Computer code programming
  • Subject matter: Python AI infrastructure, LLM tooling, and automation
What You'll Do

You will create reliable, reusable infrastructure for researchers and developers working on AI agent evaluation. The role combines production Python engineering, environment automation, security hardening, testing, and hands-on technical support.

  • Design and maintain secure sandboxes and task environments
  • Build reusable repositories, scoring pipelines, and developer environments
  • Write clean, production-grade, test-driven Python code
  • Create unit, integration, and functional tests with pytest
  • Containerize services with Docker and debug multi-stage Dockerfiles
  • Design CI/CD pipelines that lint, test, build, and deploy
  • Support researchers who rely on the infrastructure
  • Explain tools, write concise documentation, and pair-program when needed
  • Use AI coding assistants safely, including tools such as Cursor, Claude Code, or Copilot
Required Qualifications

Applicants should bring strong professional Python experience and the ability to own infrastructure that must be secure, testable, and dependable. The project specifically calls for at least five years of professional Python work and practical experience across development environments, services, containers, and automation.

  • 5+ years of professional Python experience, including production code, packaging, async I/O, and legacy-module refactoring
  • Strong testing mindset with pytest, coverage, and reliability practices
  • Advanced Linux command-line skills using tools such as bash, grep, curl, jq, and systemd
  • Understanding of basic networking, permissions, and Linux operations
  • Hands-on Docker expertise, including multi-stage Dockerfiles and docker-compose
  • Experience designing GitHub Actions or similar CI/CD workflows
  • Knowledge of CI/CD secrets management and caching
  • FastAPI or Flask proficiency for modular REST or asynchronous services
  • Experience with authentication, validation using Pydantic, and service logging
  • Ability to create devcontainer.json files, Makefiles, .env workflows, and pre-commit hooks
  • Exposure to LLM or agent infrastructure, such as sandboxes, scoring pipelines, or evaluation frameworks
  • Security awareness, including least-privilege design, image hardening, and CI scanners such as Trivy or Snyk
  • Version-control discipline, including semantic commits, PR templates, and code-review practices
Preferred Experience And Collaboration

The strongest candidates will be comfortable supporting technical researchers and adopting modern engineering workflows. Experience training or evaluating AI systems, particularly on coding-focused tasks, is a bonus rather than a stated prerequisite.

  • Comfort mentoring teammates and explaining infrastructure clearly
  • Ability to write concise technical documentation and pair-program
  • Hands-on experience with Cursor, Claude Code, Copilot, or similar AI coding tools
  • Prior experience training or evaluating AI systems is a bonus
  • Coding-focused AI work on platforms such as OpenTrain or Alignerr is a bonus
Compensation And Screening

The structured listing shows a payment rate of $12.50 USD per hour. The project description also specifies experience-based hourly tiers of $9 for Junior, $12 for Middle, and $16 for Senior contributors; final placement follows the project’s applicable tiering.

Selected candidates must complete a quick HackerRank assessment and platform coding test before recruiter interviews. The screening checklist asks candidates to complete the timed HackerRank and coding test within 48 hours of receiving an invitation.

  • Payment type: Hourly
  • Listed rate: $12.50 USD per hour
  • Project tier rates: Junior $9, Middle $12, Senior $16 USD per hour
  • Screening: HackerRank assessment and platform coding test
  • Interview stage: Recruiter interview after successful screening
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI Forward Deployed Engineer
AI Forward Deployed Engineer

OpenTrain AI • Northern (KY)

Hybrid
USD 213,000 - 288,000
GitHub Contributor for AI Code Evaluation
GitHub Contributor for AI Code Evaluation

OpenTrain AI • Northern (KY)

Hybrid
USD 69,000 - 207,000
Remote Python Infra for LLM Training and Agent Tooling
Remote Python Infra for LLM Training and Agent Tooling

OpenTrain AI • Northern (KY)

Hybrid
USD 12,000 - 22,000
Applied Machine Learning Task Auditor
Applied Machine Learning Task Auditor

OpenTrain AI • Northern (KY)

Hybrid
USD 96,000 - 124,000
Senior Python Engineer - AI Testing Project (Freelance, Mindrift)
Senior Python Engineer - AI Testing Project (Freelance, Mindrift)

Mindrift • Town of Florida (NY)

Remote
Freelance project-based collaboration
Fully remote and flexible participation
Task-based compensation up to $80/hour
+2
Software Engineer, Internal Platforms
Software Engineer, Internal Platforms

OpenTrain AI • Northern (KY)

Hybrid
USD 120,000 - 155,000
Senior Python Engineer - AI Testing Project (Freelance, Mindrift)
Senior Python Engineer - AI Testing Project (Freelance, Mindrift)

Mindrift • North Carolina

On-site
Freelance project-based collaboration
Fully remote and flexible participation
Task-based compensation
Senior Python Engineer - AI Testing Project (Freelance, Mindrift)
Senior Python Engineer - AI Testing Project (Freelance, Mindrift)

Mindrift • Dallas (TX)

On-site
Fully remote participation
Task-based compensation
Opportunity to work on AI projects
+1
Senior Engineering and Software Domain Expert
Senior Engineering and Software Domain Expert

OpenTrain AI • California (MO), Northern (KY)

Hybrid
USD 90,000 - 145,000
MCP & Tools Python Developer - Agent Evaluation Infrastructure
MCP & Tools Python Developer - Agent Evaluation Infrastructure

Mindrift • New York (NY)

Remote
Up to $80/hour depending on skills
Flexible remote work
Valuable experience in AI project