Remote | Senior Software Engineer – LLM Evaluation (US/Canada/WEU based)

24-Mag Llc

United States

Remote

USD 14,000 - 55,000

Part time

2 days ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Benefits offered by this job

Fully remote
Flexible hours
Contract-based

Job summary

24-MAG LLC is seeking experienced software engineers for a part-time, fully remote consulting role focused on evaluating advanced language models and coding benchmarks. You will curate code, verify solutions, and assess AI-generated software across languages, contributing to research collaboration and reproducible evaluation processes.

Ideal candidates have 3+ years in software engineering, strong full-stack skills, and familiarity with Python, JavaScript, Java, C++, Rust, or Go.

Qualifications

  • 3+ years of professional software-engineering experience.
  • Strong full-stack development capabilities.
  • Experience building scalable, production-grade software.
  • Deep knowledge of development, debugging, and code-quality assessment.

Responsibilities

  • Curate high-quality code examples for model training and benchmarking.
  • Develop precise solutions to software-engineering tasks across multiple languages.
  • Evaluate AI-generated code for technical correctness and efficiency.
  • Assess software design, architecture, and production implementation decisions.
  • Collaborate with research and cross-functional teams on evaluation strategies.

Skills

Python
JavaScript
Java
C++
Rust
Go
ReactJS

Job description

We are sharing a specialised part-time consulting opportunity for experienced software engineers to contribute to advanced large language model evaluation, coding benchmark development, and AI-assisted software-engineering research.

Selected professionals will curate and evaluate code, develop verification mechanisms, assess AI-generated software across multiple programming languages, and help research teams understand how advanced models perform throughout realistic software-development workflows.

Key Responsibilities
Code Curation & Solution Development
  • Curate high-quality code examples for model training and benchmarking
  • Develop precise solutions to software-engineering tasks
  • Correct and improve code across multiple programming languages
  • Work with Python, JavaScript, ReactJS, C/C++, Java, Rust, and Go
  • Maintain strong standards for correctness and maintainability
AI-Generated Code Evaluation
  • Evaluate AI-generated code for technical correctness
  • Assess solutions for efficiency, scalability, and reliability
  • Identify implementation weaknesses and recurring error patterns
  • Review code quality against professional engineering standards
  • Provide structured rationales supporting evaluation decisions
Verification & Automated Assessment
  • Build agents that assess code quality
  • Design mechanisms for automatically verifying software solutions
  • Identify recurring model-generated coding errors
  • Develop reliable checks for engineering tasks
  • Support reproducible evaluation across repeated assignments
Software Engineering Lifecycle Evaluation
  • Evaluate model capabilities across the software-development lifecycle
  • Assess reasoning around prototyping and architecture design
  • Review API design and production implementation decisions
  • Evaluate launch, experimentation, monitoring, and maintenance scenarios
  • Identify areas where models struggle with real-world engineering workflows
Research & Benchmark Collaboration
  • Collaborate with research and cross-functional technical teams
  • Contribute to datasets used for training and benchmarking
  • Help define engineering evaluation strategies
  • Compare model performance against professional engineering expectations
  • Support iterative improvements to coding-focused evaluation systems
Ideal Profile
  • 3+ years of professional software-engineering experience
  • Strong full-stack development capabilities
  • Experience building scalable, production-grade software
  • Strong understanding of software architecture and system design
  • Deep knowledge of development, debugging, and code-quality assessment
  • Experience reviewing and improving complex software implementations
  • Proficiency in one or more of Python, JavaScript, Java, C++, Rust, or related languages
  • ReactJS, C, or Go experience may also be relevant to project assignments
  • Strong understanding of API design and production implementation
  • Familiarity with software monitoring and operational maintenance
  • Ability to reason across the complete software-engineering lifecycle
  • Strong analytical and problem-solving capabilities
  • Excellent written and verbal communication skills
  • Ability to provide clear, structured evaluation rationales
  • Comfortable collaborating remotely with research and technical teams
Engagement Details
  • Part-time independent contractor engagement
  • Fully remote
  • Candidates must be based in the United States, Canada, or eligible Western European (WEU) countries
  • Source examples of WEU locations include Austria, Belgium, France, and Germany
  • Minimum commitment: 10 hours per week
  • Flexible workload of up to 40 hours per week
  • Initial project duration is approximately 1 month
  • Extension may be available depending on performance and project fit
  • No medical or paid-leave benefits are included under the contractor arrangement
  • Application process takes approximately 15–30 minutes
  • Completion of an AI video interview is required
  • Compensation is not specified in the source materials
  • Work must be completed without using confidential, proprietary, unreleased, employer-restricted, client-restricted, or otherwise protected code, datasets, architecture materials, or technical information belonging to any employer, client, institution, or other third party

By submitting this application, you acknowledge that your information may be processed by 24-MAG LLC for recruitment and opportunity matching in accordance with our Privacy Policy: https://www.24-mag.com/privacy-policy

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Remote | Senior Software Engineer – LLM Evaluation
Remote | Senior Software Engineer – LLM Evaluation

24-Mag Llc • New York (NY)

Remote
USD 14,000 - 55,000
Remote | Senior Software Engineer – LLM Evaluation & Repository Validation
Remote | Senior Software Engineer – LLM Evaluation & Repository Validation

24-Mag Llc • New York (NY)

Remote
USD 83,000 - 138,000
Fully remote
Remote | Senior Software Engineer — $50–$100/hour
Remote | Senior Software Engineer — $50–$100/hour

engineeringjobs.net, Inc. • New York (NY)

Remote
USD 69,000 - 138,000
Remote Software Engineer - C++
Remote Software Engineer - C++

turing • San Francisco (CA)

Hybrid
USD 83,000 - 124,000
Remote
Remote

24-MAG • New York (NY)

Remote
USD 138,000 - 207,000
Senior AI Software Engineer (Remote)
Senior AI Software Engineer (Remote)

Quik Hire Staffing • United States

Remote
GBP 90,000 - 130,000
Remote Senior Python Engineer – LLM Evaluation (US-based)
Remote Senior Python Engineer – LLM Evaluation (US-based)

Turing • Chicago (IL)

On-site
USD 68,880 - 110,208
Remote Software Engineer (C++)
Remote Software Engineer (C++)

turing • San Francisco (CA)

Remote
USD 83,000 - 117,000
Remote Software Engineer - Java
Remote Software Engineer - Java

turing • San Francisco (CA)

Remote
USD 83,000 - 152,000
Contractor—no medical/paid leave
Remote Senior Software Engineer: AI Code Evaluation
Remote Senior Software Engineer: AI Code Evaluation

24-Mag Llc • United States

Remote
USD 14,000 - 55,000
Fully remote
Flexible hours
Contract-based