What this role asks for
- Python
- TypeScript
- Java
- C++
Read from the posting's own words: show the exact sentences
- Python: “…answer leakage / reward hacking • Fluency across common ecosystems (Python and at least one of Java / Go / TypeScript / C++) Preferred Qualifications •…”
- TypeScript: “…across common ecosystems (Python and at least one of Java / Go / TypeScript / C++) Preferred Qualifications • Familiarity with SWE-Bench (Verified) or similar…”
- Java: “…• Fluency across common ecosystems (Python and at least one of Java / Go / TypeScript / C++) Preferred Qualifications • Familiarity with SWE-Bench…”
- C++: “…ecosystems (Python and at least one of Java / Go / TypeScript / C++) Preferred Qualifications • Familiarity with SWE-Bench (Verified) or similar repository…”
The role, as Mercor describes it
Evaluate the quality, correctness, and reproducibility of software-engineering benchmark tasks used to train and evaluate a frontier AI lab's models. You'll assess repository-level tasks, reference patches, test harnesses, and grading integrity — and provide clear, rubric-based written feedback.
Basic Qualifications
• 3+ years professional software engineering
• Real open-source contribution or maintainer experience (merged PRs, committer / maintainer roles)
• Strong ability to audit reference patches, test runners, and Docker isolation, and to detect answer leakage / reward hacking
• Fluency across common ecosystems (Python and at least one of Java / Go / TypeScript / C++)
Preferred Qualifications
• Familiarity with SWE-Bench (Verified) or similar repository benchmarks
Read the rest of the description (2 more paragraphs)
• Maintainer history on major Python OSS (Django, Flask, scikit-learn, sympy, pytest, etc.)
• Prior code-review or task-grading experience
Posted by Mercor, reproduced here so you can judge the role before clicking. Original posting ↗
Source: platform job feed · first seen 9d ago · ID list_AAABoEpzuF806Dv0eAZAV5rw
More like this
- Apply ↗Kubernetes Task Auditorup to $90/hrMercorOther AI workWorldwideposted 12d ago
- Apply ↗Meridial (Invisible)Rating & model evaluationWorldwideposted 4h ago
- Apply ↗Senior Software Engineer: AI Evaluation & Benchmarksup to $100/hrAlignerrCoding & software evalWorldwide
- Apply ↗AI Developer Trace Task Auditorup to $90/hrMercorCoding & software evalWorldwideposted 12d ago
- Apply ↗ML Challenge Task Auditorup to $90/hrMercorCoding & software evalWorldwideposted 12d ago
up to $90/hrMercor · SWE-Bench Task Auditor
Apply on Mercor ↗