Remote | Academic Research & Python Task Consultant - $70-$100/hour
About the job Remote | Academic Research & Python Task Consultant - $70-$100/hourWe are sharing a specialised part-time consulting opportunity for professors, PhD students, and advanced academic researchers experienced in domain-specific problem design, Python-based evaluation, benchmark task development, and structured reasoning assessment.This role supports current and upcoming remote consulting opportunities focused on academic benchmark task design, Python-based evaluation workflows, domain-specific problem development, golden solution preparation, model behavior analysis, and high-quality project execution. Selected professionals will apply their academic expertise to create challenging real-world tasks, define precise expected outputs, develop executable tests, and evaluate reasoning or problem-solving performance across advanced subject areas.Key ResponsibilitiesProfessionals in this role may contribute to:Academic Task Design & DevelopmentDesign challenging, real-world problems drawn from your academic or professional domainCreate tasks across areas such as machine learning, coding, data science, computer science, physics, mathematics, engineering, statistics, biology, chemistry, finance, accounting, economics, law, or businessBuild tasks that test reasoning, problem solving, instruction following, and domain-specific judgmentEnsure task prompts are clear, rigorous, realistic, and aligned with expert-level expectationsPython-Based Solutions & Evaluation AssetsPrepare task specifications, golden solutions, and supporting evaluation components using PythonDevelop executable tests or structured checks that support objective evaluationTranslate complex domain problems into clear, testable workflows with measurable success criteriaReview task materials for correctness, completeness, reproducibility, and technical clarity Model Behavior Analysis & Failure ClassificationEvaluate model or agent performance on domain-specific tasksIdentify tasks where outputs fail to satisfy tests, instructions, or expected reasoning standardsClassify failure modes involving logical reasoning, problem decomposition, technical execution, or domain understandingWrite clear analysis explaining where and why a task response succeeds or fails Rubric Development & Structured ReviewDevelop detailed rubrics and evaluation frameworks for academic and technical benchmark tasksApply consistent evaluation standards across tasks, outputs, and solution materialsProvide clear written feedback explaining quality, reasoning gaps, and improvement areasCollaborate with other subject matter experts to support consistency and accuracy across review workflows Ideal Profile Strong candidates may have:Current or retired professor status, or current PhD student status, in a relevant academic or professional fieldAcademic expertise in STEM, quantitative, professional, or research-intensive domainsWorking proficiency in Python applied through research, industry work, GitHub projects, coursework, or technical task developmentAbility to design rigorous domain-specific problems and evaluate solutions with precisionStrong reasoning, written communication, problem-solving, and independent work skillsAbility to manage time effectively and contribute reliably in a remote project-based environmentAvailability for high-commitment project work, potentially 30+ hours per week during weekdays depending on project scope Educational BackgroundA completed or in-progress PhD from a strong university program is highly relevantAcademic backgrounds may include machine learning, coding, data science, computer science, physics, mathematics, engineering, statistics, biology, chemistry, finance, accounting, economics, law, business, or related fieldsTeaching, research, publication, technical writing, benchmark design, coding, or evaluation experience may be especially valuable Nice to HaveExperience in AI training, model evaluation, benchmark development, data annotation, or structured task reviewExperience writing Python tests, executable checks, golden solutions, or reproducible research codeFamiliarity with agentic task design, model behavior analysis, reasoning evaluation, or failure-mode classificationExperience developing academic assessments, problem sets, rubrics, grading criteria, or research evaluation materialsStrong ability to turn complex academic or professional problems into clear, testable tasks Why This OpportunityApply academic expertise to structured remote benchmark and evaluation workContribute to high-quality task design, Python-based solution development, and reasoning assessmentWork on flexible assignments aligned with your research field, domain knowledge, and technical strengthsUse your ability to identify reasoning gaps, design rigorous problems, and evaluate outputs with precisionRemote structure with competitive hourly compensation Contract DetailsIndependent contractor roleFully remote with flexible schedulingEligible professionals should be based in the United States depending on project needsHigh-commitment project availability may be required, potentially 30+ hours per week during weekdays depending on project scopeCompetitive rates between $70-$100 per hour depending on expertise and project scopeWeekly payments via Stripe or WiseProjects may be extended, shortened, or adjusted depending on scope and performanceWork will not involve access to confidential or proprietary information from any employer, client, or institution About the Platform This opportunity is available through 24-MAG LLC. We connect experienced professionals with remote consulting opportunities across technical, evaluation, and project-based workstreams. By submitting this application, you acknowledge that your information may be processed by 24-MAG LLC for recruitment and opportunity matching in accordance with our Privacy Policy: https://www.24-mag.com/privacy-policy.