Remote | AI Evaluation Guidelines & Rubric Specialist — $40–$60/hour

24-Mag

  • New York, New York
  • 13 days ago
  • Remote
    Want to know if you’re a fit?
    Upload your resume and let our AI show you.

    Skills

    • Artificial Intelligence (AI)unmatched
    • Artificial Intelligence (AI) Programming Languagesunmatched
    • Calibrationunmatched
    • Concreteunmatched
    • Consultingunmatched
    • Document Managementunmatched
    • Documentation Reviewunmatched
    • Financeunmatched
    • Human-Computer Interactionunmatched
    • Information Architectureunmatched
    • Information Designunmatched
    • Instructional Designunmatched
    • Insuranceunmatched
    • Legalunmatched
    • Linguisticsunmatched
    • Modeling Languagesunmatched
    • Onboardingunmatched
    • Problem Solving Skillsunmatched
    • Project Evaluationunmatched
    • Project/Program Managementunmatched
    • Quality Assuranceunmatched
    • Quality Metricsunmatched
    • Reinforcement Learningunmatched
    • Reliability Analysisunmatched
    • Requirements Managementunmatched
    • Retailunmatched
    • Source Code/Configuration Management (SCM)unmatched
    • Sportsunmatched
    • Taxonomiesunmatched
    • Technical Consultingunmatched
    • Technical Writingunmatched
    • Technical/Engineering Designunmatched
    • Training/Teaching Materialsunmatched
    • Usability Testingunmatched
    • User Interface Designunmatched
    • Workflow Analysisunmatched

    Description

    We are sharing a specialised full-time consulting opportunity for US-based linguists, instructional designers, and technical writers experienced in developing clear evaluation guidelines, structured rubrics, and human-rating instructions for generative AI programmes.

    This role supports a high-impact generative AI initiative focused on translating complex and potentially ambiguous programme requirements into precise, practical guidance for human evaluators. Selected professionals will develop rater-ready instructions across domains such as finance, retail, insurance, legal, and sports while resolving contradictions, defining edge cases, and improving consistency throughout evaluation workflows.

    Key Responsibilities

    Rater Guideline Development

    • Translate programme requirements into clear, structured, and actionable instructions for human evaluators
    • Develop guidelines that can be applied consistently across standard scenarios and complex edge cases
    • Define terminology, rating criteria, decision rules, exceptions, and escalation pathways
    • Ensure instructions are accessible to raters while preserving necessary domain-specific precision

    Rubric & Evaluation Framework Design

    • Design detailed scoring rubrics for evaluating generative AI outputs
    • Establish measurable criteria covering correctness, relevance, reasoning quality, completeness, and instruction adherence
    • Create examples and counterexamples illustrating different performance levels
    • Align evaluation frameworks with programme objectives and quality standards

    Ambiguity & Consistency Review

    • Review draft specifications for ambiguity, contradiction, missing information, and inconsistent terminology
    • Identify instructions that may lead to conflicting interpretations across raters
    • Revise guideline sets until they can be applied reliably with minimal escalation
    • Document concrete before-and-after improvements to written requirements and evaluation instructions

    Cross-Domain Instructional Translation

    • Convert specifications from finance, retail, insurance, legal, sports, and other specialist domains into rater-ready guidance
    • Collaborate with subject matter experts to understand domain-specific terminology and professional judgment
    • Preserve important technical nuance while making instructions clear to non-specialist evaluators
    • Maintain consistent structure and quality across multiple domain-specific guideline sets

    Ideal Profile

    Strong candidates may have:

    • At least 3 years of professional experience in linguistics, instructional design, technical writing, content design, or a closely related field
    • Direct experience developing or refining guidelines and rubrics for human evaluators in generative AI, RLHF, or model-assessment programmes
    • Demonstrated ability to resolve ambiguity and contradiction in complex written specifications
    • Experience translating specialist requirements into clear and practical instructions
    • Ability to work effectively across multiple subject-matter domains
    • A portfolio or concrete examples showing measurable improvements to guidelines, rubrics, or instructional materials
    • Demonstrable professional growth and increasing responsibility
    • Reliable availability for at least 35 hours per week during weekdays

    Educational Background

    • A degree in linguistics, instructional design, education, communications, technical writing, language studies, or a related field is highly relevant
    • Graduate-level education in applied linguistics, learning design, human-computer interaction, or information design may be helpful
    • Equivalent professional experience in AI evaluation, technical documentation, or guideline development may also be considered
    • Training in assessment design, taxonomy development, content strategy, or quality assurance may be valuable

    Nice to Have

    • Experience supporting large language model evaluation, reinforcement learning from human feedback, or AI training-data programmes
    • Familiarity with annotation platforms, human-feedback workflows, and rater calibration processes
    • Experience developing domain-specific guidance for finance, insurance, retail, legal, sports, or comparable fields
    • Knowledge of controlled language, information architecture, taxonomy design, or content governance
    • Experience conducting guideline usability tests or analysing inter-rater consistency
    • Familiarity with version control, documentation systems, and structured authoring tools
    • Previous collaboration with researchers, programme managers, engineers, and subject matter experts

    Why This Opportunity

    • Apply linguistic and instructional-design expertise to an advanced generative AI initiative
    • Influence the clarity and reliability of human evaluation processes
    • Develop guidelines used across a wide range of professional subject-matter domains
    • Solve complex problems involving ambiguity, edge cases, and evaluation consistency
    • Join a full-time remote engagement with competitive hourly compensation

    Contract Details

    • Full-time W-2 contingent employment arrangement
    • Fully remote role available to candidates based in the United States
    • Expected commitment of at least 35 hours per week during weekdays
    • Competitive rates between $40–$60 per hour depending on expertise and project scope
    • Direct experience developing rater guidelines or rubrics for generative AI or RLHF programmes is required
    • Applicants should be prepared to provide concrete examples of guideline or specification improvements
    • Immediate availability is preferred
    • Work may include onboarding, calibration, documentation review, and ongoing guideline refinement
    • Project scope and duration may be adjusted according to programme requirements and performance

    About the Platform

    This opportunity is available through 24-MAG LLC. We connect experienced professionals with remote consulting opportunities across technical, evaluation, and project-based workstreams.

    By submitting this application, you acknowledge that your information may be processed by 24-MAG LLC for recruitment and opportunity matching in accordance with our Privacy Policy: https://www.24-mag.com/privacy-policy.

    Numbers & Facts

    LocationNew York, New York (
    Remote
    )
    Website4-mag.com/privacy-policy

    Similar Jobs