Annotation Judging

Multilingual AI Response And Rubric Evaluator

Evaluate multilingual AI reference responses and scoring rubrics, identify critical errors, and improve evaluation quality. Flexible remote work across seven selected language and location options.

Flexible Brazil · Indonesia Available in 7 countries Brazil Indonesia Malaysia Mexico Thailand United Kingdom Vietnam

Evaluate multilingual AI reference responses and scoring rubrics for Project Meteor. You’ll use detailed guidelines to identify quality issues, correct critical errors, and help improve how advanced AI systems are assessed.

About this project

Meteor focuses on the reference answers and evaluation rubrics used to train and assess multilingual AI models. The work includes both open-ended and closed-ended tasks and requires careful, consistent judgment.

What you’ll do

  • Review AI-generated reference responses for quality and accuracy.
  • Evaluate and improve scoring rubrics.
  • Identify and correct critical P0 issues.
  • Follow detailed annotation guidelines and meet project quality requirements.

Requirements

  • Native-level fluency in one of the available project languages.
  • Strong reading comprehension, written communication, and attention to detail.
  • Basic familiarity with medical, financial, and legal topics.
  • Ability to follow detailed instructions carefully.
  • A computer with a reliable internet connection.

Project details

  • Location: Remote within the seven eligible language and location options.
  • Schedule: Flexible; task availability may vary.

Compensation

Compensation is calculated at a fixed hourly rate. The rate shown on the OneForma platform depends on the language and location you select.

Similar Projects

AI/ML Training Annotation

E-Commerce Search Relevance Annotator

Review search queries and e-commerce products and judge how well they match. Remote annotation project, available in French, Spanish and Swedish.

Belgium · Canada Available in 8 countries Belgium Canada Colombia Mexico Morocco Sweden Tunisia United States Flexible freelance contract, with a minimum commitment of 3 hours of production work per day.
Languages Judging

Audio Transcription Quality Reviewer

Review long-form audio transcriptions for accuracy, check speaker labels, timestamps and segmentation, and flag audio issues against project guidelines. Remote, hourly work for native speakers across 31 locales.

Australia · Brazil Available in 31 countries Australia Brazil Canada China Denmark Finland France Germany Hong Kong India Indonesia Israel Italy Japan Malaysia Mexico Netherlands Norway Poland Portugal Saudi Arabia South Korea Spain Sweden Taiwan Thailand Turkey United Arab Emirates United Kingdom United States Vietnam Flexible; weekly task releases
Fixed rate per hour Apply
Annotation

Multilingual Search Relevance Contributor

Evaluate how well retrieved content matches a user's search intent and assign relevance scores against project guidelines. Remote annotation work supporting multilingual search across several European languages.

Remote

Equal Opportunity At OneForma

OneForma welcomes experts from every background. Eligibility and selection are based on the requirements of each project, without discrimination based on race, ethnicity, gender, religion, disability, age, sexual orientation, cultural background, or any other characteristic protected by applicable law.

Choose your language or location

Select the option that matches your language and location.