Judging

Arabic AI Safety Evaluator

Review Arabic AI responses for harmful content, safety risks, and policy violations. Open to experienced native speakers in Switzerland, Saudi Arabia, or the UAE who have a personal MacBook.

Saudi Arabia · Switzerland Available in 3 countries Saudi Arabia Switzerland United Arab Emirates Remote

About This Project

OneForma is looking for native Arabic linguists based in Switzerland, Saudi Arabia or the United Arab Emirates to join an exciting remote AI research project focused on AI safety evaluation and quality assurance.

You will review AI-generated responses and perform structured linguistic assessments to identify harmful content, safety concerns, and policy violations. The work involves evaluating large language model (LLM) outputs, classifying potential risks, documenting findings, and validating improvements after issues have been addressed.

This is a flexible, remote opportunity for experienced Arabic language professionals who are interested in helping improve the safety, reliability, and accuracy of next-generation AI systems.

Why This Work Matters

The purpose of this project is to improve the safety and quality of AI chatbot systems by using human expertise to evaluate model responses, identify potential risks, and provide high-quality feedback that supports ongoing AI development.

Requirements

  • Must be a native Arabic speaker based in Switzerland, Saudi Arabia or the United Arab Emirates.
  • Linguistic background or professional language experience is preferred.
  • Must be familiar with Apple products and AppleCare content.
  • Experience evaluating Generative AI (GenAI) models or large language model (LLM) outputs is preferred.
  • Must own a personal MacBook running the latest version of macOS Tahoe.
  • Device must be personal (not shared) and have administrator privileges for software installation.
  • Must confirm whether you have an AppleConnect account.
  • Must be available to receive immediate, ad hoc task assignments.
  • Must demonstrate professionalism, attention to detail, and the ability to follow evaluation guidelines.
  • Strong analytical and written communication skills are required.

Project Details

  • This is a fully remote project.
  • You will review AI-generated conversations for safety, quality, and policy compliance.
  • Tasks include identifying harmful or unsafe content, classifying issues by risk level, documenting findings, and validating fixes after remediation.
  • The use of AI tools, automation, or speech-to-text software is strictly prohibited.
  • Completion of onboarding and certification is mandatory before production work begins.
  • Direct communication with the client point of contact will be required throughout the project.

Help shape the future of trustworthy AI by using your language expertise to make AI systems safer, more accurate, and more reliable for users worldwide.

Similar Projects

AI/ML Training Annotation

E-Commerce Search Relevance Annotator

Review search queries and e-commerce products and judge how well they match. Remote annotation project, available in French, Spanish and Swedish.

Belgium · Canada Available in 8 countries Belgium Canada Colombia Mexico Morocco Sweden Tunisia United States Flexible freelance contract, with a minimum commitment of 3 hours of production work per day.
Languages Judging

Audio Transcription Quality Reviewer

Review long-form audio transcriptions for accuracy, check speaker labels, timestamps and segmentation, and flag audio issues against project guidelines. Remote, hourly work for native speakers across 31 locales.

Australia · Brazil Available in 31 countries Australia Brazil Canada China Denmark Finland France Germany Hong Kong India Indonesia Israel Italy Japan Malaysia Mexico Netherlands Norway Poland Portugal Saudi Arabia South Korea Spain Sweden Taiwan Thailand Turkey United Arab Emirates United Kingdom United States Vietnam Flexible; weekly task releases
Fixed rate per hour Apply
AI/ML Training Annotation

Multilingual AI Response And Rubric Evaluator

Evaluate multilingual AI reference responses and scoring rubrics, identify critical errors, and improve evaluation quality. Flexible remote work across seven selected language and location options.

Brazil · Indonesia Available in 7 countries Brazil Indonesia Malaysia Mexico Thailand United Kingdom Vietnam Flexible
Fixed rate per hour Apply

Equal Opportunity At OneForma

OneForma welcomes experts from every background. Eligibility and selection are based on the requirements of each project, without discrimination based on race, ethnicity, gender, religion, disability, age, sexual orientation, cultural background, or any other characteristic protected by applicable law.

Choose your language or location

Select the option that matches your language and location.