Paid, remote work authoring and validating research-grade ML benchmark tasks. For Master's or PhD-level ML researchers with at least one first-author paper. Domain: Machine Learning / AI Research Requirements Experience: 1+ years of hands-on ML research; no upper limit…
PhD required
Apply on AfterQuery ↗We may earn a fee if you sign up.We are hiring skilled software engineers with strong full-stack experience. The ideal candidate is comfortable architecting and working across multi-file codebases, structuring realistic development environments, and building reproducible setups supported by clear documentation…
Apply on AfterQuery ↗We may earn a fee if you sign up.Mercor is looking for US-based people with strong baseball knowledge to evaluate AI assistants live, while real MLB games are being played. On scheduled game days, you'll ask AI assistants the questions you'd naturally ask while following a game, then rate how helpful, accurate…
English
Apply on Mercor ↗We may earn a fee if you sign up.Mercor is seeking Clinical Mental Health Experts to support an AI research initiative focused on evaluating the realism and quality of simulated clinical conversations. This is a short-term, fully remote evaluation task designed for mental health professionals who can apply…
Apply on Mercor ↗We may earn a fee if you sign up.micro1 is engaging Gmail & Google Calendar AI Assistant Evaluators to collaborate on a customer-driven project enhancing AI assistant quality through real-world task evaluation. In this role, you'll use your everyday experience managing your digital life — email, calendars…
Apply on Micro1 ↗We may earn a fee if you sign up.micro1 is engaging Generalist — U.S. Tax Workflow Evaluations to participate in a short-term customer project focused on evaluating U.S. tax and financial workflows. In this role, you'll apply your expertise to help train next-generation AI systems. Your work will shape how…
Apply on Micro1 ↗We may earn a fee if you sign up.We are seeking elite Content Writers & AI Evaluation Specialists with expertise in copywriting, content creation and LLM evaluation. In this role, you will apply your deep expertise in writing, content strategy, and AI evaluation to create, review, and evaluate creative and…
Arabic, Chinese, French, German, Portuguese, Spanish
Apply on AfterQuery ↗We may earn a fee if you sign up.About Turing Turing’s mission is to accelerate superintelligence to drive real economic progress. Headquartered in San Francisco, Turing works with frontier AI labs to generate high-quality data, evaluations, and reinforcement learning environments that improve model…
Bachelor's required
Apply on Turing ↗We may earn a fee if you sign up.Mercor is hiring board-certified Dermatologists for non-clinical work developing and evaluating clinical AI systems. You will apply your diagnostic expertise to image interpretation, annotation, and evaluation work that helps define what a clinically sound dermatological…
English
Apply on Mercor ↗We may earn a fee if you sign up.Mercor is hiring board-certified Radiologists for non-clinical work developing and evaluating clinical AI systems in medical imaging. You will apply your diagnostic expertise to annotation, reference-report authoring, and evaluation work that sets the standard these systems are…
English
Apply on Mercor ↗We may earn a fee if you sign up.We are looking for Speech AI Evaluation Specialist to support the improvement of AI-generated content in Bengali. Job Type: Freelance Location: India (work from home) Work Schedule: Part-time - 10+ hours per week. Flexible - work whenever you want. Start Date: Immediately…
Bengali, English, Japanese
Apply on RWS ↗Opens the platform's own listing.We are seeking candidates to develop evaluation scenarios that test cutting-edge autonomous AI systems across multiple domains. You'll create complex research tasks that challenge AI capabilities in sports analytics, financial analysis, patent research, academic planning, and…
Apply on AfterQuery ↗We may earn a fee if you sign up.micro1 is engaging Cantonese Language Evaluators to contribute to a language-focused project supporting a valued customer. In this role, you'll apply your expertise to help train next-generation AI systems. Your work will shape how models learn, reason, and perform through…
ChineseMaster's required
Apply on Micro1 ↗We may earn a fee if you sign up.Scope of the Role: We are looking for detail-oriented Voice Specialists to evaluate the performance of AI models through real-time, voice-based conversations. In this role, you will interact with two different AI models using the same assigned scenario, carefully compare their…
Apply on Innodata ↗Opens the platform's own listing.Outlier helps the world’s most innovative companies improve their AI models through human feedback. Are you fluent in Korean and comfortable using your own Android phone to test AI experiences? About the opportunity: Outlier is looking for Korean-speaking contributors to help…
Korean
Apply on Outlier ↗Opens the platform's own listing.Outlier helps the world’s most innovative companies improve their AI models through human feedback. Are you fluent in Japanese and comfortable using your own Android phone to test AI experiences? About the opportunity: Outlier is looking for Japanese-speaking contributors to…
Japanese
Apply on Outlier ↗Opens the platform's own listing.Outlier helps the world’s most innovative companies improve their AI models by providing human feedback. Are you a strong communicator with fluency in English who would like to lend your expertise to train AI models? About the opportunity: Outlier is looking for talented…
English
Apply on Outlier ↗Opens the platform's own listing.Outlier helps the world’s most innovative companies improve their AI models through human feedback. Are you fluent in Spanish, living in Mexico, and comfortable using your own Android phone to test AI experiences? About the opportunity: Outlier is looking for Spanish-speaking…
Spanish
Apply on Outlier ↗Opens the platform's own listing.Handshake is seeking motivated Associates, Bachelor’s, and Master’s students and grads to work remotely to test Large Language Models (LLMs) in collaboration with top AI labs. As part of this program, you’ll apply your overall skills and academic knowledge to help improve the…
Apply on Handshake ↗We may earn a fee if you sign up.Qualifications Ideal candidates meet all of the following: Location — Based in South Korea and able to work remotely in an asynchronous, self-directed capacity Language — Fluent in Korean, both written and spoken, with strong written communication skills for explaining your…
English, KoreanBachelor's required
Apply on Handshake ↗We may earn a fee if you sign up.Qualifications Ideal candidates meet all of the following: Location — Based in Japan and able to work remotely in an asynchronous, self-directed capacity Language — Fluent in Japanese, both written and spoken, with strong written communication skills for explaining your…
English, JapaneseBachelor's required
Apply on Handshake ↗We may earn a fee if you sign up.Qualifications Ideal candidates meet all of the following: Location — Based in Mexico and able to work remotely in an asynchronous, self-directed capacity Language — Fluent in Spanish, both written and spoken, with strong written communication skills for explaining your…
English, SpanishBachelor's required
Apply on Handshake ↗We may earn a fee if you sign up.We are looking for Generative Audio Evaluation Specialists! A great entry point into ongoing work within one of our most active AI markets! Job Type: Freelance Location: Remote (The Netherlands) Work Schedule: Part-time - 10+ hours per week. Flexible - work whenever you want!…
Dutch, English
Apply on RWS ↗Opens the platform's own listing.We are looking for Generative Audio Evaluation Specialists! A great entry point into ongoing work within one of our most active AI markets! Job Type: Freelance Location: Remote (Poland) Work Schedule: Part-time - 10+ hours per week. Flexible - work whenever you want! Your work…
English, Polish
Apply on RWS ↗Opens the platform's own listing.AI models improve when native speakers judge their output carefully. That is the work here. You receive items in Danish, each with a prompt and two responses generated by our client's language model. For every item you evaluate both responses on instruction following…
Danish, English
Apply on Sovrano ↗We may earn a fee if you sign up.AI models improve when native speakers judge their output carefully. That is the work here. You receive items in Norwegian, each with a prompt and two responses generated by our client's language model. For every item you evaluate both responses on instruction following…
English, Norwegian
Apply on Sovrano ↗We may earn a fee if you sign up.Role Overview We are seeking detail-oriented evaluators to conduct human quality evaluations for an enterprise AI customer support product on Instagram, WhatsApp, and Messenger. You will evaluate and benchmark AI model responses against complex evaluation rubrics using provided…
English, Vietnamese
Apply on Innodata ↗Opens the platform's own listing.CVE & Application Security AI Task Auditor - Freelance AI Trainer Project — a Meridial (Invisible Technologies) freelance AI training project.
Apply on Meridial ↗Opens the platform's own listing.Machine Learning (ML) AI Task Auditor - Freelance AI Trainer Project — a Meridial (Invisible Technologies) freelance AI training project.
Apply on Meridial ↗Opens the platform's own listing.AWS Serverless & IaC AI Task Auditor - Freelance AI Trainer Project — a Meridial (Invisible Technologies) freelance AI training project.
Apply on Meridial ↗Opens the platform's own listing.Mercor is hiring generalist annotators to evaluate conversations between users and a Health & Fitness Assistant AI. You will judge how well the AI communicates: whether it answered the actual question, whether a non-expert could follow it, and whether it was honest even when…
EnglishBachelor's required
Apply on Mercor ↗We may earn a fee if you sign up.Mercor is hiring practicing outpatient physicians (MD) on behalf of a healthcare AI partner building advanced clinical documentation and decision-support tools. In this role you will apply your day-to-day ambulatory expertise to review, annotate, and evaluate clinical…
Arabic, Bulgarian, Chinese, Czech, Danish, Dutch, English, Finnish, French, German, Hungarian, Indonesian, Italian, Japanese, Korean, Polish, Portuguese, Romanian, Russian, Spanish, Tagalog, Turkish, Ukrainian, Vietnamese
Apply on Mercor ↗We may earn a fee if you sign up.Outlier helps the world’s most innovative companies improve their AI models through human feedback. Are you fluent in French and comfortable using your own Android phone to test AI experiences? About the opportunity: Outlier is looking for French-speaking contributors to help…
French
Apply on Outlier ↗Opens the platform's own listing.Handshake is seeking motivated Associates, Bachelor’s, and Master’s students, as well as individuals who have already completed their degrees residing in Canada, to contribute to AI research initiatives in collaboration with top AI labs. As part of this program, you’ll apply…
Apply on Handshake ↗We may earn a fee if you sign up.Qualifications Ideal candidates meet all of the following: Location — Based in Indonesia and able to work remotely in an asynchronous, self-directed capacity Language — Fluent in Bahasa Indonesian, both written and spoken, with strong written communication skills for explaining…
English, IndonesianBachelor's required
Apply on Handshake ↗We may earn a fee if you sign up.Mercor is sourcing Sales experts for a structured judge-calibration study. You will apply professional judgment to realistic Sales work and evaluate model-generated outputs. What You'll Do Sign an NDA before receiving study materials. Complete one realistic Sales task using…
English
Apply on Mercor ↗We may earn a fee if you sign up.Mercor is hiring practicing Inpatient Hospitalists (MD) on behalf of a healthcare AI partner building advanced clinical documentation and decision-support tools. In this role you will apply your active inpatient expertise to review, annotate, and evaluate hospital clinical…
Arabic, Bulgarian, Chinese, Czech, Danish, Dutch, English, Finnish, French, German, Hungarian, Indonesian, Italian, Japanese, Korean, Polish, Portuguese, Romanian, Russian, Spanish, Tagalog, Turkish, Ukrainian
Apply on Mercor ↗We may earn a fee if you sign up.AI-Assisted Developer Workflows (Trace) AI Task Auditor - Freelance AI Trainer Project — a Meridial (Invisible Technologies) freelance AI training project.
Apply on Meridial ↗Opens the platform's own listing.GPU Kernels AI Task Auditor - Freelance AI Trainer Project — a Meridial (Invisible Technologies) freelance AI training project.
Apply on Meridial ↗Opens the platform's own listing.AWS Trainium/NKI (TRN) AI Task Auditor - Freelance AI Trainer Project — a Meridial (Invisible Technologies) freelance AI training project.
Apply on Meridial ↗Opens the platform's own listing.SWE-Bench AI Task Auditor - Freelance AI Trainer Project — a Meridial (Invisible Technologies) freelance AI training project.
Apply on Meridial ↗Opens the platform's own listing.About the Program: Innodata's Federal Practice builds the trusted data layer for critical infrastructure Trust & Safety work. Partnering with a leading systems integrator, we're delivering a modern, governed data services platform in a secure federal (IL4) environment. Over an…
Bachelor's required
Apply on Innodata ↗Opens the platform's own listing.Mindrift connects specialists with project-based AI opportunities for leading tech companies, focused on testing, evaluating, and improving AI systems. Participation is project-based, not permanent employment. What this opportunity involves We're building a dataset to evaluate…
reposted 13x
Apply on Mindrift ↗Opens the platform's own listing.Mindrift connects specialists with project-based AI opportunities for leading tech companies, focused on testing, evaluating, and improving AI systems. Participation is project-based, not permanent employment. What this opportunity involves We're building a dataset to evaluate…
Apply on Mindrift ↗Opens the platform's own listing.Mindrift connects specialists with project-based AI opportunities for leading tech companies, focused on testing, evaluating, and improving AI systems. Participation is project-based, not permanent employment. What this opportunity involves We're building a dataset to evaluate…
reposted 8x
Apply on Mindrift ↗Opens the platform's own listing.Mindrift connects specialists with project-based AI opportunities for leading tech companies, focused on testing, evaluating, and improving AI systems. Participation is project-based, not permanent employment. What this opportunity involves We're building a dataset to evaluate…
reposted 2x
Apply on Mindrift ↗Opens the platform's own listing.We are looking for Speech AI Evaluation Specialists living in Malaysia to support the improvement of AI-generated content in Chinese Simplified. Job Type: Freelance Location: Malaysia (work from home) Work Schedule: Part-time - 10+ hours per week. Flexible - work whenever you…
Chinese, English, Malay
Apply on RWS ↗Opens the platform's own listing.We are looking for Speech AI Evaluation Specialists living in Malaysia to support the improvement of AI-generated content in Vietnamese. Job Type: Freelance Location: Malasya (work from home) Work Schedule: Part-time - 10+ hours per week. Flexible - work whenever you want. Start…
English, Malay, Vietnamese
Apply on RWS ↗Opens the platform's own listing.We are looking for Speech AI Evaluation Specialist to support the improvement of AI-generated content in Vietnamese (USA). Job Type: Freelance Location: USA (work from home) Work Schedule: Part-time - 10+ hours per week. Flexible - work whenever you want. Start Date: Immediately…
English, Vietnamese
Apply on RWS ↗Opens the platform's own listing.We are looking for Speech AI Evaluation Specialist to support the improvement of AI-generated content in Chinese Simplified (USA). Job Type: Freelance Location: USA (work from home) Work Schedule: Part-time - 10+ hours per week. Flexible - work whenever you want. Start Date…
Chinese, English
Apply on RWS ↗Opens the platform's own listing.
No RLHF and evaluation roles are open today. The last one closed Oct 5. New ones show up here the morning after they're posted.
Seen open in the last 48 hours Last seen open 2 or more days ago