On this page · 8 sections
Summary. Amazon Mechanical Turk, the crowdwork platform Amazon launched in 2005, stops accepting new customers on 30 July 2026. This is not a full shutdown: Amazon Web Services says existing customers "can continue to use the service as normal," but it will add no new features, which puts the platform in maintenance mode. The reason is pointed. A 2023 study estimated that 33% to 46% of MTurk workers used large language models on a text-summarisation task, so the service built to supply human judgement was quietly filling with machine output. If you were about to build on MTurk, you now need a plan. The main alternatives are Prolific, Surge AI, Scale AI, Toloka, Clickworker and CloudResearch, and they are not interchangeable: expert RLHF work runs $85 to $200-plus per hour on Surge AI, while Toloka bills pay-as-you-go for high-volume tasks. This guide compares them by job, price and quality, for teams outside and inside India.
What Amazon actually announced
Read the announcement precisely, because the headlines oversold it. As TechCrunch reported on 5 July 2026, Amazon will stop accepting new customers, not shut the service down. From 30 July 2026, new requesters and new workers cannot register. Existing accounts keep working, and AWS says it will continue investing in security and availability while adding no new features. Amazon framed the move as coming "after careful consideration."
For anyone already running production pipelines on MTurk, that buys time but not comfort: a platform on maintenance mode with no new features and no committed roadmap is a dependency to migrate off, not to build more on. For anyone who was about to start, the door closes at the end of this month. Either way, the practical question is the same. Where do you get reliable human data now?
Why MTurk is fading: AI ate the human layer
The irony is the story. MTurk's original pitch was "artificial artificial intelligence," humans doing tasks that software could not. By 2023, software could, and the workers noticed. Researchers who reran an abstract-summarisation task on the platform estimated, using keystroke logging and synthetic-text classification, that between 33% and 46% of crowd workers used LLMs to complete it. A follow-up in Communications of the ACM studied how to prevent this, testing a "request" strategy (asking workers not to use LLMs) against a "hurdle" strategy (converting text to images or disabling copy-paste). Roughly 30% to 40% of crowdworkers rely on LLMs for text production, and about 34% of Prolific participants self-report using them for open-ended questions.
That is the real lesson for buyers. The risk is no longer slow or sloppy humans; it is undisclosed AI contaminating the very data you are paying humans to produce. Any replacement has to be judged on how well it detects and prevents that, not just on price per task. If you are building evaluation sets, this is the same problem we cover in why AI agent evals fail silently in CI/CD: unverified data quietly poisons the metric.
What to look for in a replacement
Six criteria separate a good fit from an expensive mistake.
- Worker verification. Are participants identity-checked and vetted, or anonymous? This is the direct defence against undisclosed LLM use.
- Job type. Surveys and behavioural research, bulk labeling, and expert RLHF are three different markets. No single vendor wins all three.
- Pricing model. Pay-as-you-go suits spiky volume; custom enterprise contracts suit sustained frontier work.
- Quality controls. Attention checks, gold questions, and reviewer layers matter more than headline worker counts.
- Compliance. For personal data, India's Digital Personal Data Protection Act 2023 (DPDP) and cross-border rules shape who you can use and where data sits.
- Migration effort. Some platforms offer MTurk-compatible tooling; others need a rebuild.
The six alternatives, compared
| Platform | Best for | Worker model | Pricing | Quality control |
|---|---|---|---|---|
| Prolific | Research, surveys, evals | Verified, representative participants | Per-response, transparent | Identity checks, screening |
| Surge AI | Frontier RLHF, expert labeling | Vetted, better-paid experts | $85-$200+/expert hour | Elite, educated pool |
| Scale AI | Large enterprise labeling | Managed global workforce | Custom enterprise | Structured review layers |
| Toloka | High-volume, multimodal labeling | Global crowd | Pay-as-you-go | Configurable QC workflows |
| Clickworker | Microtasks at scale | Large general crowd | Per-task | Standard checks |
| CloudResearch | MTurk-style studies | Curated MTurk-derived pool | Per-response | Vetting on top of MTurk |
Research, surveys and evaluations: Prolific. Prolific is the quality-and-evaluation specialist that labs reach for when they need verified, representative humans rather than the cheapest labelers. If you ran academic-style studies or model evaluations on MTurk, this is the closest philosophical replacement, with better identity guarantees.
Frontier RLHF and expert judgement: Surge AI. Surge uses a more heavily vetted, better-paid pool and specialises in RLHF, at $85 to $200-plus per expert hour. Mercor and Turing compete here; all three have grown by taking work from customers who left Scale AI after Meta acquired a 49% stake in June 2025.
Large-scale labeling: Scale AI or Toloka. For classification, transcription and multimodal labeling with structured quality controls, Scale AI offers a managed enterprise workforce, while Toloka gives you pay-as-you-go flexibility for cost-sensitive, high-volume runs.
Migrating existing MTurk studies: CloudResearch or Toloka. If you need the least rebuild, CloudResearch layers vetting and tooling on an MTurk-derived participant pool, and Toloka's configurable projects map cleanly onto microtask workflows.
If you strip it down to the switching decision, the contrast with MTurk itself is clearest on the dimensions that actually failed: worker verification and defence against undisclosed AI.
| Dimension | MTurk (closing to new users) | Prolific | Surge AI | Toloka |
|---|---|---|---|---|
| Worker verification | Weak, largely anonymous | Strong, identity-checked | Strong, vetted experts | Moderate, crowd-based |
| Defence against undisclosed LLM use | Poor, the reason for decline | Screening and verified pool | Vetted, better-paid pool | Configurable QC checks |
| Primary job | General microtasks | Research and evals | RLHF and expert labeling | High-volume labeling |
| Pricing shape | Per-task, low | Per-response | $85-$200+/expert hour | Pay-as-you-go |
| Migration effort from MTurk | n/a | Medium | Medium to high | Low to medium |
India-specific considerations
For Indian teams, two forces matter. First, cost: rupee-denominated budgets favour pay-as-you-go models like Toloka for volume work and reserve expert-hour pricing for the RLHF that genuinely needs it. Second, compliance: if tasks involve personal data, DPDP obligations around consent and cross-border transfer shape which platform and region you can use, so a vendor with clear data-residency controls is worth more than a marginally cheaper one. Teams building retrieval or evaluation systems can pair a labeling vendor with our RAG knowledge assistant service and treat human data as a governed input, not an afterthought. Our DPDP engineering playbook for Indian startups covers the consent and residency mechanics.
FAQ
How eCorpIT can help
eCorpIT is a Gurugram-based, senior-led engineering organisation, founded in 2021 and assessed at CMMI Level 5, that builds AI systems where human data quality decides the outcome. We help teams choose and integrate the right human-data vendor, design evaluation and labeling pipelines that detect undisclosed AI, and keep the whole flow aligned with DPDP requirements. As an AWS, Microsoft and Google technology partner, we treat human data as a governed input to your models. To plan a labeling or evaluation pipeline, talk to our engineering team or explore our AI evaluations and observability service.
References
_Last updated: 23 July 2026._