Amazon Mechanical Turk closes for good on September 30, 2026. Not a freeze on new signups. A full shutdown, for workers and requesters both, twenty-one years after it launched.
For twenty years MTurk answered one question. How do I get a lot of small human judgments, fast and cheap? If you were planning to use it for AI training data, you now need somewhere else to go.
Looking for paid task work rather than a data supplier? This page will not help you. Prolific, CloudResearch, Appen and clickworker all take worker signups directly.
One disclosure first. LXT is the parent company of clickworker. Two options on this page are ours: clickworker’s self-service marketplace, and LXT’s managed programmes. We have marked both, and we have left our competitors on the list.
What happened, and when
- July 30, 2026. Amazon closed MTurk to new requester accounts. The service moved to AWS’s maintenance list: security and availability fixes only, no new features.
- August 25, 2026. Amazon announced the platform will close entirely on September 30. News broke via CNBC before most users saw anything in-product.
- August 27, 2026. A bulk email went out to workers and requesters, thanking them for being “part of Amazon Mechanical Turk’s journey” and asking them to update payment information before the close.
- September 30, 2026. Mechanical Turk shuts down. Amazon’s own banner puts it plainly: the platform will permanently close on that date, following what the company describes as a routine assessment of its programs and services.
That is a five-week window from public announcement to full closure. If you are running anything on MTurk today, this is not a “no roadmap” problem anymore. It is a hard deadline.
AI took the cheap tasks, and quality slipped
Models now do the easy work. Much of MTurk’s volume was basic sorting, simple transcription and straightforward tagging. A model does that for a fraction of the cost. What is left needs a real human, and that is the work an anonymous marketplace handles worst.
The quality problem is documented. A 2023 study in PLOS ONE by Benjamin Douglas, Patrick Ewell and Markus Brauer compared five platforms used for online research: MTurk, Prolific, CloudResearch, Qualtrics and SONA. MTurk was not the worst on every measure, but it sat near the bottom on most. Its workers passed fewer attention checks than workers on Prolific or CloudResearch. When the researchers repeated the same measurement on the same people, MTurk’s answers held together far less well, with test-retest reliability “considerably smaller on MTurk than on the other four platforms.” Scale reliabilities were lower too. Each high-quality response also cost more: $4.36 on MTurk against $1.90 and $2.00 on Prolific and CloudResearch.
Some of the human work was a model. In a 2023 paper called “Artificial Artificial Artificial Intelligence”, Veniamin Veselovsky, Manoel Horta Ribeiro and Robert West gave MTurk workers a text summarising task. Using keystroke detection and a synthetic-text classifier, they estimate that 33 to 46% of those workers used a large language model to do it. That is one task, not a census of the platform. But if you pay for human judgment and up to half of it might be a model, you are buying an expensive detour to a worse model than you would have picked yourself.
Once you cannot trust what comes back, cheap stops being cheap. The hours your team spends finding the bad work are a real cost. They just do not show up on the invoice.
The pay structure worked against quality. MTurk’s published pricing takes 20% of what you pay a worker, plus another 20% on tasks with ten or more assignments. Some tasks paid a cent. Pay a cent a judgment to anonymous workers with no career path and you buy speed, not care.
MTurk was good at the job it was built for. Cheap parallel microtasks at volume, twenty years before anyone said “AI training data.” The requirements moved and it did not follow.
Researchers say the bottleneck is data, not compute
We run the AI Data Bottleneck Index. It reads arXiv papers in machine learning, language, vision and AI, then sorts each one by a keyword rule: does the paper name a data problem, or a compute problem? Across 62,637 papers from two full quarters, 17.5% named a data problem, and the figure barely moved between the two quarters. Evaluation data is the fastest-rising complaint in the set.
Two limits worth stating. This counts what papers mention, not what provably held them back. And a keyword rule is not human review.
The part that matters here: the complaint is rarely “we could not get enough labels.” It is quality, coverage and evaluation data. Those are the three things a penny-per-task anonymous marketplace is worst at.
Ten alternatives, and who each one suits
| Platform | How you buy | Who runs quality control | Best for | Public pricing |
|---|---|---|---|---|
| clickworker (ours) | Self-serve marketplace | You | Running your own survey tasks, no sales call | Yes |
| Prolific | Self-serve panel | You | Academic and market research surveys | Yes |
| CloudResearch | Self-serve panel | You | Behavioural research, moving off MTurk | Partly |
| Toloka | Self-serve crowd | You | High-volume microtasks | Yes |
| Appen | Managed | Provider | Large-scale annotation, many languages | No |
| CloudFactory | Managed teams | Provider | Sustained annotation with dedicated staff | No |
| Sama | Managed | Provider | Annotation with a stated labour-ethics model | No |
| Scale AI | Platform plus managed | Shared | Large-scale ML labelling, computer vision | No |
| Surge AI | Managed expert | Provider | RLHF and expert-level judgment | No |
| LXT (ours) | Managed | Provider | Specialist fields, rare languages, controlled collection | No |
Check current details with each provider before you commit. This is a starting map, not a procurement document.
If you want to run the tasks yourself
This is the closest thing to what you lost.
clickworker’s self-service marketplace lets you set up and launch work against a crowd without a sales call. You define the task, set the fee per participant, and run it. Pricing is public. clickworker is part of the LXT group, so this option is ours.
One limit, because getting it wrong wastes your time. The self-service marketplace covers online surveys. Data labelling, annotation, categorisation and collection go through managed service instead. So it replaces MTurk’s survey and respondent-recruitment work, and not its labelling work. If labelling is what you came for, skip to the training data section.
We’re not the only ones pointing there. Writing about where to collect data now that MTurk is closing, researcher Elena Brandt splits her recommendations by audience: Besample, CloudResearch Connect and Prolific for academic researchers, and clickworker for everyone else. Her pick for “general-purpose crowdsourcing at scale,” she calls it one of the closest relatives of MTurk still standing. AIMultiple’s comparison of MTurk alternatives backs that up from a different angle: clickworker carries the highest worker rating of any platform in their review, 4.4 out of 5 across 2,454 Trustpilot reviews.
If you need survey respondents for research
clickworker’s marketplace (ours) handles this self-serve. You recruit the participants and run the survey yourself, at your own fee per participant. It suits standard cases with a clear target group.
For academic and behavioural research, Prolific and CloudResearch are the specialists, and both are worth your time. They were built for research, they keep vetted participant pools, and they handle sampling and screening that MTurk left you to work out. If you are publishing peer-reviewed work, start there rather than with us. Managed enterprise data collection, which is what LXT does, is the wrong tool for a small academic survey.
If you need training data for an AI or ML model
A like-for-like swap will disappoint you.
MTurk gave you a crowd, not a workforce. You wrote the task, hoped it was clear, and took whatever came back. Quality control was your job: your gold-standard questions, your attention checks, your redundant labelling, your review pipeline. At a cent a task that trade made sense. For data a production model trains on, it does not.
This work now needs four things.
- Annotators who know the subject. Radiology images need clinical knowledge. Legal annotation needs legal literacy. Dialect transcription needs native speakers of that dialect.
- Quality checks inside the pipeline, not bolted on afterwards: multi-pass review, measured agreement between annotators, and feedback that corrects an annotator’s work over time.
- Provenance you can evidence. In a regulated field you need to show who collected the data, under what consent, in which country. Anonymous marketplaces cannot answer that, and your auditors will ask.
- The same people for months. An annotator who knows your edge cases in month four is worth several fresh workers reading your guidelines for the first time. That only happens where people are managed, paid properly, and stay.
That is the work we do at LXT. We build and run annotation and data collection programmes across text, speech, image and video, in over 1,000 language locales.
LXT is not a self-serve marketplace where you post a task and get results in an hour. If that is what you want, use the clickworker marketplace instead, same group, built for exactly that. Managed work fits when the data is hard: a rare language pair, a specialist field, a controlled collection setup, an evaluation programme that has to stay consistent for months. That is the work that broke on MTurk.
If you used MTurk for moderation or QA at volume
Look at managed BPO providers or specialist trust and safety vendors. Same reasons as above, plus one more: moderation carries duty-of-care obligations you should not hand to anonymous piecework.
If you already have an MTurk account
Every earlier version of this advice assumed you could choose to stay a while longer. That option is gone. September 30 is a hard close for every requester account, regardless of how long you have run on the platform or how well your pipeline works today.
Four things to check:
- Check what is still running. Any HITs in flight, any longitudinal or multi-wave study, needs to either finish or migrate before September 30. Do not start anything new on MTurk that cannot complete in that window.
- Export everything. Task results, worker qualifications, your custom quality logic. AWS has not said how long data stays retrievable after the close, and “permanently close” is not language that leaves room for a grace period.
- Decide your replacement now, not in week four. Self-service survey work maps to clickworker’s marketplace. Labelling, annotation and evaluation work at production quality maps to managed providers like LXT. Moderation and QA at volume maps to BPO or trust and safety vendors. Match your actual workload to the option, not the other way round.
- Budget migration time, not just migration cost. Standing up quality control with a new provider, even a strong one, takes longer than people expect in week one. Five weeks is tight. Start the conversation this week if you have not already.
Managed work costs more per label and often less per usable label
Managed data work does cost more per label than MTurk did. There is no point pretending otherwise.
But cost per label is the wrong number. What you pay for is cost per usable label, which includes the labels you throw away, the reviewer time spent finding them, and the training runs wasted on data that turned out to be wrong.
Here is the arithmetic. The numbers are assumptions, not quotes. Say a cheap crowd charges $0.05 a label and 30% fail review. Your real cost is $0.05 divided by 0.70, so about $0.07 a label, plus the reviewer time to find that bad 30%, plus whatever a contaminated training run costs you. Say a managed provider charges $0.15 and 3% fail review. That is about $0.155 a label, with far less review work and far less risk of a quality problem reaching your model.
The shape is the point. The gap between sticker price and real price is much wider on cheap crowd labels, and the largest cost never appears on an invoice.
The cheap-crowd era ended before MTurk did
Anonymous piecework at a cent a task could once produce data good enough to train systems people relied on. The quality bar that arrived with generative AI is what ended that. If you were about to start on MTurk for AI training data, the closure probably saved you a detour. If you were already running on it, the clock is now five weeks, not indefinite.



