p.enthalabs

Mechanical Turk shutting down September 30

mturk.com · Read Story HN original

Comments

As AMT's largest requester for the past 10 years, this news was relayed to requesters at the same time as respondents. It's also worth noting that our lead contact, the Sr Program Manager at AWS leading AMT, transitioned to Amazon Bedrock and SageMaker Model Evaluations a ~2-3 years ago.. Leaving behind essential zero team managing the project after they migrated over the stored value accounts to native AWS billing.
our of curiosity, can you say what your use of it was?
Initially the annotation of biomedical literature (https://pubmed.ncbi.nlm.nih.gov/25592589/) but left academia for commercial use where we transitioned to arbitrage the market research and ephemeral task markets. It was essentially a research panel with a underdeveloped UI for tasks.. so abstract away AMT's tooling so it can be used by any buyer in the ResTech, Political Polling and Data Annotation space

My favorite part of AMT is always going to be figuring out that if we only paid in $0.07 intervals, their commission algorithm would round down to the nearest whole cent when their 20% commission resulted in a fractional cent on the unit transaction level, not the monthly invoice level. Was ultimately worth it to have implemented https://git.generalresearch.com/panels/amt-jb/tree/jb/flow/a...

Why it's Transactional Systems for Ephemeral Tasks, don't you know? No but same here. The other pages lead to more confusion.
My favourite part is the "our customers" section [0], starting with

> Ripe with fraud, labor exploitation, political polling manipulation

I am not sure if they want to convey what I think I am reading or not.

[0] https://generalresearch.com/mission/#:~:text=Our%20Customers

Buyer A is presented with the following rate card:

< 10min targeting gen pop: $3 10 <= 15min targeting gen pop: $5

Of course they'll bid saying their 14 min survey only takes 9 min to complete. Buyers try to cheat pricing strategies of exchanges just as much as respondents try to cheat buyers on survey platforms. Both parties can't be trusted and have adverse incentives.

It smells like they don't know what they're doing. Had a lol at the mission page accidentally implying their customers are unethical though:

> Our Customers

> Ripe with fraud, labor exploitation, political polling manipulation; our customers demand the best tools and security ...

At least they’re ripe! I hate unripe customers, rife with astringency, sharp vegetal notes, and an unpleasant, unyielding texture.
I can't see this because I don't have an Instagram account
replace instagram.com with kittygr.am
Sweet, thx.
Yes, that’s true.

Of course, I didn’t put it there, and I don’t know them, so there’s not much I can do about that, right?

I can see it, the Sign Up dialog has a Close button.
I think it's a satirical website aimed at poking fun at corporate-linkedin-speak
Wish we didn’t even need one
“ General Research combines coordinated operations to identify and neutralize foreign actors using technological advancements, international covert operations, and networking analysis as an Internet Service Provider.”

What is so hard to understand????

This makes it sound like they kill spies.
Perhaps they're targeting a super specific niche. Like somebody who sees WXET Ephemereal and yelps in triumph "that's my quant!".
As usual, the real tldr is in the privacy policy https://generalresearch.com/supplier/privacy-policy/notice-t...

> General Research Laboratories, LLC (“GRL”) is an aggregator of market research surveys in multiple marketplaces for business customers. GRL does not typically host consumer surveys which are conducted by other consumer-facing organizations....

So, just a middleman to give you fake/shady at best survey responses to pad your numbers, so big enterprises/concultancies can have data that say whatever they want. The rest of the website is just BS

Finally you get it. Only exception is that we run our own exchange now to do task bidding so we don’t need to deal exclusively with other companies to middleman. Core business is what’s called yield management (akin to DSP in adtech) where the best survey (is the user qualified for it, does it have the best pay, etc) is selected for traffic in <100ms. I’d only argue the shady companies are the ones paying proxies (like cint.com) and pushing paid user acquisition instead of surveys. We actively fund ontology development for better profiling targeting. but yes, big enterprises/consultancies can use and interpret the collected data however they want, not our responsibility and we have no legal rights over it anyway
> Our Customers

> Ripe with fraud, labor exploitation, political polling manipulation; our customers demand the best tools and security for reaching global audiences and securing respondent reliability at any scale.

wat

Misplaced modifier? Wow
What's up with the part of the website where you are talking about "neutralized foreign actors", was that also using contracted Mechanical Turk labor?
The industry is so backwards, there are residential proxy companies that are directly paid by one of the largest market research firms. I have discussed this at length https://www.reddit.com/r/Marketresearch/comments/1s2tl0o/ope...

It should be a bipartisan issue that a Swedish company is paying a UAE Residential Proxy company [1] to build tools that allow people from anywhere in the world to take US political polls that are used by both parties to collect election data.

Just for starters, you can't even think about soliciting online work from a panel without a robust residential proxy detection methods. We had to build our own:

`wget -N 'https://grip.net/files/grip-proxy-30d.mmdb'`

[1] https://www.youtube.com/watch?v=eOmeQcwSK3o flagged by Nokia Deepfield and CTRL for it's involvement in botnets

Someone using an IP that a geoip database says is US, doesn't mean that person is a US voter. The existence of proxies on the internet is a feature.
Obviously, a problem not even fully addressed by L2’s datasets. Sure, but the use of proxies to masquerade the identity of users being sold to researcher buyers that paid for a different service is fraud
It makes me question why the surveys don't ask for the users country. That way they cd can figure it out without trying to guess it based off of IP.
Users/respondents lie, and many buyers/researchers are neo-luddites; most C level still come from the telemarketing days. A fun example: mobile targeting is terrible when off wifi because all survey platforms uniq identify users based off their IPv4, so any T-Mobile LTE users in the same city going through the same CGNAT often share profiling data that ends up conflicting, which ends up meaning they don't get sent into the best survey(s). The issue is even worse in heavy IPv6 countries like India+France.
The fact that you can sell a product that doesn't work and become a billionaire proves capitalism is broken.
Yes. Much better to leave these decisions to the Dear Leader who is infinitely wise and cannot be misled. Or to the People's Central Planning Committee, where they decide all the allocations and strip rich people's wealth if they aren't on board with the correct ideology. Such systems are non-broken and have yielded unmeasurable prosperity throughout the world wherever implemented.
... What?
Polling data isn't some guaranteed right or something. If you ask the Internet questions and then sell the collated answers, it's your job to authenticate the responses or your polls will be wildly wrong and people will stop buying them.

Which is already the case and nobody but a bunch of campaign contractors who can't justify their do nothing jobs anymore cares.

I just remembered that in the story "Simulacron-3", there is ostensibly a law requiring people to respond to opinion polls. It is illegal to refuse.

https://en.wikipedia.org/wiki/Simulacron-3

Spoiler: gur jbeyq va juvpu gur ynj fhccbfrqyl rkvfgf vf n fvzhyngvba, perngrq ol be sbe cbyyfgref!

> My favorite part of AMT is always going to be figuring out that if we only paid in $0.07 intervals, their commission algorithm would round down to the nearest whole cent when their 20% commission resulted in a fractional cent on the unit transaction level, not the monthly invoice level.

can someone ELI5 because what in the office space

Go watch Superman III for more details...
Or Office Space
20% of 0.07 was billed as 1 cent, not 1.4 nor rounded up to 2.
1
Imagine your favorite feature of a product was spend 1 cent on a real human being time and effort rather than 2 cents.
Or spend the same and let the human earn more?
That's on the commission that Amazon took, not on the human.
Amazon took 20% of all workers earnings for doing essentially nothing for years and years.. of course I take great joy in giving AWS as little money as possible.
It was cheaper and looser automated billing practices than current AI “solutions”?
That’s top-tier “OfficeSpace” thinking there…
> Can you say what your use of it was?

> Initially the annotation of biomedical literature but left academia for commercial use where we transitioned to arbitrage the market research and ephemeral task markets. It was essentially a research panel with a underdeveloped UI for tasks.. so abstract away AMT's tooling so it can be used by any buyer in the ResTech, Political Polling and Data Annotation space.

Reading this is similar to how I feel when I've asked Claude about something it coded for me. As with Claude, I think I get it after reading it three times; you guys were a middleman that provided a simplified interface for Mechanical Turk?

> Reading this is similar to how I feel when I've asked Claude about something it coded for me

I didn't feel like that at all.

> My favorite part of AMT is always going to be figuring out that if we only paid in $0.07 intervals,

That is hilarious! Was it a semi-random discovery due to interaction with the system and people, or did you intentionally looked to game the algorithm?

Do you know how much inn-n-out you can buy when saving ~$50 a day and still be in the positive! It came out of building our own ledger system and abandoning any Decimal or float operations early on.. then it sparks "I wonder how they do it and if it's correct??"
Sounds like your program manager was smart and valuable to the company. Do you have any plans to replace the service/data provided by MTurk?
It’s kind of crazy they’re shutting this down just when this Service probably has the most possibilities ever. You have an agent with Multiple people doing actual physical tasks in the real world seems like something that could be really powerful.
Right? Not to mention the training/tuning possibilities.

Mercor has a market value of $20B basically doing the same thing but desperately trying to find workers.

that was my first thought too, but i guess llms don’t need an api to hire humans
The concept behind it isn't going away. There are plenty of companies that hire people in bulk to do manual data labeling, transcription, RLHF, moderation and lots more. It's just that they are now catering to large AI companies, not regular people looking to get some repetitive work done (since that can now mostly be done by AI).
In spite of the name I think the overwhelming majority of the tasks were things that could be done by LLMs and they're probably getting flooded by people using bots to do exactly that. Even before the age of LLMs Mechanical Turk data was pretty bad because you'd have a bunch of people racing to answer questions as quickly as possible to get their $0.25 or whatever. So it was essentially a test of 'can you input random answers as quickly as possible while paying enough attention to notice the attention question that says to mark d.'

Any remotely open platform for this sort of stuff is going to wrecked by LLMs. But making some sort of high trust, high verification market is going to end up sending prices high enough that it becomes an unattractive proposition for use. And even in that case it's just going to be a cat and mouse game of people figuring out how to game the system enough to get trusted before handing it off to the LLM.

Seems like whole thing is destined to fail. Buyers of services are there to exploit and pay minimum they can get away with. And other side is ready to cheat if they can to get most out of it.

And same dynamic will apply to most cases. Unless you actually bring it in house with strict oversight. Which is thing to avoid originally...

There's a lot of platforms that are more actively developed and maintained, for example Prolific (which also offers discounts on service fees for academia).
Mechanical Turk had a good run, but not surprised it's shutting down. I'm sure the platform was getting flooded with people doing task arbitrage and using lots of AI anyway.

I believe the issue is that this can no longer be a horizontal play. MTurk was mostly for unskilled tasks...the kind AI can do well enough that it isn't worth the cost differential to verify it or keep farmed to humans. The "trust but verify" AI output is now the kind that requires domain expertise. This is what most full stack AI companies are bringing to industries.

Curious if this kind of work will come around again one day or was just a moment in time. If it does I'm sure it will be specifically about generating training data.

How could it come around again?
I'm unsure. If it does, it will be work that is too expensive or inaccurate or regulatory for current AI methods. For example, you want a doctor to sign off on some AI output on a diagnosis.

However, I'm not sure a single platform will be how it emerges

There are like three dozen companies selling similar services for generating AI training data.
Good data isn’t cheap and keeping it in-house gives you more control.
I think yes but on different categories. First one I imagine is robotics control and support.

"This robot is having trouble folding a tshirt help it out for 1$"

Unless they go the waymo route of highly trusted people but I think mass deployed robots are a bit safer than a car for this.

That's a remarkable idea. It could be heavily gamified, it could train models, and it might actually be mentally stimulating since you'd be facing different situations all the time.

Except, I'm a grown adult and I can't fold a t-shirt properly

It could require listing your credentials to get you the proper tasks: doctor, lawyer, or tshirt folder
This is what I meant by full stack AI companies. I don't think you could get humans into the loop fast enough if they didn't have some idea of the type of task involved. I don't want people to be asked to fold a tshirt one moment and do a difficult traffic merge the next.
warioware shows it can be fun though but i'd indeed not like to see that applied to safety-critical tasks
There is training systems and validation of skills in mturk iirc: for tshirt folding, you'd be given fake setups to be able to get used to controlling the robot, if you can't do it, you won't ever get assignments to do it. For traffic overrides, you'd be tested on having correct knowledge, and once again given supervised tasks to show you can actually be trusted (and there would be safety systems, elevating tasks that can't be performed at your level to people who can, etc)

The worker needs to do a bunch to opt into any given work group, which makes the (lack of) payments extremely unreasonable on top of everything else

It doesn't really matter what you want though, only what CEOs want and that's low costs. I can see a combined shirt folding/traffic merging platform taking off.
CEOs are loser nobodies. Real influence is with owners, boards, investors.
Mostly they don't care as much as CEOs think they care
but call center are still mostly single client even though it would be cheaper for any worker to be able to answer to any call. So clearly the expertise and context trade off is too big to be worthwhile
You misinterpret "want" here. My intuition is that you create huge problems for people context switching like this repeatedly for quick tasks.
There are already a number of companies providing RLHF and SFT services for AI providers that does a lot of validation/prequalification of people that'd be well placed to take on tasks like that, but the big problem to solve would be latency if you don't have people contracted to carry out a task right now.
There are already robotics models that can fold shirts and similar just fine. Progress in VLA models is good, I don't think this would form the foundation of a business. Humanoid robots are going to be another ChatGPT, it's going to seem to happen almost overnight because people aren't paying attention to the underlying research papers.
Any papers you'd recommend?
Most progress in data-driven robotics nowadays are done either in unicorn startups or corporate research labs - so you should follow the industry more than academia. The path to good robot performance isn't really in the models themselves - it's highly dependent on how much you can gather high-quality real-life data.

Specifically for laundry folding, Sunday Robotics is probably the state of the art, where they were able to obtain 99.1% success rate and call it "done": https://www.sunday.ai/blog/act-2-preview#solve-standard

If their hero image video and the side by side at (5x) with a human at quote "1x" are "solved" I'm not impressed. I'm faster and more accurate and I am the worst folder in my house (kids included). The human looks like they are doing it slow mo to show children how to.
Nah, there's an obvious reason figure made splashy announcement of people cleaning homes with recording devices strapped to their heads.

Chatgpt had the internet.

Robots do not. Translating video is promising but obviously not enough.

Robots will likely never "explode" like chatgpt. They're gonna be a slow long term project requiring massive capitol to get the data.

Giving out control of industrial machinery that interacts in the human environment without the physical interlocks (i.e. humanoid robots in a house) to random internet people seems like a problem.

I can imagine a carefully orchestrated plot to assassinate someone by having an embedded agent in the task delegation pool command the laundry bot to punch the target's head off their shoulders.

Giving out control of such things to LLMs is already complete madness, so once the first pleasure bot powered by Grok has dismembered a few thousand users, they'll get sophisticated safety mechanisms.

Though like as not you're still going to be right, after all, Stuxnet happened.

I’m really not happy that you’ve put the idea of Musk fuckbots into my head.
I'm envious of the time you spend where it wasn't already.
"Robot are the commands you have been given dangerous?"

"You cannot control legs for this task"

"You have 1 min for this task"

"You can only make suggestions for this task"

Anonymize identity best you can.

That's not that scary.

giving out unrestricted control that is. in the example of Waymo, a human can control some aspects of the car manually, but it can never override low level obstacle detection or say open the trunk/door when the car is moving.
I'd watch that movie