Crowdsourcing

Crowdsourcing

Crowdsourcing means assigning a task not to individual employees but publicly to a large group of volunteers or paid casual workers. Many small contributions together add up to a result that a single company could hardly manage alone.

Crowdsourcing is made up of two English words: “crowd” means a large group of people, “sourcing” means procuring. So what is meant is: work is procured from a large crowd of people. Instead of tasking ten employees with something, one calls out to thousands of strangers on the internet, each contributing a tiny piece. Wikipedia is the best-known example of this. No publisher wrote this encyclopedia; rather, millions of volunteers each added or corrected a few sentences. Sometimes participation is voluntary and unpaid, sometimes it is compensated with a few cents per task.

Why the crowd is often smarter than the individual

Some tasks cannot be solved with more expertise, only with more people. A single expert cannot walk down every street in Germany and enter it into a map. Tens of thousands of volunteers can. This is exactly how OpenStreetMap came into being, a freely usable world map.

There is also a statistical effect at play. When many people estimate or judge independently of one another, their errors partially cancel each other out. The average of many mediocre answers is often better than the answer of a single person. Experts call this the wisdom of the crowd. However, it only works if the participants do not influence one another.

For the AI industry, crowdsourcing is even a basic prerequisite. Modern models learn from huge amounts of labeled examples. These labels have to be created by humans. Without distributed micro-work by tens of thousands of people, many of today’s AI systems simply would not exist.

From the big task to a thousand micro-jobs

The crucial step is breaking things down. A large task is cut into very small, uniform units that anyone can complete without training. Instead of “label 100,000 images,” the instruction then reads: “Is there a pedestrian in this image? Yes or no.” Such micro-tasks take seconds and require no training.

This is mediated through platforms. Amazon Mechanical Turk or Appen are marketplaces where clients post their micro-tasks and workers complete them for payment. A few cents per task is typical. The name Mechanical Turk alludes to a supposed 18th-century chess-playing automaton, which in reality had a human hidden inside.

Because you cannot blindly trust strangers, quality assurance is needed. Usually several people receive the same task, and only the majority result counts. In addition, test questions with known answers are mixed in to identify inattentive workers. A common misconception is that crowdsourcing is automatically cheap and fast. Organizing, controlling, and filtering out poor contributions requires considerable effort.

Between Wikipedia, map services, and click work

In everyday life, people often use crowdsourcing without noticing. The traffic jam display in Google Maps is generated from the movement data of many phones. Ratings on Amazon or Google reviews are collected judgments from strangers. And anyone who solves an image puzzle to prove they are human is, incidentally, helping to label training data.

In business news, the term usually appears in two contexts. First, in crowdfunding, i.e., financing a project through many small contributions via platforms like Kickstarter. This is a special form in which money, not labor, is collected. Second, in debates about the working conditions of so-called click workers.

This criticism should be taken seriously. A large part of the data labeling for AI companies takes place in countries with low wages, sometimes for just a few dollars a day. The workers are rarely permanently employed and have little protection. Anyone reading about the costs of AI systems should also keep this invisible human contribution in mind.

Subscribe free. Unsubscribe the second it sucks.

High-signal news across AI, business, UX, and tech. Every morning.