
Human Intelligence Tasks
Human Intelligence Tasks are small, clearly defined work assignments that people complete online for pay because computers are poor at solving them. They supply a large share of the labeled data from which AI systems learn.
Human Intelligence Tasks are tiny work assignments distributed to people over the internet. The name is meant literally: these are tasks that require human judgment. Typical examples include marking all pedestrians in a photo, transcribing an illegible receipt, or judging whether a comment is offensive. A single task often takes only a few seconds and is paid a few cents. They are brokered through online platforms where clients and workers never meet in person. The term originates from Amazon's platform Mechanical Turk, but today it is used generally for this type of micro-task.
The invisible raw material behind AI models
An AI system learns from examples. For it to learn what a traffic sign is, someone must first have shown it, across thousands of images, where traffic signs are. This labeling of data happens to a large extent through Human Intelligence Tasks. Without this human groundwork, most of today’s models would not exist.
The demand is enormous. A single image recognition system can require millions of individual labels. The famous ImageNet dataset contains over 14 million labeled images. Starting in 2007, they were sorted by tens of thousands of people on Mechanical Turk, spread across many countries.
This topic is also tied to a serious debate. Many workers earn the equivalent of only a few euros per hour and have no permanent employment. In content moderation, people must also view violent videos or depictions of abuse so that a filter can later recognize them automatically. Critics therefore speak of hidden labor behind supposedly fully automated systems.
From assignment to finished label
A client uploads their data to a platform and breaks the work down into uniform chunks. Each chunk comes with a short instruction and a price. Workers choose what they want to take on from a long list. After submission, the client reviews the result and pays out or rejects it.
Because individual people make mistakes, the same task is usually completed multiple times. Three or five people judge the same image, and the majority opinion counts as the result. In addition, control tasks are interspersed for which the correct answer is already known. Anyone who gets these wrong too often is filtered out. This principle is called quality assurance through redundancy.
A related but narrower term is Reinforcement Learning from Human Feedback, or RLHF for short. Here, people do not evaluate images but instead compare two answers from a language model and choose the better one. Technically, these are also Human Intelligence Tasks, just with considerably higher demands on the workers. For sensitive fields such as medicine or law, providers now employ trained specialists instead of anonymous casual workers.
Where these micro-tasks show up in everyday life
The most direct encounter with this principle is through captchas. When a website asks which tiles show a crosswalk, it is not only checking whether you are human. For years, the answers were also used to improve map data and image recognition. In doing so, you perform for free a task for which others get paid.
In business news, the topic appears under terms like data annotation or clickwork. Companies such as Scale AI, Appen, or Sama make their living organizing such assignments. Their valuations have risen sharply in recent years because AI providers urgently need high-quality training data. At the same time, there have repeatedly been reports of poor working conditions in Kenya and the Philippines.
A common misconception is the assumption that AI will soon make this work unnecessary. So far, the opposite has happened: the more capable the models become, the more demanding and expensive the human evaluations needed for the next step become.