Dreistufiges Ablaufschema eines AI Notetakers: links das Mikrofon einer Videokonferenz, daraus ein Pfeil zur Spracherkennung mit wörtlichem Transkript und Sprecherzuordnung, daraus ein Pfeil zum Sprachmodell, das rechts ein kurzes Protokoll mit Aufgabenliste ausgibt.

AI Notetaker

An AI Notetaker is a program that joins video conferences, transcribes what is said, and automatically generates minutes with a task list from it. Tools like this are now firmly built into services such as Zoom, Microsoft Teams, or Google Meet.

In many meetings, one person takes notes so that everyone can later read up on what was discussed. An AI Notetaker takes on exactly this task, but as software. The program listens in on a video conference, converts the spoken words into text, and then summarizes that text. In the end, there’s a set of minutes in your inbox: the key points, decisions made, and open tasks with names attached. Some of these helpers appear as a separate participant in the attendee list, others run invisibly in the background of the conferencing app. The English name has become common usage; what’s meant is simply an AI meeting scribe.

What automatic minutes change in everyday work

Meetings cost a lot of work time, and a large part of that goes into follow-up work. Anyone taking notes finds it harder to think along and contribute to the discussion. This is exactly where the benefit comes in: everyone involved can focus on the conversation, while the note-taking happens on the side. It also creates a searchable archive. Months later, you can look up when a decision was made and on what grounds.

For companies, this is also a business. Providers like Otter, Fireflies, or Read sell subscriptions per user and month, often in the range of ten to thirty euros. Zoom, Microsoft, and Google have built the feature into their own packages, putting pressure on the smaller providers. That’s why the term frequently comes up in tech news when it comes to software companies' revenue figures or acquisitions.

But there is a serious downside: data protection. In Germany, you’re not allowed to record conversations without the consent of those involved. Many companies therefore ban third-party notetakers from their meetings, because confidential content could end up on other providers' servers. Anyone using such a tool must inform participants beforehand.

From microphone to task list

The process consists of three steps. First, the software records the audio of the conference. Then a speech recognition system translates that audio into written text, word for word. Experts call the result a transcript, i.e. a verbatim record. Good systems also recognize who is speaking at any given moment and assign each sentence to a person.

In the third step, a language model comes into play, that is, an AI system trained on very large amounts of text that can continue or shorten texts. It receives the transcript along with the task of building a summary from it. Since an hour of conversation quickly adds up to 10,000 words, this is tough condensing work. The raw text becomes a few paragraphs and a list of tasks.

This last step is also the weak point. Language models can invent things that were never actually said; this is referred to as hallucination. Irony, dialect, and technical jargon lead to additional errors. A common mistake, therefore, is to blindly trust the minutes. It remains a draft that someone should briefly proofread.

Where these helpers are already taking notes today

You most often encounter them in work-related video conferences. If you see a participant named “Notetaker” or “Meeting Assistant” in a Zoom meeting, you’re looking at such a program. Schools and universities also use the technology, for instance to provide lectures as text for students with hearing impairments.

Related applications exist in medicine and law. Doctors have patient conversations transcribed so they don’t have to type during the examination. Journalists use the same technology for interviews. An AI Notetaker differs from a plain dictation app in that it doesn’t just transcribe, but also analyzes and organizes the content itself.

On a small scale, you may know this principle from your smartphone. Voice messages in messenger apps have been able to be displayed as text for some time now. That’s exactly the second step of a notetaker, just without the summary. So the technology is the same, only the scope is different.

Subscribe free. Unsubscribe the second it sucks.

High-signal news across AI, business, UX, and tech. Every morning.