
Smart Speaker
A smart speaker is a speaker with a built-in microphone that responds to spoken commands and reads out answers aloud. Well-known devices include Amazon's Echo with Alexa, Google Nest, and Apple's HomePod with Siri.
A smart speaker is a speaker that additionally has a microphone and an internet connection. You address it with a fixed word, for example “Alexa” or “Hey Google”. After that, you can ask it something or give it a task: play music, set a timer, read out the weather. The speaker replies with an artificially generated voice. Well-known models include Amazon Echo, Google Nest, and Apple HomePod. Several hundred million units have been sold worldwide.
The speaker as a gateway into the household
Until a few years ago, people typed commands into devices. The smart speaker was the first mass-market product to make voice a normal way of operating a device. This is especially handy when your hands are busy, for instance while cooking. It’s also a genuine benefit for people with poor eyesight or who struggle with keyboards.
For manufacturers, it’s about more than just speakers. Whoever places the assistant in the living room also has a say in which music service starts first and which retailer gets the reorder. That’s why Amazon sold the devices very cheaply for years, sometimes below manufacturing cost. The idea was to make money later through purchases and subscriptions. This calculation only partly worked out: many people use the devices almost exclusively for music and timers.
At the same time, the devices are a constant topic in data protection discussions. A microphone that is always on sits right in the middle of someone’s home. It has repeatedly come to light that employees of the manufacturers listened to voice recordings in order to improve recognition. Today, it is legally required that users consent to such evaluations and are able to delete recordings.
From wake word to answer
Inside the device, a small program constantly runs that listens for just one single word: the so-called wake word. This recognition happens directly on the speaker, without the internet. Only once the wake word has been detected does recording begin and get sent to the manufacturer’s data centers. Think of it like a doorman who only opens the door upon hearing his specific cue. False alarms still occur, though: similar-sounding words from the TV occasionally wake the device.
In the cloud, meaning on the provider’s servers, the recording is converted into text. This step is called speech recognition. After that, another system checks what the sentence is actually meant to accomplish. “Set a timer for ten minutes” is thereby turned into a clear command with a numeric value. The service then carries out the action or fetches the information from a suitable source.
In the end, the answer is converted back into speech and played through the speaker. The whole process usually takes less than two seconds. Since 2023, manufacturers have been building in large language models, i.e. AI systems capable of formulating free-form sentences. As a result, the assistant can also understand nested questions that used to simply get the response “I’m sorry, I don’t know that”.
Echo, Nest, and the whole smart home business
You most commonly encounter smart speakers in the living room and kitchen. There, they serve as a music system, kitchen clock, and news radio. They often also control lights, radiator thermostats, or blinds — a so-called smart home. They’re now also showing up in hotel rooms and offices.
Related devices come with a screen, such as the Echo Show. These can additionally display recipes, front-door camera feeds, or video calls. In principle, it’s the same technology, just with a display. The same assistants are also built into smartphones, TVs, and cars.
In business news, smart speakers usually come up in two contexts. First, in the dispute over market share between Amazon, Google, and Apple. Second, when data protection authorities impose fines or manufacturers announce cost-cutting programs in their device divisions. A common misconception, by the way, is that these devices permanently store everything that’s spoken: normally, only what is said after the wake word gets stored.