Schematische Darstellung einer LSTM-Zelle mit den drei Toren (Vergessenstor, Eingangstor, Ausgangstor) und dem Zellzustand als separatem horizontalem Kanal, der von links nach rechts durch mehrere Zeitschritte verläuft.

LSTM

An LSTM is a type of artificial neural network specifically built to remember important information across long sequences. For years it was the standard solution for tasks like language translation or speech recognition, before transformer models replaced it.

An LSTM — short for Long Short-Term Memory — is a blueprint for an artificial neural network, that is, a computing system loosely modeled on the structure of the human brain. It was developed in 1997 by researchers Sepp Hochreiter and Jürgen Schmidhuber. The particular purpose of an LSTM is processing sequences — that is, data in which order matters, for example sentences, time series, or melodies. An ordinary network processes each input independently of the previous one. An LSTM, by contrast, maintains an internal state that stores information from earlier steps and passes it on to later ones.

Why the “forgetting problem” made LSTMs necessary

Before the LSTM, so-called recurrent neural networks, or RNNs for short, already existed. These networks pass information from each step on to the next step. That sounds good, but it has a decisive catch: the longer the sequence becomes, the more the early information fades. The network “forgets” what came at the beginning. This problem is known as the vanishing gradient problem.

Imagine you’re reading a long essay and are supposed to answer a question at the end whose answer was in the first sentence. A classic RNN would barely take that first sentence into account anymore. For tasks like translation or text comprehension, this is fatal, because a word at the end of a sentence often depends on the beginning of the sentence. This is exactly the problem the LSTM solved — and in doing so, it made an entire class of AI applications possible in the first place.

Gates instead of a through-flow: the mechanism inside the LSTM

The heart of an LSTM consists of three so-called gates: the forget gate, the input gate, and the output gate. Each gate is a small computational unit that decides how much of a piece of information is allowed through. The forget gate determines which old memories are erased. The input gate decides which new information is added. The output gate determines what is passed on to the next step.

In addition to these gates, the LSTM maintains what is called a cell state — a kind of separate memory channel that runs through the entire sequence. This channel changes only where the gates allow it to and otherwise remains stable. This prevents the fading of early information. All of these weights — that is, the numbers that control how the gates react — are learned automatically by the network during training from example data.

An important distinction: an LSTM processes a sequence step by step, one entry after another. It cannot therefore look at all parts of a sequence at the same time. This is exactly what later transformer models can do, which makes them more capable on very long texts.

LSTMs in products and in the news

LSTMs were built into almost every language product up until around 2018. Google Translate used them before the system was switched over to transformer architecture. Siri and Alexa used LSTM-based components for speech recognition. LSTMs were also deployed early on in music generation — for example, Google’s Magenta project — because music has a temporal structure that an LSTM can represent well.

Today the term mostly comes up in comparison with transformers. Media reports explaining the rise of GPT or other large language models often mention LSTMs as predecessors. In specialized fields such as the analysis of financial time series or medical signal data — for example, ECG evaluation — LSTMs, however, are still in use. The reason: for short, structured sequences, they are often faster to train and more resource-efficient than a large transformer model.

So anyone who reads headlines about AI models and understands what an LSTM is will also better understand why AI research didn’t stand still — and which specific problem the transformer solved that the LSTM still had.

Latest News

Subscribe free. Unsubscribe the second it sucks.

High-signal news across AI, business, UX, and tech. Every morning.