
DisCo
DisCo is a training method for AI language models in which two models work together: one generates responses, the other evaluates them and provides feedback. The goal is to improve the quality and reliability of outputs without requiring large amounts of human evaluation.
DisCo stands for “Discriminator-Cooperator” and describes a process in which two AI models are trained together. One model — the generator — formulates responses or texts. The other — the discriminator, i.e. an evaluation model — checks these outputs and judges whether they are good or bad. This judgment flows directly back to the generator, which then improves accordingly. Both models therefore do not learn independently of one another, but through their cooperation. The special feature: this feedback loop can run largely automatically, without a human needing to intervene at every step.
Why the feedback loop is so valuable
The classic problem with improving language models is effort. Humans must read and evaluate thousands of model responses — this costs time and money. The larger the model, the more evaluations are needed to steer it reliably. DisCo attempts to circumvent this bottleneck.
Because the discriminator delivers feedback automatically, the generator can be corrected much more frequently than through human evaluators alone. This is comparable to a student who doesn’t just get a grade once a month, but receives feedback after every single sentence. Learning thus proceeds faster and more precisely. At the same time, the question remains how good the discriminator itself is — a poor evaluator gives poor feedback, and the generator then learns in the wrong direction.
The interplay between generator and discriminator
During training, the process runs in rounds. The generator produces a text. The discriminator assigns this text a score — for instance, for how correct, coherent, or helpful it is. This score is converted into a signal that tells the generator which of its decisions were good and which it should avoid in the future. Then the next round begins.
Importantly, the discriminator itself is also trained — often on a smaller set of actual human judgments. It thus learns what humans would consider a good response, and then transfers this knowledge to the generator at large scale. This is sometimes called a “learned reward signal,” because the reward is not assigned by hand but calculated by a model.
DisCo is thus related to a more widely known technique called Reinforcement Learning from Human Feedback — RLHF for short — in which human evaluations are likewise used to train a reward model. The difference lies in the details: DisCo places greater emphasis on the direct cooperation of both models during training and, in certain variants, can manage entirely without external reward models.
DisCo in practice and in the news
DisCo appears primarily in research surrounding large language models — that is, the kind of models behind chatbots like ChatGPT or Gemini. Labs are constantly searching for ways to make such models safer and more useful without needing unlimited amounts of human evaluation data. DisCo is one approach in this field of research.
In tech reports, the term is often encountered in connection with so-called “alignment” — the effort to train models so that they do what humans expect of them, rather than merely delivering statistically plausible answers. DisCo is not a finished product one can buy, but rather a method that companies and universities incorporate into their own training pipelines.
Anyone reading reports about new model versions and wondering how manufacturers improved their models so significantly between two versions often comes across training procedures of this kind. DisCo is one of the answers to the question of how to raise a model to a higher quality level quickly and cost-effectively.