YouTube: Zhichao Wei Reads Everything →
Zhichao Wei: Books·Psychology·Film

Daniel Kahneman's Noise: A Flaw in Human Judgment | A 10,000-Character Deep Dive [Text-Polished Edition]

Daniel Kahneman's Noise: A Flaw in Human Judgment | A 10,000-Character Deep Dive [Text-Polished Edition]

This piece runs through the knowledge framework of Noise from beginning to end. Think of it as a complete guide to the book; if you read it with this overview in mind, the book may go down a little more smoothly.

I’ve divided my reading of Noise into five parts:

1. What is noise?

2. What kinds of noise are there?

3. What are the noise amplifiers?

4. How can we reduce noise?

5. What obstacles will measures to reduce noise run into?

There’s quite a lot here, so I’ve made a mind map. Save it if you think it might be useful:

https://www.mubucm.com/doc/5LQ2nZpti3Q


If you’d rather watch than read, head over to the video:

https://www.bilibili.com/video/BV1hg411c773/


1. What is noise?

Let’s use a shooting-range example to explain it.

There are four targets in the picture, showing the results for teams A, B, C, and D. Each team has five people, and they all use the same rifle.

Team A: every one of them is a crack shot, hitting the bull’s-eye every time. This is the perfect team.

Team B: everyone misses, but they miss in a remarkably tight, uniform cluster—all of them hit the lower-left corner of the target. Let’s call Team B the bias team. Bias is a uniform error like this. The reason for such a uniform error isn’t hard to guess: most likely, something is wrong with the sights on their rifle.

Team C: again, everyone misses, but they miss in all sorts of random directions. The bullet holes are scattered, with almost no pattern to be found. Team C is the noise team. Those wildly varied errors are noise.

Teams B and C represent the two kinds of errors humans make when judging. The first is Team B’s bias: when different people make judgments, they veer uniformly in the same direction. Why does this happen? Because all human beings share certain psychological tendencies. For example, we are all hypersensitive to negative information. So when we read a sensational, fear-inducing article online, it can trigger fear in all of us, leading us, without even coordinating, to make remarkably similar irrational decisions.

The second kind of error humans make is Team C’s “noise.” Noise is essentially fluctuation, or inconsistency: one person may make a different decision this time than next time, and a group of people judging the same thing may disagree with one another. That’s noise.

Noise is everywhere. There is noise in job interviews: interviewers often disagree sharply. One interviewer may decide you’re exactly right for the job, while another kicks you out because you may not fit the company culture. There is noise in medical diagnoses: one doctor says you need to be hospitalized and operated on immediately, while another says some medication will do. There is noise in stock-market forecasts: half the institutions say stocks will rise tomorrow, and the other half say they’ll fall. That’s business as usual. There is noise when teachers grade papers, too: in a good mood in the morning, they give higher scores; tired in the afternoon, they grade more harshly. Noise is extraordinarily common.

When making judgments, of course we’d all like to be crack shots like Team A. In reality, though, we may have bias like Team B, noise like Team C, or both. We keep missing the bull’s-eye and struggle to make accurate judgments. Bias and noise often appear together—that’s Team D in the picture: the whole group leans in one direction, showing bias at work, while the individual shots are still scattered, showing noise at work.

In one sentence—human error in judgment = bias + noise.

Kahneman’s previous book, Thinking, Fast and Slow, was essentially about Team B: bias. This book, Noise, is of course about Team C. With Noise, the problem of “error in human judgment” is finally complete.

2. What kinds of noise are there?

There are three:

1) Level noise

2) Stable pattern noise

3) Occasion noise

Level noise is the difference between people.

Gangzi and Dazhu are two judges. Gangzi is stern and impartial, extremely strict. Dazhu is more lenient. The same criminal gets 30 years from Gangzi, but only 5 from Dazhu. The sentencing of these two judges is wildly inconsistent. That’s level noise.

The second kind is called stable pattern noise. Stable pattern noise is personal preference.

For example, Judge Gangzi is generally harsh with criminals, but he has always had a soft spot for white-collar workers, so he is more lenient with white-collar offenders. Two people commit crimes of roughly the same severity, and both cases go to Gangzi—but one gets a lighter sentence simply because he’s a white-collar worker. Personal preference leads to inconsistent judgments. That’s stable pattern noise.

If there is stable pattern noise, naturally there is unstable pattern noise too. Gangzi likes white-collar workers; that’s a fairly stable tendency—once he’s like that, he stays like that. But some tendencies are unstable and constantly shifting. Take mood. Gangzi is a night owl and is always especially sleepy in the morning. When he’s sleepy, he’s in a worse mood, so criminals who appear before him in the morning get heavier sentences. By afternoon, Gangzi is more alert and in a better mood, so the afternoon criminals get lighter sentences. That’s unstable pattern noise. Unstable pattern noise is really noise produced by different settings and situations, so it’s also called “occasion noise.”

To sum up:

The disagreement between Gangzi and Dazhu is level noise;

Gangzi’s difference in treatment of white-collar and non-white-collar workers is stable pattern noise;

The difference between Gangzi in the morning and Gangzi in the afternoon is occasion noise.

—Those are the three types of noise.

One thing to note: these three types aren’t mutually exclusive. They can show up at the same time. Gangzi might encounter a white-collar criminal while in a bad mood and also disagree with Dazhu’s sentence.

3. Noise amplifiers

Differences between individuals always exist, everyone has their own quirks and preferences, and people’s states rise and fall with the situation. So noise is almost impossible to escape. To make matters worse, some factors can increase its magnitude. The book doesn’t give these factors a collective name, but I think we can call them “noise amplifiers.”

There are three common noise amplifiers.

The first is called “objective ignorance.”

Objective ignorance means that the person making the judgment genuinely has no way of knowing some necessary information. When is objective ignorance most likely to appear?—When we’re predicting the future.

Will stocks rise or fall tomorrow? How will international relations change over the next three years? What will the real-estate market do over the next ten? Noise is enormous when we make predictions like these. Stockbrokers, political commentators, and economists often disagree dramatically, sometimes making completely opposite predictions. But we can’t entirely blame them for being incompetent: the future is determined by the present together with many things that haven’t happened yet, and those future events can’t be predicted today. That’s objective ignorance. So predicting the future inevitably involves some blind guessing. Different experts guess blindly in different directions at different times—how could the noise not be huge?

Objective ignorance is what we encounter when predicting the future. It’s the first noise amplifier.

The second noise amplifier is called “the matching problem.”

What is the matching problem? After watching a movie, you want to review it. Generally, you won’t write a long review laying out every single one of your feelings. You’ll open Douban, go to the movie’s page, and give it a rating from one to five stars. That’s matching: you’re matching your subjective response to the movie with five levels, from one star to five. We call a one-to-five-star rating system a scale. Whenever we use a scale to make a judgment, the matching problem is inevitably involved.

So how does the matching problem amplify noise?

First, it amplifies level noise—the disagreement between people. Not long ago I gave Free Guy five stars. My standard for five stars was “Was it surprising?” I thought it would be a silly, chaotic movie, but it turned out to be unexpectedly brilliant. I was pleasantly surprised, so I gave it five stars. But five stars may mean something completely different to someone else. For another movie fan, five stars are reserved for timeless classics. So even if that fan’s subjective response to Free Guy is actually pretty similar to mine, they may give it only three stars. To them, three stars is already a very high score. We understand the one-to-five-star scale completely differently, so our ratings produce level noise, just like the disagreement between Gangzi and Dazhu.

Second, the matching problem amplifies occasion noise as well. Take performance reviews at a company: a manager has to rate a subordinate somewhere between 0 and 100. Matching a person’s performance to a number between 0 and 100 is actually very difficult, so the manager’s ratings can be pretty erratic. Dazhu gets 80 today and might get 90 tomorrow. His performance hasn’t changed in the manager’s subjective assessment, but once it’s turned into a number, it starts wobbling. The matching problem has amplified occasion noise—in other words, people’s understanding of a scale can change from one situation to another.

The matching problem is the second noise amplifier.

The third noise amplifier is the harmful influence of groups.

We often say, “Three cobblers together can outsmart Zhuge Liang.” Pooling everyone’s efforts seems as if it should improve the quality of decisions. But we’ll soon see that when a group makes a judgment or decision together, improving decision quality actually requires some stringent conditions. If those conditions aren’t met, the group may increase decision noise instead. Three cobblers are often worse than one.

So how do groups amplify noise? There are two ways it can happen.

The first situation is that whoever gets the first-mover advantage gains an overwhelming edge. The book mentions an experiment in which scientists gave volunteers a playlist and let them listen to any songs they wanted; if they came across one they especially liked, they could download it. By counting the number of downloads, the researchers could tell which songs were the most popular and best liked. But the researchers had quietly rigged things a little: when they sent the playlist to the volunteers, some songs were already shown as having been downloaded many times. The volunteers assumed those songs had been downloaded by other people who took part in the experiment earlier. In reality, the supposedly popular songs had been chosen at random by the researchers. The surprising thing was that, by the end of the experiment, the songs randomly designated as popular at the beginning had genuinely become the popular ones. The volunteers really did download them more often. In other words, if you gain even a slight advantage at the start, that advantage will automatically grow.

This is also the principle behind the way traffic celebrities manipulate online opinion, or “comment control.” Once the top comments are praising the celebrity, everyone who comes later is influenced by them.

Why does this effect happen? Because how good most songs are is somewhat ambiguous, so people are easily influenced by other people’s judgments. If someone decides at the beginning that these few songs are good, everyone else starts thinking they are good too.

And that creates a great deal of occasion noise. Next time, if a different batch of songs happens to get the first-mover advantage, those songs will become the popular ones.

The experiment also turned up an interesting side finding: the very best and very worst songs were unaffected by this dirty trick. Even if the best song started with zero downloads, people would eventually discover it. As for the worst songs, even if they topped the charts at the beginning, people would still end up hating them. So operations like comment control have their limits too. If the quality is truly awful, the eventual crash is only a matter of time.

There is another situation similar to that experiment. After interviewing each candidate, for example, interviewers may sit together to discuss whom to choose. The interviewer Gangzi is especially fond of Little Li. He speaks first and explains how good Little Li is. Next it is Dazhu’s turn. Dazhu had no particular fondness for Little Li, but seeing how certain Gangzi sounds, he assumes Gangzi must have noticed some special advantage in Little Li, so he echoes him a little: “Little Li really is pretty good.” Then Tiedan speaks. Tiedan’s initial impression of Little Li was bad, though he did not have much evidence for it. Now that he sees both Gangzi and Dazhu supporting Little Li, he decides not to say much. And so Little Li, who may not have been particularly outstanding in the first place, ends up passing the interview unanimously and by a landslide.

But in reality, this is entirely the result of Little Li happening to get the first-mover advantage. If, in another interview, the interviewer Tiedan—who likes another candidate, Little Gui—had spoken first, Little Li would have had no chance. That is how occasion noise is formed.

Whoever happens to get the first-mover advantage will then expand that advantage and influence everyone else’s choices in the group. This is the first reason groups increase noise.

The second reason groups increase noise is called “group polarization.” Polarization means becoming more extreme. Group polarization is what happens when people who already hold roughly similar views on an issue gather to discuss it, and afterward their views often become extremely polarized: people who already liked something like it even more, while people who already disliked it hate it even more. That is group polarization.

For example, a bunch of fans spend all day discussing the traffic celebrity they idolize. The result is that, although they initially only kind of liked the celebrity, afterward they have all become obsessive fans.

——You might say, “That doesn’t sound right. Doesn’t this reduce noise? Everyone has become equally extreme, so they are more consistent with one another.” But the problem is that the world contains many different groups. A group that initially only somewhat disliked a traffic celebrity may, after group polarization, all become bitterly hostile toward them. The disagreement between fans and nonfans therefore grows larger. So, overall, group polarization still increases noise.

In short, whether because someone got the first-mover advantage or because of group polarization, groups often amplify noise.

Objective ignorance, the matching problem, and the harmful influence of groups: these are the three common noise amplifiers.

4. How can we reduce noise? What methods are available?

I’ve distilled the book’s content a little. There are four common methods:

1) Give the decision-making process to a group in which each individual can make an independent judgment;

2) Replace matching with ranking;

3) Give the decision-making process to a model;

4) Give the decision-making process to “decision experts.”

The first method for reducing noise is to give the decision-making process to a group in which every individual can make an independent judgment.

The keyword here is “independent.”

We just said that groups often lower the quality of decisions. The key reason is that the judgments of people in the group are not independent enough—the judgment of a song is influenced by its initial download count, and an interviewer’s judgment is influenced by the other interviewers.

But once every individual in a group makes a completely independent judgment, the situation reverses by 180 degrees. Combining many independent judgments usually reduces noise significantly. Three independent cobblers really can equal Zhuge Liang.

For example, Noise mentions that, in the legal field, fingerprint identification is particularly vulnerable to the harmful influence of groups.

I was pretty surprised when I got to this part. I had assumed fingerprint identification worked like it does in so many movies: fingerprints from the crime scene are entered into a computer, a program automatically compares them, and it eventually spits out one matching suspect. But Kahneman tells us that reality is completely different. Fingerprints collected at crime scenes are generally very poor quality—either incomplete or blurry—so comparing fingerprints depends to a great extent on the experience and subjective judgment of identification experts. That makes fingerprint identification prone to exactly the kind of meeting-room dynamic we just saw with the interviewers.

The identification expert Gangzi compared a fingerprint collected at the crime scene with the suspect Little Gui’s fingerprint and concluded that it belonged to Little Gui.

The detective handling the case was uneasy about the result, so he took Gangzi’s report to another identification expert, Dazhu, and said, “Dazhu, this is Gangzi’s report. Could you please take another look?” After examining it, Dazhu also said it was Little Gui’s fingerprint. But this was not because Dazhu and Gangzi had genuinely reached a consensus. Dazhu already knew Gangzi’s conclusion, so he was very likely to agree with it. Dazhu had not made an independent judgment.

Yet this identification procedure was actually the standard procedure used in American forensic identification for a long time. Each expert later in the process received the earlier experts’ results before conducting their own examination. The result was a great many wrongful convictions and other unjust cases, because in practice the identification became a one-person show led by the first expert.

How can we avoid this error? It is actually simple: give the fingerprint to different experts at the same time and have them make independent judgments. If 10 experts independently examine the fingerprint and most of them—for example, seven—say it belongs to Little Gui, then the detective has reason to strongly suspect that Little Gui is the culprit.

Make independent judgments, then aggregate the results of those independent judgments. This principle can be applied to all kinds of decision-making situations. When making forecasts, for example, you can use a procedure called the “Delphi method.”

The Delphi method works like this: if you want a group of experts to forecast economic trends, first have each expert make an individual forecast. Then average their forecasts of the economic data and feed the average back to every expert, allowing them to revise their forecasts on that basis over several rounds. In other words, experts can receive statistical information about the other experts’ forecasts, but they cannot discuss things with one another. The average forecast produced this way is much more accurate than one produced by having the experts gather and discuss things freely.

The core of this procedure is independent judgment.

Giving the decision-making process to a group in which every individual can make an independent judgment: that is the first method for reducing noise.

The second method for reducing noise is to replace matching with ranking.

The independent judgment we just discussed is a countermeasure against the harmful influence of groups. Replacing matching with ranking is a countermeasure against another noise amplifier mentioned earlier—the matching problem.

As I said earlier, if a leader wants to use a score from 0 to 100 to match employees’ performance, that kind of matching is difficult, and the standard will inevitably be pretty slippery. So how can we make it easier? First divide the employees into simpler levels—for example, excellent, good, average, poor, and very poor. That is much easier than assigning each person a score out of 100. Then rank the employees within each level. Little Li and Little Gui are both in the excellent level, the 81–100 range. If Little Li performs better than Little Gui, Little Li ranks higher. Once the ranking is complete, assign scores in order based on the ranking.

The principle behind this operation is that we can often judge more accurately when making pairwise comparisons. It is hard to say exactly how many points Little Li and Little Gui’s job performance is worth, but it is relatively easy to determine which of them performs a little better. So replacing matching with ranking can reduce noise.

The two methods we just discussed are targeted methods. The next two are general-purpose methods.

The third method for reducing noise is to give the decision-making process to a model.

A model is a formula, a set of rules, or an algorithm.

For example, when hiring, instead of deciding whom to hire based on the interviewers’ subjective judgments, use a formula. Add up the candidate’s scores for conscientiousness, work ability, and teamwork, and hire the person with the highest total.

Likewise, when forecasting the economy, instead of gathering a group of economists, feed all kinds of current economic data into an artificial-intelligence algorithm and let the algorithm produce an economic forecast.

The gaokao is actually a model too: it uses the model of an exam to select students in place of human judgment.

The reason human judgment is noisy is, put bluntly, that people are too flexible and too changeable, while models are very rigid. So models have no noise, because at bottom they are just fixed mathematical formulas. As long as the input is the same each time, the output is guaranteed to be exactly the same. When it comes to reducing noise, models have an overwhelming advantage over people.

So, if we start from the goal of reducing noise, humans should hand over as much judgment as possible to models.

Even if humans cannot step out completely, they should let the model go first whenever possible and intervene as late as possible. We should follow the principle of “model first, humans second.” Google does exactly this when hiring. When Google interviews applicants, the process has two major stages. The first asks applicants to complete a series of fully standardized tests and interviews. Each assessment is scored and evaluated independently, and the details of every assessment are fixed—including strict rules about which questions may be asked in the interview. The results are then aggregated into an applicant profile. That profile is essentially generated by a completely rigid interview model.

But Google does not eliminate human judgment entirely. In the second stage, the profile is handed to a hiring committee, whose members read it and give the final hiring recommendation.

Google’s hiring process does not eliminate those subtle human intuitions and judgments. It simply delays human intervention as much as possible: let the model judge first, and hand things over to humans only at the end.

One point needs emphasizing here: using a model to make decisions does indeed eliminate noise, but that does not mean the model’s judgment is more accurate. Remember, error in judgment = bias + noise. Very often, a model is merely changing Team D on the target diagram into Team B. The noise is gone, but the bias may still be there: the guns are now firing in perfect unison, just at the wrong spot. So handing decisions to a model does not magically solve everything. Humans are responsible for continually optimizing the model, nudging its judgments as close as possible from Team B toward Team A.

Here is the fourth way to reduce noise: hand the decision-making process over to “decision masters.”

The last way to reduce noise is simple and brutal: find the sharpshooters on Team A in that target diagram. Find those “decision masters” who can hit the bull’s-eye almost every time, and let them make the decisions. Wouldn’t that solve the noise problem?

So how do we find these “decision masters”?

Decision masters generally have three notable characteristics.

First, they may be experts in a particular field. Lawyers at top law firms and doctors at top hospitals make judgments far more accurately than ordinary people, so we can give priority to their judgments.

But be very careful. There are two kinds of experts: genuine experts and honorary experts.

Lawyers and doctors are generally genuine experts, because they earned their status through real results. A famous lawyer became famous by actually winning cases; a famous doctor became famous through genuinely outstanding medical skill. Their performance can be objectively verified. These are the experts whose judgments are worth trusting.

But there is another kind of expert: the “honorary expert.” An honorary expert’s performance cannot be verified. Take some management consultants or political analysts, for example. Their track records are difficult to test against objective standards. If a company improves, it is credited to the management consultant; if the company collapses, that is blamed on other objective causes—in short, never the consultant’s fault. Experts like these build their status on the respect of clients and peers. Put bluntly, it is a little flimsy. The judgments of honorary experts are no better than those of ordinary people, so you should keep your eyes open when listening to them.

Besides expert status, the second notable characteristic of “decision masters” is that they are generally smart people.

In almost every field, intelligence is associated with better performance. More intelligent people are also more likely to have good judgment. So if we have to choose between two opinions and have no other information to go on, we should give priority to the more intelligent person.

The third notable characteristic of “decision masters” is that they have very open minds.

Intelligence is only one side of the story; the way people think matters a lot too. Some smart people are stubborn and self-righteous. They are convinced their own judgments are right, and people like that usually do not have very good judgment.

What is truly impressive is someone with a very open mind. Such people are not only smart but also humble and practical. They do not mind hearing opinions that contradict theirs. They may even actively seek out new information that conflicts with their original views. They do not reject the task of integrating new information with their current opinions; they may even hope that new knowledge will change their minds. Judgments made by people with this kind of open mind are often more accurate.

So if we have to choose between the advice of two smart people, we should prioritize the one with the more open mind.

Listen to experts, listen to smart people, and listen to people with open minds. That is the fourth way to reduce noise.

We have said a lot. We now know that noise is widespread, and we know many ways to reduce it. So why not have every industry spring into action and use these methods to reduce noise as much as possible? In reality, measures to reduce noise face a lot of resistance.

And that brings us to the final question—

5. What obstacles will measures to reduce noise face?

The first obstacle is that people do not trust algorithms.

As I said earlier, handing decisions over to models can reduce noise dramatically, and the most popular models today are of course algorithms—especially artificial-intelligence algorithms based on deep neural networks. The trouble is that people generally do not trust algorithms. Humans have a very subtle psychological tendency: we often expect algorithms to be perfect. Making mistakes is a human privilege; machines are not allowed to make them. When humans cause a car accident, we find it understandable. But an autonomous-driving system must have 0 accidents, or it is untrustworthy. Naturally, if people do not trust algorithms, they will resist measures that use algorithms to reduce noise.

The second obstacle measures to reduce noise face is that people worry models will crush motivation and creativity.

If the entire decision-making process were handed over to a rigid model, would people feel like cogs in a machine, with no room for agency at all? If every employee in a company felt they could not decide anything, would their morale take a serious hit? And although human decisions can be erratic, they also produce a lot of creativity. If the decision-making process were handed over to a rigid model, would that crush human creativity?

This actually concerns a major question: in what situations do we really need to reduce noise? We hope a doctor will not get creative or exercise personal discretion when diagnosing whether someone has high blood pressure; in situations like that, the less noise, the better. But in a company, if we want employees to be happier and more inspired, should we allow some noise? Or could we learn from Google’s hiring method: standardize the creative process and follow the model’s steps at the broad structural level, while still giving people plenty of flexibility at each step to exercise their creativity? Is that workable?

Which situations require noise to be reduced? How can we reduce noise without damaging people’s motivation and creativity? These too are questions well worth thinking about.

The two obstacles above are really aimed at models. The final obstacle is the universal one.

The last obstacle to reducing noise is cost. Reducing noise is certainly good, but is the cost-effectiveness high enough?

For example, if one teacher grades a set of papers, there will inevitably be noise. The most appropriate solution would be to have five teachers grade the papers independently and then average their scores. But who pays for all that extra work? Even if someone does pay, is the expense worth it? How do we measure the cost-effectiveness here? These are also questions that deserve further thought and research.

Although Kahneman repeatedly emphasizes in the book that measures to reduce noise are necessary because noise creates a great deal of unfairness, after reading this section I felt that questions such as how algorithms can win human trust and how to measure the cost-effectiveness of reducing noise still lack fairly deep research at this stage. Compared with the earlier sections, this part is more open-ended, with many details still worth thinking about and exploring.

That completes the outline of the knowledge framework of Noise: A Flaw in Human Judgment.

That’s all.

Written by: Zhichao Wei

Original article link

Leave a Reply

Your email address will not be published. Required fields are marked *