Zhihu question: Which films and TV works make especially good use of the uncanny valley theory?
As someone who is hopeless with technology, I’ll boldly use the questioner’s prompt as an excuse to talk a little about the technical side of Avatar. My examples are limited to my own moviegoing experience, so there’s a very good chance I’m spectacularly wrong.
There’s a famous uncanny valley hypothesis in theories about robots, proposed in 1970 by Japanese robotics expert Masahiro Mori. What little I know about it comes from Wikipedia, so I have no idea how much scientific evidence supports it, but it does seem pretty plausible. The hypothesis says that with robots that imitate humans, “because they are quite similar to humans in both appearance and movement, humans will also develop positive feelings toward robots.” But once “they look more than 95% human, even the tiniest difference will seem glaringly obvious, making the whole robot look stiff and terrifying, and giving people the feeling that they’re facing the walking dead.”

A simple example: when a doll doesn’t look much like a real person, you find it cute. But once it becomes realistic to a certain degree, even if its design is still cute, you’ll find it eerie and frightening.

This hypothesis naturally applies to CG as well. Unfortunately, when CG creators charge toward the goal of simulating real human characters perfectly, they fall into the uncanny valley nine times out of ten. The bloodiest example was probably Final Fantasy: The Spirits Within from years ago. Hironobu Sakaguchi and the other creative leads poured their hearts and souls into trying to raise the realism of CG characters to a new level. Technically, they surely did reach a new level for the time—but the result was a screen full of blank-faced zombies, and the film ended in total critical and box-office disaster. Another figure who could practically be called a living sacrifice to the uncanny valley hypothesis is Robert Zemeckis. Ever since he started tinkering with motion-capture technology, the great director who once made Back to the Future and Forrest Gump suddenly seemed to have been possessed by the wooden people in his Beowulf, completely losing his spark. Still, he kept trying after falling into the valley again and again. You have to admire the courage.
So to this day, even when CG works can technically get fairly close to real people, creators often lower the realism on purpose because of the uncanny valley. They either cartoonize the characters, like Pixar (the characters in Up have heads as big as their bodies), or, as in some Japanese CG, capture only movement and not facial expressions. Take Final Fantasy VII: Advent Children, the half-brother of Final Fantasy: The Spirits Within. Its body movement and background textures aim for maximum realism, while its treatment of faces goes in the opposite direction: the textures look plastic, and the expressions are basically the kinds of expressions traditional anime characters are expected to have, complete with the anime label. Or take Appleseed, where the characters’ movements are mostly created through motion capture and are almost indistinguishable from reality, while the faces are basically simple line drawings.
Of course, a forbidden zone is just waiting for someone to cross it. Efforts to approach the uncanny valley have never been abandoned. The story goes that Cameron finally decided to make Avatar because the success of Gollum, the CG character in The Lord of the Rings, showed him that CG characters might be able to cross the uncanny valley. But I think Gollum’s design still avoided the valley deliberately, to some extent. First, the character was defined as a monster, neither human nor ghost, which psychologically removed the obstacle of treating Gollum as a human. Second, when CG simulates a real person, apart from skin texture, the hardest thing to simulate is facial expression (the cannon fodder mentioned above actually did an acceptable job with texture; what truly made people uncomfortable was the stiff expression). Gollum didn’t make much of a breakthrough here, because his character design called for exaggerated or extreme expressions. The reason viewers dislike CG characters’ expressions is that they are intimately familiar with the delicate muscle movements on their own—real human—faces. Exaggerated expressions are distorted and abnormal; compared with bland, ordinary expressions, they can slip past the audience’s fussy scrutiny more easily.

In my view, it was King Kong, which burst onto the scene in 2005, that made an effective attempt to approach the uncanny valley with a CG character. You’ll probably say, isn’t that a contradiction? Kong is even less of a “human.” Actually, it isn’t contradictory. Gollum is human, but was designed as a monster; Kong is a monster, but was designed as human. Kong’s design sidestepped the problem of simulating human skin, but as I said, the hardest part of crossing the uncanny valley is simulating facial expressions and mannerisms. And Kong’s expressions and mannerisms—especially those big, dewy, soulful eyes—are truly heartbreaking.

Last year (this article was written in early 2010), another CG masterpiece that tried to challenge the uncanny valley was born: The Curious Case of Benjamin Button. Before seeing the making-of footage, it’s hard to imagine that almost every scene featuring old Benjamin was made by combining a CG head (capturing Pitt’s expressions) with the body of a live-action stand-in. In still images, these half-real, half-fake people could absolutely pass for real; in motion, though, the film’s treatment of expression was still a little crafty. The stiffness and lag that are hardest to overcome in facial capture were very cleverly neutralized by the character design. Benjamin was already a stiff, slow-moving old man in his youth (it feels awkward even saying that), so even though the technology still couldn’t eliminate the stiffness of CG expressions, the character’s own stiffness cleverly concealed it. While watching the film, I didn’t find old Benjamin’s scenes awkward in the slightest. Instead, I thought, well, old people are supposed to be this adorably dazed. The underlying idea is the same as with Gollum, though the method is exactly the opposite. Still, that judgment is too harsh on Benjamin Button. After all, old Benjamin was vastly more advanced than Gollum had been back then.

And then, finally, our masterpiece takes the stage. While watching Avatar, I’m sure many viewers felt as I did: the most incredible thing about the film’s dynamic CG was the perfect rendering of the Na’vi’s facial expressions and mannerisms. (I hear Cameron specially invented a “facial-capture system,” filming the actors’ expressions with tiny cameras and digitizing them to achieve this nearly perfect result.) Their expressions and mannerisms were so natural and believable that I completely forgot the figures displaying every shade of joy, anger, sorrow, and delight in front of me were actually CG characters.
Put it this way: when I watched them, my illusion was that these were real actors in makeup performing!
Wait a minute—isn’t that strange? After all these years of progress in CG anthropomorphic technology, the final result is supposedly to achieve the effect of real actors performing in makeup? If that’s the ultimate goal, why go to so much trouble to take the long way around? Why not just find some actors and paint them blue? Is CG merely a misguided, self-abusive form of performance art?
Of course not. The greatness of CG still lies in creating images that real photography cannot achieve—or, more precisely, images that real photography cannot achieve cheaply. Don’t be fooled by how much CG costs; to achieve the same result, CG can be far cheaper than shooting on location. Take a historical film: normally we would build a film set town. In The Curious Case of Benjamin Button, by contrast, many scenes used real sets only for the small areas the characters actually touched, while the rest of the background was drawn in CG. Those sets could have been built, but the director chose CG anyway; it wasn’t a case of having money to burn. A more extreme example is a film like Avatar: Pandora’s natural world is entirely virtual CG. If it had to be shot on location, not even a whole nation’s resources could build a “Pandora film-set town.”

So here we have environmental backgrounds that are entirely CG. If we shot with real blue people, we would inevitably run into problems joining the live actors to the CG where they touch—at points and across surfaces. One side is real and the other fake; when they touch and interact, especially while making irregular movements together, coordinating them is extraordinarily difficult. You have to simulate the real contact between two physical materials (for example, how should the skin and muscles on a CG face react when a real finger pokes it?) while also coordinating both movement paths with absolute precision. In my view, this is another notoriously difficult problem in CG technology, and for now a perfect result remains hard to achieve. In scenes like these, we often seem to see live-action characters moving with the unnatural, stop-motion-like quality of animation—for example, the scene in Spider-Man 2 where Doctor Octopus uses CG tentacles to kidnap Aunt May. The joining of the real and the virtual may have reached a major breakthrough in The Curious Case of Benjamin Button: in many cases, you can detect a little unnaturalness between the CG head and the real body only after someone reveals how it was made. But Benjamin Button still couldn’t go further—for example, old Benjamin rarely makes large, full-body movements relative to the camera.
Since perfect interaction between CG and live actors is so difficult, turning the characters into CG as well is actually a good way to get around the problem while saving time and effort. The barrier between character and background disappears, and interaction becomes much simpler. If we shot with real blue people, how could we get such natural scenes of taming horses and riding dragons? (Think of the bear-riding scene in The Golden Compass.)
Since Avatar had no choice but to use CG characters, and since it raised the rendering of their facial expressions to an almost perfect level, has it successfully crossed the uncanny valley? I think it has passed through the bottom of the valley and is climbing up the cliff on the side of perfect simulation—but saying it has crossed all the way would still be going too far. Because we’ve returned to the other side of CG character simulation: skin texture.
I think simulating materials is, to a large extent, a case of the simpler the material, the harder it is to reproduce. This is really a matter of psychological processing: the more obvious an object’s distinguishing features are, the fuzzier the mental image we need to store in our brains to process and recognize it; the less distinctive the features, the more detailed that stored image has to be. Say you get robbed by a scar-faced man—you can identify him without much trouble. But if you get robbed by someone with an utterly ordinary face, you have to picture what his eyes, eyebrows, and nose looked like before you can recognize him. I wonder if this is why The Curious Case of Benjamin Button stopped using CG to create Benjamin’s head once he became a child. An old man’s heavily wrinkled skin is an object with very obvious features. Our perceptual habit when it comes to things like an “old man’s face” is that we don’t need to form a highly detailed mental image; we only need to pick out obvious features such as the wrinkles, without paying much attention to the other details. That makes simulating the material in CG much easier. With the smooth, taut skin of a young person, though, our perceptual habit is to demand a great deal of detail: tiny shifts in the skin’s color and texture all fall within the viewer’s field of attention. It’s easy to imagine how sharply the difficulty increases when CG has to render this kind of material. When I watched the final stretch of The Curious Case of Benjamin Button, I couldn’t get emotionally absorbed at all. The main reason was that teenage and toddler Benjamin were no longer created with CG, so the film lost its spirit right at the emotional climax where it should have hit hardest.
On this particular challenge, Avatar cleverly sidestepped the problem—like King Kong, the Na’vi aren’t “human,” after all. Cameron may have boldly claimed that audiences wouldn’t be able to tell whether these CG characters were real when standing alongside real people, but I suspect that if the CG characters had been Earth humans rather than Na’vi, telling real from fake wouldn’t have been difficult. In many close-up shots, the Na’vi’s skin texture is still rendered perfectly: the pores, patterns, and spots look astonishingly lifelike. But in certain shots, a bit of a “plastic” quality is still unavoidable. I’m sure that feeling would be even more obvious if the CG characters were Earth humans. Avatar’s clever move was to give the audience the preconception that the Na’vi are “aliens,” so its weaknesses no longer really count as weaknesses.

Anyway, I still think Avatar is the only CG work I’ve seen so far that has almost crossed the uncanny valley. That’s a major step forward. But it isn’t the reason Avatar became a masterpiece in my mind. I’ll save that reason for another post.
Leave a Reply