As a person who can’t read lips, I’ve always been unsettled with bad lip reading videos, and it’s because as I’m listening to the audio and watching the speaker’s lips, I can’t confidently say that the voice over is fake. It calls into question how reliable our senses are, particularly sight and hearing, and the relationship between those senses.
We’ve always been told that “seeing is believing,” and for some people, that’s led to a confidence in our sense of sight. In fact, if someone were to ask me whether I felt my sight or hearing was more “accurate,” I’d tell them that my sight was. Our ability to see things provides us with something that feels more concrete, something more tangible than soundwaves. But bad lip reading videos take away this sense of concreteness and accuracy. Suddenly what we are “seeing” as a person speaks is not what they are actually saying. The only “truth” is in the fake audio that is playing, and we start to “see” what we are physically hearing, or at least that is true in my case. The more I concentrate on the person’s lips, the more I can convince myself that what I’m hearing is what’s being said. And what makes the experience even more strange is that now we have this voice from a unknown source being put over these clips. The voice overs take listeners out of the context of the clip. Although after talking to Andrew, it seems like the ways this is accomplished is different for live videos versus cinematic footage. In the Star Wars bad lip reading, it was probably much easier to find scenes or clips that fit the song, especially because most of the video consisted of Yoda’s puppet speaking. As for the presidential debate, the person behind these videos most likely had to see what word he/she could fit into Trump’s or Clinton’s mouth.
In a way, bad lip reading makes me think of lyrebird and other forms of artificial voice. While the person behind the dubbing is a real human with a real voice, it is not the voice of the original speaker; it’s taking away what Aristotle saw as one of the distinguishing features we have as humans: our voice. This distortion brings in the long relationship that sound and sight have developed over the years of sound studies.