Bad Lip Reading, although meant for the mere comical amusement of its listeners, can bring perspective to questions surrounding sound and voice. In these dubbed clips of culturally recognizable events/scenes, we are presented an outlandish rendition which is far what is known to be the true recording. However, as we listen into the recording for long enough, we find that we cannot imagine the characters to be saying anything else. Especially if the video is our first exposure, the combination of visual imagery with this audio lock in this version as somewhat viable. This is due to the fact that linguistically, much of the oral movement we make are similar across varied speech – and therefore, can be easily fabricated.
Through our discussions surrounding Aristotle’s thoughts on sound and voice, we have come to define voice as a utterance, or vibration of the body in which the soul is the catalyst. In this way, the sounds we create in communication not only meaning, but the will, emotion, and intentions of the speaker. The introduction of technology that has made Bad Lip Reading and Lyrebird make us to hypothesize if the “grain” of the voice, as Cavarero defines it, could ever be synthetically manufactured. Could the fervor at which Martin Luther King Jr. delivers a speech or “spunk” in Michael Jackson’s viral songs be replicated? Could “live” footage be dubbed without an audience’s knowledge? Would we welcome such in our society? Our current technological advancements in the audio-visual demand that we look at these questions proactively as a society and draw resultant boundaries
The Uncanny in Bad Lip Reading
Bad Lip Reading Clips, while comedic, also makes one feel uneasy. The voice in the video looks like it is coming from the individual when it is understood that the video had been dubbed to make it seem like they are saying comical things. Voice is something personal to the individual, whether one believes in Barthe’s idea of voice coming from “deep down in the cavities, the muscles, the membranes, the cartilage” of the individual— also known as the grain—or one believes in Aristotle’s idea that voice comes the breath hitting “the soul resident” of the body (Barthes 505, Aristotle 670). It becomes unsettling to hear a computerized robot voice coming from a human’s mouth. The robotic sounding voice leads us to fall into the uncanny valley. The uncanny valley occurs when a humanoid object, which appears almost truly human or strangely familiar to a human, causes observers to feel strange and repulsed. The voice of a computer gives one a very different feeling then when “someone in flesh and bone who emits it” (Cavarero 522). The voice of human is strong but limited. Society has trained people to think, act, and speak in certain ways that are suitable for certain situations, making the bad lip reading live footage clips feel more uncomfortable. However, the cinematic lip reading to Star Wars, including nonhumans and music, makes is easier to watch and trust since one is not listening for the flaws; the setting, often used comedy allows for it because of viewer expectations. While the political stage adds comedic value, it also makes the sound untrustworthy. With sight being very linear and direct, sound is more nebulous; so while the video is easy to trust the words are not, making the viewer second guess the presentation. Overall, while the videos are entertaining, the uncanny value and setting makes the viewer feel uncomfortable.
The Essence of Our Voice
The Bad Lip Reading clips are quite amazing if one realizes the amount of work that must have been put into them. First, the automated voices sound exactly like the actual voices of the people in the clips, especially for the voices of Trump and Yoda. Second, though what is being said is absurd, the lip movement is scarily accurate. In both the live footage and the cinematic footage, the lip reading and automated voice looked valid and there was not much of a difference between the footages. That was the unsettling part; the videos were hilarious, but also showed insight on fake news. While Bad Lip Reading clips are used for comedic purposes not all clips using automated voices will be and some might actually be able to influence politics, especially with the current rise of fake news in social media.
This relates to what we have spent time discussing in class about the uniqueness of someone’s voice. Cavarero discusses this idea in her essay “Multiple Voices”, in which she makes the claim that the essence of someone’s voice is not in what they say but how it sounds. For instance, “… The human condition of uniqueness resounds in the register of the voice” (Cavarero525). Cavarero questions the idea of the voice being an intellectual concept for human uniqueness and, instead, supports the sound of voice as being what makes a human unique. In relation to the Bad Lip Reading clips, Cavarero’s claim is put into question. Now that the ‘uniqueness’ of any voice can be made into an automated copy, is the sound of our voice really unique? These Bad Lip Reading clips suggest that they aren’t as special as Cavarero claims.
Sight vs Sound
As a person who can’t read lips, I’ve always been unsettled with bad lip reading videos, and it’s because as I’m listening to the audio and watching the speaker’s lips, I can’t confidently say that the voice over is fake. It calls into question how reliable our senses are, particularly sight and hearing, and the relationship between those senses.
We’ve always been told that “seeing is believing,” and for some people, that’s led to a confidence in our sense of sight. In fact, if someone were to ask me whether I felt my sight or hearing was more “accurate,” I’d tell them that my sight was. Our ability to see things provides us with something that feels more concrete, something more tangible than soundwaves. But bad lip reading videos take away this sense of concreteness and accuracy. Suddenly what we are “seeing” as a person speaks is not what they are actually saying. The only “truth” is in the fake audio that is playing, and we start to “see” what we are physically hearing, or at least that is true in my case. The more I concentrate on the person’s lips, the more I can convince myself that what I’m hearing is what’s being said. And what makes the experience even more strange is that now we have this voice from a unknown source being put over these clips. The voice overs take listeners out of the context of the clip. Although after talking to Andrew, it seems like the ways this is accomplished is different for live videos versus cinematic footage. In the Star Wars bad lip reading, it was probably much easier to find scenes or clips that fit the song, especially because most of the video consisted of Yoda’s puppet speaking. As for the presidential debate, the person behind these videos most likely had to see what word he/she could fit into Trump’s or Clinton’s mouth.
In a way, bad lip reading makes me think of lyrebird and other forms of artificial voice. While the person behind the dubbing is a real human with a real voice, it is not the voice of the original speaker; it’s taking away what Aristotle saw as one of the distinguishing features we have as humans: our voice. This distortion brings in the long relationship that sound and sight have developed over the years of sound studies.
BLR and the displacement of voice
In all forms of writing, from academic to artistic, there is a large emphasis placed on the strength of the writer’s “voice,” and that the strength of this voice generally correlates to the quality and power of the writing. From a writer’s voice we interpolate their personality, who they are, what their message is, and what makes them unique, and, in general, we assume the voice and the author are connected: if I’m watching a clip of Bernie Sanders give a speech, I believe that it is Bernie Sanders’ voice coming through my speakers. To frame this in terms of Aristotle’s work, we assume that the “impact of the inbreathed air against the windpipe” is originating from the windpipe of the individual we see is talking. The “Bad Lip Reading” clips play off of that idea by robbing and distorting an individual’s voice while still satisfying the visual component of listening – we can match the words of the voiceover with the movement of the subject’s mouth, their facial expressions, and their gestures. Particularly with the presidential debate editions of BLR, it becomes extremely satisfying and funny to see individuals “say” things that we know they would never say. However, it is not hard to envision this robbing and replacing of the voice (or, the displacement of the origin of the voice from one windpipe to a different windpipe) being used in ways that are disconcerting and dangerous. We are already approaching a point in our society of truth not being truth. What will happen when voice is no longer voice — In other words, when an individual voice can be put in anyone’s mouth?