Hardware FixRecommendedDevice not working? Your driver may be the problemCheck updates for common hardware issues.Fix DriversOctober DealsAmazon USOctober deal check: compare before you payAmazon US: current deals, useful picks and tech finds.Check DealsSlow PC?RecommendedPC slow today? Run a repair scan before it gets worseResolve common Windows issues and optimize system performance.Scan Now×
Skip to content
DocumentaryTube
AI history

AI Created a Convincing Fake Obama Video in 2017—How It Worked

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

“AI Creates Fake Obama” refers to a July 2017 research demonstration, not a real Obama statement. Researchers at the University of Washington generated a video in which Barack Obama appears to deliver audio that was not recorded with the footage. Their system learned how his speech sounds corresponded to visible mouth movements, synthesized a matching mouth, and composited it into existing video.

The project, Synthesizing Obama: Learning Lip Sync from Audio, was an early, tightly constrained example of audio-to-video synthesis. It showed why video could no longer be treated as automatic proof that a public figure had actually said the words heard in a clip.

What “AI Creates Fake Obama” actually describes

The phrase is the headline of an IEEE Spectrum article published on July 12, 2017: “AI Creates Fake Obama.” The underlying work was presented at SIGGRAPH 2017 by Supasorn Suwajanakorn, Steven M. Seitz and Ira Kemelmacher-Shlizerman of the University of Washington’s Graphics and Imaging Laboratory. The team’s official project page is available here, with the paper at this PDF.

It was not a new recording of Obama, a digital clone created from nothing, or evidence that he delivered the newly paired speech. It was an altered video: authentic Obama footage served as the visual base, while the mouth region was generated to match a different audio track.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
What is authentic What is synthesized What the result does not establish
Obama’s body, face, camera view and original performance in the source footage Mouth shapes and mouth texture timed to supplied speech audio That Obama actually uttered the words in the new audio

How the audio-to-video system worked

The central contribution was audio-driven lip synthesis, not a conventional face swap. In simplified form, the pipeline was:

  1. Collect training footage. The researchers assembled approximately 17 hours of Obama’s weekly-address footage—nearly two million frames spanning eight years, according to the paper.
  2. Learn speech-to-motion relationships. A recurrent neural network learned how audio features corresponded to the shapes and positions of Obama’s mouth while he spoke.
  3. Predict a mouth for new audio. Given an audio track, the model generated the mouth appearance that would plausibly accompany each moment of speech.
  4. Match the target video. The synthesized region was adjusted for the target frame’s facial pose, geometry and timing.
  5. Composite the frame. The generated mouth was blended into the original face so the result retained the surrounding skin, lighting and expression from the source video.

The result could look photorealistic in the mouth area because the system was not trying to invent an unconstrained person. It was reusing a carefully selected archive of one person’s face and visible speech.

Why Obama was a practical research subject

Obama’s public addresses offered an unusually useful dataset. The archive contained many high-definition recordings, often with his face large and near the center of the frame and with relatively controlled framing. Weekly addresses also provided repeated examples of speech, facial motion and changing expressions over multiple years. The footage was publicly available for academic research and publication.

Published summaries sometimes cite about 14 hours of material, but the primary paper reports approximately 17 hours and nearly two million frames. The larger figure is the one to use when describing the researchers’ dataset.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

What the demonstrations showed

The project page lists demonstrations using weekly-address footage, unrelated Obama recordings, speech from Steve Harvey, a 60 Minutes interview, audio from The View, Obama speaking from an earlier period, an impressionist’s audio and a speech-summarization example. These pairings mattered because they showed the system was not simply replaying the original soundtrack.

They did not show that the system could make Obama perform any sentence under any conditions. Quality depended on the training archive, the visibility and pose of the face, lighting, framing and how compatible the supplied audio was with the learned visual speech patterns.

What the system did—and did not—make Obama say

The audio could be genuine speech by someone else, while the accompanying image was a synthetic performance. The model generated visible mouth movements that matched the sound; it did not uncover a hidden recording of Obama saying those words.

That distinction is essential when evaluating a clip. An authentic voice recording and an authentic video performance are separate claims. The first may be real even when the second has been altered, and neither claim can be proved merely by watching a convincing mouth movement.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Intended uses beyond deception

The researchers described applications that were not inherently fraudulent:

  • Lower-bandwidth communication: transmit audio and reconstruct a visual representation instead of sending full video.
  • Videoconferencing recovery: synthesize a talking face when a camera feed is frozen, degraded or low resolution.
  • Virtual and augmented reality: create digital humans whose visible speech follows live audio.
  • Entertainment and visual effects: produce controlled mouth animation for media production.
  • Accessibility: explore video synthesis that could help infer visible speech for lip-reading from telephone audio.

The same capability becomes harmful when a fabricated performance is presented as documentary evidence, news footage or a direct statement.

Why the demonstration raised a misinformation alarm

A synthetic clip could make a public figure appear to announce a policy, confess to misconduct or endorse a false claim. Possible consequences include fabricated political statements, fraudulent testimony, reputational damage and manipulation of public opinion. The 2017 demonstration was a research prototype, not evidence of a documented attack; its warning was that existing public footage could be repurposed into misleading audiovisual “evidence.” The AI Incident Database entry records that broader concern.

The milestone also changed the evidentiary status of video. A clip can remain visually persuasive while no longer being a reliable record of what happened at the time and place implied by its soundtrack.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Limitations of the original system

“Photorealistic” described the quality of the demonstration, not perfection. The paper and contemporary reporting identify several failure modes:

  • Head turns and unfavorable poses could expose errors in the facial model.
  • The generated mouth could spill beyond the face boundary into the background.
  • The system modeled the mouth more strongly than the full range of facial emotion.
  • Expression could conflict with the emotional tone of the supplied audio.
  • Results depended on favorable source footage, visible facial features and compatible lighting and framing.
  • The work focused on mouth-region synthesis rather than unconstrained generation of an entire person.

Can you spot one of these fakes?

IEEE Spectrum reported that the researchers noticed possible softness or blur around the mouth and teeth compared with the rest of the frame. That was an interesting research-era clue, not a universal detector. Compression, focus, motion blur and ordinary editing can produce the same appearance, while newer systems may leave different artifacts.

Visual inspection can suggest a question, but it cannot authenticate a high-stakes video. A clip that looks clean is not thereby genuine, and a soft mouth is not by itself proof of manipulation.

A practical verification checklist

  1. Find the earliest known upload and preserve the complete file if possible.
  2. Search for the full, unedited source recording from an official archive or broadcaster.
  3. Compare the soundtrack with an official transcript and independent recordings of the event.
  4. Check whether reputable news organizations or subject-matter fact-checkers authenticated the clip.
  5. Inspect mouth boundaries, teeth, lighting, reflections and audio continuity only as clues.
  6. Prefer provenance, corroboration and original files over an intuitive judgment based on a few frames.

Why this 2017 project still matters

The Obama demonstration was an influential early audio-to-video, deepfake-style milestone because it turned a large public archive into a reusable model of visible speech. It was narrower than later synthetic-media systems: the face, pose and source footage imposed real constraints. Even so, it proved a crucial point in a form viewers could immediately understand—an existing video of a famous person could be made to perform a different soundtrack convincingly enough to challenge casual observation.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

When encountering the phrase “AI Creates Fake Obama,” the accurate reading is therefore historical: a University of Washington SIGGRAPH 2017 experiment in lip-synced synthetic video, not a current Obama statement and not proof that Obama spoke the generated words.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Leave a Reply

Your email address will not be published. Required fields are marked *

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Read next

Recommended PC Tool
Recommended PC Tool
Windows Errors? Fix Them Before They SpreadFree repair scan
Crashes, No Sound, or Screen Glitches?Free driver scan

Two free Windows tools

One Free Minute Could Fix That PC

Before you go - each of these free tools takes about a minute and tackles what quietly slows a Windows PC down.

Special offer. View Outbyte info, uninstall instructions, EULA, and Privacy Policy.