Definition of deepfakes

Basic terminology of audio and video deepfakes

Deepfakes are synthetic or manipulated media produced with AI models. We detect them in two places: in the speech of a recording, and in the picture of a video.

🔉 Audio deepfakes

Audio deepfakes are synthetic speech samples generated using AI models to mimic human voices. They are increasingly used for fraud, impersonation, and misinformation, making robust detection essential.

Key challenges:

  • High realism: Modern models (e.g., ElevenLabs, PlayHT) produce near-human speech.
  • Cross-lingual threats: Fake audio spans multiple languages.
  • Speaker variability: Need detection independent of specific speaker profiles.

🗣️ Types of Deepfake Generators & Attacks

Deepfake audio can be created using two main approaches:

  • Voice Cloning Generators: Replicate a specific person’s voice using a short audio sample as reference. Commonly used by commercial and open-source tools to mimic real individuals.
  • Text-to-Speech (TTS) Generators: Convert written text into speech with synthetic voices that are generic or prebuilt.

Open-Source Generators

These are models that are open-sourced to the community. Examples include:

  • F5-TTS: Multilingual zero-shot voice cloning (English, Chinese, Russian, Arabic)
  • xTTS-v2: Coqui's multilingual zero-shot voice cloning

Commercial API Services

These are some commercial models that may be misused by bad actors, for cloning voices:

  • ElevenLabs: High-quality commercial voice cloning
  • PlayHT: Enterprise-grade TTS services

🎥 Video deepfakes

Video deepfakes are recordings in which a person's face, or the whole scene, has been generated or altered by AI. They are used to defeat remote identity checks, to impersonate executives on video calls, and to put words in the mouth of a public figure.

Key challenges:

  • Compression: re-encoding for social platforms strips out the fine detail a detector relies on.
  • Partial manipulation: only the face region may be fake, and only for part of the clip.
  • Framing: small faces, motion blur and heavy filters all reduce the signal available.
  • Pace of change: video generators improve faster than any fixed set of artifacts can track.

🎭 Types of Video Manipulation

  • Face swap: the face of one person replaces another's, frame by frame, while the body and setting stay real.
  • Reenactment and lip sync: a real face is driven by new audio or by another performance, so the person appears to say something they never said.
  • Fully generated video: no camera was involved. The person, or the entire scene, comes from a prompt or from an avatar template.

🎬 Types of Video Generators

Open-Source Generators

These are models and tools open-sourced to the community. Examples include:

  • DeepFaceLab: the long-standing face-swap toolkit behind much of the public deepfake material
  • SimSwap: single-image face swapping
  • Wav2Lip: drives the lips of an existing video from an audio track

Commercial API Services

These are commercial services for synthetic video and avatars that may be misused by bad actors:

  • HeyGen, Synthesia, D-ID: avatar and talking-head video generation from a script
  • Text-to-video models such as Sora, Veo and Kling, which generate a scene, and the people in it, from a prompt

🌐 Online Examples

Below are real-world demonstrations of deepfakes. While the exact tools or models used to create these samples are not disclosed, they illustrate the realism and potential risks of the technology.

Anderson Cooper, 4K Original/(Deep)Fake Example

Deepfake Example. Original/Deepfake Elon Musk

🧑‍💻 Our Technology

Our detectors follow the same method whether the input is speech or picture. We train models on a mix of authentic (bonafide) and synthetic (spoofed) material, so the system learns the general patterns that separate real from fake rather than the fingerprints of any one generator. That is what lets it hold up against generators it has never seen.

Detection runs in stages. The input is cut into short segments, utterances for audio and fixed-length windows for video. Key features are extracted from each segment, and the model scores every segment as bonafide or spoofed with a probability, so you see where in a file the manipulation sits rather than a single verdict for the whole file.

For speech, the features are behavioral: intonation, emotional tone and hesitation patterns, which synthetic voices tend to flatten or distort. The system is speaker-agnostic and does not rely on voice-prints or spectral artifacts alone. For video, the same principle applies to the picture, and the speech track goes through the speech detector as well. The two verdicts are independent, which is what catches the mixed cases: a genuine recording with a cloned voice dubbed over it, and a face swap that carries the real person's audio.

You can learn more on how to try our demo on your own, and see Submit file to send a file for analysis.


Did this page help you?