Video summary

КАК РАСПОЗНАТЬ ИИ В ПЕСНЕ? ДИПФЕЙКИ, ОСНОВНЫЕ МАРКЕРЫ И ЧТО БУДЕТ ДАЛЬШЕ?

Main summary

Key takeaways

Technology

Overview

The video focuses on how to recognize music that may be generated—or heavily processed—by neural networks / deepfakes. The speaker provides “listening markers” (audio cues) and then applies them to critique specific popular tracks on Russian streaming charts, especially Yandex Music.


Platform / ecosystem claims

  • The speaker cites Yandex Music statistics:
    • ~100,000 artists
    • ~140,000 tracks released monthly (claimed to be ~40% of all releases)
  • He claims every 10th track on the Yandex Music chart is neural-network-generated.
  • He says Yandex Music does not plan to label AI-generated tracks.
  • He also claims artists who try to expose AI tracks are met with harassment, including:
    • spam attacks
    • account hacking
  • He frames himself as helping users identify “bad recommendations.”

Listening guide: markers to detect AI tracks

In the first major block, the speaker proposes a practical method: if a recommended song “sounds like crap,” listen for the following audio cues.

  1. Unnatural breaths / inhalations

    • Loud, separate “accented” breaths that don’t connect naturally to the subsequent vocal phrasing.
  2. “Sand/metallic” artifacts in vocals and music

    • A gritty, sandy texture; metallic “taste”; robotic or compressed timbre.
  3. Compression and mismatched voice quality

    • Examples described include:
      • male/female mismatch
      • overly compressed sound
      • vocals sounding like a not-fully-finished “trans” / “therapy-like” voice
  4. Hissing / wheezing / crackling consonants

    • Audible hiss or whistling (or crackle-like artifacts), especially on S/TS/SH/SHCH sounds.
  5. Missing human imperfections

    • Lack of micro-sighs, minor vocal “tears,” natural intonation drift, and live performance imperfections.
    • Vocals described as “mathematically smooth” and “plasticine.”
  6. Stress / articulation errors

    • Incorrect placement of word stress.
    • Mispronunciations.
  7. Technical vocal moments handled poorly

    • Neural nets may struggle with extreme techniques (e.g., vocal splitting / overly processed rasp), leading to oversuppressed or artifact-ridden passages.
  8. Frequency artifacts / spectral “dirt”

    • He claims that when analyzing spectrally, above ~18 kHz there may be no clean content—only dirt/robotic flaws.

Review / analysis of specific chart tracks (examples)

The speaker repeatedly applies these markers to tracks he believes are AI-related.

#1 on Yandex Music: “Sade” (Mod, HCHO, Bandiya)

  • Described as low-quality AI output:
    • dirty choir (“egg chorus”)
    • volume imbalance
    • harsh artifacts
    • wheezing / crackling
  • He claims the backing suppression process (“minus-track” handling) can leave frequency artifacts.
  • He interprets the lyrics as implying a “not-you / voices-in-head” narrative (framed as narrative horror).
  • He references controversies in comments and alleged spam attacks around the track.

#3 on Yandex Music: “Malbora” (Sayan)

  • He credits the artist with not denying neural-network use:
    • singing recorded by the artist
    • followed by neural-network voice transformation/deepfake to create a “girl” chorus
  • Still, he notes the recurring “sand” in transformed vocals:
    • especially around drawn-out letters (e.g., “A”) and high notes
  • He argues consonants (s/ts/sha/shch) reveal AI limitations due to high-frequency artifacts.

Other referenced viral / AI-associated cases

  • Mentions neural-network covers and “stadium/cover” formats (including examples influenced by “D(z)heg-like ‘Mama’” and a Kanye West-type viral cover).
  • Mentions “Tell me, Snow Maiden” (Sasha Kamovich) and other tracks as having AI markers: sand, hissing, compressed smoothness.
  • Mentions tracks attributed to additional AI-leaning performers/producers, including:
    • an SDP + Harmonica track
    • “Russian Beauty” by Hanna, described as using “GPT-like” prompt structure themes and containing rhythmic / arrhythmic artifacts

“Celebrities using AI” discussion (second analytical block)

The speaker examines mainstream artists and argues they also use AI, often without full transparency.

Philip Kirkorov – “Pink Wine” (cover)

  • Claims “cyborg vocals” / deepfake identification based on:
    • mismatch with the artist’s usual articulation and stress patterns
  • Notes absence of the “smacking/excess articulation errors” he associates with older vocal performance, suggesting the voice was swapped.

Hanna – “Russian Beauty”

  • Claims the song follows a “prompt for Russian beauty” checklist:
    • sarafan, braids, lace, candles, snow, flame, fairy-tale patterns, berries
  • Cites issues such as:
    • synth sound choice resembling popular AI tools
    • rhythmic breaks
    • odd metaphor lines

Shaman – “Ros’ Mama”

  • Discusses controversy involving AI-generated animated agents/portraits and claims the choir and vocalizations may also be AI.
  • Notes structural/dynamic behavior:
    • verse-to-chorus dynamics not increasing as expected
    • sounds “overcompressed/smooth,” implying AI-tuned production
  • Mentions testing a platform called “Erosenet/Narosset” (spelling uncertain) and getting AI output that “sounds like Shaman,” implying prompt-to-style generation.

Overall position: labeling, ethics, and future impact

  • Core ethical/market claim: AI usage should be disclosed/labelled.
    • Without labeling, recommendations and music feeds can’t be meaningfully controlled, and users can’t reliably filter out AI creativity.
  • He contrasts AI “prompt-writing” with human artistic performance and implies AI lacks human emotion.
  • He predicts:
    • the industry will initially remain partly dominated by “AI slop”
    • once audiences learn to ignore it, real human performance may preserve a niche
  • He also suggests an analogy to other creative/design markets:
    • AI lowers the barrier to entry
    • non-adopters may become less competitive for some clients

Proposed positive / neutral uses (not purely critical)

The speaker allows neural networks for:

  • demos
  • inspiration
  • turning sketches into rough song structures
  • using AI as a writing/arrangement assistant rather than deceptive final authorship

He also cites a Grimes idea (as quoted): more low-quality content may make truly unique real art more valuable.


Main speakers / sources

  • Speaker/author: Daniil Zhelyanin — YouTube channel “Propet”
  • Main external source discussed:
    • Yandex Music (chart statistics and claims about lack of labeling)
  • Specific Russian chart tracks/artists mentioned throughout:
    • Mod / HCHO / Bandiya (“Sade”)
    • Sayan (“Malbora”)
    • Sasha Kamovich (“Tell me, Snow Maiden”)
    • Hanna (“Russian Beauty”)
    • Philip Kirkorov (“Pink Wine”)
    • Shaman (“Ros’ Mama”)
    • SDP (plus an unspecified Harmonica track)
    • plus other referenced producers/performers and cover-format examples

Original video