AT-ADD: All-Type Audio Deepfake Detection Challenge Evaluation Plan
A new benchmark challenge aims to push audio deepfake detection beyond speech to music, sound effects, and singing — exposing how narrow today's defenses really are.

The Thesis
Most audio deepfake detectors today are trained almost entirely on synthetic speech, which means they are blind to faked music, sound effects, and singing voices. The AT-ADD Grand Challenge, proposed for ACM Multimedia 2026, sets up a formal competition to fix that by testing detectors against a much wider range of audio types and real-world distortions. This matters because Audio Large Language Models — AI systems capable of generating convincing non-speech audio, not just voice clones — have made synthetic soundscapes, fake musical performances, and fabricated ambient audio easy to produce at scale. The commercial and legal stakes are high: media verification, insurance fraud, evidence authentication, and platform content moderation all depend on tools that don't yet exist in robust form. The catch is that this paper is an evaluation plan, not a solved problem — it describes the competition framework, not a detector that already works.
Catalyst
Audio Large Language Models capable of synthesizing music, sound effects, and singing (not just speech) have matured rapidly in the past two years, with models like AudioCraft, Stable Audio, and Suno reaching consumer-grade quality. Existing detection benchmarks were designed for earlier, speech-only synthesis systems and have not kept pace with this generative diversity. The ACM Multimedia 2026 timeline reflects the research community recognizing this gap has become practically urgent.
What's New
Earlier detection benchmarks — such as ASVspoof and ADD Challenge — focused almost exclusively on synthetic or replayed speech, evaluating systems on whether a voice clip was human or machine-generated. Those systems exploit artifacts specific to speech synthesis pipelines (vocoders, text-to-speech models) and perform poorly when confronted with music or ambient sound generation. AT-ADD proposes two new tracks: one that stress-tests speech detection against real-world degradation and newer generation methods, and a second that demands type-agnostic detection across speech, sound, singing, and music simultaneously — a combination no prior public benchmark has formally evaluated.
The Counter
This paper is a competition proposal, not a result. There is no detector, no dataset released yet, and no baseline performance numbers to evaluate. The history of ML challenges is littered with benchmarks that attracted academic participation but produced systems that never transferred to deployment — ASVspoof itself has been running for over a decade without producing commercially robust speech deepfake detectors. The framing assumes that a unified model can generalize across speech, music, sound, and singing, but these audio modalities have fundamentally different statistical structures; a detector trained to spot artifacts in a vocoder may have nothing in common with one that catches seams in a generative music model. The challenge is also reactive by design — it evaluates against 'unseen' generation methods, but the generation side is moving faster than the detection side has historically managed to follow. Finally, the paper acknowledges that existing countermeasures lack robustness to real-world distortions, which is a known open problem — proposing a benchmark does not solve it.
Longs
- ZETA — digital media authentication and communications security
- IRDM (iRhythm, proxy for AI signal processing in sensor domains)
- AKAM (Akamai) — CDN and media delivery platforms with deepfake filtering interest
- FTNT (Fortinet) — security vendors building audio/media integrity into enterprise tools
- BOTZ (Global Robotics & AI ETF) — broad AI application exposure
Shorts
- ASVspoof benchmark maintainers and associated detection vendors — their speech-only framing becomes visibly insufficient if AT-ADD demonstrates broad generalization failures
- Platform trust-and-safety teams at YouTube, Spotify, and SoundCloud — currently lack validated tools for non-speech audio authentication, and AT-ADD highlights this gap publicly
- Voice biometric vendors (Nuance, now Microsoft) — if audio deepfakes extend reliably beyond voice to all audio types, the threat surface of their authentication products widens significantly
Enablers (Picks & Shovels)
- Hugging Face Hub — hosting and distributing baseline models and datasets for the challenge
- AudioCraft (Meta, open source) — one of the generative systems whose outputs the challenge must detect
- Suno and Udio (private) — commercial music generation platforms creating real-world test cases
- ACM Multimedia conference infrastructure — provides the formal evaluation and reproducibility scaffolding
- VoxCeleb and AISHELL datasets — foundational speech corpora that inform the speech track baselines
Private Watchlist
- Resemble AI — synthetic voice detection and watermarking tools
- Pindrop Security — voice fraud detection for call centers, directly threatened by audio deepfakes
- AI Foundation — synthetic media authentication research
- Sievert Larson Cyber — audio forensics consultancy space
Resources
The Paper
The rapid advancement of Audio Large Language Models (ALLMs) has enabled cost-effective, high-fidelity generation and manipulation of both speech and non-speech audio, including sound effects, singing voices, and music. While these capabilities foster creativity and content production, they also introduce significant security and trust challenges, as realistic audio deepfakes can now be generated and disseminated at scale. Existing audio deepfake detection (ADD) countermeasures (CMs) and benchmarks, however, remain largely speech-centric, often relying on speech-specific artifacts and exhibiting limited robustness to real-world distortions, as well as restricted generalization to heterogeneous audio types and emerging spoofing techniques. To address these gaps, we propose the All-Type Audio Deepfake Detection (AT-ADD) Grand Challenge for ACM Multimedia 2026, designed to bridge controlled academic evaluation with practical multimedia forensics. AT-ADD comprises two tracks: (1) Robust Speech Deepfake Detection, which evaluates detectors under real-world scenarios and against unseen, state-of-the-art speech generation methods; and (2) All-Type Audio Deepfake Detection, which extends detection beyond speech to diverse, unknown audio types and promotes type-agnostic generalization across speech, sound, singing, and music. By providing standardized datasets, rigorous evaluation protocols, and reproducible baselines, AT-ADD aims to accelerate the development of robust and generalizable audio forensic technologies, supporting secure communication, reliable media verification, and responsible governance in an era of pervasive synthetic audio.