Google DeepMind Lyria
Lyria (deepmind.google/models/lyria/) is Google DeepMind’s flagship family of generative music and audio foundation models. Built using latent diffusion over temporal audio latents and integrated into the Gemini API (Interactions API), Vertex AI, and Google AI Studio, the Lyria model suite—spanning Lyria 3 Clip, Lyria 3 Pro, Lyria 3.5, and Lyria RealTime—generates full-fidelity stereo tracks with vocals, timed lyrics, and section arrangements from text or multimodal image inputs, all protected with imperceptible SynthID audio watermarking.
Early AI music synthesis often suffered from structural incoherence, muffled vocal artifacts, and limited track lengths. Lyria provides a musically grounded audio generation architecture capable of modeling complex musical phrasing (verses, choruses, bridges), multi-instrument arrangements, and real-time interactive performance parameters.
Model Lineup & Architectural Profile
| Model Variant | Duration / Output | Audio Quality | Core Modality & Input | Primary Use Case |
Lyria 3 Clip (lyria-3-clip-preview) |
30 Seconds (MP3) | 44.1 kHz High-Fidelity Stereo | Text prompts or up to 10 reference images | Hooks, intro loops, ads, social clips |
Lyria 3 Pro (lyria-3-pro-preview) |
Up to 184 Seconds / 3 Min (MP3/WAV) | 44.1 kHz High-Fidelity Stereo | Multimodal inputs, BPM, key, section tags | Full songs, video soundtrack beds, orchestral scores |
| Lyria 3.5 | Variable length (up to 3 min) | High-Fidelity Stereo | Advanced text/lyric conditioning | Nuanced vocal performance & musicality |
| Lyria RealTime | Continuous Live Streaming | 48 kHz Professional Stereo | Interactive parameter sliders (BPM, mood, genre) | Live DJ performance, MusicFX DJ, adaptive game audio |
Core Breakthroughs & Capabilities
-
Multimodal Conditioning (Text & Vision): Developers and creators can condition Lyria 3 on up to 10 visual images alongside text descriptions, enabling the model to score video soundtracks, mood boards, or artwork dynamically.
-
Structured Composition & Lyric Synchronization: Supports explicit song structure control using standard musical tags (
[Verse],[Chorus],[Bridge]) alongside BPM, tempo, scale, and density specifications. Lyrics are automatically aligned to the generated vocal rhythm. -
SynthID Provenance & Responsible Guardrails: Every piece of audio generated by the Lyria family embeds an imperceptible, tamper-resistant SynthID watermark, ensuring verifiable AI provenance while enforcing safety guardrails against artist voice impersonation.
-
Lyria RealTime Interactive Jamming: Unlike turn-based models that require waiting for a render, Lyria RealTime functions like an adaptive live instrument—letting users morph styles, adjust density, and shift genres on the fly in real time.
-
Developer API Integration: Accessible via the standard Gemini API and Google AI Studio, developers can integrate high-fidelity audio generation directly into applications, games, and creative automation pipelines.
