Google DeepMind Lyria

Lyria (deepmind.google/models/lyria/) is Google DeepMind’s flagship family of generative music and audio foundation models. Built using latent diffusion over temporal audio latents and integrated into the Gemini API (Interactions API), Vertex AI, and Google AI Studio, the Lyria model suite—spanning Lyria 3 Clip, Lyria 3 Pro, Lyria 3.5, and Lyria RealTime—generates full-fidelity stereo tracks with vocals, timed lyrics, and section arrangements from text or multimodal image inputs, all protected with imperceptible SynthID audio watermarking.

Early AI music synthesis often suffered from structural incoherence, muffled vocal artifacts, and limited track lengths. Lyria provides a musically grounded audio generation architecture capable of modeling complex musical phrasing (verses, choruses, bridges), multi-instrument arrangements, and real-time interactive performance parameters.

Model Lineup & Architectural Profile

Model Variant Duration / Output Audio Quality Core Modality & Input Primary Use Case
Lyria 3 Clip (lyria-3-clip-preview) 30 Seconds (MP3) 44.1 kHz High-Fidelity Stereo Text prompts or up to 10 reference images Hooks, intro loops, ads, social clips
Lyria 3 Pro (lyria-3-pro-preview) Up to 184 Seconds / 3 Min (MP3/WAV) 44.1 kHz High-Fidelity Stereo Multimodal inputs, BPM, key, section tags Full songs, video soundtrack beds, orchestral scores
Lyria 3.5 Variable length (up to 3 min) High-Fidelity Stereo Advanced text/lyric conditioning Nuanced vocal performance & musicality
Lyria RealTime Continuous Live Streaming 48 kHz Professional Stereo Interactive parameter sliders (BPM, mood, genre) Live DJ performance, MusicFX DJ, adaptive game audio

Core Breakthroughs & Capabilities

  • Multimodal Conditioning (Text & Vision): Developers and creators can condition Lyria 3 on up to 10 visual images alongside text descriptions, enabling the model to score video soundtracks, mood boards, or artwork dynamically.

  • Structured Composition & Lyric Synchronization: Supports explicit song structure control using standard musical tags ([Verse], [Chorus], [Bridge]) alongside BPM, tempo, scale, and density specifications. Lyrics are automatically aligned to the generated vocal rhythm.

  • SynthID Provenance & Responsible Guardrails: Every piece of audio generated by the Lyria family embeds an imperceptible, tamper-resistant SynthID watermark, ensuring verifiable AI provenance while enforcing safety guardrails against artist voice impersonation.

  • Lyria RealTime Interactive Jamming: Unlike turn-based models that require waiting for a render, Lyria RealTime functions like an adaptive live instrument—letting users morph styles, adjust density, and shift genres on the fly in real time.

  • Developer API Integration: Accessible via the standard Gemini API and Google AI Studio, developers can integrate high-fidelity audio generation directly into applications, games, and creative automation pipelines.