Suno is expanding beyond AI-generated songs with a new tool for spoken audio. The Verge reported on October 2 that the company’s Speech feature has entered public beta on its web and mobile platforms, allowing users to generate a voiceover from a description or a written script.

The product’s central pitch is that narration and background music can be created together rather than assembled in separate tools. Suno chief product officer Jack Brody described Speech as the company’s first audio model designed to produce voice and music as one cohesive track, according to The Verge. That description is Suno’s own characterization of the system.

Illustration of prompt-based and script-based paths leading to generated speech.
Simple mode starts from a description, while Advanced mode accepts a custom script and additional voice controls.

Speech offers two creation paths. Simple mode accepts a general description of the desired result, while Advanced mode lets a user provide a custom script. The Verge said the advanced controls include choices for the voice’s gender and speaking style, plus a setting that changes how much variation appears between generations. Outputs can run for roughly eight minutes.

Background music is optional, so the same feature can produce a clean spoken track. With music enabled, Suno is positioning it for material such as poems with calm accompaniment, energetic speeches, dramatic voiceovers, meditations, pep talks, and bedtime stories. Those examples indicate the company is aiming at a broader range of audio creation than its original song-focused product.

Illustration of generated voice and music expanding into a broader audio landscape.
Suno is entering an established speech market with a model designed to create narration and music together.

Generated speech itself is not new. The Verge noted that DeepMind has worked on deep-learning speech synthesis for about a decade, Adobe offers text-to-speech technology, and ElevenLabs has become a prominent company in the category since its 2023 launch. Suno’s differentiator is the attempt to generate the spoken performance and its musical bed together.

The launch also broadens Suno’s identity at a time when its music generator has drawn multiple lawsuits, The Verge observed. The report suggested diversification may be one reason for entering the established speech market, but that motive was presented as an interpretation rather than a confirmed explanation from the company.

Suno is explicitly labeling Speech as unfinished. Brody said accents can shift unexpectedly and dramatic pauses may become overly dramatic, and the company plans to improve the feature using feedback from the public beta. The initial release therefore presents a wider creative canvas, but not a promise of consistently controlled narration.