Skip to main content

AI audio models

Try leading AI voice, text-to-speech, and sound-generation models online · 6 models currently available · 4 models available for reference only

Compare models for text-to-speech, voiceover, multi-speaker dialogue, music drafts, and sound effects.

Suno Music is Kyeo AI's text-to-music entry. The active form exposes a prompt, version selector, and instrumental switch. Eligible successful results can launch extension, WAV conversion, or vocal separation. Lyrics, title, style weights, Persona, Voices, Custom Models, My Taste, and other controls from Suno's own product are not exposed here.

Text to audio

Suno Sounds is Kyeo AI's text-to-sound entry. Choose V5 or V5.5, request a one-shot or loop-oriented result, and optionally set 1–300 BPM and a key. Features described for Suno's official beta product, sample counts, and subscription rights do not automatically apply to results generated on Kyeo.

Text to audio

Suno Lyrics is Kyeo AI's text-only lyric generator. Submit a topic or brief of up to 200 characters, then review the parsed titles and lyric text. The current entry has no ReMi or Classic selector and produces no melody, performance, or audio.

Write lyrics
Draft lyrics

ElevenLabs Dialogue V3 provides line-by-line dialogue generation: enter each turn, choose a voice, and pay by total dialogue length. It is not a real-time voice agent, and ElevenLabs notes that users may need multiple generations to find a usable result. Start with a sample to check speaker changes, audio tags, language code, output format, and long-script completeness.

Dialogue to audio

ElevenLabs Turbo 2.5 is a retained single-speaker TTS workflow on Kyeo AI. It exposes a voice choice plus stability, similarity boost, style, speed, timestamps, surrounding text, and language code. The workflow remains selectable, but ElevenLabs now recommends Flash v2.5 instead. The Turbo name does not make low latency, language support, timestamp accuracy, voice authorization, or output quality a Kyeo guarantee.

Text to speech

ElevenLabs Multilingual V2 is Kyeo AI's current single-speaker, high-naturalness voiceover page. It shares a similar form with Turbo 2.5, but the task boundary differs: Turbo is a fast voiceover baseline, while Multilingual V2 is aimed at long-form narration, cross-language content, and sustained brand voice. Kyeo still uses a fixed voice list and a 5,000-character entry, so the practical decision is whether the listening experience justifies twice Turbo's per-character cost.

Text to speech
Currently unavailable

The official ElevenLabs Sound Effect V2 interface can generate sound effects from text. Kyeo AI has disabled this model's runtime, so this page provides no generation entry and creates no task or credit charge.

Full model documentation remains available, but online generation is currently unavailable.

Text to audio
Currently unavailable

ElevenLabs Audio Isolation covers audio cleanup that separates speech from background noise. Generation is currently closed, so no task or credit charge is created. This page keeps the verified context and points to active alternatives.

Full model documentation remains available, but online generation is currently unavailable.

Clean recording
Isolate voice
Currently unavailable

Gemini 3.1 Flash TTS supports documented multi-speaker dialogue and voice controls, but a reliable maximum output-audio estimate is not yet available. Its details remain available while generation stays disabled.

Full model documentation remains available, but online generation is currently unavailable.

Text to speech
dialogue-to-speech
Currently unavailable

Gemini 2.5 Pro TTS targets detailed multi-character voice design. Because no reliable output-audio limit is available, this page provides current capability details and practical alternatives while generation remains disabled.

Full model documentation remains available, but online generation is currently unavailable.

Text to speech
dialogue-to-speech