AI Speech Reference

Use useAISpeech to convert text into natural-sounding audio. Generated speech returns a hosted URL for playback and persistence.

useAISpeech

Hook Options

  • onError ((error: Error) => void): Custom error handler.
  • onFinish (function): Called with the generated audio object when generation completes.

Returns

  • generateSpeech (function): Generate speech from text.
  • isLoading (boolean): Whether generation is in progress.
  • error (Error | null): Current error.
  • clearError (() => void): Clear the current error.

Result Shape

  • url (string): Hosted URL for playback and persistence.

generateSpeech Input

  • text (string, required): Text to convert to speech.
  • voice (string): Model-specific voice ID. Omit it by default.
  • instructions (string): Additional voice direction for supported OpenAI models.
  • language (string, default "en"): ISO 639-1 language code.
  • model (string): Model ID for the request.

generateSpeech resolves to the result shape above on success or undefined on failure. It never rejects.

Example

Import the hook from the AI hook module:

import { useAISpeech } from '@/hooks/use-ai'
import { usePersistentItem } from '@/hooks/use-persistent-item'

export default function App() {
  const [text, setText] = React.useState('')
  const [audioUrl, setAudioUrl] = usePersistentItem<string | null>(
    'audioUrl',
    null,
  )
  const { generateSpeech, isLoading, error } = useAISpeech()

  const handleGenerate = async () => {
    if (!text.trim()) return

    const audio = await generateSpeech({ text: text.trim() })
    if (audio?.url) {
      setAudioUrl(audio.url)
    }
  }

  return (
    <div>
      <textarea
        value={text}
        onChange={(event) => setText(event.target.value)}
        placeholder="Enter text to convert to speech..."
      />
      <button onClick={handleGenerate} disabled={isLoading}>
        {isLoading ? 'Generating...' : 'Generate speech'}
      </button>
      {error ? <p className="text-destructive">{error.message}</p> : null}
      {audioUrl ? (
        <audio controls src={audioUrl} className="w-full">
          Your browser does not support the audio element.
        </audio>
      ) : null}
    </div>
  )
}

Use the returned hosted URL for playback and persistence.

Voices and Providers

Omit voice for portable apps so the platform can choose a compatible model and default voice. OpenAI voices suit most narration, accessibility, and high-volume use cases. ElevenLabs voices are best when expressive or branded delivery is central to the experience.

For model-specific voices:

import { listSpeechModels, listSpeechVoices, useAISpeech } from '@/hooks/use-ai'

const models = listSpeechModels()
const [modelId, setModelId] = React.useState(models[0]?.id)
const voices = modelId ? listSpeechVoices(modelId) : []
const [voiceId, setVoiceId] = React.useState<string | undefined>(undefined)
const { generateSpeech } = useAISpeech()

await generateSpeech({
  text: 'Welcome to our application!',
  model: modelId,
  voice: voiceId,
})

Model Selection

When no model is specified, the platform chooses an efficient default speech model for the user's plan. To expose model choice, pass a model ID and build selectors from listSpeechModels():

import { listSpeechModels, useAISpeech } from '@/hooks/use-ai'

const models = listSpeechModels()
const [modelId, setModelId] = React.useState<string | undefined>(undefined)
const { generateSpeech } = useAISpeech()

await generateSpeech({
  text: 'Hello world',
  model: modelId,
})

listSpeechModels() returns model objects with id, displayName, provider, tier, description, and capabilities. Speech model capabilities include tts.

listSpeechVoices(modelId) returns voice objects:

  • id: Voice ID to pass as voice.
  • name: Human-readable voice name.
  • description: Guidance about the voice character.

All models are available on every plan; more capable models cost more credits. Deprecated models automatically fall back to a recommended replacement.

Best Practices

  • Omit voice unless the app specifically needs a voice picker or branded voice.
  • Use the url field for playback and persistence.
  • Store the URL with usePersistentItem if the audio should survive reloads.
  • Split long content into chunks because maximum text length varies by model.
  • If you require a specific voice, choose the model first and load voices with listSpeechVoices(modelId).