Voice
Build speech input and output around an agent conversation.
- Libraries
Runtime.VoiceRuntime.VoiceErrors - npm
@cloudflare/voice 0.4.0 - Upstream/API Experimental
Speech is a separate pipeline
Voice connects speech providers and audio handling to the rest of an agent. Choose transcription and synthesis providers, define when a turn begins and ends, and decide how partial text and interruptions affect the conversation.
The example below covers text-to-speech only: it turns part of an article into audio using Workers AI. It can be used independently of a chat agent. The pinned Voice package contains the wider provider and pipeline interfaces.
Audio Preview
The Worker responds to a POSTed article with MPEG audio of its first two sentences, which SentenceChunker splits out and WorkersAITTS synthesizes with Workers AI. If the synthesize call throws, logVoiceError writes one structured log entry and the caller receives the error text, cut to at most 300 characters.
open Fable.Core
module Workers = FSharp.CloudEdge.Runtime.Workers
module Voice = FSharp.CloudEdge.Runtime.Voice
module Errors = FSharp.CloudEdge.Runtime.VoiceErrors.Errors
type Env =
abstract AI: Voice.AiLike
let opening (article: string) =
let chunker = Voice.Exports.SentenceChunker()
Array.append (chunker.add article) (chunker.flush ()) |> Array.truncate 2 |> String.concat " "
let failed (message: string) =
Workers.Exports.Response.json ({| error = message |}, U2.Case2(Workers.ResponseInit.Create(status = 502.)))
[<ExportDefault>]
let worker: Workers.ExportedHandler<Env, obj, obj, obj> =
Workers.ExportedHandler.Create(
fetch = fun request env _ ->
async {
try
let! article = request.text () |> Async.AwaitPromise
let tts = Voice.Exports.WorkersAITTS env.AI
let! audio = tts.synthesize (opening article) |> Async.AwaitPromise
match audio with
| Some bytes ->
let headers = [| [| "content-type"; "audio/mpeg" |] |]
return Workers.Exports.Response.Create(bytes, Workers.ResponseInit.Create(headers = headers))
| None -> return failed "No audio"
with caught ->
let error = Errors.Exports.toVoiceError (caught, "Speech is unavailable")
Errors.Exports.logVoiceError (Errors.VoiceErrorLogOptions.Create(``component`` = "AudioPreview", stage = "synthesize", message = "Preview failed", error = error))
return failed (Errors.Exports.voiceErrorMessage (error, "Speech is unavailable"))
}
|> Async.StartAsPromise
|> U2.Case1
)
Needs a Workers AI binding named AI. Voice.AiLike is the Voice package's own type for that binding, and WorkersAITTS calls the @cf/deepgram/aura-1 model unless its options set another. The free plan includes 10,000 Workers AI Neurons a day.
Integrating a conversation
Keep conversation identity and history in the chat layer. Connect text from speech recognition to that conversation, then stream or synthesize the reply through the chosen output provider. Audio encoding, chunking, cancellation, and provider errors belong to this integration and should be handled explicitly.
Related Pages
Testing and feedback
Useful test cases include audio and stream types, provider callbacks, synthesis results, cancellation, and error handling.
Verify bindings explains how to run checks and report an issue. Include a small reproduction and the package versions used; working examples are welcome too.
NuGet packages
Runtime.Voice 0.1.0, Runtime.VoiceErrors 0.1.0, Runtime.Workers 0.1.0.