Skip to content

Voice

Build speech input and output around an agent conversation.

  • Libraries Runtime.Voice Runtime.VoiceErrors
  • npm @cloudflare/voice 0.4.0
  • Upstream/API Experimental

Speech is a separate pipeline

Voice connects speech providers and audio handling to the rest of an agent. Choose transcription and synthesis providers, define when a turn begins and ends, and decide how partial text and interruptions affect the conversation.

The example below covers text-to-speech only: it turns part of an article into audio using Workers AI. It can be used independently of a chat agent. The pinned Voice package contains the wider provider and pipeline interfaces.

Audio Preview

The Worker responds to a POSTed article with MPEG audio of its first two sentences, which SentenceChunker splits out and WorkersAITTS synthesizes with Workers AI. If the synthesize call throws, logVoiceError writes one structured log entry and the caller receives the error text, cut to at most 300 characters.

open Fable.Core

module Workers = FSharp.CloudEdge.Runtime.Workers
module Voice = FSharp.CloudEdge.Runtime.Voice
module Errors = FSharp.CloudEdge.Runtime.VoiceErrors.Errors

type Env =
    abstract AI: Voice.AiLike

let opening (article: string) =
    let chunker = Voice.Exports.SentenceChunker()
    Array.append (chunker.add article) (chunker.flush ()) |> Array.truncate 2 |> String.concat " "

let failed (message: string) =
    Workers.Exports.Response.json ({| error = message |}, U2.Case2(Workers.ResponseInit.Create(status = 502.)))

[<ExportDefault>]
let worker: Workers.ExportedHandler<Env, obj, obj, obj> =
    Workers.ExportedHandler.Create(
        fetch = fun request env _ ->
            async {
                try
                    let! article = request.text () |> Async.AwaitPromise
                    let tts = Voice.Exports.WorkersAITTS env.AI
                    let! audio = tts.synthesize (opening article) |> Async.AwaitPromise
                    match audio with
                    | Some bytes ->
                        let headers = [| [| "content-type"; "audio/mpeg" |] |]
                        return Workers.Exports.Response.Create(bytes, Workers.ResponseInit.Create(headers = headers))
                    | None -> return failed "No audio"
                with caught ->
                    let error = Errors.Exports.toVoiceError (caught, "Speech is unavailable")
                    Errors.Exports.logVoiceError (Errors.VoiceErrorLogOptions.Create(``component`` = "AudioPreview", stage = "synthesize", message = "Preview failed", error = error))
                    return failed (Errors.Exports.voiceErrorMessage (error, "Speech is unavailable"))
            }
            |> Async.StartAsPromise
            |> U2.Case1
    )

Needs a Workers AI binding named AI. Voice.AiLike is the Voice package's own type for that binding, and WorkersAITTS calls the @cf/deepgram/aura-1 model unless its options set another. The free plan includes 10,000 Workers AI Neurons a day.

Integrating a conversation

Keep conversation identity and history in the chat layer. Connect text from speech recognition to that conversation, then stream or synthesize the reply through the chosen output provider. Audio encoding, chunking, cancellation, and provider errors belong to this integration and should be handled explicitly.

Testing and feedback

Useful test cases include audio and stream types, provider callbacks, synthesis results, cancellation, and error handling.

Verify bindings explains how to run checks and report an issue. Include a small reproduction and the package versions used; working examples are welcome too.

NuGet packages

Runtime.Voice 0.1.0, Runtime.VoiceErrors 0.1.0, Runtime.Workers 0.1.0.

See installation and release availability.

Edit this page