Acme
Speech

Speech input

Configure the independent speech package and its transcription providers.

@repo/speech is the shared microphone-to-text package. It has separate browser and server entry points and does not depend on the agent runtime, chat history, database, or storage. The consuming interface owns what happens to the text.

In Agent Chat, Stop-to-send uses the current conversational model and thinking choices. Changing them during recording does not change the speech provider. If the selected model cannot accept pending images, finalization keeps the text and attachments for review instead of submitting. The browser workflow covers model changes during recording and unchanged admission retries in both placements.

Configuration

Speech is disabled by default. Set enabled in packages/speech/config.ts to enable it. That file also owns recording duration, connection/finalization deadlines, audio frame and buffer sizes, transcript limits, and concurrency limits. Invalid limits fail validation rather than creating an unbounded recording.

Choose the provider with the re-export in packages/speech/provider/index.ts, following the same convention as the payments package. Each provider's config.ts owns its endpoint, model, and language. Speech language is independent of the UI locale; customize the selected provider's supported language settings there.

ProviderServer credentialSettings
ElevenLabsELEVENLABS_API_KEYpackages/speech/provider/elevenlabs/config.ts
DeepgramDEEPGRAM_API_KEYpackages/speech/provider/deepgram/config.ts

Set only the selected provider's credential in .env.local or the API deployment environment. Credentials resolve lazily on the server. Missing credentials leave speech unavailable without preventing unrelated application imports.

getSpeechCapabilities from @repo/speech/server reports availability and public input limits without contacting the provider. Import shared types from @repo/speech/types; browser consumers use @repo/speech/client and must never import the server entry point.

Package checks

pnpm --filter @repo/speech test
pnpm --filter @repo/speech type-check

The default tests use controlled settings and credentials; they do not require a live transcription account. Changing provider defaults should not require updating tests that exercise the same behavior with their own settings.

Integrate another input

The browser package exports useSpeech for React and createSpeechRecorder for other browser interfaces. Neither submits text to an agent or form. A consumer chooses whether completed text fills a field, runs a search, or submits a message.

"use client";

import { useSpeech } from "@repo/speech/client";

export function Dictation({ onText }: { onText: (text: string) => void }) {
	const speech = useSpeech();
	const recording = speech.status === "recording";
	const busy = ["connecting", "recording", "finalizing"].includes(speech.status);
	return (
		<>
			<button
				disabled={busy && !recording}
				onClick={async () => {
					try {
						if (recording) onText(await speech.stop());
						else {
							const url = new URL("/api/speech/stream", location.href);
							url.protocol = url.protocol === "https:" ? "wss:" : "ws:";
							await speech.start({
								url: url.href,
								workletUrl: "/speech/pcm-worklet.js",
								maxTranscriptCharacters: 1000,
							});
						}
					} catch {
						// Render a localized message from speech.error below the field.
					}
				}}
			>
				{recording ? "Finish dictation" : "Dictate"}
			</button>
			{busy && <button onClick={speech.cancel}>Cancel dictation</button>}
			<p>{[speech.finalized, speech.partial].filter(Boolean).join(" ")}</p>
			{speech.error && <p role="alert">{speech.error}</p>}
		</>
	);
}

Localize control labels and safe error codes in the consuming application. Check speech.capabilities through the application's oRPC helpers before offering recording. A denied microphone request, unsupported browser, or disconnected provider must leave the original input usable. Microphone access requires a secure context (HTTPS, or localhost during development).

SaaS development and build scripts run scripts/prepare-speech.mjs, copying @repo/speech/pcm-worklet.js to the same-origin public asset path above. Other hosts can serve that exported asset themselves and pass their own URL. The worklet produces mono signed 16-bit PCM at the actual AudioContext sample rate. The browser flushes its trailing samples before asking the server to finish.

API ownership and deployment

apps/api/src/speech.ts attaches /api/speech/stream to the existing Node HTTP server. speech-auth.ts validates the real login and optional organization membership using the existing auth helpers, including current session, ban, and impersonation checks. The package itself imports no application auth or agent runtime. Availability is a separate protected procedure under packages/api/modules/speech.

Route same-origin WebSocket upgrades to the API process and preserve cookies and the browser's Origin header. The built Next.js rewrite is covered by the browser workflow. For external ingress, configure HTTP/1.1 upgrade forwarding, for example with nginx (replace the upstream name with your deployment's API service):

location = /api/speech/stream {
    proxy_pass http://api:3004;
    proxy_http_version 1.1;
    proxy_set_header Upgrade $http_upgrade;
    proxy_set_header Connection "upgrade";
    proxy_set_header Host $host;
    proxy_read_timeout 360s;
}

Configure NEXT_PUBLIC_SAAS_URL to the exact public SaaS origin. Providers' permanent keys stay in the API process and are never placed in WebSocket URLs. Restart the API after changing speech enablement, provider selection, or server credentials.

The host checks access before upgrade, during recording, and before returning a completed result. The default periodic authorization interval is 30 seconds; session expiry closes the connection independently. Concurrency is enforced per user per API process. Multiple API replicas need shared admission counters if a deployment requires a cluster-wide quota. Shutdown closes active recordings.

Completion, limits, and diagnostics

Stop waits for final words rather than submitting a provisional transcript. ElevenLabs uses serialized manual commits in segments of at most 20 seconds to stay below its documented automatic-commit boundary; a bounded buffer holds the next segment until acknowledgement. Deepgram uses CloseStream and collects the remaining results and closing metadata, without depending on the optional from_finalize flag. Neither adapter automatically reconnects, replays audio, or falls back to another provider after failure.

Recording duration, connection/finalization deadlines, transcript length, chunk size, and queued audio are bounded. Limit and finalization failures leave recovered text for review instead of auto-sending a truncated message. Raw audio is transient; the speech transport stores neither recordings nor transcripts. A consumer such as chat persists the resulting text through its own ordinary submission flow.

API process diagnostics contain recording correlation, provider identity, safe failure categories, recognized provider error codes, HTTP rejection status, and provider identifiers when available. Transport and provider records share the recording ID. Audio and transcript payloads are not logged. @repo/logs bounds and redacts these diagnostic fields.

Browser and live-provider verification

The development routing check starts a temporary Next development server for the SaaS project. Stop an existing SaaS dev server before running it, then restart your normal development command afterward:

pnpm --filter saas e2e:speech-dev

It verifies cookies, Origin, PCM, final transcript delivery, and the worklet asset through the actual development rewrite. The built browser workflows are:

pnpm --filter saas e2e:agent speech.spec.ts
AGENT_CHAT_PLACEMENT=floating pnpm --filter saas e2e:agent speech.spec.ts

These commands use the existing agent browser fixture and require its PostgreSQL and private S3-compatible test storage setup; see agent testing. Speech itself needs neither storage nor an agent. A standalone browser consumer is also exercised by this workflow. Tests feed a real synthetic MediaStream through the capture worklet and authenticated API to a local provider fixture. They verify PCM delivery and final-word ordering without a physical microphone or paid account.

Live verification is separate: configure each provider in turn and record short speech, silence, a longer multi-segment request, and a final word immediately before Stop. Measure time to the first provisional text and Stop-to-completion; confirm the final words are preserved. Repeat physical-microphone permission and cleanup checks in the browsers your deployment supports. Fixture success alone does not establish live latency, recognition accuracy, or physical-device support.

On this page