Speech input
Configure the independent speech package and its transcription providers.
@repo/speech is the shared microphone-to-text package. It has separate browser
and server entry points and does not depend on the agent runtime, chat history,
database, or storage. The consuming interface owns what happens to the text.
In Agent Chat, Stop-to-send uses the current conversational model and thinking choices. Changing them during recording does not change the speech provider. If the selected model cannot accept pending images, finalization keeps the text and attachments for review instead of submitting. The browser workflow covers model changes during recording and unchanged admission retries in both placements.
Configuration
Speech is disabled by default. Set enabled in packages/speech/config.ts to
enable it. That file also owns recording duration, connection/finalization
deadlines, audio frame and buffer sizes, transcript limits, and concurrency limits.
Invalid limits fail validation rather than creating an unbounded recording.
Choose the provider with the re-export in packages/speech/provider/index.ts,
following the same convention as the payments package. Each provider's config.ts
owns its endpoint, model, and language. Speech language is independent of the UI
locale; customize the selected provider's supported language settings there.
| Provider | Server credential | Settings |
|---|---|---|
| ElevenLabs | ELEVENLABS_API_KEY | packages/speech/provider/elevenlabs/config.ts |
| Deepgram | DEEPGRAM_API_KEY | packages/speech/provider/deepgram/config.ts |
Set only the selected provider's credential in .env.local or the API deployment
environment. Credentials resolve lazily on the server. Missing credentials leave
speech unavailable without preventing unrelated application imports.
getSpeechCapabilities from @repo/speech/server reports availability and public
input limits without contacting the provider. Import shared types from
@repo/speech/types; browser consumers use @repo/speech/client and must never
import the server entry point.
Package checks
pnpm --filter @repo/speech test
pnpm --filter @repo/speech type-checkThe default tests use controlled settings and credentials; they do not require a live transcription account. Changing provider defaults should not require updating tests that exercise the same behavior with their own settings.
Integrate another input
The browser package exports useSpeech for React and createSpeechRecorder for
other browser interfaces. Neither submits text to an agent or form. A consumer
chooses whether completed text fills a field, runs a search, or submits a message.
"use client";
import { useSpeech } from "@repo/speech/client";
export function Dictation({ onText }: { onText: (text: string) => void }) {
const speech = useSpeech();
const recording = speech.status === "recording";
const busy = ["connecting", "recording", "finalizing"].includes(speech.status);
return (
<>
<button
disabled={busy && !recording}
onClick={async () => {
try {
if (recording) onText(await speech.stop());
else {
const url = new URL("/api/speech/stream", location.href);
url.protocol = url.protocol === "https:" ? "wss:" : "ws:";
await speech.start({
url: url.href,
workletUrl: "/speech/pcm-worklet.js",
maxTranscriptCharacters: 1000,
});
}
} catch {
// Render a localized message from speech.error below the field.
}
}}
>
{recording ? "Finish dictation" : "Dictate"}
</button>
{busy && <button onClick={speech.cancel}>Cancel dictation</button>}
<p>{[speech.finalized, speech.partial].filter(Boolean).join(" ")}</p>
{speech.error && <p role="alert">{speech.error}</p>}
</>
);
}Localize control labels and safe error codes in the consuming application. Check
speech.capabilities through the application's oRPC helpers before offering
recording. A denied microphone request, unsupported browser, or disconnected
provider must leave the original input usable. Microphone access requires a
secure context (HTTPS, or localhost during development).
SaaS development and build scripts run scripts/prepare-speech.mjs, copying
@repo/speech/pcm-worklet.js to the same-origin public asset path above. Other
hosts can serve that exported asset themselves and pass their own URL. The worklet
produces mono signed 16-bit PCM at the actual AudioContext sample rate. The browser
flushes its trailing samples before asking the server to finish.
API ownership and deployment
apps/api/src/speech.ts attaches /api/speech/stream to the existing Node HTTP
server. speech-auth.ts validates the real login and optional organization
membership using the existing auth helpers, including current session, ban, and
impersonation checks. The package itself imports no application auth or agent
runtime. Availability is a separate protected procedure under
packages/api/modules/speech.
Route same-origin WebSocket upgrades to the API process and preserve cookies and the browser's Origin header. The built Next.js rewrite is covered by the browser workflow. For external ingress, configure HTTP/1.1 upgrade forwarding, for example with nginx (replace the upstream name with your deployment's API service):
location = /api/speech/stream {
proxy_pass http://api:3004;
proxy_http_version 1.1;
proxy_set_header Upgrade $http_upgrade;
proxy_set_header Connection "upgrade";
proxy_set_header Host $host;
proxy_read_timeout 360s;
}Configure NEXT_PUBLIC_SAAS_URL to the exact public SaaS origin. Providers' permanent
keys stay in the API process and are never placed in WebSocket URLs. Restart the
API after changing speech enablement, provider selection, or server credentials.
The host checks access before upgrade, during recording, and before returning a completed result. The default periodic authorization interval is 30 seconds; session expiry closes the connection independently. Concurrency is enforced per user per API process. Multiple API replicas need shared admission counters if a deployment requires a cluster-wide quota. Shutdown closes active recordings.
Completion, limits, and diagnostics
Stop waits for final words rather than submitting a provisional transcript.
ElevenLabs uses serialized manual commits in segments of at most 20 seconds to
stay below its documented automatic-commit boundary; a bounded buffer holds the
next segment until acknowledgement. Deepgram uses CloseStream and collects the
remaining results and closing metadata, without depending on the optional
from_finalize flag. Neither adapter automatically reconnects, replays audio, or
falls back to another provider after failure.
Recording duration, connection/finalization deadlines, transcript length, chunk size, and queued audio are bounded. Limit and finalization failures leave recovered text for review instead of auto-sending a truncated message. Raw audio is transient; the speech transport stores neither recordings nor transcripts. A consumer such as chat persists the resulting text through its own ordinary submission flow.
API process diagnostics contain recording correlation, provider identity, safe
failure categories, recognized provider error codes, HTTP rejection status, and
provider identifiers when available. Transport and provider records share the
recording ID. Audio and transcript payloads are not logged. @repo/logs bounds
and redacts these diagnostic fields.
Browser and live-provider verification
The development routing check starts a temporary Next development server for the SaaS project. Stop an existing SaaS dev server before running it, then restart your normal development command afterward:
pnpm --filter saas e2e:speech-devIt verifies cookies, Origin, PCM, final transcript delivery, and the worklet asset through the actual development rewrite. The built browser workflows are:
pnpm --filter saas e2e:agent speech.spec.ts
AGENT_CHAT_PLACEMENT=floating pnpm --filter saas e2e:agent speech.spec.tsThese commands use the existing agent browser fixture and require its PostgreSQL and private S3-compatible test storage setup; see agent testing. Speech itself needs neither storage nor an agent. A standalone browser consumer is also exercised by this workflow. Tests feed a real synthetic MediaStream through the capture worklet and authenticated API to a local provider fixture. They verify PCM delivery and final-word ordering without a physical microphone or paid account.
Live verification is separate: configure each provider in turn and record short speech, silence, a longer multi-segment request, and a final word immediately before Stop. Measure time to the first provisional text and Stop-to-completion; confirm the final words are preserved. Repeat physical-microphone permission and cleanup checks in the browsers your deployment supports. Fixture success alone does not establish live latency, recognition accuracy, or physical-device support.