Speech-to-Text Engine
Choose how audio is converted to text
AWS Transcribe runs in the cloud — instant, no GPU cost, but no speaker diarization. WhisperLiveKit runs on a dedicated GPU with diarization and multilingual support, but requires ~60s to start and incurs GPU costs.
Engine
Engine Status
GPU Instance:
Stopped
WLK Engine:
Stopped
Pipeline Tuning
Pick a style — the same numbers apply to every language pair
Style
Advanced — raw pipeline values
Pre-Buffer
Min characters
Floor for sentence-boundary and pause triggers. Ignored when eager-fire is ON.
Min words
Companion to min characters — both must be met.
Silence timeout (seconds)
How long a pause must last to release. Resets every time new text arrives.
Max buffer time (seconds)
Hard ceiling. Fires this long after the first word, no matter what. Never resets.
Smart chunking
Holds release when buffer ends mid-phrase (after
and, the, は, が…). Uses the word list below. Capped by max buffer time.Eager punctuation fire
Fires on any sentence-end punctuation, ignoring min chars / min words. Snappier but can emit short fragments.
Post-Stitcher
Controls how the operator-UI rows are sized. Each row holds at least
this many characters of translated output before firing on the next
sentence-end. Lower = shorter rows, fires often. Higher = longer rows,
more clustering. The post-stitcher always fires on sentence-end once
past the floor, never mid-sentence.
Row length floor:
—
characters
Drag to adjust. 10–30 = short, sentence-paced rows.
40–80 = balanced. 100+ = paragraph
clustering for lecture-style content.
Max wait on pause:
—
After the speaker pauses, how long to wait for more speech
before force-firing the row regardless of the floor above.
0 = wait forever (only fires on sentence-end
above the floor or when recording stops). Higher values cluster
more text per row at the cost of latency on isolated short utterances.
Save Changes covers the numeric Pre-Buffer and Post-Stitcher fields above. The two sub-sections below are language-scoped and have their own Save buttons.
Sentence-End Pattern — per target language
Regex used by the post-stitcher to detect a finished sentence in the translated output. Defaults are Latin
[.!?], CJK [。!?], Mixed [。!?.!?] for Chinese.
Editing pattern for: Japanese
Smart Chunking — Hold Words — per source language
When smart chunking is enabled, the buffer won't fire if the text ends with any of these words.
Editing list for: English
Translation
Translator chain (per language pair) and Nova prompts
Translator Chain (hedged race)
Not a fallback chain. The 1st choice fires immediately. If it hasn't returned by the hedge delay, the 2nd choice is launched in parallel (the 1st keeps running). Same for the 3rd. First non-empty result wins — the rest are cancelled. If nothing returns by the chain deadline, the engine emits the marker (default
…) — TTS skips it and the UI greys it out.
Why hedge? When the 1st choice is healthy, you pay nothing extra. When it's slow, the 2nd choice catches up and wins after just the hedge delay — no full timeout penalty. Putting the fastest model first means most chunks come back from it; the rest are quietly absorbed by the hedge.
1st choice runs at t = 0s
The translator that handles every chunk under healthy conditions.
Hard timeout (s)
Stop waiting on this call entirely (it never wins after this).
2nd choice joins at hedge delay
Launches in parallel if the 1st choice is slow. First good answer wins.
Hard timeout (s)
Stop waiting on this call. (Race continues with whatever's left.)
3rd choice joins at 2× hedge delay
Last contestant. Lower-quality fallback (e.g. Amazon Translate) is fine here.
Hard timeout (s)
Stop waiting on this call.
Hedge delay (s)
If the 1st choice hasn't returned by this point, the 2nd joins. The 3rd joins at 2× this. Lower = more aggressive hedging (more parallel calls); higher = leaner but slower fallback. Default
0.8.Chain deadline (s)
If no choice has returned by this hard cap, the marker is emitted and all calls are cancelled. Default
4.Marker text
Placeholder emitted when no tier returns in time. Audience sees a greyed row; TTS skips it.
Nova / Claude System Prompts
Shared by every Bedrock-backed tier (Nova, Haiku, Sonnet). Amazon Translate ignores prompts.
Connection
Service endpoints and authentication
Documentation
For developers integrating with this service or operators tuning the engine.
- → gRPC Integration Guide — for clients writing a gRPC consumer
- → API Reference — full proto and HTTP endpoint reference
- → Operator Guide — runtime configuration and troubleshooting
- → Deployment Guide — CDK modes and infrastructure
- → Architecture — how this engine fits with hibiki-stage and kizuna-tts