Audience: people configuring a hibiki-translate deployment — choosing STT engines, tuning translation profiles, managing GPU lifecycle, customising prompts.
This is about running the engine, not consuming it. If you're writing a gRPC client, see grpc-integration.md. If you're deploying it, see deployment.md.
The settings page at /translate/settings is the main operator UI. It surfaces these knobs:
| Section | What it controls |
|---|---|
| STT engine | WhisperLiveKit (GPU, diarization) vs. AWS Transcribe (no GPU) |
| Pipeline tuning | Pre-buffer thresholds, smart-chunking word lists, post-stitcher behavior — per language pair |
| Translation model | Nova system prompt (with and without context); model ID display |
| GPU lifecycle | Manual start / stop of the WLK GPU instance |
| Connection | Display-only: WebSocket URL and auth mode |
All settings persist to DynamoDB (the HibikiTranslateProfiles table) where applicable. Restarting the server preserves them.
Settings that are NOT in the UI:
- Region (deploy-time only; AWS_REGION env var)
- Cognito pool / client ID (deploy-time)
- Default target language at server start (TGT_LANG env var)
- gRPC port (hardwired to 50053)
The big choice. Two backends, very different tradeoffs.
Runs on GPU. Uses faster-whisper + Sortformer for transcription with built-in speaker diarization.
Pros:
- Multilingual: handles all 10 supported source languages
- Speaker diarization (speaker_id per word, gender estimate via F0)
- Large-v3 model is the most accurate option
- Streaming partial transcripts work cleanly
Cons:
- Requires a GPU instance (g6e.xlarge / L40S recommended)
- ~60s cold start to load the model
- More expensive — GPU is ~$1–4/hour vs. Transcribe's per-second-of-audio billing
- Single language per WLK process; switching WLK_LANGUAGE requires restarting the WLK service
When to use: events with multiple speakers where diarization matters; when accuracy is the priority; when you have a sustained workload that justifies keeping the GPU warm.
Cloud API. No GPU required.
Pros: - No GPU instance to manage - Instant start (no cold-start delay) - Pay per second of audio actually transcribed (no idle cost) - Multilingual within the supported set
Cons:
- No speaker diarization — speaker_id is always 0
- No gender estimate
- Network latency adds ~50–100ms per chunk
- AWS-region-specific (use the same region as your deploy)
When to use: single-speaker scenarios; cost-sensitive deployments; bursty workloads where keeping a GPU warm doesn't pay off; development and testing.
In the operator UI: Settings → STT Engine → choose → Apply. The server immediately switches, and:
The change persists. Subsequent server restarts use the saved engine choice.
WLK_LANGUAGE mid-session?Today: not possible without restarting the WLK service (the underlying Whisper model loads with a fixed language). The /translate/restart-wlk endpoint exists to do this gracefully.
This is a fundamental limitation of the WLK model loading, not an architectural choice. To work around it, deploy multiple WLK services with different language configs and route to the right one — but that's hibiki-stage's problem, not the engine's.
Translation behavior depends on the source-target pair. Different language pairs have different optimal pre-buffer thresholds, smart-chunking rules, and post-stitcher behavior.
| Category | Languages | Why |
|---|---|---|
| SVO→SVO | en, fr, es, pt, it, de — when both src and tgt are Subject-Verb-Object | Word order is similar; small chunks translate well. Lower latency. |
| SVO→SOV | en→ja, en→ko (and similar) — translating to Subject-Object-Verb | Verb at end means you need most of the sentence before translation makes sense. Higher latency, larger buffers. |
| SOV→SVO | ja→en, ko→en | Inverse of above; reasonable buffers but smart-chunking helps. |
| SOV→SOV | ja→ko, ko→ja | Both verb-final; smaller buffers fine. |
| SVO→Mixed | →zh | Chinese is technically SVO but has unique chunking rules. |
| Setting | Effect |
|---|---|
| Min characters | Don't fire the pre-buffer until at least N characters have accumulated. Higher = bigger chunks, more context for translation, more latency. |
| Min words | Same idea, but word-count gate. Both must be met. |
| Silence timeout (s) | After speaker pauses this long, fire the buffer regardless of size. Lower = more responsive; higher = more grouping. |
| Max buffer time (s) | Hard ceiling: fire after this many seconds even if other thresholds aren't met. Prevents indefinite stalling. |
| Smart chunking | When on, the buffer doesn't fire if the text ends with a "function word" (e.g. "the", "and", particles in JA). Reduces mid-thought translations. |
| Eager punctuation fire | When on, sentence-end punctuation (., 。, ?, etc.) fires the buffer even before thresholds are met. Lower latency for clean speech; can produce fragmenty translations on hesitant speech. |
Defaults are tuned per category. The settings UI lets you override per language pair.
| Pair | Min chars | Min words | Timeout | Max buffer | Smart chunking | Eager punct fire |
|---|---|---|---|---|---|---|
| en→fr (SVO→SVO) | 30 | 4 | 2.5s | 5.0s | on | off |
| en→ja (SVO→SOV) | 40 | 5 | 3.0s | 6.0s | on | off |
| ja→en (SOV→SVO) | 15 | 3 | 3.0s | 6.0s | on | off |
| en→zh (SVO→Mixed) | 25 | 4 | 2.5s | 5.0s | on | off |
These are the post-revert defaults from June 2026. The event-translator branch had briefly used aggressive settings for SVO→SOV (timeout=2.5s, min_chars=20, min_words=3) — those produced fragmenty Japanese and were reverted. If you need to test the aggressive settings again (e.g. for very short utterances like tour guides), the override is preserved in README.md.
Each source language has a list of "function words" — if the buffer's last word is on the list AND smart chunking is on, the buffer waits for more text before firing.
Examples:
- English: the, and, or, but, of, to, in, ...
- Japanese: は, が, を, に, で, と, も, の, ... (case particles)
- French: le, la, les, et, ou, mais, ...
Edit via Settings → Smart Chunking Words. The list is per source language, not per pair.
When STT_ENGINE=whisperlivekit, the engine needs a GPU. Lifecycle is manual.
| Action | UI | API |
|---|---|---|
| Start GPU | Settings → GPU → Start | POST /translate/gpu/start |
| Stop GPU | Settings → GPU → Stop | POST /translate/gpu/stop |
| Status | Settings → GPU → Status badge | GET /translate/gpu/status |
The GPU keeps running indefinitely once started. There's no idle auto-stop and no max-runtime cap. Both used to exist; we removed them after they triggered mid-event terminations.
If you want time-based auto-stop (e.g. for cost control on dev environments), implement it externally — a scheduled Lambda calling /translate/gpu/stop outside business hours, for example. The engine deliberately doesn't make this decision.
When the GPU goes from stopped to running:
1. ASG sets desired capacity → 1 (~30s for instance launch)
2. ECS schedules the WLK task on the new instance (~30s)
3. Whisper model loads (~30–60s for large-v3)
4. WLK service registers in Cloud Map; gRPC server can now reach it
Total: ~1.5–2 minutes of "GPU is starting" before translations work. Plan accordingly for events.
There are two ways to control the GPU instance:
- gpu_manager.start/stop (ASG-based, current default in the UI)
- wlk_service.start/stop (ECS service-based, legacy)
Both manipulate the same physical instance via different AWS APIs. Stopping one doesn't necessarily stop the other. If they get out of sync, the operator UI may show contradictory state. To resolve: stop both, then start the one you want.
This dual-control-plane situation is a known issue. Captured in CHANNEL_FEATURE_REVIEW.md for resolution as part of the consumer-app extraction.
Amazon Nova Lite via Bedrock Converse Stream API. Region-specific model IDs:
| Region | Model ID |
|---|---|
ap-northeast-1 (Tokyo) |
jp.amazon.nova-2-lite-v1:0 |
us-west-2 (Oregon) |
us.amazon.nova-lite-v1:0 |
us-east-1 (Virginia) |
us.amazon.nova-lite-v1:0 |
Streaming output (token-by-token) is enabled, which lets the operator UI show typewriter-style live translation.
Two prompts, one with conversation context, one without. Both are heavily templated — they include:
--, ..., …, repeated commas in the input are treated as speech-recognition hesitation artifacts, not punctuation, and are NOT preserved in the output (Japanese specifically: never use —— or …… to mirror them)"you-- you can extract"), output ONLY the recovered version, never both[SKIP] rule for fillers, garbled input, and refusalsEXCITED|, HAPPY|, ..., NEUTRAL|)Edit prompts in Settings → Translation Model. Both prompts have a "Reset to default" button. Custom prompts persist to DynamoDB.
Be careful editing prompts. The post-processing logic depends on the prompt producing output in the expected format (EMOTION|text or [SKIP]). Breaking that contract leads to malformed translations. If you need radically different output, the right path is the tool-use refactor in TOOLUSE_REFACTOR.md, not prompt acrobatics.
The defaults are tuned for live speech where the speaker hesitates and corrects themselves frequently. If you customise the prompt, keep the disfluency / false-start sections — without them the translator faithfully renders WLK's -- artifacts as Japanese em-dashes, which reads as broken speech.
Common failure modes:
| Symptom | Likely cause | Fix |
|---|---|---|
Translations contain [SKIP] literally |
Prompt isn't strong enough on the [SKIP] rule |
Reset prompt to default; default has stronger language |
Output has EMOTION|EMOTION|text |
Prompt has the format described twice | Reset prompt |
| Output in wrong language | Prompt's "output entirely in {tgt_lang}" rule was edited out | Reset prompt |
| Refusals leak through ("I cannot translate...") | Nova ignored the prompt; was a known issue, fixed by stronger [SKIP] instruction |
Reset prompt; if it persists, check Nova model ID is correct for region |
/translate).end=True to the WS / gRPC).The session record (segments, timing, language, etc.) persists in memory until the server restarts. There's no automatic cleanup; this is captured as a known issue and will move to hibiki-stage with proper persistence.
/health or container logs) will show "Nova translator: credentials validated" or an error.enable_emotion and enable_diarization aren't set on a non-supporting backend. Diarization on Transcribe is silently no-op; emotion on a stripped prompt produces empty strings.tgt_lang in the request. The settings UI shows the current default but per-stream gRPC config overrides it.GPU_ASG_NAME is set in env. Without it, the GPU manager is disabled and shows "unmanaged".This was a known bug fixed in June 2026: /translate/config was reading the env var instead of the runtime state. If you see this on a deploy older than that, redeploy.
Profile changes apply on the next stream open. Existing streams continue with the old config. Reconnect to pick up changes.
grpc-integration.md — for clients consuming the gRPC APIapi-reference.md — full proto and HTTP endpoint referencedeployment.md — CDK deploy modes, infrastructurearchitecture.md — how the engine relates to hibiki-stage and hibiki-tts../HIBIKI_STAGE_PLAN.md — the broader plan to extract the operator/audience UX into its own app../CHANNEL_FEATURE_REVIEW.md — known issues being deferred to that extraction