Audience: new maintainers, security reviewers, future architects, anyone trying to understand how hibiki-translate fits in the broader system.
This is the high-level orientation. For specifics, follow the cross-references at the end.
A streaming translation microservice. Audio in, translated text + emotion + speaker diarization out.
It is not: - An end-user application - A TTS engine - An audience-broadcast system - A session-storage / history system
The audience-facing UX, channel switching, session export, TTS audio synthesis, and similar product workflow are the responsibility of consumer apps (hibiki-stage, future HibikiVoice, HibikiDub, etc.). hibiki-translate is the engine they consume.
This separation is recent — see CHANNEL_FEATURE_REVIEW.md and HIBIKI_STAGE_PLAN.md for the architectural reasoning and the extraction plan.
┌─────────────────────────────────────────────────────────────────┐
│ │
│ ╔═══════════════════════════════╗ │
│ ║ hibiki-infra (CDK) ║ │
│ ║ ║ │
│ ║ - Cognito user pool ║ │
│ ║ - VPC ║ │
│ ║ - Route53 zone ║ │
│ ║ - ACM SAN cert (us-east-1) ║ │
│ ║ - Cloud Map namespace ║ │
│ ║ ║ │
│ ╚═══════════════════════════════╝ │
│ △ △ │
│ │ refs │ refs │
│ │ │ │
│ ┌───────────────────┴───────┐ ┌─────┴─────────────────────┐│
│ │ hibiki-translate │ │ hibiki-stage ││
│ │ (this engine) │ │ (operator + audience UX) ││
│ │ │ │ ││
│ │ - STT │ │ - operator UI ││
│ │ - pre-buffer │◀───│ - audience UI ││
│ │ - Nova translation │gRPC│ - channel switching ││
│ │ - post-stitcher │ │ - session storage ││
│ │ - emotion │ │ - history / CSV / SRT ││
│ │ - diarization │ │ - Polly TTS (Phase 1) ││
│ │ - settings UI │ │ → hibiki-tts (Phase 3) ││
│ │ │ │ ││
│ └───────────────────────────┘ └────────────────────────────┘│
│ │ │
│ │ gRPC │
│ ▼ (Phase 3+) │
│ ┌─────────────────────────────────┐ │
│ │ hibiki-tts │ │
│ │ (multi-engine TTS) │ │
│ │ - Qwen3 (GPU) │ │
│ │ - Polly │ │
│ │ - MiniMax (planned) │ │
│ └─────────────────────────────────┘ │
│ │
└─────────────────────────────────────────────────────────────────┘
Tokyo (testing) and LA (production)
share this architecture; differ only in
CDK -c region context var
Three repositories, plus the shared-infra repo:
| Repo | Role | Public surface |
|---|---|---|
hibiki-infra |
Foundation. Owns Cognito, VPC, DNS, certs. Deploys first. | None. CfnOutputs only. |
hibiki-translate (this repo) |
Translation engine. Audio → translated text events. | gRPC port 50053 (internal). HTTP for settings UI (CloudFront-fronted). |
hibiki-stage |
Operator + audience UX. The "live event" application. | HTTPS via CloudFront on stage.liveprod.cloud. |
hibiki-tts |
Multi-engine TTS microservice. | gRPC (when production-ready; currently WIP). |
hibiki-translate/
├── pipeline/ # The actual engine
│ ├── stt.py # WhisperLiveKit wrapper
│ ├── stt_transcribe.py # AWS Transcribe wrapper
│ ├── pre_buffer.py # Language-aware audio buffering
│ ├── translator_nova.py # Bedrock Nova Lite client
│ ├── post_stitcher.py # Sentence assembly
│ ├── language_profiles.py # Per-language-pair config
│ ├── profile_store.py # DynamoDB-backed profile persistence
│ ├── emotion.py # Emotion classifier (rule-based backup)
│ ├── hallucination_filter.py # STT artifact + filler word filtering
│ └── gpu_manager.py # ASG-based GPU lifecycle
│
├── server.py # gRPC TranslationService implementation
├── session.py # Per-stream pipeline lifecycle
├── web_server.py # FastAPI: HTTP endpoints + legacy WS path
│
├── protos/translation.proto # gRPC service definition (source of truth)
├── protos/generated/ # protoc-generated stubs
│
├── web/
│ ├── live_transcription.html # Operator UI (legacy; will trim during Phase 2)
│ ├── live_transcription.js # Operator UI logic
│ ├── live_transcription.css # Operator UI styles
│ └── settings.html # Pipeline tuning UI (engine-layer)
│
├── docs/ # This documentation set
│
└── cdk/
├── app.py # CDK app entry point
└── stacks/
├── shared.py # Shared CDK constructs (Cognito ref, VPC, etc.)
├── fargate_stack.py # No-GPU deploy mode
├── ecs_gpu_stack.py # Single-instance GPU deploy mode
└── split_stack.py # Production deploy mode
The pipeline/ directory is the heart. Everything else either exposes the pipeline (server.py, web_server.py, protos), tunes it (settings.html, language_profiles.py), or deploys it (cdk/).
The data flow inside the engine:
audio (16kHz float32 PCM)
│
▼
┌─────────────────┐
│ STT │ WhisperLiveKit (GPU, with diarization)
│ │ OR AWS Transcribe (no GPU)
└────────┬────────┘
│ confirmed words + speaker_id + timing
▼
┌─────────────────┐
│ Pre-buffer │ Accumulate until min_chars/min_words AND
│ │ smart-chunking allows fire (or timeout/max-time)
└────────┬────────┘
│ chunk text, speaker, timing
▼
┌─────────────────┐
│ Hallucination │ Filter Whisper artifacts and pure filler words
│ filter │
└────────┬────────┘
│ chunk text (filtered)
▼
┌─────────────────┐
│ Nova Lite │ Bedrock Converse Stream API
│ (translator) │ Output: EMOTION|translation_text or [SKIP]
└────────┬────────┘
│ translated text + emotion
▼
┌─────────────────┐
│ Post-stitcher │ Assemble target-language complete sentences
│ │ before emitting (TTS-ready)
└────────┬────────┘
│ complete sentence
▼
gRPC TranslatedSentence event
(or WebSocket JSON for legacy clients)
Each stage is configurable per language pair. See operator-guide.md for details on each.
It does a surprising amount of work for a "translation service":
| Concern | Why the engine handles it |
|---|---|
| STT | Translation has to start somewhere; coupling STT to the buffer/translation logic gives consistent latency |
| Language-pair-aware buffering | Naive translation of partial sentences produces bad output (esp. SVO→SOV); this is genuinely the engine's job |
| Smart chunking | Same reason: language-specific knowledge |
| Post-stitching | Naive output is fragmented; consumers need TTS-ready sentences |
| Diarization | Optional; produced by WLK; passed through transparently |
| Emotion classification | Optional; produced by Nova inline; passed through transparently |
| GPU lifecycle | Required because WLK runs in this engine's deployment |
What's NOT the engine's job (and we've been removing where it crept in):
- ❌ Audience broadcast (channel feature → moves to hibiki-stage)
- ❌ Session storage / history / export (→ moves to hibiki-stage)
- ❌ TTS synthesis (→ moves to hibiki-stage Phase 1, then hibiki-tts Phase 3)
- ❌ Operator UI for orchestrating events (→ moves to hibiki-stage)
- ❌ QR code sharing, audience auth, channel routing (→ moves to hibiki-stage)
The engine emits structured translation events. What the consumer does with them — display, broadcast, persist, synthesize, route — is the consumer's call.
Two ways to consume the engine:
TranslateStream (recommended for new consumers)Bidirectional streaming RPC. Audio in, translation events out. See grpc-integration.md.
This is what hibiki-stage will use. It's the future. New consumers should always start here.
/translate/asr (legacy)The pre-gRPC contract. Audio in over a binary WS, translation events out as JSON. Used by:
- The current live_transcription.html operator UI
- HibikiVoice's existing client (likely)
This will be removed once hibiki-stage is the production consumer (Phase 2 of HIBIKI_STAGE_PLAN.md). New consumers should not use it.
| Surface | Auth |
|---|---|
| Operator UI (HTTPS) | Cognito JWT validation in _validate_token() (currently only enforced on the WS — soft spot) |
WebSocket /translate/asr |
JWT in ?token= query param; validated against shared Cognito pool |
gRPC :50053 |
None. Trust boundary is the network (VPC + NLB security groups). Internal-only. |
| HTTP REST endpoints | None at code level. Settings page relies on the login overlay; this is a soft spot to be fixed when extracting to hibiki-stage (which will do per-endpoint validation). |
The shared Cognito pool means an operator who logs into one app (e.g. hibiki-stage) has a JWT that works at another (e.g. hibiki-translate's settings page) without re-authenticating.
See HIBIKI_STAGE_PLAN.md "How auth actually works end-to-end" for the full diagram.
Every step is configurable, and most settings are per-language-pair:
| Layer | Configurable via | Persistence |
|---|---|---|
| STT engine | /translate/set-stt-engine API or settings UI |
DynamoDB |
| Language pair | /translate/set-languages API or settings UI |
DynamoDB |
| Pre-buffer thresholds | /translate/update-profile API or settings UI |
DynamoDB |
| Smart-chunking word lists | /translate/incomplete-endings API or settings UI |
DynamoDB |
| Post-stitcher | /translate/update-profile API or settings UI |
DynamoDB |
| Nova prompts | /translate/update-prompt API or settings UI |
DynamoDB |
| GPU lifecycle | /translate/gpu/* API or settings UI |
Live state (ASG) |
All persisted settings live in HibikiTranslateProfiles DynamoDB table. Consumer apps don't read or write this directly — they pass per-stream config via gRPC StreamConfig.
A handful of architectural decisions are explicitly captured as "we considered this and chose not to do it":
| Considered | Not done | Where documented |
|---|---|---|
| Tool-use refactor for Nova output | Deferred until current prompt-based contract starts breaking | TOOLUSE_REFACTOR.md |
| Single shared CloudFront for both apps | Per-app CloudFront chosen instead | HIBIKI_STAGE_PLAN.md "Why per-app CloudFront" |
Cross-stack Fn.import_value for inter-app references |
Convention + override context vars instead | HIBIKI_STAGE_PLAN.md "Cross-app references" |
| In-process region switching | Removed entirely; region is deploy-time | Engine code (deleted) + CHANNEL_FEATURE_REVIEW.md |
| Speculative translation (UI typewriter) | Removed; bypassed pipeline contract | CHANNEL_FEATURE_REVIEW.md item 4 |
| GPU auto-stop on idle | Removed; manual lifecycle only | CHANNEL_FEATURE_REVIEW.md (and engine code) |
| Foreign-script post-strip | Removed; trust the prompt | TOOLUSE_REFACTOR.md decision log |
| Refusal-pattern regex | Removed; trust the prompt | TOOLUSE_REFACTOR.md decision log |
If you find yourself wanting to add one of these back, read the linked doc first — there's likely a reason it's gone.
grpc-integration.md — for clients consuming the gRPC contractapi-reference.md — proto and HTTP endpoint referenceoperator-guide.md — runtime configuration and troubleshootingdeployment.md — CDK modes, infrastructure, deploy order../HIBIKI_STAGE_PLAN.md — extraction plan for the operator/audience UX../CHANNEL_FEATURE_REVIEW.md — known issues being deferred to that extraction; architectural patterns we've removed../TOOLUSE_REFACTOR.md — Nova output-format refactor, when it becomes worth doing../CLAUDE.md — high-level orientation../README.md — public-facing description and quickstart