← Documentation index

Architecture

Audience: new maintainers, security reviewers, future architects, anyone trying to understand how hibiki-translate fits in the broader system.

This is the high-level orientation. For specifics, follow the cross-references at the end.


What hibiki-translate is

A streaming translation microservice. Audio in, translated text + emotion + speaker diarization out.

It is not: - An end-user application - A TTS engine - An audience-broadcast system - A session-storage / history system

The audience-facing UX, channel switching, session export, TTS audio synthesis, and similar product workflow are the responsibility of consumer apps (hibiki-stage, future HibikiVoice, HibikiDub, etc.). hibiki-translate is the engine they consume.

This separation is recent — see CHANNEL_FEATURE_REVIEW.md and HIBIKI_STAGE_PLAN.md for the architectural reasoning and the extraction plan.


The system as a whole

┌─────────────────────────────────────────────────────────────────┐
│                                                                 │
│              ╔═══════════════════════════════╗                  │
│              ║  hibiki-infra (CDK)    ║                  │
│              ║                               ║                  │
│              ║  - Cognito user pool          ║                  │
│              ║  - VPC                        ║                  │
│              ║  - Route53 zone               ║                  │
│              ║  - ACM SAN cert (us-east-1)   ║                  │
│              ║  - Cloud Map namespace        ║                  │
│              ║                               ║                  │
│              ╚═══════════════════════════════╝                  │
│                       △                  △                      │
│                       │ refs             │ refs                 │
│                       │                  │                      │
│   ┌───────────────────┴───────┐    ┌─────┴─────────────────────┐│
│   │  hibiki-translate         │    │  hibiki-stage              ││
│   │  (this engine)            │    │  (operator + audience UX)  ││
│   │                           │    │                            ││
│   │  - STT                    │    │  - operator UI             ││
│   │  - pre-buffer             │◀───│  - audience UI             ││
│   │  - Nova translation       │gRPC│  - channel switching       ││
│   │  - post-stitcher          │    │  - session storage         ││
│   │  - emotion                │    │  - history / CSV / SRT     ││
│   │  - diarization            │    │  - Polly TTS (Phase 1)     ││
│   │  - settings UI            │    │    → hibiki-tts (Phase 3)  ││
│   │                           │    │                            ││
│   └───────────────────────────┘    └────────────────────────────┘│
│                                                  │               │
│                                                  │ gRPC          │
│                                                  ▼ (Phase 3+)    │
│                              ┌─────────────────────────────────┐ │
│                              │  hibiki-tts                      │ │
│                              │  (multi-engine TTS)              │ │
│                              │  - Qwen3 (GPU)                   │ │
│                              │  - Polly                         │ │
│                              │  - MiniMax (planned)             │ │
│                              └─────────────────────────────────┘ │
│                                                                 │
└─────────────────────────────────────────────────────────────────┘

                  Tokyo (testing) and LA (production)
                  share this architecture; differ only in
                  CDK -c region context var

Three repositories, plus the shared-infra repo:

Repo Role Public surface
hibiki-infra Foundation. Owns Cognito, VPC, DNS, certs. Deploys first. None. CfnOutputs only.
hibiki-translate (this repo) Translation engine. Audio → translated text events. gRPC port 50053 (internal). HTTP for settings UI (CloudFront-fronted).
hibiki-stage Operator + audience UX. The "live event" application. HTTPS via CloudFront on stage.liveprod.cloud.
hibiki-tts Multi-engine TTS microservice. gRPC (when production-ready; currently WIP).

What lives inside hibiki-translate

hibiki-translate/
├── pipeline/                  # The actual engine
│   ├── stt.py                 # WhisperLiveKit wrapper
│   ├── stt_transcribe.py      # AWS Transcribe wrapper
│   ├── pre_buffer.py          # Language-aware audio buffering
│   ├── translator_nova.py     # Bedrock Nova Lite client
│   ├── post_stitcher.py       # Sentence assembly
│   ├── language_profiles.py   # Per-language-pair config
│   ├── profile_store.py       # DynamoDB-backed profile persistence
│   ├── emotion.py             # Emotion classifier (rule-based backup)
│   ├── hallucination_filter.py # STT artifact + filler word filtering
│   └── gpu_manager.py         # ASG-based GPU lifecycle
│
├── server.py                  # gRPC TranslationService implementation
├── session.py                 # Per-stream pipeline lifecycle
├── web_server.py              # FastAPI: HTTP endpoints + legacy WS path
│
├── protos/translation.proto   # gRPC service definition (source of truth)
├── protos/generated/          # protoc-generated stubs
│
├── web/
│   ├── live_transcription.html # Operator UI (legacy; will trim during Phase 2)
│   ├── live_transcription.js   # Operator UI logic
│   ├── live_transcription.css  # Operator UI styles
│   └── settings.html           # Pipeline tuning UI (engine-layer)
│
├── docs/                      # This documentation set
│
└── cdk/
    ├── app.py                 # CDK app entry point
    └── stacks/
        ├── shared.py          # Shared CDK constructs (Cognito ref, VPC, etc.)
        ├── fargate_stack.py   # No-GPU deploy mode
        ├── ecs_gpu_stack.py   # Single-instance GPU deploy mode
        └── split_stack.py     # Production deploy mode

The pipeline/ directory is the heart. Everything else either exposes the pipeline (server.py, web_server.py, protos), tunes it (settings.html, language_profiles.py), or deploys it (cdk/).


The translation pipeline

The data flow inside the engine:

audio (16kHz float32 PCM)
   │
   ▼
┌─────────────────┐
│  STT            │  WhisperLiveKit (GPU, with diarization)
│                 │  OR AWS Transcribe (no GPU)
└────────┬────────┘
         │ confirmed words + speaker_id + timing
         ▼
┌─────────────────┐
│  Pre-buffer     │  Accumulate until min_chars/min_words AND
│                 │  smart-chunking allows fire (or timeout/max-time)
└────────┬────────┘
         │ chunk text, speaker, timing
         ▼
┌─────────────────┐
│  Hallucination  │  Filter Whisper artifacts and pure filler words
│  filter         │
└────────┬────────┘
         │ chunk text (filtered)
         ▼
┌─────────────────┐
│  Nova Lite      │  Bedrock Converse Stream API
│  (translator)   │  Output: EMOTION|translation_text or [SKIP]
└────────┬────────┘
         │ translated text + emotion
         ▼
┌─────────────────┐
│  Post-stitcher  │  Assemble target-language complete sentences
│                 │  before emitting (TTS-ready)
└────────┬────────┘
         │ complete sentence
         ▼
   gRPC TranslatedSentence event
   (or WebSocket JSON for legacy clients)

Each stage is configurable per language pair. See operator-guide.md for details on each.


Why is the engine so heavy?

It does a surprising amount of work for a "translation service":

Concern Why the engine handles it
STT Translation has to start somewhere; coupling STT to the buffer/translation logic gives consistent latency
Language-pair-aware buffering Naive translation of partial sentences produces bad output (esp. SVO→SOV); this is genuinely the engine's job
Smart chunking Same reason: language-specific knowledge
Post-stitching Naive output is fragmented; consumers need TTS-ready sentences
Diarization Optional; produced by WLK; passed through transparently
Emotion classification Optional; produced by Nova inline; passed through transparently
GPU lifecycle Required because WLK runs in this engine's deployment

What's NOT the engine's job (and we've been removing where it crept in): - ❌ Audience broadcast (channel feature → moves to hibiki-stage) - ❌ Session storage / history / export (→ moves to hibiki-stage) - ❌ TTS synthesis (→ moves to hibiki-stage Phase 1, then hibiki-tts Phase 3) - ❌ Operator UI for orchestrating events (→ moves to hibiki-stage) - ❌ QR code sharing, audience auth, channel routing (→ moves to hibiki-stage)

The engine emits structured translation events. What the consumer does with them — display, broadcast, persist, synthesize, route — is the consumer's call.


Translation contract

Two ways to consume the engine:

Bidirectional streaming RPC. Audio in, translation events out. See grpc-integration.md.

This is what hibiki-stage will use. It's the future. New consumers should always start here.

2. WebSocket /translate/asr (legacy)

The pre-gRPC contract. Audio in over a binary WS, translation events out as JSON. Used by: - The current live_transcription.html operator UI - HibikiVoice's existing client (likely)

This will be removed once hibiki-stage is the production consumer (Phase 2 of HIBIKI_STAGE_PLAN.md). New consumers should not use it.


Auth model

Surface Auth
Operator UI (HTTPS) Cognito JWT validation in _validate_token() (currently only enforced on the WS — soft spot)
WebSocket /translate/asr JWT in ?token= query param; validated against shared Cognito pool
gRPC :50053 None. Trust boundary is the network (VPC + NLB security groups). Internal-only.
HTTP REST endpoints None at code level. Settings page relies on the login overlay; this is a soft spot to be fixed when extracting to hibiki-stage (which will do per-endpoint validation).

The shared Cognito pool means an operator who logs into one app (e.g. hibiki-stage) has a JWT that works at another (e.g. hibiki-translate's settings page) without re-authenticating.

See HIBIKI_STAGE_PLAN.md "How auth actually works end-to-end" for the full diagram.


Translation pipeline customization

Every step is configurable, and most settings are per-language-pair:

Layer Configurable via Persistence
STT engine /translate/set-stt-engine API or settings UI DynamoDB
Language pair /translate/set-languages API or settings UI DynamoDB
Pre-buffer thresholds /translate/update-profile API or settings UI DynamoDB
Smart-chunking word lists /translate/incomplete-endings API or settings UI DynamoDB
Post-stitcher /translate/update-profile API or settings UI DynamoDB
Nova prompts /translate/update-prompt API or settings UI DynamoDB
GPU lifecycle /translate/gpu/* API or settings UI Live state (ASG)

All persisted settings live in HibikiTranslateProfiles DynamoDB table. Consumer apps don't read or write this directly — they pass per-stream config via gRPC StreamConfig.


What's deliberately not here

A handful of architectural decisions are explicitly captured as "we considered this and chose not to do it":

Considered Not done Where documented
Tool-use refactor for Nova output Deferred until current prompt-based contract starts breaking TOOLUSE_REFACTOR.md
Single shared CloudFront for both apps Per-app CloudFront chosen instead HIBIKI_STAGE_PLAN.md "Why per-app CloudFront"
Cross-stack Fn.import_value for inter-app references Convention + override context vars instead HIBIKI_STAGE_PLAN.md "Cross-app references"
In-process region switching Removed entirely; region is deploy-time Engine code (deleted) + CHANNEL_FEATURE_REVIEW.md
Speculative translation (UI typewriter) Removed; bypassed pipeline contract CHANNEL_FEATURE_REVIEW.md item 4
GPU auto-stop on idle Removed; manual lifecycle only CHANNEL_FEATURE_REVIEW.md (and engine code)
Foreign-script post-strip Removed; trust the prompt TOOLUSE_REFACTOR.md decision log
Refusal-pattern regex Removed; trust the prompt TOOLUSE_REFACTOR.md decision log

If you find yourself wanting to add one of these back, read the linked doc first — there's likely a reason it's gone.


Cross-references