Audience: developers needing the raw, exhaustive reference for every gRPC field, HTTP endpoint, and error mode.
For a tutorial-style walk-through with examples and gotchas, see grpc-integration.md. This doc is the field-by-field reference; nothing here is omitted for brevity.
TranslationServiceDefined in protos/translation.proto. Package: hibikivoice.translation.
TranslateStream(stream TranslateStreamRequest) → stream TranslateStreamResponseBidirectional streaming RPC. Audio in, translation events out. The main contract.
TranslateStreamRequestA oneof payload — exactly one of these per message:
| Field | Type | Description |
|---|---|---|
config |
StreamConfig |
Stream configuration. Must be the first message. |
audio |
bytes |
DEPRECATED. Raw float32 PCM. Kept for HibikiVoice back-compat. New clients use audio_chunk. |
control |
StreamControl |
Flush or end signal. |
audio_chunk |
AudioChunk |
PCM with PTS metadata. Preferred. |
StreamConfigmessage StreamConfig {
string src_lang = 1;
string tgt_lang = 2;
string session_id = 3;
bool enable_diarization = 4;
bool enable_emotion = 5;
string whisper_model = 6;
bool skip_post_stitcher = 7;
}
| Field | Required | Default | Notes |
|---|---|---|---|
src_lang |
yes | — | ISO code: en, ja, fr, de, es, it, pt, ko, zh, ru. |
tgt_lang |
yes | — | Same set as src_lang. Server resolves a language-pair profile from the pair. |
session_id |
recommended | empty | Opaque to server. Logged. Use a UUID for traceability. |
enable_diarization |
optional | false |
Populates speaker_id in output if true AND STT engine supports it (WLK does, Transcribe does not). |
enable_emotion |
optional | false |
Populates emotion and emotion_instruction in output. |
whisper_model |
optional | large-v3 |
One of large-v3, medium, small. WLK only; ignored by Transcribe. |
skip_post_stitcher |
optional | false |
Bypass sentence assembly. Lower latency, less natural output. |
AudioChunkmessage AudioChunk {
bytes pcm = 1;
int64 pts_start_ms = 2;
int64 pts_end_ms = 3;
}
| Field | Format |
|---|---|
pcm |
float32 little-endian, mono, 16 kHz. 4 bytes per sample. Recommended chunk: 100ms = 1600 samples = 6400 bytes. |
pts_start_ms |
Source-stream clock (not wall clock), milliseconds. |
pts_end_ms |
Source-stream clock, milliseconds. Should be pts_start_ms + chunk_duration_ms. |
Both PTS values may be 0 to indicate "unknown timing"; output will then have src_pts_*_ms = 0 too.
StreamControlmessage StreamControl {
bool flush = 1;
bool end = 2;
}
| Field | Effect |
|---|---|
flush |
Drain pre-buffer immediately; emit any pending translation. Useful at speaker pauses. |
end |
End of session. Server emits remaining translations and closes the response stream. Always send before closing the channel for graceful shutdown. |
TranslateStreamResponseA oneof payload:
| Field | Type | When emitted |
|---|---|---|
sentence |
TranslatedSentence |
One per committed sentence (or chunk if skip_post_stitcher=true). |
partial |
PartialTranscript |
In-progress preview. Frequency varies by STT engine and translator. |
live |
LiveTranscriptUpdate |
Per-WLK-tick (~5Hz) full snapshot of operator-display state. Optional for consumers that don't render a live UI. |
TranslatedSentencemessage TranslatedSentence {
string translated_text = 1;
string src_text = 2;
int32 sequence = 3;
int32 speaker_id = 4;
string speaker_gender = 5;
string emotion = 6;
string emotion_instruction = 7;
float latency_ms = 8;
string model_used = 9;
float start_time = 10;
float end_time = 11;
int64 src_pts_start_ms = 12;
int64 src_pts_end_ms = 13;
bool is_replacement = 14;
string translation_status = 15;
}
| Field | Type | Range / values |
|---|---|---|
translated_text |
string | Cleaned translation in tgt_lang. No emotion prefix, no [SKIP], no Nova meta-commentary. |
src_text |
string | Original source text the translator received. Useful for debugging and display. |
sequence |
int32 | Monotonic per stream, starts at 1, increments per sentence. |
speaker_id |
int32 | 1, 2, ... if diarization on; 0 otherwise. |
speaker_gender |
string | "male", "female", or "". |
emotion |
string | One of excited, happy, sad, angry, fearful, surprised, neutral, or "" if enable_emotion=false. |
emotion_instruction |
string | Free-text TTS instruction. Pipe to hibiki-tts's instruct field. |
latency_ms |
float | Server-side translation latency (Nova call duration). Excludes audio capture, STT, pre-buffer wait. |
model_used |
string | E.g. nova-2-lite. |
start_time |
float | STT-relative seconds. Back-compat. Prefer src_pts_start_ms. |
end_time |
float | STT-relative seconds. Back-compat. Prefer src_pts_end_ms. |
src_pts_start_ms |
int64 | Echoed-back PTS where this sentence's source audio starts. 0 if unknown. |
src_pts_end_ms |
int64 | Echoed-back PTS where this sentence's source audio ends. 0 if unknown. |
is_replacement |
bool | When true, this revises the previous sentence (post-stitcher carry-over case). Display clients should replace the prior row in place rather than appending. |
translation_status |
string | Outcome of the translator chain race. One of "ok", "fallback_secondary", "fallback_tertiary", "marker", "cache". See grpc-integration.md "Translation chain status" for rendering recommendations. |
PartialTranscriptmessage PartialTranscript {
string text = 1;
int32 speaker_id = 2;
}
Treat as advisory, not authoritative. Inconsistent across STT engines.
LiveTranscriptUpdatemessage LiveTranscriptUpdate {
repeated TranscriptLine lines = 1;
string buffer_transcription = 2;
string buffer_diarization = 3;
string buffer_translation = 4;
string status = 5;
bool speculative = 6;
float remaining_time_transcription = 7;
float remaining_time_diarization = 8;
}
message TranscriptLine {
int32 speaker = 1;
string text = 2;
string translation = 3;
string speaker_gender = 4;
float start = 5;
float end = 6;
string detected_language = 7;
}
| Field | Type | Range / values |
|---|---|---|
lines[] |
TranscriptLine[] |
Ordered: every committed sentence as a line with both text and translation populated, followed by a single trailing line representing the in-flight English. The trailing line is always present (even when empty) so downstream renderers can keep a stable DOM element across commits. |
buffer_transcription |
string | WLK's not-yet-confirmed trailing word(s). May oscillate; suppress for visual stability. |
buffer_diarization |
string | Speaker label for buffer_transcription. |
buffer_translation |
string | Streaming Japanese tokens from the translator's on_partial callback. |
status |
string | "active_transcription", "no_audio_detected", or "ready_to_stop". |
speculative |
bool | True when a non-primary tier's partial is being rendered. |
remaining_time_transcription |
float | WLK's transcription pipeline lag, in seconds. |
remaining_time_diarization |
float | WLK's diarization pipeline lag, in seconds. |
TranslateText(TranslateTextRequest) → TranslateTextResponseUnary text-only translation. No audio. Useful for testing and one-shot use cases.
TranslateTextRequestmessage TranslateTextRequest {
string text = 1;
string src_lang = 2;
string tgt_lang = 3;
bool enable_emotion = 4;
}
TranslateTextResponsemessage TranslateTextResponse {
string translated_text = 1;
string emotion = 2;
string emotion_instruction = 3;
float latency_ms = 4;
string model_used = 5;
}
Identical shape to TranslatedSentence minus the audio-related fields (no sequence, speaker_id, speaker_gender, timestamps, PTS).
HealthCheck(HealthRequest) → HealthResponseLiveness check. No-op request, returns server status.
message HealthResponse {
string status = 1; // "ok", "initializing", or error string
string models_loaded = 2; // comma-separated list of loaded models
string device = 3; // "cuda:0", "cpu", etc.
}
Use this in load balancer health checks and deploy-readiness probes.
grpc.RpcError code() values you may see:
| Code | Meaning | Client action |
|---|---|---|
OK |
Stream completed normally | — |
UNAVAILABLE |
Server unreachable | Reconnect with exponential backoff (cap ~30s). |
FAILED_PRECONDITION |
Config not sent first, or server in bad state | Don't retry. Surface error. |
CANCELLED |
Stream cancelled by either side | Usually no action. |
DEADLINE_EXCEEDED |
RPC deadline exceeded | Increase deadline or break into smaller streams. |
RESOURCE_EXHAUSTED |
Server overloaded | Backoff + retry. |
INTERNAL |
Server bug | Log, surface, file an issue. |
INVALID_ARGUMENT |
Malformed request (e.g. invalid src_lang) |
Don't retry. Fix the request. |
The server doesn't currently distinguish many of these; most failure modes return either UNAVAILABLE (network), FAILED_PRECONDITION (config order), or INTERNAL (anything else).
These are non-gRPC HTTP endpoints exposed by web_server.py. They serve the operator UI and runtime configuration. Most of them are not relevant to a gRPC client like hibiki-stage — but documenting here for completeness.
| Method | Path | Purpose |
|---|---|---|
GET |
/health |
Liveness check. Returns service uptime, version, metrics. No auth. |
GET |
/translate/config |
Returns current config: src_lang, tgt_lang, auth_mode, cognito_*, stt_engine. Used by the frontend to populate UI. |
GET |
/translate/metrics |
Counters: sessions_total, sessions_active, translations_total, translations_timeouts, translations_errors. |
GET |
/translate/chain-metrics |
Per-tier translator chain stats: attempts, successes, timeouts, errors, p50/p99 latency. Useful for monitoring whether Sonnet vs. Nova vs. AWS Translate is currently winning the hedged race. |
GET |
/translate/languages |
Lists supported language pairs and their categories (svo_svo, svo_sov, etc.). |
GET |
/translate/presets |
Returns the three built-in profile presets (Snappy / Conversation / Patient) for the current (src_lang, tgt_lang) pair. Settings UI uses this to populate the preset chip buttons. |
GET |
/translate/stt-status |
Current STT engine and GPU service status. |
GET |
/translate/gpu/status |
Current GPU instance status (stopped, starting, running, etc.). |
| Method | Path | Purpose |
|---|---|---|
POST |
/translate/set-languages |
Change current src_lang / tgt_lang globally. Body: {"src_lang": "en", "tgt_lang": "ja"}. |
POST |
/translate/set-stt-engine |
Switch STT engine: whisperlivekit or transcribe. Triggers GPU lifecycle. |
POST |
/translate/gpu/start |
Manual GPU start. |
POST |
/translate/gpu/stop |
Manual GPU stop. |
POST |
/translate/wlk-service/start |
Start WLK service via ECS (legacy parallel control plane to gpu/start; same physical resource). |
POST |
/translate/wlk-service/stop |
Stop WLK service via ECS. |
GET |
/translate/pipeline-config |
Returns full pipeline config for all language pairs (pre-buffer, post-stitcher, prompts). |
POST |
/translate/update-profile |
Update pipeline profile for a language pair. Body: {"src_lang": "en", "tgt_lang": "ja", "pre_buffer": {...}, "post_stitcher": {...}}. |
POST |
/translate/reset-profile |
Reset profile to category default for a language pair. |
GET |
/translate/incomplete-endings |
Get incomplete-phrase endings for smart-chunking, by source language. Query param: ?lang=en. |
POST |
/translate/incomplete-endings |
Update endings list. |
POST |
/translate/reset-incomplete-endings |
Reset to defaults. |
GET |
/translate/sentence-end |
Get the sentence-end regex for a target language. Query param: ?lang=ja. |
POST |
/translate/sentence-end |
Override the sentence-end regex for a target language. |
POST |
/translate/reset-sentence-end |
Reset sentence-end regex to default. |
POST |
/translate/update-prompt |
Update the translator system prompt. Body: {"system_prompt": "...", "system_prompt_with_context": "..."}. Both Sonnet (primary) and Nova (secondary) read these. |
POST |
/translate/reset-prompt |
Reset prompt to default. |
POST |
/translate/test-backend |
Smoke-test the configured Bedrock backend. Returns latency + sample translation. |
| Path | Purpose |
|---|---|
WS /translate/asr |
Audio in, translation events out. Pre-gRPC contract. Kept for development and the existing operator UI. New consumers should use the gRPC TranslateStream instead. Will be removed once hibiki-stage is the production consumer (Phase 2 of HIBIKI_STAGE_PLAN.md). |
| Path | Purpose |
|---|---|
GET / or GET /translate |
Operator UI (web/live_transcription.html). |
GET /translate/settings |
Settings panel (web/settings.html). |
GET /translate/docs |
Renders the markdown docs from /docs/. |
Set on the server, not the client. Documented here for clients that may want to detect server config (via /translate/config).
| Variable | Default | Purpose |
|---|---|---|
STT_ENGINE |
whisperlivekit |
whisperlivekit (GPU) or transcribe (no GPU). |
TRANSLATOR |
nova |
Translation backend. |
NOVA_MODEL_ID |
region-dependent | Bedrock model ID. |
AUTH_MODE |
dev |
dev (no auth) or normal (Cognito JWT validation on /translate/asr). |
COGNITO_USER_POOL_ID |
— | Cognito pool for auth. |
COGNITO_CLIENT_ID |
— | Cognito app client. |
AWS_REGION |
ap-northeast-1 |
Bedrock + Transcribe + Polly + ECS region. |
WLK_MODEL |
large-v3 |
Whisper model size (WLK mode). |
WLK_LANGUAGE |
en |
Default source language for STT. |
TGT_LANG |
ja |
Default target language. |
GPU_ASG_NAME |
— | ASG name for GPU lifecycle management. |
See deployment.md for which CDK context vars produce these env vars at deploy time.
The proto file is the source of truth. Treat field numbers as immutable; never reuse a deprecated number. Adding new optional fields is backward-compatible; renaming or removing fields is not.
Current proto version: see git log -- protos/translation.proto for history.
When breaking changes are necessary, mark the old field with // DEPRECATED (as audio already is) and add a new one rather than mutating the existing field.