← Documentation index

API Reference

Audience: developers needing the raw, exhaustive reference for every gRPC field, HTTP endpoint, and error mode.

For a tutorial-style walk-through with examples and gotchas, see grpc-integration.md. This doc is the field-by-field reference; nothing here is omitted for brevity.


gRPC service: TranslationService

Defined in protos/translation.proto. Package: hibikivoice.translation.

TranslateStream(stream TranslateStreamRequest) → stream TranslateStreamResponse

Bidirectional streaming RPC. Audio in, translation events out. The main contract.

Request: TranslateStreamRequest

A oneof payload — exactly one of these per message:

Field Type Description
config StreamConfig Stream configuration. Must be the first message.
audio bytes DEPRECATED. Raw float32 PCM. Kept for HibikiVoice back-compat. New clients use audio_chunk.
control StreamControl Flush or end signal.
audio_chunk AudioChunk PCM with PTS metadata. Preferred.

StreamConfig

message StreamConfig {
    string src_lang = 1;
    string tgt_lang = 2;
    string session_id = 3;
    bool enable_diarization = 4;
    bool enable_emotion = 5;
    string whisper_model = 6;
    bool skip_post_stitcher = 7;
}
Field Required Default Notes
src_lang yes — ISO code: en, ja, fr, de, es, it, pt, ko, zh, ru.
tgt_lang yes — Same set as src_lang. Server resolves a language-pair profile from the pair.
session_id recommended empty Opaque to server. Logged. Use a UUID for traceability.
enable_diarization optional false Populates speaker_id in output if true AND STT engine supports it (WLK does, Transcribe does not).
enable_emotion optional false Populates emotion and emotion_instruction in output.
whisper_model optional large-v3 One of large-v3, medium, small. WLK only; ignored by Transcribe.
skip_post_stitcher optional false Bypass sentence assembly. Lower latency, less natural output.

AudioChunk

message AudioChunk {
    bytes pcm = 1;
    int64 pts_start_ms = 2;
    int64 pts_end_ms = 3;
}
Field Format
pcm float32 little-endian, mono, 16 kHz. 4 bytes per sample. Recommended chunk: 100ms = 1600 samples = 6400 bytes.
pts_start_ms Source-stream clock (not wall clock), milliseconds.
pts_end_ms Source-stream clock, milliseconds. Should be pts_start_ms + chunk_duration_ms.

Both PTS values may be 0 to indicate "unknown timing"; output will then have src_pts_*_ms = 0 too.

StreamControl

message StreamControl {
    bool flush = 1;
    bool end = 2;
}
Field Effect
flush Drain pre-buffer immediately; emit any pending translation. Useful at speaker pauses.
end End of session. Server emits remaining translations and closes the response stream. Always send before closing the channel for graceful shutdown.

Response: TranslateStreamResponse

A oneof payload:

Field Type When emitted
sentence TranslatedSentence One per committed sentence (or chunk if skip_post_stitcher=true).
partial PartialTranscript In-progress preview. Frequency varies by STT engine and translator.
live LiveTranscriptUpdate Per-WLK-tick (~5Hz) full snapshot of operator-display state. Optional for consumers that don't render a live UI.

TranslatedSentence

message TranslatedSentence {
    string translated_text = 1;
    string src_text = 2;
    int32 sequence = 3;
    int32 speaker_id = 4;
    string speaker_gender = 5;
    string emotion = 6;
    string emotion_instruction = 7;
    float latency_ms = 8;
    string model_used = 9;
    float start_time = 10;
    float end_time = 11;
    int64 src_pts_start_ms = 12;
    int64 src_pts_end_ms = 13;
    bool is_replacement = 14;
    string translation_status = 15;
}
Field Type Range / values
translated_text string Cleaned translation in tgt_lang. No emotion prefix, no [SKIP], no Nova meta-commentary.
src_text string Original source text the translator received. Useful for debugging and display.
sequence int32 Monotonic per stream, starts at 1, increments per sentence.
speaker_id int32 1, 2, ... if diarization on; 0 otherwise.
speaker_gender string "male", "female", or "".
emotion string One of excited, happy, sad, angry, fearful, surprised, neutral, or "" if enable_emotion=false.
emotion_instruction string Free-text TTS instruction. Pipe to hibiki-tts's instruct field.
latency_ms float Server-side translation latency (Nova call duration). Excludes audio capture, STT, pre-buffer wait.
model_used string E.g. nova-2-lite.
start_time float STT-relative seconds. Back-compat. Prefer src_pts_start_ms.
end_time float STT-relative seconds. Back-compat. Prefer src_pts_end_ms.
src_pts_start_ms int64 Echoed-back PTS where this sentence's source audio starts. 0 if unknown.
src_pts_end_ms int64 Echoed-back PTS where this sentence's source audio ends. 0 if unknown.
is_replacement bool When true, this revises the previous sentence (post-stitcher carry-over case). Display clients should replace the prior row in place rather than appending.
translation_status string Outcome of the translator chain race. One of "ok", "fallback_secondary", "fallback_tertiary", "marker", "cache". See grpc-integration.md "Translation chain status" for rendering recommendations.

PartialTranscript

message PartialTranscript {
    string text = 1;
    int32 speaker_id = 2;
}

Treat as advisory, not authoritative. Inconsistent across STT engines.

LiveTranscriptUpdate

message LiveTranscriptUpdate {
    repeated TranscriptLine lines = 1;
    string buffer_transcription = 2;
    string buffer_diarization = 3;
    string buffer_translation = 4;
    string status = 5;
    bool speculative = 6;
    float remaining_time_transcription = 7;
    float remaining_time_diarization = 8;
}

message TranscriptLine {
    int32 speaker = 1;
    string text = 2;
    string translation = 3;
    string speaker_gender = 4;
    float start = 5;
    float end = 6;
    string detected_language = 7;
}
Field Type Range / values
lines[] TranscriptLine[] Ordered: every committed sentence as a line with both text and translation populated, followed by a single trailing line representing the in-flight English. The trailing line is always present (even when empty) so downstream renderers can keep a stable DOM element across commits.
buffer_transcription string WLK's not-yet-confirmed trailing word(s). May oscillate; suppress for visual stability.
buffer_diarization string Speaker label for buffer_transcription.
buffer_translation string Streaming Japanese tokens from the translator's on_partial callback.
status string "active_transcription", "no_audio_detected", or "ready_to_stop".
speculative bool True when a non-primary tier's partial is being rendered.
remaining_time_transcription float WLK's transcription pipeline lag, in seconds.
remaining_time_diarization float WLK's diarization pipeline lag, in seconds.

TranslateText(TranslateTextRequest) → TranslateTextResponse

Unary text-only translation. No audio. Useful for testing and one-shot use cases.

TranslateTextRequest

message TranslateTextRequest {
    string text = 1;
    string src_lang = 2;
    string tgt_lang = 3;
    bool enable_emotion = 4;
}

TranslateTextResponse

message TranslateTextResponse {
    string translated_text = 1;
    string emotion = 2;
    string emotion_instruction = 3;
    float latency_ms = 4;
    string model_used = 5;
}

Identical shape to TranslatedSentence minus the audio-related fields (no sequence, speaker_id, speaker_gender, timestamps, PTS).


HealthCheck(HealthRequest) → HealthResponse

Liveness check. No-op request, returns server status.

message HealthResponse {
    string status = 1;          // "ok", "initializing", or error string
    string models_loaded = 2;   // comma-separated list of loaded models
    string device = 3;          // "cuda:0", "cpu", etc.
}

Use this in load balancer health checks and deploy-readiness probes.


gRPC error codes

grpc.RpcError code() values you may see:

Code Meaning Client action
OK Stream completed normally —
UNAVAILABLE Server unreachable Reconnect with exponential backoff (cap ~30s).
FAILED_PRECONDITION Config not sent first, or server in bad state Don't retry. Surface error.
CANCELLED Stream cancelled by either side Usually no action.
DEADLINE_EXCEEDED RPC deadline exceeded Increase deadline or break into smaller streams.
RESOURCE_EXHAUSTED Server overloaded Backoff + retry.
INTERNAL Server bug Log, surface, file an issue.
INVALID_ARGUMENT Malformed request (e.g. invalid src_lang) Don't retry. Fix the request.

The server doesn't currently distinguish many of these; most failure modes return either UNAVAILABLE (network), FAILED_PRECONDITION (config order), or INTERNAL (anything else).


HTTP API (engine operator surface)

These are non-gRPC HTTP endpoints exposed by web_server.py. They serve the operator UI and runtime configuration. Most of them are not relevant to a gRPC client like hibiki-stage — but documenting here for completeness.

Public

Method Path Purpose
GET /health Liveness check. Returns service uptime, version, metrics. No auth.
GET /translate/config Returns current config: src_lang, tgt_lang, auth_mode, cognito_*, stt_engine. Used by the frontend to populate UI.
GET /translate/metrics Counters: sessions_total, sessions_active, translations_total, translations_timeouts, translations_errors.
GET /translate/chain-metrics Per-tier translator chain stats: attempts, successes, timeouts, errors, p50/p99 latency. Useful for monitoring whether Sonnet vs. Nova vs. AWS Translate is currently winning the hedged race.
GET /translate/languages Lists supported language pairs and their categories (svo_svo, svo_sov, etc.).
GET /translate/presets Returns the three built-in profile presets (Snappy / Conversation / Patient) for the current (src_lang, tgt_lang) pair. Settings UI uses this to populate the preset chip buttons.
GET /translate/stt-status Current STT engine and GPU service status.
GET /translate/gpu/status Current GPU instance status (stopped, starting, running, etc.).

Engine configuration (auth required when AUTH_MODE=normal)

Method Path Purpose
POST /translate/set-languages Change current src_lang / tgt_lang globally. Body: {"src_lang": "en", "tgt_lang": "ja"}.
POST /translate/set-stt-engine Switch STT engine: whisperlivekit or transcribe. Triggers GPU lifecycle.
POST /translate/gpu/start Manual GPU start.
POST /translate/gpu/stop Manual GPU stop.
POST /translate/wlk-service/start Start WLK service via ECS (legacy parallel control plane to gpu/start; same physical resource).
POST /translate/wlk-service/stop Stop WLK service via ECS.
GET /translate/pipeline-config Returns full pipeline config for all language pairs (pre-buffer, post-stitcher, prompts).
POST /translate/update-profile Update pipeline profile for a language pair. Body: {"src_lang": "en", "tgt_lang": "ja", "pre_buffer": {...}, "post_stitcher": {...}}.
POST /translate/reset-profile Reset profile to category default for a language pair.
GET /translate/incomplete-endings Get incomplete-phrase endings for smart-chunking, by source language. Query param: ?lang=en.
POST /translate/incomplete-endings Update endings list.
POST /translate/reset-incomplete-endings Reset to defaults.
GET /translate/sentence-end Get the sentence-end regex for a target language. Query param: ?lang=ja.
POST /translate/sentence-end Override the sentence-end regex for a target language.
POST /translate/reset-sentence-end Reset sentence-end regex to default.
POST /translate/update-prompt Update the translator system prompt. Body: {"system_prompt": "...", "system_prompt_with_context": "..."}. Both Sonnet (primary) and Nova (secondary) read these.
POST /translate/reset-prompt Reset prompt to default.
POST /translate/test-backend Smoke-test the configured Bedrock backend. Returns latency + sample translation.

WebSocket endpoint (legacy / dev)

Path Purpose
WS /translate/asr Audio in, translation events out. Pre-gRPC contract. Kept for development and the existing operator UI. New consumers should use the gRPC TranslateStream instead. Will be removed once hibiki-stage is the production consumer (Phase 2 of HIBIKI_STAGE_PLAN.md).

Pages

Path Purpose
GET / or GET /translate Operator UI (web/live_transcription.html).
GET /translate/settings Settings panel (web/settings.html).
GET /translate/docs Renders the markdown docs from /docs/.

Environment variables

Set on the server, not the client. Documented here for clients that may want to detect server config (via /translate/config).

Variable Default Purpose
STT_ENGINE whisperlivekit whisperlivekit (GPU) or transcribe (no GPU).
TRANSLATOR nova Translation backend.
NOVA_MODEL_ID region-dependent Bedrock model ID.
AUTH_MODE dev dev (no auth) or normal (Cognito JWT validation on /translate/asr).
COGNITO_USER_POOL_ID — Cognito pool for auth.
COGNITO_CLIENT_ID — Cognito app client.
AWS_REGION ap-northeast-1 Bedrock + Transcribe + Polly + ECS region.
WLK_MODEL large-v3 Whisper model size (WLK mode).
WLK_LANGUAGE en Default source language for STT.
TGT_LANG ja Default target language.
GPU_ASG_NAME — ASG name for GPU lifecycle management.

See deployment.md for which CDK context vars produce these env vars at deploy time.


Versioning

The proto file is the source of truth. Treat field numbers as immutable; never reuse a deprecated number. Adding new optional fields is backward-compatible; renaming or removing fields is not.

Current proto version: see git log -- protos/translation.proto for history.

When breaking changes are necessary, mark the old field with // DEPRECATED (as audio already is) and add a new one rather than mutating the existing field.