← Documentation index

gRPC Integration Guide

Audience: developers writing a client (e.g. hibiki-stage) that consumes hibiki-translate over gRPC.

This is the contract document. For the underlying proto definition see api-reference.md. For an executable walk-through, see tests/spike_grpc_streaming_client.py — every code snippet here corresponds to a section of that file.


TL;DR

from protos.generated import translation_pb2, translation_pb2_grpc
import grpc

channel = grpc.aio.insecure_channel("translate.hibiki.local:50053")
stub = translation_pb2_grpc.TranslationServiceStub(channel)

async def request_iter():
    yield translation_pb2.TranslateStreamRequest(
        config=translation_pb2.StreamConfig(src_lang="en", tgt_lang="ja", session_id="...")
    )
    # ... yield AudioChunk messages with float32 16kHz PCM ...
    yield translation_pb2.TranslateStreamRequest(control=translation_pb2.StreamControl(end=True))

async for response in stub.TranslateStream(request_iter()):
    if response.WhichOneof("payload") == "sentence":
        print(response.sentence.translated_text)

That's the shape. Everything below explains the details.


Endpoint

The hibiki-translate gRPC server listens on port 50053. The address depends on the deployment mode:

Deploy mode Address Notes
Local dev localhost:50053 Run python server.py directly
split (production-ish) translate.hibiki.local:50053 Cloud Map private DNS, VPC-internal NLB
ecs-gpu EC2 instance public IP / DNS, port 50053 Public — not recommended for production
fargate Not exposed Fargate mode doesn't run the gRPC tier

In the split deploy, the service is internal-only — accessible from any task in the same VPC, not from the public internet. There is no authentication on the gRPC layer; the trust boundary is the network (VPC + security groups). Don't expose port 50053 publicly.

Use grpc.aio.insecure_channel(...) (no TLS for the internal hop). The shared infra layer terminates TLS at the per-app CloudFront edge for HTTP traffic; gRPC-internal stays plaintext.


RPC methods

The service definition (protos/translation.proto) exposes three RPCs:

RPC Type Purpose
TranslateStream bidi-streaming The main contract. Audio in, translation events out.
TranslateText unary Translate a single text string. Useful for testing without audio.
HealthCheck unary Returns server status, loaded models, device. Use as a liveness check.

99% of the integration work is TranslateStream. The rest of this doc focuses on it.


TranslateStream — the streaming contract

The message ordering is strict

Every TranslateStream call must follow this order:

  1. First message: StreamConfig. If audio arrives before config, the server returns FAILED_PRECONDITION and closes the stream.
  2. Then: zero or more AudioChunk messages (or the deprecated audio field — see below).
  3. Optionally: StreamControl messages to signal flush or end.
yield TranslateStreamRequest(config=StreamConfig(...))
yield TranslateStreamRequest(audio_chunk=AudioChunk(pcm=..., pts_start_ms=..., pts_end_ms=...))
yield TranslateStreamRequest(audio_chunk=AudioChunk(...))
# ...
yield TranslateStreamRequest(control=StreamControl(flush=True))   # end-of-utterance
yield TranslateStreamRequest(control=StreamControl(end=True))     # end-of-session

StreamConfig fields

Field Type Required? Notes
src_lang string yes ISO code: en, ja, fr, etc. Determines STT language and translation source.
tgt_lang string yes Target language. The server resolves a language-pair profile (pre-buffer thresholds, smart-chunking flags) from this pair.
session_id string recommended Opaque to server; included in logs. Use a UUID per stream so you can correlate gRPC traffic with your own traces.
enable_diarization bool optional Default false. When true (and STT engine supports it), TranslatedSentence.speaker_id is populated. WLK STT supports this; AWS Transcribe does not — silently ignored if Transcribe is the configured engine.
enable_emotion bool optional Default false. When true, TranslatedSentence.emotion is populated by the translator.
whisper_model string optional large-v3 (default), medium, small. Only honored by the WLK STT engine.
skip_post_stitcher bool optional Default false. When true, the post-stitcher is bypassed and translations are emitted in chunks rather than complete sentences. Lower latency; less natural output. Recommended only for display-only consumers (e.g. live-captions display) where final-sentence fidelity isn't required.

Per-call profile config (pre-buffer thresholds, smart-chunking, prompts) is not in StreamConfig. The server resolves a profile from (src_lang, tgt_lang) via the language-pair profile registry. To customise per-pair tuning, use POST /translate/update-profile (HTTP; persisted to DynamoDB) before opening the gRPC stream — or accept the defaults.

AudioChunk — the audio format

message AudioChunk {
    bytes pcm = 1;          // float32 PCM samples, mono, 16 kHz
    int64 pts_start_ms = 2; // PTS of the first sample (source-stream clock)
    int64 pts_end_ms = 3;   // PTS of the last sample
}

Required format:

PTS values: - Express timestamps in milliseconds. - The clock source is your stream, not the server's wall clock. Whatever monotonic timeline your audio comes from (a media file's timestamp, a live capture's elapsed time, etc.) is what should be in pts_start_ms / pts_end_ms. - These are echoed back in TranslatedSentence.src_pts_start_ms / src_pts_end_ms, which is how downstream consumers (e.g. a video dub mixer) correlate translation output with the source stream timing. - Setting both to 0 is acceptable but loses timing fidelity. Always populate them if you have them.

Deprecated: the audio field

TranslateStreamRequest.audio (a raw bytes field) is the legacy path. It still works for backward compatibility with HibikiVoice. New clients should always use audio_chunk so PTS information flows through.

StreamControl — flush and end

message StreamControl {
    bool flush = 1;   // end of utterance, drain buffers
    bool end = 2;     // end of session, close the stream
}

Always send end=True before closing the channel — without it, the server may drop the last partial translation.


Receiving translations

The response stream emits TranslateStreamResponse messages with a oneof payload:

message TranslateStreamResponse {
    oneof payload {
        TranslatedSentence sentence = 1;
        PartialTranscript partial = 2;
        LiveTranscriptUpdate live = 3;
    }
}

TranslatedSentence — the main event

Emitted once per committed sentence (or chunk if skip_post_stitcher=True). All fields:

Field Type Meaning
translated_text string The translation, in the target language. Already cleaned (no Nova meta-commentary, no emotion prefix, no [SKIP] markers).
src_text string The original source text the translator saw. Useful for display or debugging.
sequence int32 Monotonic counter starting at 1. Increments per sentence, per stream.
speaker_id int32 1, 2, ... if enable_diarization=True and STT supports it; 0 otherwise.
speaker_gender string "male", "female", or "" if unknown. WLK estimates from F0; Transcribe does not provide it.
emotion string excited, happy, sad, angry, fearful, surprised, neutral if enable_emotion=True; "" otherwise.
emotion_instruction string Free-text TTS instruction (e.g. "speak with warm excitement") — generated alongside the emotion label. Pipe to hibiki-tts's instruct field if doing TTS.
latency_ms float Server-side translation latency (Nova call duration). Excludes audio capture, STT, pre-buffer wait.
model_used string The translator that produced this — e.g. nova-2-lite.
start_time float STT-relative seconds. Kept for back-compat. Prefer src_pts_start_ms.
end_time float STT-relative seconds. Kept for back-compat. Prefer src_pts_end_ms.
src_pts_start_ms int64 Echoed-back PTS of where this sentence's source audio starts. 0 if you didn't set it on AudioChunk.
src_pts_end_ms int64 Echoed-back PTS of where this sentence's source audio ends. 0 if unknown.
is_replacement bool When true, this sentence revises the previously emitted one (the post-stitcher's carry-over case). Display clients should replace the prior row in place rather than appending a new one. Audience-side broadcast uses the same field.
translation_status string One of "ok", "fallback_secondary", "fallback_tertiary", "marker", "cache". Reflects how the translator chain produced this row (see "Translation chain status" below). Lets consumers decide rendering / TTS behavior without parsing model_used.

Translation chain status

The translator chain runs a hedged race: primary tier (default Sonnet) at t=0; secondary (Nova) joins at hedge_delay_s; tertiary (AWS Translate) at 2 × hedge_delay_s. The first non-empty result wins; the rest are cancelled. translation_status tells you which won:

Status Meaning Recommended display
ok Primary tier produced the translation. Highest quality. Default rendering.
fallback_secondary Primary timed out / errored / returned empty; secondary won. Default rendering — translation is still good, just from a faster fallback.
fallback_tertiary Both primary and secondary failed; tertiary won. Render normally; optionally surface degradation hint to operator.
marker All hedges failed within chain_deadline_s. translated_text is the configured marker (default "…"). Render with degraded styling (grey). Do NOT TTS-speak this.
cache Translation came from the per-session cache (identical source seen earlier in the stream). No tier was contacted. Default rendering. Useful for cost accounting (no Bedrock call billed).

PartialTranscript — the in-progress preview

message PartialTranscript {
    string text = 1;
    int32 speaker_id = 2;
}

Partial transcripts are emitted while a translation is being assembled — for example, the streaming Nova on_partial callback fires per token. They're useful for typewriter-style UI but should not be persisted as final output.

In current code, partials are emitted from the WLK STT path consistently and from Transcribe inconsistently. Don't rely on them for your client logic; treat them as a UX nicety.

LiveTranscriptUpdate — operator-UI display state

Emitted on every WLK tick (~5Hz). Carries the full snapshot of what the operator's screen should show — committed rows plus an in-flight row for content currently being recognised / translated. Consumer apps that just want committed translations can ignore this; it exists to drive the operator's live-display UX without requiring the client to reconstruct row state from sentence + partial events.

message LiveTranscriptUpdate {
    repeated TranscriptLine lines = 1;
    string buffer_transcription = 2;
    string buffer_diarization = 3;
    string buffer_translation = 4;
    string status = 5;
    bool speculative = 6;
    float remaining_time_transcription = 7;
    float remaining_time_diarization = 8;
}

message TranscriptLine {
    int32 speaker = 1;
    string text = 2;
    string translation = 3;
    string speaker_gender = 4;
    float start = 5;
    float end = 6;
    string detected_language = 7;
}

Field meanings:

Field Meaning
lines[] Ordered rows: every committed sentence as a TranscriptLine with both text (source) and translation populated, followed by always-present trailing line holding the current in-flight English. The trailing line's text may be empty during genuine speaker pauses; rendering clients should hide such empty rows (e.g. via display:none) but keep the DOM element in place to avoid CSS-transition restart on the next blue-text update.
buffer_transcription WLK's not-yet-confirmed word at the very tail of the in-flight line. May oscillate as WLK re-recognises the trailing word(s). Most clients suppress this for visual stability.
buffer_diarization Speaker label for buffer_transcription.
buffer_translation Streaming Japanese tokens from the translator's on_partial callback. Renders next to the in-flight English on the trailing line; finalised translation arrives via the next TranslatedSentence event.
status One of "active_transcription", "no_audio_detected", "ready_to_stop".
speculative True when a non-primary tier's partial is being rendered (display hint only).
remaining_time_transcription / remaining_time_diarization WLK's internal pipeline lag, in seconds. UI can show as a "translator catching up" indicator.

The trailing-row contract was tightened in the 2026-06 line of work to fix a chunk-by-chunk flicker. The engine maintains _uncommitted_english — a running buffer of WLK-confirmed text since the last post-stitcher commit — which becomes the trailing row's text. It grows monotonically as the speaker continues; commits drain it by removing the matching prefix. Clients that mirror this behavior will get smooth blue-text rendering for free.


Error handling and reconnect

gRPC errors come through as grpc.RpcError. Common codes:

Code Meaning Suggested response
UNAVAILABLE Server unreachable (network down, port closed, service stopped) Reconnect with exponential backoff. Cap at ~30s between attempts.
FAILED_PRECONDITION Config wasn't sent first, or server in bad state Don't retry. Surface to the user / log.
CANCELLED The stream was cancelled (you did it, or the server did) Usually no action — graceful shutdown.
DEADLINE_EXCEEDED RPC took longer than the deadline Increase deadline or break into smaller streams.
RESOURCE_EXHAUSTED Server overloaded Backoff + retry.
INTERNAL Server bug Log, surface, file an issue.

No session resumption. The server keeps no state between streams. If your stream drops mid-utterance, the partial audio is lost. Reconnects start from scratch with a new session_id.

For a UX where dropping audio mid-utterance is unacceptable, mitigate at the client: buffer the source audio for a short window (1–2 seconds), reconnect on UNAVAILABLE, and replay buffered audio with the same PTS values. The server will dedupe nothing, but you'll get a clean restart at sentence boundaries.


Cancellation

Two ways:

  1. Stop the request iterator. async for ... exits, the stream closes from the client side, the server sees the half-close and stops emitting.
  2. Close the channel. await channel.close() aborts immediately. The server raises CANCELLED on its side.

Always send StreamControl(end=True) before close if you want a graceful end. Skipping it is fine for emergencies but you lose any in-flight translation.


Worked example

See tests/spike_grpc_streaming_client.py — a complete working client. Run it against a local server:

# Start hibiki-translate locally (one terminal)
AUTH_MODE=dev STT_ENGINE=transcribe python web_server.py

# In another terminal — health check first
python tests/spike_grpc_streaming_client.py health

# Synthetic audio (silence + sine tone, ~5 seconds)
python tests/spike_grpc_streaming_client.py synthetic en ja

# Real WAV file (must be 16kHz mono int16; resample with ffmpeg if not)
python tests/spike_grpc_streaming_client.py stream sample.wav en ja

To point at a deployed server:

HIBIKI_TRANSLATE_GRPC=translate.hibiki.local:50053 \
    python tests/spike_grpc_streaming_client.py synthetic en ja

The script demonstrates: - Config-first message ordering - Real-time pacing of audio chunks (sleep between sends to match audio duration) - PTS tracking (so TranslatedSentence.src_pts_* is meaningful) - Both TranslatedSentence and PartialTranscript handling - Graceful Ctrl-C cancellation - gRPC error classification and exit codes


Common gotchas


Next steps