Audience: developers writing a client (e.g. hibiki-stage) that consumes hibiki-translate over gRPC.
This is the contract document. For the underlying proto definition see api-reference.md. For an executable walk-through, see tests/spike_grpc_streaming_client.py — every code snippet here corresponds to a section of that file.
from protos.generated import translation_pb2, translation_pb2_grpc
import grpc
channel = grpc.aio.insecure_channel("translate.hibiki.local:50053")
stub = translation_pb2_grpc.TranslationServiceStub(channel)
async def request_iter():
yield translation_pb2.TranslateStreamRequest(
config=translation_pb2.StreamConfig(src_lang="en", tgt_lang="ja", session_id="...")
)
# ... yield AudioChunk messages with float32 16kHz PCM ...
yield translation_pb2.TranslateStreamRequest(control=translation_pb2.StreamControl(end=True))
async for response in stub.TranslateStream(request_iter()):
if response.WhichOneof("payload") == "sentence":
print(response.sentence.translated_text)
That's the shape. Everything below explains the details.
The hibiki-translate gRPC server listens on port 50053. The address depends on the deployment mode:
| Deploy mode | Address | Notes |
|---|---|---|
| Local dev | localhost:50053 |
Run python server.py directly |
split (production-ish) |
translate.hibiki.local:50053 |
Cloud Map private DNS, VPC-internal NLB |
ecs-gpu |
EC2 instance public IP / DNS, port 50053 | Public — not recommended for production |
fargate |
Not exposed | Fargate mode doesn't run the gRPC tier |
In the split deploy, the service is internal-only — accessible from any task in the same VPC, not from the public internet. There is no authentication on the gRPC layer; the trust boundary is the network (VPC + security groups). Don't expose port 50053 publicly.
Use grpc.aio.insecure_channel(...) (no TLS for the internal hop). The shared infra layer terminates TLS at the per-app CloudFront edge for HTTP traffic; gRPC-internal stays plaintext.
The service definition (protos/translation.proto) exposes three RPCs:
| RPC | Type | Purpose |
|---|---|---|
TranslateStream |
bidi-streaming | The main contract. Audio in, translation events out. |
TranslateText |
unary | Translate a single text string. Useful for testing without audio. |
HealthCheck |
unary | Returns server status, loaded models, device. Use as a liveness check. |
99% of the integration work is TranslateStream. The rest of this doc focuses on it.
Every TranslateStream call must follow this order:
StreamConfig. If audio arrives before config, the server returns FAILED_PRECONDITION and closes the stream.AudioChunk messages (or the deprecated audio field — see below).StreamControl messages to signal flush or end.yield TranslateStreamRequest(config=StreamConfig(...))
yield TranslateStreamRequest(audio_chunk=AudioChunk(pcm=..., pts_start_ms=..., pts_end_ms=...))
yield TranslateStreamRequest(audio_chunk=AudioChunk(...))
# ...
yield TranslateStreamRequest(control=StreamControl(flush=True)) # end-of-utterance
yield TranslateStreamRequest(control=StreamControl(end=True)) # end-of-session
StreamConfig fields| Field | Type | Required? | Notes |
|---|---|---|---|
src_lang |
string | yes | ISO code: en, ja, fr, etc. Determines STT language and translation source. |
tgt_lang |
string | yes | Target language. The server resolves a language-pair profile (pre-buffer thresholds, smart-chunking flags) from this pair. |
session_id |
string | recommended | Opaque to server; included in logs. Use a UUID per stream so you can correlate gRPC traffic with your own traces. |
enable_diarization |
bool | optional | Default false. When true (and STT engine supports it), TranslatedSentence.speaker_id is populated. WLK STT supports this; AWS Transcribe does not — silently ignored if Transcribe is the configured engine. |
enable_emotion |
bool | optional | Default false. When true, TranslatedSentence.emotion is populated by the translator. |
whisper_model |
string | optional | large-v3 (default), medium, small. Only honored by the WLK STT engine. |
skip_post_stitcher |
bool | optional | Default false. When true, the post-stitcher is bypassed and translations are emitted in chunks rather than complete sentences. Lower latency; less natural output. Recommended only for display-only consumers (e.g. live-captions display) where final-sentence fidelity isn't required. |
Per-call profile config (pre-buffer thresholds, smart-chunking, prompts) is not in StreamConfig. The server resolves a profile from (src_lang, tgt_lang) via the language-pair profile registry. To customise per-pair tuning, use POST /translate/update-profile (HTTP; persisted to DynamoDB) before opening the gRPC stream — or accept the defaults.
AudioChunk — the audio formatmessage AudioChunk {
bytes pcm = 1; // float32 PCM samples, mono, 16 kHz
int64 pts_start_ms = 2; // PTS of the first sample (source-stream clock)
int64 pts_end_ms = 3; // PTS of the last sample
}
Required format:
[-1.0, 1.0]PTS values:
- Express timestamps in milliseconds.
- The clock source is your stream, not the server's wall clock. Whatever monotonic timeline your audio comes from (a media file's timestamp, a live capture's elapsed time, etc.) is what should be in pts_start_ms / pts_end_ms.
- These are echoed back in TranslatedSentence.src_pts_start_ms / src_pts_end_ms, which is how downstream consumers (e.g. a video dub mixer) correlate translation output with the source stream timing.
- Setting both to 0 is acceptable but loses timing fidelity. Always populate them if you have them.
audio fieldTranslateStreamRequest.audio (a raw bytes field) is the legacy path. It still works for backward compatibility with HibikiVoice. New clients should always use audio_chunk so PTS information flows through.
StreamControl — flush and endmessage StreamControl {
bool flush = 1; // end of utterance, drain buffers
bool end = 2; // end of session, close the stream
}
flush=True: tells the server "I've finished an utterance; drain the pre-buffer and emit any pending translation now." Useful when the speaker pauses noticeably or you have a VAD-detected speech-end event.end=True: tells the server "I'm done sending audio; close gracefully." After end=True, no more audio chunks will be accepted; the server emits any final translations and closes the response stream.Always send end=True before closing the channel — without it, the server may drop the last partial translation.
The response stream emits TranslateStreamResponse messages with a oneof payload:
message TranslateStreamResponse {
oneof payload {
TranslatedSentence sentence = 1;
PartialTranscript partial = 2;
LiveTranscriptUpdate live = 3;
}
}
TranslatedSentence — the main eventEmitted once per committed sentence (or chunk if skip_post_stitcher=True). All fields:
| Field | Type | Meaning |
|---|---|---|
translated_text |
string | The translation, in the target language. Already cleaned (no Nova meta-commentary, no emotion prefix, no [SKIP] markers). |
src_text |
string | The original source text the translator saw. Useful for display or debugging. |
sequence |
int32 | Monotonic counter starting at 1. Increments per sentence, per stream. |
speaker_id |
int32 | 1, 2, ... if enable_diarization=True and STT supports it; 0 otherwise. |
speaker_gender |
string | "male", "female", or "" if unknown. WLK estimates from F0; Transcribe does not provide it. |
emotion |
string | excited, happy, sad, angry, fearful, surprised, neutral if enable_emotion=True; "" otherwise. |
emotion_instruction |
string | Free-text TTS instruction (e.g. "speak with warm excitement") — generated alongside the emotion label. Pipe to hibiki-tts's instruct field if doing TTS. |
latency_ms |
float | Server-side translation latency (Nova call duration). Excludes audio capture, STT, pre-buffer wait. |
model_used |
string | The translator that produced this — e.g. nova-2-lite. |
start_time |
float | STT-relative seconds. Kept for back-compat. Prefer src_pts_start_ms. |
end_time |
float | STT-relative seconds. Kept for back-compat. Prefer src_pts_end_ms. |
src_pts_start_ms |
int64 | Echoed-back PTS of where this sentence's source audio starts. 0 if you didn't set it on AudioChunk. |
src_pts_end_ms |
int64 | Echoed-back PTS of where this sentence's source audio ends. 0 if unknown. |
is_replacement |
bool | When true, this sentence revises the previously emitted one (the post-stitcher's carry-over case). Display clients should replace the prior row in place rather than appending a new one. Audience-side broadcast uses the same field. |
translation_status |
string | One of "ok", "fallback_secondary", "fallback_tertiary", "marker", "cache". Reflects how the translator chain produced this row (see "Translation chain status" below). Lets consumers decide rendering / TTS behavior without parsing model_used. |
The translator chain runs a hedged race: primary tier (default Sonnet) at t=0; secondary (Nova) joins at hedge_delay_s; tertiary (AWS Translate) at 2 × hedge_delay_s. The first non-empty result wins; the rest are cancelled. translation_status tells you which won:
| Status | Meaning | Recommended display |
|---|---|---|
ok |
Primary tier produced the translation. Highest quality. | Default rendering. |
fallback_secondary |
Primary timed out / errored / returned empty; secondary won. | Default rendering — translation is still good, just from a faster fallback. |
fallback_tertiary |
Both primary and secondary failed; tertiary won. | Render normally; optionally surface degradation hint to operator. |
marker |
All hedges failed within chain_deadline_s. translated_text is the configured marker (default "…"). |
Render with degraded styling (grey). Do NOT TTS-speak this. |
cache |
Translation came from the per-session cache (identical source seen earlier in the stream). No tier was contacted. | Default rendering. Useful for cost accounting (no Bedrock call billed). |
PartialTranscript — the in-progress previewmessage PartialTranscript {
string text = 1;
int32 speaker_id = 2;
}
Partial transcripts are emitted while a translation is being assembled — for example, the streaming Nova on_partial callback fires per token. They're useful for typewriter-style UI but should not be persisted as final output.
In current code, partials are emitted from the WLK STT path consistently and from Transcribe inconsistently. Don't rely on them for your client logic; treat them as a UX nicety.
LiveTranscriptUpdate — operator-UI display stateEmitted on every WLK tick (~5Hz). Carries the full snapshot of what the operator's screen should show — committed rows plus an in-flight row for content currently being recognised / translated. Consumer apps that just want committed translations can ignore this; it exists to drive the operator's live-display UX without requiring the client to reconstruct row state from sentence + partial events.
message LiveTranscriptUpdate {
repeated TranscriptLine lines = 1;
string buffer_transcription = 2;
string buffer_diarization = 3;
string buffer_translation = 4;
string status = 5;
bool speculative = 6;
float remaining_time_transcription = 7;
float remaining_time_diarization = 8;
}
message TranscriptLine {
int32 speaker = 1;
string text = 2;
string translation = 3;
string speaker_gender = 4;
float start = 5;
float end = 6;
string detected_language = 7;
}
Field meanings:
| Field | Meaning |
|---|---|
lines[] |
Ordered rows: every committed sentence as a TranscriptLine with both text (source) and translation populated, followed by always-present trailing line holding the current in-flight English. The trailing line's text may be empty during genuine speaker pauses; rendering clients should hide such empty rows (e.g. via display:none) but keep the DOM element in place to avoid CSS-transition restart on the next blue-text update. |
buffer_transcription |
WLK's not-yet-confirmed word at the very tail of the in-flight line. May oscillate as WLK re-recognises the trailing word(s). Most clients suppress this for visual stability. |
buffer_diarization |
Speaker label for buffer_transcription. |
buffer_translation |
Streaming Japanese tokens from the translator's on_partial callback. Renders next to the in-flight English on the trailing line; finalised translation arrives via the next TranslatedSentence event. |
status |
One of "active_transcription", "no_audio_detected", "ready_to_stop". |
speculative |
True when a non-primary tier's partial is being rendered (display hint only). |
remaining_time_transcription / remaining_time_diarization |
WLK's internal pipeline lag, in seconds. UI can show as a "translator catching up" indicator. |
The trailing-row contract was tightened in the 2026-06 line of work to fix a chunk-by-chunk flicker. The engine maintains _uncommitted_english — a running buffer of WLK-confirmed text since the last post-stitcher commit — which becomes the trailing row's text. It grows monotonically as the speaker continues; commits drain it by removing the matching prefix. Clients that mirror this behavior will get smooth blue-text rendering for free.
gRPC errors come through as grpc.RpcError. Common codes:
| Code | Meaning | Suggested response |
|---|---|---|
UNAVAILABLE |
Server unreachable (network down, port closed, service stopped) | Reconnect with exponential backoff. Cap at ~30s between attempts. |
FAILED_PRECONDITION |
Config wasn't sent first, or server in bad state | Don't retry. Surface to the user / log. |
CANCELLED |
The stream was cancelled (you did it, or the server did) | Usually no action — graceful shutdown. |
DEADLINE_EXCEEDED |
RPC took longer than the deadline | Increase deadline or break into smaller streams. |
RESOURCE_EXHAUSTED |
Server overloaded | Backoff + retry. |
INTERNAL |
Server bug | Log, surface, file an issue. |
No session resumption. The server keeps no state between streams. If your stream drops mid-utterance, the partial audio is lost. Reconnects start from scratch with a new session_id.
For a UX where dropping audio mid-utterance is unacceptable, mitigate at the client: buffer the source audio for a short window (1–2 seconds), reconnect on UNAVAILABLE, and replay buffered audio with the same PTS values. The server will dedupe nothing, but you'll get a clean restart at sentence boundaries.
Two ways:
async for ... exits, the stream closes from the client side, the server sees the half-close and stops emitting.await channel.close() aborts immediately. The server raises CANCELLED on its side.Always send StreamControl(end=True) before close if you want a graceful end. Skipping it is fine for emergencies but you lose any in-flight translation.
See tests/spike_grpc_streaming_client.py — a complete working client. Run it against a local server:
# Start hibiki-translate locally (one terminal)
AUTH_MODE=dev STT_ENGINE=transcribe python web_server.py
# In another terminal — health check first
python tests/spike_grpc_streaming_client.py health
# Synthetic audio (silence + sine tone, ~5 seconds)
python tests/spike_grpc_streaming_client.py synthetic en ja
# Real WAV file (must be 16kHz mono int16; resample with ffmpeg if not)
python tests/spike_grpc_streaming_client.py stream sample.wav en ja
To point at a deployed server:
HIBIKI_TRANSLATE_GRPC=translate.hibiki.local:50053 \
python tests/spike_grpc_streaming_client.py synthetic en ja
The script demonstrates:
- Config-first message ordering
- Real-time pacing of audio chunks (sleep between sends to match audio duration)
- PTS tracking (so TranslatedSentence.src_pts_* is meaningful)
- Both TranslatedSentence and PartialTranscript handling
- Graceful Ctrl-C cancellation
- gRPC error classification and exit codes
int16_value / 32768.0 → float32.session_id: while optional, an empty session_id makes server logs harder to correlate. Always send a UUID.pts_start_ms = sample_index * 1000 // 16000.enable_diarization requires WLK: with the AWS Transcribe STT engine, the flag is silently ignored. If diarization is critical, deploy in ecs-gpu or split mode with WLK enabled.whisper_model only affects WLK: ignored under Transcribe.python -m grpc_tools.protoc -I=protos --python_out=. --grpc_python_out=. protos/translation.proto.@grpc/grpc-js plus the proto file directly, or pre-generate stubs with protoc-gen-ts.hibiki-tts's gRPC contract (separate microservice; consume TranslatedSentence.translated_text + emotion_instruction and pipe to hibiki-tts.SynthesizeStream).api-reference.md for the raw proto field-by-field.architecture.md for where this fits in the broader system.