← Documentation index

HibikiTranslate Runbook

For: the operator on-call during a live event, or anyone investigating an alarm. Goal: every named scenario below has concrete commands. No interpretation required.

This document is terse on purpose. If you're reading it during an event, you don't have time for narrative — you need the next command.

Quick reference

Symptom Section
Operator session has no Japanese rows for >1 minute Operator session is silent
Audience sees … placeholders instead of translation Translation chain failing
Many [STT] line[N] re-anchored log lines WLK quality degraded
Operator sees a Cognito login form mid-event Auth bounce
GPU instance won't start GPU instance won't start
Bedrock returning errors Bedrock throttling or outage
Memory or CPU climbing on a task Resource exhaustion
Alarm received but service looks healthy Alarm triage
Need to restart a service Force-restart commands

Operator session is silent

Alarm: OperatorSilentAlarm — no [Session ...] Output # log lines for 5 minutes during an active session.

Diagnosis (in order — stop when you find the culprit):

  1. Audio source is the most common cause. Look at the operator UI's silence banner (added 2026-06-14). If it's showing, the speaker has stopped or the source died.
  2. YouTube tab paused → resume it
  3. Mic muted → unmute
  4. Operator stepped away from a recorded source → resume
  5. Network blip on operator's machine → check their connection

  6. Bridge → engine connectivity: ```bash # Is the bridge healthy? curl -sf https://tokyo-stage.liveprod.cloud/health # Expect: {"status":"ok"}

# Is the engine reachable from the bridge's network? aws ecs execute-command --region ap-northeast-1 \ --cluster HibikiStage-Tokyo-ClusterEB0386A7-xG8E4DMSUkhn \ --task $(aws ecs list-tasks --region ap-northeast-1 \ --cluster HibikiStage-Tokyo-ClusterEB0386A7-xG8E4DMSUkhn \ --service-name $(aws ecs list-services --region ap-northeast-1 \ --cluster HibikiStage-Tokyo-ClusterEB0386A7-xG8E4DMSUkhn \ --query 'serviceArns[0]' --output text) \ --query 'taskArns[0]' --output text) \ --command "/bin/sh -c 'curl -sf http://translate.hibiki.local:50053 || echo UNREACHABLE'" \ --interactive ```

  1. Engine → WLK connectivity: bash # Tail the gRPC log for recent errors aws logs tail HibikiTranslate-Split-SharedLogsBEDC897C-Ax81o4KVjOcP \ --region ap-northeast-1 \ --filter-pattern '?ERROR ?Traceback ?"connection refused" ?"ws_connect"' \ --since 10m

  2. WLK process itself: bash # Tail the WLK log for crashes / hallucination warnings aws logs tail HibikiTranslate-Split-SharedLogsBEDC897C-Ax81o4KVjOcP \ --region ap-northeast-1 \ --log-stream-name-prefix hibiki-wlk \ --since 10m

  3. Last resort — restart the gRPC service. See Force-restart commands. This kills the active session; the operator will need to re-click record.

If none of the above: the alarm may be a false positive (session ended naturally and the alarm self-clears within 5 min). Wait one more minute before escalating.

Translation chain failing

Alarm: ChainFailureAlarm — 5+ all hedges failed log lines in 5 minutes.

Diagnosis:

  1. Bedrock service health: check https://health.aws.amazon.com/health/status for ap-northeast-1 (Tokyo). Sonnet-4.6 and Nova-Lite both run there for us.

  2. Recent error sample: bash aws logs tail HibikiTranslate-Split-SharedLogsBEDC897C-Ax81o4KVjOcP \ --region ap-northeast-1 \ --filter-pattern '?"all hedges failed" ?"tier nova" ?"tier claude-sonnet"' \ --since 5m Look for the pattern: are all three tiers timing out? Is one tier specifically failing? If only Sonnet is failing, Bedrock for that model is degraded; the chain should still hedge to Nova. If both Sonnet AND Nova fail, that's region-wide.

  3. Per-tier metrics: bash curl -sf https://tokyo-translate.liveprod.cloud/translate/chain-metrics Look at errors, timeouts, successes per tier.

Mitigation options (escalating impact):

WLK quality degraded

Alarm: LineReanchorAlarm — >30 re-anchor / drift events per minute, 2+ minutes running.

This is not a service-down alarm. The pipeline is still emitting translations; quality may be off because WLK's transcription has degraded.

Diagnosis:

  1. Audio source quality is the usual culprit:
  2. Speaker too far from mic
  3. Loud ambient noise (HVAC, audience chatter)
  4. Mic muted intermittently (Zoom, Bluetooth, etc)

  5. Look at recent re-anchor types: bash aws logs tail HibikiTranslate-Split-SharedLogsBEDC897C-Ax81o4KVjOcP \ --region ap-northeast-1 \ --filter-pattern '?"re-anchored after shrink" ?"drift detected"' \ --since 5m | head -20 Many "shrink" events with Hallucination loop collapsed reason → WLK is hallucinating, audio is too quiet or too noisy. Many "drift detected" → WLK is rewriting earlier text, often a sign of a fresh segmentation boundary (multi-speaker audio confusing WLK's diarization).

  6. No alarm action needed if the speaker is currently active and audience is seeing correct translations — quality may simply be borderline. Note for post-event review.

Mitigation: - Talk to the operator: improve mic placement / reduce background noise. - This alarm doesn't trigger a service restart. Quality recovers when the source improves.

Auth bounce

Symptom: operator (or panelist) sees the Cognito login form mid-event when they were already logged in.

Recent fix (2026-06-14): silent JWT refresh + 401-retry was added. If this is happening AFTER 2026-06-14, the refresh path failed.

Diagnosis:

  1. Browser console — look for [auth] log lines. The refresh path now logs its failure mode explicitly (e.g. getSession failed, no current Cognito user).

  2. Cognito user pool health: bash aws cognito-idp describe-user-pool --region ap-northeast-1 \ --user-pool-id ap-northeast-1_DF0shsAHl \ --query 'UserPool.[Status,LambdaConfig]'

  3. Token validity (sanity check): bash aws cognito-idp describe-user-pool-client --region ap-northeast-1 \ --user-pool-id ap-northeast-1_DF0shsAHl \ --client-id a524o4kd9u2v2ujglulnhgo7d \ --query 'UserPoolClient.{IdTokenValidity:IdTokenValidity,RefreshTokenValidity:RefreshTokenValidity}' ID/access tokens are 12h, refresh is 30d. Anything different is a config drift.

Mitigation: - Tell the operator to log in again. Their refresh-token will get a fresh ID-token; recovery is silent thereafter. - If multiple operators are bouncing simultaneously, that's a Cognito issue; check service health.

GPU instance won't start

Symptom: WLK service is desired_count=1 but no task is running, or task keeps cycling.

Diagnosis:

  1. ASG capacity: bash aws autoscaling describe-auto-scaling-groups --region ap-northeast-1 \ --query 'AutoScalingGroups[?contains(AutoScalingGroupName, `Wlk`)].[AutoScalingGroupName,DesiredCapacity,Instances[*].LifecycleState]' If DesiredCapacity=0, no instance — the operator UI's "Start GPU" button needs to be clicked, OR start it manually: bash aws autoscaling set-desired-capacity --region ap-northeast-1 \ --auto-scaling-group-name <ASG_NAME_FROM_ABOVE> --desired-capacity 1

  2. Capacity Reservation / SKU availability: bash aws ec2 describe-instance-status --region ap-northeast-1 \ --filters "Name=instance-state-name,Values=pending,running" \ --query 'InstanceStatuses[*].[InstanceId,InstanceState.Name]' If new instances are stuck pending or get terminated quickly, AWS may be out of g6e.xlarge in this AZ. Mitigation: swap to a different SKU temporarily — see cdk.json context vars or the deploy command's -c instanceType=.

  3. Recent ECS events: bash aws ecs describe-services --region ap-northeast-1 \ --cluster HibikiTranslate-Split-ClusterEB0386A7-FVYLRaqHrllo \ --services <WlkServiceName> \ --query 'services[0].events[0:5]'

Bedrock throttling or outage

Symptom: chain-failure alarm + per-tier metrics show high error/timeout rates from Bedrock.

Diagnosis:

# Recent Bedrock error patterns
aws logs tail HibikiTranslate-Split-SharedLogsBEDC897C-Ax81o4KVjOcP \
  --region ap-northeast-1 \
  --filter-pattern '?"ThrottlingException" ?"ServiceUnavailable" ?"ModelTimeoutException"' \
  --since 10m

Mitigation order:

  1. Wait 2-3 min — most Bedrock blips are transient.
  2. Bump timeouts in settings UI — Sonnet 4→6s, hedge_delay 1.5→2.5s. No deploy needed.
  3. Region swap to us-west-2 — see Translation chain failing above.

Resource exhaustion

Symptom: memory or CPU on a task is climbing without bound, or the task keeps OOMing.

Diagnosis:

# Per-task CPU/memory
aws cloudwatch get-metric-statistics --region ap-northeast-1 \
  --namespace AWS/ECS \
  --metric-name MemoryUtilization \
  --dimensions Name=ClusterName,Value=HibikiTranslate-Split-ClusterEB0386A7-FVYLRaqHrllo \
               Name=ServiceName,Value=<service-name> \
  --start-time $(date -u -d '1 hour ago' +%Y-%m-%dT%H:%M:%S) \
  --end-time $(date -u +%Y-%m-%dT%H:%M:%S) \
  --period 60 --statistics Average,Maximum

Mitigation:

Alarm triage

When an alarm fires but everything looks fine:

  1. Check if a session was just ending — OperatorSilentAlarm self-clears within 5 minutes after a session naturally ends. False positives at session-end are expected for v1.0.

  2. Check the alarm's metric history: bash aws cloudwatch describe-alarm-history --region ap-northeast-1 \ --alarm-name <alarm-name> --max-records 5

  3. If the alarm is genuinely noisy across many sessions, post-event we adjust the threshold in cdk/stacks/split_stack.py. Don't change thresholds during the event itself — that's deploy-time work.

Force-restart commands

Restart the gRPC service (kills any active session):

aws ecs update-service --region ap-northeast-1 \
  --cluster HibikiTranslate-Split-ClusterEB0386A7-FVYLRaqHrllo \
  --service <GrpcServiceName> \
  --force-new-deployment

Restart the bridge service (hibiki-stage):

aws ecs update-service --region ap-northeast-1 \
  --cluster HibikiStage-Tokyo-ClusterEB0386A7-xG8E4DMSUkhn \
  --service <WebServiceName> \
  --force-new-deployment

Restart WLK (after engine + bridge are confirmed healthy; this is the most disruptive):

aws ecs update-service --region ap-northeast-1 \
  --cluster HibikiTranslate-Split-ClusterEB0386A7-FVYLRaqHrllo \
  --service <WlkServiceName> \
  --force-new-deployment

With Phase 1's drain logic in place, these commands DO NOT lose in-flight sessions on the other tasks of the same service. They only kill the task being replaced.

Post-event review

After every event, check:

  1. Alarm history — what fired, when, was it real? bash aws cloudwatch describe-alarms --region ap-northeast-1 \ --query 'MetricAlarms[?starts_with(AlarmName, `HibikiTranslate-Split`)].[AlarmName,StateValue,StateUpdatedTimestamp]'

  2. [POST-STITCH] Fire distribution — sanity-check that row sizes match the configured min_chars: bash aws logs filter-log-events --region ap-northeast-1 \ --log-group-name HibikiTranslate-Split-SharedLogsBEDC897C-Ax81o4KVjOcP \ --start-time <event-start-ms> --end-time <event-end-ms> \ --filter-pattern '"POST-STITCH" "Fire"' \ --query 'events[*].message' --output text | grep -oE '\([0-9]+ chars' | sort | uniq -c Look for too many short rows (< min_chars) — those are timeout-fire artifacts.

  3. Re-anchor distribution — if LineReanchorAlarm fired, did it correlate with a known audio issue?

  4. Subscribe new on-call to the SNS topic if the team has changed: bash aws sns subscribe --region ap-northeast-1 \ --topic-arn <OpsAlarmsTopicArn from CFN outputs> \ --protocol email --notification-endpoint <email>