For: the operator on-call during a live event, or anyone investigating an alarm. Goal: every named scenario below has concrete commands. No interpretation required.
This document is terse on purpose. If you're reading it during an event, you don't have time for narrative — you need the next command.
| Symptom | Section |
|---|---|
| Operator session has no Japanese rows for >1 minute | Operator session is silent |
Audience sees … placeholders instead of translation |
Translation chain failing |
Many [STT] line[N] re-anchored log lines |
WLK quality degraded |
| Operator sees a Cognito login form mid-event | Auth bounce |
| GPU instance won't start | GPU instance won't start |
| Bedrock returning errors | Bedrock throttling or outage |
| Memory or CPU climbing on a task | Resource exhaustion |
| Alarm received but service looks healthy | Alarm triage |
| Need to restart a service | Force-restart commands |
Alarm: OperatorSilentAlarm — no [Session ...] Output # log lines for 5 minutes during an active session.
Diagnosis (in order — stop when you find the culprit):
Network blip on operator's machine → check their connection
Bridge → engine connectivity: ```bash # Is the bridge healthy? curl -sf https://tokyo-stage.liveprod.cloud/health # Expect: {"status":"ok"}
# Is the engine reachable from the bridge's network? aws ecs execute-command --region ap-northeast-1 \ --cluster HibikiStage-Tokyo-ClusterEB0386A7-xG8E4DMSUkhn \ --task $(aws ecs list-tasks --region ap-northeast-1 \ --cluster HibikiStage-Tokyo-ClusterEB0386A7-xG8E4DMSUkhn \ --service-name $(aws ecs list-services --region ap-northeast-1 \ --cluster HibikiStage-Tokyo-ClusterEB0386A7-xG8E4DMSUkhn \ --query 'serviceArns[0]' --output text) \ --query 'taskArns[0]' --output text) \ --command "/bin/sh -c 'curl -sf http://translate.hibiki.local:50053 || echo UNREACHABLE'" \ --interactive ```
Engine → WLK connectivity:
bash
# Tail the gRPC log for recent errors
aws logs tail HibikiTranslate-Split-SharedLogsBEDC897C-Ax81o4KVjOcP \
--region ap-northeast-1 \
--filter-pattern '?ERROR ?Traceback ?"connection refused" ?"ws_connect"' \
--since 10m
WLK process itself:
bash
# Tail the WLK log for crashes / hallucination warnings
aws logs tail HibikiTranslate-Split-SharedLogsBEDC897C-Ax81o4KVjOcP \
--region ap-northeast-1 \
--log-stream-name-prefix hibiki-wlk \
--since 10m
Last resort — restart the gRPC service. See Force-restart commands. This kills the active session; the operator will need to re-click record.
If none of the above: the alarm may be a false positive (session ended naturally and the alarm self-clears within 5 min). Wait one more minute before escalating.
Alarm: ChainFailureAlarm — 5+ all hedges failed log lines in 5 minutes.
Diagnosis:
Bedrock service health: check https://health.aws.amazon.com/health/status for ap-northeast-1 (Tokyo). Sonnet-4.6 and Nova-Lite both run there for us.
Recent error sample:
bash
aws logs tail HibikiTranslate-Split-SharedLogsBEDC897C-Ax81o4KVjOcP \
--region ap-northeast-1 \
--filter-pattern '?"all hedges failed" ?"tier nova" ?"tier claude-sonnet"' \
--since 5m
Look for the pattern: are all three tiers timing out? Is one tier specifically failing? If only Sonnet is failing, Bedrock for that model is degraded; the chain should still hedge to Nova. If both Sonnet AND Nova fail, that's region-wide.
Per-tier metrics:
bash
curl -sf https://tokyo-translate.liveprod.cloud/translate/chain-metrics
Look at errors, timeouts, successes per tier.
Mitigation options (escalating impact):
bash
# Override Bedrock region via SSM parameter (existing path, not a code change)
aws ssm put-parameter --region ap-northeast-1 \
--name "/hibiki-translate/bedrock-region-override" \
--value "us-west-2" --overwrite
# Engine reads this on session start; operator must restart the active sessionAlarm: LineReanchorAlarm — >30 re-anchor / drift events per minute, 2+ minutes running.
This is not a service-down alarm. The pipeline is still emitting translations; quality may be off because WLK's transcription has degraded.
Diagnosis:
Mic muted intermittently (Zoom, Bluetooth, etc)
Look at recent re-anchor types:
bash
aws logs tail HibikiTranslate-Split-SharedLogsBEDC897C-Ax81o4KVjOcP \
--region ap-northeast-1 \
--filter-pattern '?"re-anchored after shrink" ?"drift detected"' \
--since 5m | head -20
Many "shrink" events with Hallucination loop collapsed reason → WLK is hallucinating, audio is too quiet or too noisy. Many "drift detected" → WLK is rewriting earlier text, often a sign of a fresh segmentation boundary (multi-speaker audio confusing WLK's diarization).
No alarm action needed if the speaker is currently active and audience is seeing correct translations — quality may simply be borderline. Note for post-event review.
Mitigation: - Talk to the operator: improve mic placement / reduce background noise. - This alarm doesn't trigger a service restart. Quality recovers when the source improves.
Symptom: operator (or panelist) sees the Cognito login form mid-event when they were already logged in.
Recent fix (2026-06-14): silent JWT refresh + 401-retry was added. If this is happening AFTER 2026-06-14, the refresh path failed.
Diagnosis:
Browser console — look for [auth] log lines. The refresh path now logs its failure mode explicitly (e.g. getSession failed, no current Cognito user).
Cognito user pool health:
bash
aws cognito-idp describe-user-pool --region ap-northeast-1 \
--user-pool-id ap-northeast-1_DF0shsAHl \
--query 'UserPool.[Status,LambdaConfig]'
Token validity (sanity check):
bash
aws cognito-idp describe-user-pool-client --region ap-northeast-1 \
--user-pool-id ap-northeast-1_DF0shsAHl \
--client-id a524o4kd9u2v2ujglulnhgo7d \
--query 'UserPoolClient.{IdTokenValidity:IdTokenValidity,RefreshTokenValidity:RefreshTokenValidity}'
ID/access tokens are 12h, refresh is 30d. Anything different is a config drift.
Mitigation: - Tell the operator to log in again. Their refresh-token will get a fresh ID-token; recovery is silent thereafter. - If multiple operators are bouncing simultaneously, that's a Cognito issue; check service health.
Symptom: WLK service is desired_count=1 but no task is running, or task keeps cycling.
Diagnosis:
ASG capacity:
bash
aws autoscaling describe-auto-scaling-groups --region ap-northeast-1 \
--query 'AutoScalingGroups[?contains(AutoScalingGroupName, `Wlk`)].[AutoScalingGroupName,DesiredCapacity,Instances[*].LifecycleState]'
If DesiredCapacity=0, no instance — the operator UI's "Start GPU" button needs to be clicked, OR start it manually:
bash
aws autoscaling set-desired-capacity --region ap-northeast-1 \
--auto-scaling-group-name <ASG_NAME_FROM_ABOVE> --desired-capacity 1
Capacity Reservation / SKU availability:
bash
aws ec2 describe-instance-status --region ap-northeast-1 \
--filters "Name=instance-state-name,Values=pending,running" \
--query 'InstanceStatuses[*].[InstanceId,InstanceState.Name]'
If new instances are stuck pending or get terminated quickly, AWS may be out of g6e.xlarge in this AZ. Mitigation: swap to a different SKU temporarily — see cdk.json context vars or the deploy command's -c instanceType=.
Recent ECS events:
bash
aws ecs describe-services --region ap-northeast-1 \
--cluster HibikiTranslate-Split-ClusterEB0386A7-FVYLRaqHrllo \
--services <WlkServiceName> \
--query 'services[0].events[0:5]'
Symptom: chain-failure alarm + per-tier metrics show high error/timeout rates from Bedrock.
Diagnosis:
# Recent Bedrock error patterns
aws logs tail HibikiTranslate-Split-SharedLogsBEDC897C-Ax81o4KVjOcP \
--region ap-northeast-1 \
--filter-pattern '?"ThrottlingException" ?"ServiceUnavailable" ?"ModelTimeoutException"' \
--since 10m
Mitigation order:
Symptom: memory or CPU on a task is climbing without bound, or the task keeps OOMing.
Diagnosis:
# Per-task CPU/memory
aws cloudwatch get-metric-statistics --region ap-northeast-1 \
--namespace AWS/ECS \
--metric-name MemoryUtilization \
--dimensions Name=ClusterName,Value=HibikiTranslate-Split-ClusterEB0386A7-FVYLRaqHrllo \
Name=ServiceName,Value=<service-name> \
--start-time $(date -u -d '1 hour ago' +%Y-%m-%dT%H:%M:%S) \
--end-time $(date -u +%Y-%m-%dT%H:%M:%S) \
--period 60 --statistics Average,Maximum
Mitigation:
When an alarm fires but everything looks fine:
Check if a session was just ending — OperatorSilentAlarm self-clears within 5 minutes after a session naturally ends. False positives at session-end are expected for v1.0.
Check the alarm's metric history:
bash
aws cloudwatch describe-alarm-history --region ap-northeast-1 \
--alarm-name <alarm-name> --max-records 5
If the alarm is genuinely noisy across many sessions, post-event we adjust the threshold in cdk/stacks/split_stack.py. Don't change thresholds during the event itself — that's deploy-time work.
Restart the gRPC service (kills any active session):
aws ecs update-service --region ap-northeast-1 \
--cluster HibikiTranslate-Split-ClusterEB0386A7-FVYLRaqHrllo \
--service <GrpcServiceName> \
--force-new-deployment
Restart the bridge service (hibiki-stage):
aws ecs update-service --region ap-northeast-1 \
--cluster HibikiStage-Tokyo-ClusterEB0386A7-xG8E4DMSUkhn \
--service <WebServiceName> \
--force-new-deployment
Restart WLK (after engine + bridge are confirmed healthy; this is the most disruptive):
aws ecs update-service --region ap-northeast-1 \
--cluster HibikiTranslate-Split-ClusterEB0386A7-FVYLRaqHrllo \
--service <WlkServiceName> \
--force-new-deployment
With Phase 1's drain logic in place, these commands DO NOT lose in-flight sessions on the other tasks of the same service. They only kill the task being replaced.
After every event, check:
Alarm history — what fired, when, was it real?
bash
aws cloudwatch describe-alarms --region ap-northeast-1 \
--query 'MetricAlarms[?starts_with(AlarmName, `HibikiTranslate-Split`)].[AlarmName,StateValue,StateUpdatedTimestamp]'
[POST-STITCH] Fire distribution — sanity-check that row sizes match the configured min_chars:
bash
aws logs filter-log-events --region ap-northeast-1 \
--log-group-name HibikiTranslate-Split-SharedLogsBEDC897C-Ax81o4KVjOcP \
--start-time <event-start-ms> --end-time <event-end-ms> \
--filter-pattern '"POST-STITCH" "Fire"' \
--query 'events[*].message' --output text | grep -oE '\([0-9]+ chars' | sort | uniq -c
Look for too many short rows (< min_chars) — those are timeout-fire artifacts.
Re-anchor distribution — if LineReanchorAlarm fired, did it correlate with a known audio issue?
Subscribe new on-call to the SNS topic if the team has changed:
bash
aws sns subscribe --region ap-northeast-1 \
--topic-arn <OpsAlarmsTopicArn from CFN outputs> \
--protocol email --notification-endpoint <email>