Audience: anyone deploying hibiki-translate to AWS via CDK. Covers the three deploy modes, infrastructure they create, context variables, sharing infra with hibiki-stage (and other future consumers).
For end-user operation, see operator-guide.md. For client integration, see grpc-integration.md.
cdk/app.py exposes three CDK deploy modes via the -c deployMode=... context variable. Pick based on your STT engine and operational shape.
fargate — no GPU, AWS Transcribe STTCDK_DOCKER=finch cdk deploy -c deployMode=fargate
| Component | Where |
|---|---|
| Web tier | Fargate task running web_server.py |
| STT | AWS Transcribe (cloud API, no GPU) |
| gRPC | Not exposed. Port 50053 mapped on the task but unreachable from outside. |
| Public ingress | CloudFront → ALB → Fargate, port 8080 |
Use when: - You don't need diarization - Single-speaker scenarios - Cost-sensitive (no idle GPU) - Quick deploy for dev / testing
Limitations:
- gRPC is unavailable. If you need a downstream consumer (hibiki-stage), you must use split mode.
- No multilingual diarization.
ecs-gpu — GPU on ECS, WhisperLiveKit STT, public gRPCCDK_DOCKER=finch cdk deploy -c deployMode=ecs-gpu -c instanceType=g6e.xlarge
| Component | Where |
|---|---|
| Web + WLK | Single ECS-on-EC2 task with HOST networking on a g6e.xlarge instance |
| STT | WhisperLiveKit on GPU |
| gRPC | Exposed on port 50053, publicly accessible via the EC2 security group |
| Public ingress | CloudFront → ALB for HTTP, plus direct EC2 IP for gRPC |
Use when:
- You're testing the GPU/WLK path without setting up split mode infrastructure
- Single-instance is enough (no autoscaling)
Don't use for production. Public gRPC has no authentication; the trust boundary is a security-group rule, which is too coarse-grained for production. Move to split mode before going live.
split — Fargate web + dedicated gRPC tier + on-demand GPUCDK_DOCKER=finch cdk deploy -c deployMode=split -c instanceType=g6e.xlarge
| Component | Where |
|---|---|
| Web tier | Fargate task (Transcribe STT for the WS endpoint) |
| gRPC tier | Dedicated Fargate task behind an internal NLB, registered in Cloud Map at translate.hibiki.local:50053 |
| GPU | ECS-on-EC2 ASG, capacity 0 by default (on-demand) — used by WLK STT when activated |
| Public ingress | CloudFront → ALB → web Fargate (port 8080) |
| Internal | gRPC NLB is internet-facing=False, security group restricts to VPC CIDR |
Use when:
- You're going to production
- You have a downstream consumer (hibiki-stage) that needs internal gRPC access
- You want the engine and the GPU to scale independently
This is the recommended mode for the architecture in HIBIKI_STAGE_PLAN.md. The hibiki-stage consumer app expects to find the engine at translate.hibiki.local:50053, which split mode provides.
All modes accept these context variables via -c key=value. See cdk/app.py for the full list and cdk/stacks/shared.py for how they're consumed.
hibiki-stage and hibiki-infraIf you're deploying as part of the multi-app architecture, these context vars must be passed at every cdk deploy so the apps share infrastructure correctly.
| Var | What it references | Source |
|---|---|---|
vpcId |
Existing VPC | hibiki-infra output |
cognitoUserPoolId |
Existing Cognito user pool | hibiki-infra output |
cognitoClientId |
Cognito app client | hibiki-infra output |
acmCertArn |
Wildcard ACM cert in us-east-1 for CloudFront |
hibiki-infra output |
hostedZoneId |
Route53 hosted zone for the parent domain | hibiki-infra output |
cloudMapNamespaceId |
Cloud Map private DNS namespace (hibiki.local) |
hibiki-infra output |
profileTableArn |
DynamoDB profile-table ARN | Existing engine deploy or hibiki-infra output |
domainName |
Base domain (e.g. liveprod.cloud) |
Decided once at the org level |
If you're deploying hibiki-translate standalone (no hibiki-infra yet, no hibiki-stage), the CDK can create the missing resources fresh. But that creates the migration problem documented in HIBIKI_STAGE_PLAN.md — the resources end up owned by hibiki-translate's stack, and moving them to hibiki-infra later is a cdk import exercise. Better to deploy hibiki-infra first if you anticipate needing it.
| Var | Default | Notes |
|---|---|---|
region |
ap-northeast-1 |
Single-region deploy. The stack name doesn't include region for backward compat. |
regions |
— | Multi-region deploy. Comma-separated. Stack name includes region short-name. |
account |
$CDK_DEFAULT_ACCOUNT |
AWS account ID. |
For our deployments:
- Testing: -c region=ap-northeast-1 (Tokyo)
- Production: -c region=us-west-2 (Oregon / LA)
ACM cert lives in us-east-1 regardless — see "ACM cert constraint" below.
| Var | Modes | Notes |
|---|---|---|
deployMode |
all | fargate / ecs-gpu / split. Default: fargate. |
instanceType |
ecs-gpu, split |
GPU instance type. Default: g6.xlarge. Recommended for WLK: g6e.xlarge (L40S). |
adminEmail |
when creating new Cognito pool | Creates an admin user with this email. |
CloudFront requires the ACM cert to live in us-east-1, regardless of where the Fargate / ECS workload deploys. If you're deploying everything in us-west-2 for production, you still need a cert in us-east-1.
The way hibiki-infra handles this:
DnsValidatedCertificate (or equivalent multi-region pattern) in us-east-1hibiki-translate and hibiki-stage reference it via acmCertArn context varIf you're deploying hibiki-translate standalone without hibiki-infra, the CDK creates a per-deploy cert. Workable but generates one cert per app per region — wasteful when you can share one wildcard.
The CDK supports multi-region deploys via -c regions=ap-northeast-1,us-west-2. Each region gets its own stack with the region short-name appended:
HibikiTranslate-Split-Tokyo (deployed in ap-northeast-1)HibikiTranslate-Split-Oregon (deployed in us-west-2)Each region has its own: - VPC (or a separately-imported one per region) - Cognito pool (unless you reuse via context) - DynamoDB profile table - CloudFront distribution
Cognito and DynamoDB are NOT global resources — running in two regions means duplicating user data unless you explicitly share via cognitoUserPoolId / profileTableArn.
For our setup: testing in Tokyo and production in LA share the same hibiki-infra-owned Cognito + VPC, so users created in one region work in the other. The DynamoDB profile table can also be shared if cross-region replication is acceptable, or duplicated per region for tighter consistency.
| Table | Created by | Purpose |
|---|---|---|
HibikiTranslateProfiles |
hibiki-translate (or hibiki-infra) |
Per-language-pair pipeline tuning, smart-chunking word lists, custom prompts. Partition key: pair_key. |
HibikiStageSessions (future) |
hibiki-stage |
Per-session segment storage, history, exports. |
HibikiTranslateProfiles is engine-specific. hibiki-stage doesn't read or write it. They can coexist in the same DynamoDB account or be separate.
Each app owns its own CloudFront distribution, per region. There's no shared CloudFront distribution.
hibiki-translate (production, LA): translate.liveprod.cloud → its CloudFront → its ALBhibiki-stage (production, LA): stage.liveprod.cloud → its CloudFront → its ALBhibiki-translate (testing, Tokyo): tokyo-translate.liveprod.cloud → its CloudFront → its ALBhibiki-stage (testing, Tokyo): tokyo-stage.liveprod.cloud → its CloudFront → its ALBThe wildcard ACM cert (*.liveprod.cloud) covers all four. The Route53 hosted zone has A-records pointing each subdomain at its CloudFront distribution.
See HIBIKI_STAGE_PLAN.md "Why per-app CloudFront" for the architectural reasoning.
For a from-scratch deploy of the whole architecture:
hibiki-infra — creates Cognito pool, VPC, Route53 zone, ACM cert, Cloud Map namespace.hibiki-translate — references shared-infra outputs as context vars. Pick split mode for production.hibiki-stage — references shared-infra outputs + translate.hibiki.local:50053 for gRPC.(hibiki-tts is deployed independently when needed for Phase 3 of HIBIKI_STAGE_PLAN.md.)
For a redeploy of just hibiki-translate:
cdk deploy -c deployMode=split -c region=us-west-2 -c vpcId=... -c cognitoUserPoolId=... -c acmCertArn=... -c hostedZoneId=... -c cloudMapNamespaceId=... -c profileTableArn=...The plan doc captures the full context-var lists.
Before the first deploy in any region, CDK requires bootstrapping:
cdk bootstrap aws://{account-id}/ap-northeast-1
cdk bootstrap aws://{account-id}/us-west-2
cdk bootstrap aws://{account-id}/us-east-1 # for the cert
Idempotent. Re-running on already-bootstrapped accounts is fine.
The Dockerfile builds the engine image. CDK uses Finch (or Docker if CDK_DOCKER unset) to build and push to ECR.
CDK_DOCKER=finch cdk deploy -c deployMode=split
The image includes:
- The Python application (web_server.py, server.py, pipeline/)
- The proto-generated stubs (protos/generated/)
- The web frontend (web/)
- The new docs/ directory (for the runtime docs page)
Build time on a fresh checkout: ~5–10 minutes for the first build (PyTorch, CUDA libs); ~1–2 minutes for incremental builds with cached layers.
These are set on the running container, sourced from CDK context vars and stack outputs.
| Variable | Set from | Default if not set |
|---|---|---|
STT_ENGINE |
DynamoDB-persisted setting | transcribe for fargate; whisperlivekit for ecs-gpu and split |
TRANSLATOR |
hard-coded in the stack | nova |
NOVA_MODEL_ID |
env-var override | region-derived (jp.amazon.nova-2-lite-v1:0 in Tokyo, us.amazon.nova-lite-v1:0 in Oregon) |
AUTH_MODE |
CDK context (authMode) |
dev if not set; should be normal in production |
COGNITO_USER_POOL_ID |
shared-infra output | — |
COGNITO_CLIENT_ID |
shared-infra output | — |
COGNITO_REGION |
matches AWS_REGION |
— |
AWS_REGION |
CDK env | ap-northeast-1 |
WLK_MODEL |
hard-coded | large-v3 |
WLK_LANGUAGE |
DynamoDB-persisted (default) | en |
TGT_LANG |
DynamoDB-persisted (default) | ja |
GPU_ASG_NAME |
stack output (split / ecs-gpu modes only) | — |
PROFILE_TABLE_NAME |
stack output | HibikiTranslateProfiles |
After cdk deploy completes, the stack outputs are printed. Capture:
WebUrl — the CloudFront URL for the operator UICognitoUserPoolId, CognitoClientId — for sharing with consumersWlkServiceName, ClusterName — for GPU lifecycle managementGrpcEndpoint (split mode) — internal NLB DNS for gRPCGrpcCloudMap (split mode) — translate.hibiki.local:50053DeployMode — confirms which mode actually deployedSmoke tests:
curl https://translate.liveprod.cloud/health → {"status": "ok", ...}python tests/spike_grpc_streaming_client.py healthhttps://translate.liveprod.cloud/translate in a browser; log in; click mic; speak.If gRPC is unreachable from inside the VPC, check:
- The Cloud Map service is registered (aws servicediscovery list-services)
- The NLB security group allows the VPC CIDR on port 50053
- The Fargate task is running and healthy (aws ecs describe-services)
The split-mode stack is configured to absorb mid-event deploys, task crashes, and translator-tier failures without dropping operator sessions. The behaviors below are active in code; nothing extra to configure post-deploy.
When ECS replaces a task (deploy, scale-down, health-check failure) it sends SIGTERM, waits for the container to exit, then SIGKILLs if it doesn't.
The engine's server.py and the bridge's app.py both intercept SIGTERM and:
_draining flag so new gRPC RPCs / new operator websockets are refused immediately (with UNAVAILABLE / WS close code 1013). The consumer's existing reconnect path then routes to a healthy task.Container stop_timeout is set to 90 seconds on the gRPC and WLK containers (60s drain + 30s scheduler overhead). The web container uses 30s — it doesn't hold long-running state.
NLB target deregistration delay is also 90s, matching the container budget. ALB target on the bridge tier uses the same value.
| Service | min_healthy_percent |
max_healthy_percent |
desired_count |
Notes |
|---|---|---|---|---|
| Engine gRPC (Fargate) | 100 | 200 | 1 | New task comes up before old task is stopped — no capacity dip during deploys. |
| Engine Web (Fargate) | 100 | 200 | 1 | Same. |
| Engine WLK (EC2 GPU) | 50 (default) | 200 (default) | 0 (operator-controlled) | Single-task service; the operator-controlled lifecycle means deploy-time replacements are rare. |
| hibiki-stage Web (Fargate) | 100 | 200 | 2 | Two tasks at all times. ALB load-balances; if one task crashes, the survivor immediately accepts new connections — recovery in ~10s instead of the 2-3 min Fargate cold-start a single-task service would have. |
Three alarms are created in cdk/stacks/split_stack.py and routed to a single SNS topic (OpsAlarms). The topic email subscription is set via -c opsAlarmEmail=... at deploy time; if not provided, the topic is created without a subscriber and you can subscribe manually post-deploy.
| Alarm | Trigger | Action |
|---|---|---|
LineReanchorAlarm |
More than 30 [STT] line[N] re-anchored / drift detected log lines per minute, sustained 2+ min. |
Audio source quality is degrading (loud room, mic issue, paused source). Operator should check input. |
ChainFailureAlarm |
More than 5 [CHAIN] all hedges failed log lines per 5-min window. |
Bedrock degraded. Bump translator timeouts or swap to alternate region. See runbook.md. |
OperatorSilentAlarm |
Zero [Session ...] Output # log lines for 5 consecutive minutes during an active session. |
Active session has gone silent. Catches dead sessions whether the cause is engine, bridge, or audio source. |
Metrics are derived via CloudWatch Logs metric filters — pattern-match against existing log lines. No boto3.put_metric_data instrumentation in the runtime; all observability comes from log content.
The runbook (docs/runbook.md) has diagnostic ladders for each alarm.
The operator UI (in hibiki-stage) has its own resilience layer that pairs with the engine drain:
fetch() catches 401s on auth'd calls, calls refreshToken(), retries with the fresh token. Operator sees no Cognito login form during a 12-hour session.sentence_committed events during recording surfaces a calm grey banner. Cleared instantly when audio resumes.These are bridge-side; the engine doesn't need to know about them. Documented here because operators occasionally ask why a deploy didn't drop their session.
To tear down a stack:
cdk destroy -c deployMode=split -c region=us-west-2
This removes:
- Fargate / ECS / ASG resources
- ALB + CloudFront
- DynamoDB tables (unless RemovalPolicy.RETAIN is set, which it is by default for the profile table)
- IAM roles, log groups, security groups
It does not remove: - The user pool (RETAIN policy by default — protect against accidental loss of users) - The shared-infra resources (those live in their own stack) - The Route53 zone (if shared-infra owns it)
grpc-integration.md — for clients (e.g. hibiki-stage) consuming the deployed engineapi-reference.md — proto and HTTP endpoint referenceoperator-guide.md — runtime operationarchitecture.md — how this engine relates to siblings../HIBIKI_STAGE_PLAN.md — full extraction plan including shared-infra setup../CLAUDE.md — high-level orientation