← Documentation index

Deployment Guide

Audience: anyone deploying hibiki-translate to AWS via CDK. Covers the three deploy modes, infrastructure they create, context variables, sharing infra with hibiki-stage (and other future consumers).

For end-user operation, see operator-guide.md. For client integration, see grpc-integration.md.


Three deploy modes

cdk/app.py exposes three CDK deploy modes via the -c deployMode=... context variable. Pick based on your STT engine and operational shape.

fargate — no GPU, AWS Transcribe STT

CDK_DOCKER=finch cdk deploy -c deployMode=fargate
Component Where
Web tier Fargate task running web_server.py
STT AWS Transcribe (cloud API, no GPU)
gRPC Not exposed. Port 50053 mapped on the task but unreachable from outside.
Public ingress CloudFront → ALB → Fargate, port 8080

Use when: - You don't need diarization - Single-speaker scenarios - Cost-sensitive (no idle GPU) - Quick deploy for dev / testing

Limitations: - gRPC is unavailable. If you need a downstream consumer (hibiki-stage), you must use split mode. - No multilingual diarization.

ecs-gpu — GPU on ECS, WhisperLiveKit STT, public gRPC

CDK_DOCKER=finch cdk deploy -c deployMode=ecs-gpu -c instanceType=g6e.xlarge
Component Where
Web + WLK Single ECS-on-EC2 task with HOST networking on a g6e.xlarge instance
STT WhisperLiveKit on GPU
gRPC Exposed on port 50053, publicly accessible via the EC2 security group
Public ingress CloudFront → ALB for HTTP, plus direct EC2 IP for gRPC

Use when: - You're testing the GPU/WLK path without setting up split mode infrastructure - Single-instance is enough (no autoscaling)

Don't use for production. Public gRPC has no authentication; the trust boundary is a security-group rule, which is too coarse-grained for production. Move to split mode before going live.

split — Fargate web + dedicated gRPC tier + on-demand GPU

CDK_DOCKER=finch cdk deploy -c deployMode=split -c instanceType=g6e.xlarge
Component Where
Web tier Fargate task (Transcribe STT for the WS endpoint)
gRPC tier Dedicated Fargate task behind an internal NLB, registered in Cloud Map at translate.hibiki.local:50053
GPU ECS-on-EC2 ASG, capacity 0 by default (on-demand) — used by WLK STT when activated
Public ingress CloudFront → ALB → web Fargate (port 8080)
Internal gRPC NLB is internet-facing=False, security group restricts to VPC CIDR

Use when: - You're going to production - You have a downstream consumer (hibiki-stage) that needs internal gRPC access - You want the engine and the GPU to scale independently

This is the recommended mode for the architecture in HIBIKI_STAGE_PLAN.md. The hibiki-stage consumer app expects to find the engine at translate.hibiki.local:50053, which split mode provides.


CDK context variables

All modes accept these context variables via -c key=value. See cdk/app.py for the full list and cdk/stacks/shared.py for how they're consumed.

Required-to-share with hibiki-stage and hibiki-infra

If you're deploying as part of the multi-app architecture, these context vars must be passed at every cdk deploy so the apps share infrastructure correctly.

Var What it references Source
vpcId Existing VPC hibiki-infra output
cognitoUserPoolId Existing Cognito user pool hibiki-infra output
cognitoClientId Cognito app client hibiki-infra output
acmCertArn Wildcard ACM cert in us-east-1 for CloudFront hibiki-infra output
hostedZoneId Route53 hosted zone for the parent domain hibiki-infra output
cloudMapNamespaceId Cloud Map private DNS namespace (hibiki.local) hibiki-infra output
profileTableArn DynamoDB profile-table ARN Existing engine deploy or hibiki-infra output
domainName Base domain (e.g. liveprod.cloud) Decided once at the org level

If you're deploying hibiki-translate standalone (no hibiki-infra yet, no hibiki-stage), the CDK can create the missing resources fresh. But that creates the migration problem documented in HIBIKI_STAGE_PLAN.md — the resources end up owned by hibiki-translate's stack, and moving them to hibiki-infra later is a cdk import exercise. Better to deploy hibiki-infra first if you anticipate needing it.

Region

Var Default Notes
region ap-northeast-1 Single-region deploy. The stack name doesn't include region for backward compat.
regions — Multi-region deploy. Comma-separated. Stack name includes region short-name.
account $CDK_DEFAULT_ACCOUNT AWS account ID.

For our deployments: - Testing: -c region=ap-northeast-1 (Tokyo) - Production: -c region=us-west-2 (Oregon / LA)

ACM cert lives in us-east-1 regardless — see "ACM cert constraint" below.

Mode-specific

Var Modes Notes
deployMode all fargate / ecs-gpu / split. Default: fargate.
instanceType ecs-gpu, split GPU instance type. Default: g6.xlarge. Recommended for WLK: g6e.xlarge (L40S).
adminEmail when creating new Cognito pool Creates an admin user with this email.

ACM cert constraint

CloudFront requires the ACM cert to live in us-east-1, regardless of where the Fargate / ECS workload deploys. If you're deploying everything in us-west-2 for production, you still need a cert in us-east-1.

The way hibiki-infra handles this:

  1. Creates a DnsValidatedCertificate (or equivalent multi-region pattern) in us-east-1
  2. Validates against the Route53 hosted zone (which can be in any region — Route53 is global)
  3. Exports the cert ARN as a CfnOutput
  4. hibiki-translate and hibiki-stage reference it via acmCertArn context var

If you're deploying hibiki-translate standalone without hibiki-infra, the CDK creates a per-deploy cert. Workable but generates one cert per app per region — wasteful when you can share one wildcard.


Single-region vs. multi-region

The CDK supports multi-region deploys via -c regions=ap-northeast-1,us-west-2. Each region gets its own stack with the region short-name appended:

Each region has its own: - VPC (or a separately-imported one per region) - Cognito pool (unless you reuse via context) - DynamoDB profile table - CloudFront distribution

Cognito and DynamoDB are NOT global resources — running in two regions means duplicating user data unless you explicitly share via cognitoUserPoolId / profileTableArn.

For our setup: testing in Tokyo and production in LA share the same hibiki-infra-owned Cognito + VPC, so users created in one region work in the other. The DynamoDB profile table can also be shared if cross-region replication is acceptable, or duplicated per region for tighter consistency.


DynamoDB tables

Table Created by Purpose
HibikiTranslateProfiles hibiki-translate (or hibiki-infra) Per-language-pair pipeline tuning, smart-chunking word lists, custom prompts. Partition key: pair_key.
HibikiStageSessions (future) hibiki-stage Per-session segment storage, history, exports.

HibikiTranslateProfiles is engine-specific. hibiki-stage doesn't read or write it. They can coexist in the same DynamoDB account or be separate.


CloudFront distributions

Each app owns its own CloudFront distribution, per region. There's no shared CloudFront distribution.

The wildcard ACM cert (*.liveprod.cloud) covers all four. The Route53 hosted zone has A-records pointing each subdomain at its CloudFront distribution.

See HIBIKI_STAGE_PLAN.md "Why per-app CloudFront" for the architectural reasoning.


Deploy order

For a from-scratch deploy of the whole architecture:

  1. hibiki-infra — creates Cognito pool, VPC, Route53 zone, ACM cert, Cloud Map namespace.
  2. hibiki-translate — references shared-infra outputs as context vars. Pick split mode for production.
  3. hibiki-stage — references shared-infra outputs + translate.hibiki.local:50053 for gRPC.

(hibiki-tts is deployed independently when needed for Phase 3 of HIBIKI_STAGE_PLAN.md.)

For a redeploy of just hibiki-translate:

  1. Confirm shared-infra is still up.
  2. cdk deploy -c deployMode=split -c region=us-west-2 -c vpcId=... -c cognitoUserPoolId=... -c acmCertArn=... -c hostedZoneId=... -c cloudMapNamespaceId=... -c profileTableArn=...

The plan doc captures the full context-var lists.


CDK bootstrap

Before the first deploy in any region, CDK requires bootstrapping:

cdk bootstrap aws://{account-id}/ap-northeast-1
cdk bootstrap aws://{account-id}/us-west-2
cdk bootstrap aws://{account-id}/us-east-1   # for the cert

Idempotent. Re-running on already-bootstrapped accounts is fine.


Container image build

The Dockerfile builds the engine image. CDK uses Finch (or Docker if CDK_DOCKER unset) to build and push to ECR.

CDK_DOCKER=finch cdk deploy -c deployMode=split

The image includes: - The Python application (web_server.py, server.py, pipeline/) - The proto-generated stubs (protos/generated/) - The web frontend (web/) - The new docs/ directory (for the runtime docs page)

Build time on a fresh checkout: ~5–10 minutes for the first build (PyTorch, CUDA libs); ~1–2 minutes for incremental builds with cached layers.


Environment variables

These are set on the running container, sourced from CDK context vars and stack outputs.

Variable Set from Default if not set
STT_ENGINE DynamoDB-persisted setting transcribe for fargate; whisperlivekit for ecs-gpu and split
TRANSLATOR hard-coded in the stack nova
NOVA_MODEL_ID env-var override region-derived (jp.amazon.nova-2-lite-v1:0 in Tokyo, us.amazon.nova-lite-v1:0 in Oregon)
AUTH_MODE CDK context (authMode) dev if not set; should be normal in production
COGNITO_USER_POOL_ID shared-infra output —
COGNITO_CLIENT_ID shared-infra output —
COGNITO_REGION matches AWS_REGION —
AWS_REGION CDK env ap-northeast-1
WLK_MODEL hard-coded large-v3
WLK_LANGUAGE DynamoDB-persisted (default) en
TGT_LANG DynamoDB-persisted (default) ja
GPU_ASG_NAME stack output (split / ecs-gpu modes only) —
PROFILE_TABLE_NAME stack output HibikiTranslateProfiles

Verifying a deploy

After cdk deploy completes, the stack outputs are printed. Capture:

Smoke tests:

  1. Health check: curl https://translate.liveprod.cloud/health → {"status": "ok", ...}
  2. gRPC liveness (from a VPC-internal task): python tests/spike_grpc_streaming_client.py health
  3. Operator UI: open https://translate.liveprod.cloud/translate in a browser; log in; click mic; speak.

If gRPC is unreachable from inside the VPC, check: - The Cloud Map service is registered (aws servicediscovery list-services) - The NLB security group allows the VPC CIDR on port 50053 - The Fargate task is running and healthy (aws ecs describe-services)


Operational behavior

The split-mode stack is configured to absorb mid-event deploys, task crashes, and translator-tier failures without dropping operator sessions. The behaviors below are active in code; nothing extra to configure post-deploy.

Graceful shutdown drain

When ECS replaces a task (deploy, scale-down, health-check failure) it sends SIGTERM, waits for the container to exit, then SIGKILLs if it doesn't.

The engine's server.py and the bridge's app.py both intercept SIGTERM and:

Container stop_timeout is set to 90 seconds on the gRPC and WLK containers (60s drain + 30s scheduler overhead). The web container uses 30s — it doesn't hold long-running state.

NLB target deregistration delay is also 90s, matching the container budget. ALB target on the bridge tier uses the same value.

Service deployment configuration

Service min_healthy_percent max_healthy_percent desired_count Notes
Engine gRPC (Fargate) 100 200 1 New task comes up before old task is stopped — no capacity dip during deploys.
Engine Web (Fargate) 100 200 1 Same.
Engine WLK (EC2 GPU) 50 (default) 200 (default) 0 (operator-controlled) Single-task service; the operator-controlled lifecycle means deploy-time replacements are rare.
hibiki-stage Web (Fargate) 100 200 2 Two tasks at all times. ALB load-balances; if one task crashes, the survivor immediately accepts new connections — recovery in ~10s instead of the 2-3 min Fargate cold-start a single-task service would have.

CloudWatch alarms (split mode)

Three alarms are created in cdk/stacks/split_stack.py and routed to a single SNS topic (OpsAlarms). The topic email subscription is set via -c opsAlarmEmail=... at deploy time; if not provided, the topic is created without a subscriber and you can subscribe manually post-deploy.

Alarm Trigger Action
LineReanchorAlarm More than 30 [STT] line[N] re-anchored / drift detected log lines per minute, sustained 2+ min. Audio source quality is degrading (loud room, mic issue, paused source). Operator should check input.
ChainFailureAlarm More than 5 [CHAIN] all hedges failed log lines per 5-min window. Bedrock degraded. Bump translator timeouts or swap to alternate region. See runbook.md.
OperatorSilentAlarm Zero [Session ...] Output # log lines for 5 consecutive minutes during an active session. Active session has gone silent. Catches dead sessions whether the cause is engine, bridge, or audio source.

Metrics are derived via CloudWatch Logs metric filters — pattern-match against existing log lines. No boto3.put_metric_data instrumentation in the runtime; all observability comes from log content.

The runbook (docs/runbook.md) has diagnostic ladders for each alarm.

Operator session resilience

The operator UI (in hibiki-stage) has its own resilience layer that pairs with the engine drain:

These are bridge-side; the engine doesn't need to know about them. Documented here because operators occasionally ask why a deploy didn't drop their session.


Cleanup

To tear down a stack:

cdk destroy -c deployMode=split -c region=us-west-2

This removes: - Fargate / ECS / ASG resources - ALB + CloudFront - DynamoDB tables (unless RemovalPolicy.RETAIN is set, which it is by default for the profile table) - IAM roles, log groups, security groups

It does not remove: - The user pool (RETAIN policy by default — protect against accidental loss of users) - The shared-infra resources (those live in their own stack) - The Route53 zone (if shared-infra owns it)


Reference