In Java you drop a WAR into a Tomcat that is already running and already paid for:
the server is a sunk cost, and undeploying frees nothing. Fargate has no server for the cost to
sink into. Every priced box below starts billing by the hour when terraform apply
returns, whether or not a request ever arrives. The load balancer bills the most, and it is
also the box you are most likely to forget. So this is two days of one habit. Day 15
(cloud-06) writes the module and reads the plan's cost line by line before anything
exists. Day 16 (cloud-07) applies it, runs py-10's live test against the public URL, records
three numbers, and destroys it the same session.
Fifteen Terraform resources in ap-south-1, built into
cloud-04's VPC, using only its two public subnets. No
NAT gateway is involved. An internet-facing Application Load Balancer
(aws_lb) sits in both subnets and has a listener on :80 that forwards to a
target group of IP targets on :8000. The target group health-checks /readyz.
An ECS cluster runs a service with desired_count = 1. The service
starts tasks from a task definition: 0.25 vCPU and 0.5 GB, running py-10's image from
cloud-05's repository, with /readyz as the
container health check as well. The container's stdout goes to a CloudWatch log group
with 7-day retention. A task execution role may pull from that one repository and write
to that one log group, and nothing else. Two security groups let the internet reach the
ALB on :80, and let the ALB (and only the ALB) reach the task on :8000.
Where the money goes: Fargate bills vCPU-hours and GB-hours per second, with a one-minute minimum. The ALB bills every hour or partial hour it exists, plus LCUs for traffic. And every public IPv4 address bills $0.005/h: the task has one, because in a public subnet with no NAT that address is its only route to ECR and CloudWatch. The ALB has one per subnet. AWS's ELB pricing page says: "You will incur standard public IPv4 address charges for all the addresses you consume with load balancers." The meter uses AWS's published us-east-1 rates, which I read on 2026-09-11: Fargate Linux/x86 at $0.000011244 per vCPU-second and $0.000001235 per GB-second (so $0.04048 per vCPU-hour and $0.004446 per GB-hour), an ALB at $0.0225 per hour plus $0.008 per LCU-hour, public IPv4 at $0.005 per hour, and CloudWatch Logs at $0.50 per GB ingested (5 GB free). A month is 720 hours. Prices in ap-south-1 differ, so they are approx.: verify on the pricing page today. The cold-start line is an estimate. Day 16 measures your real number.
terraform apply before that e-mail.
On Day 16 the alarm will fire, because this is the first item in the track that
costs real money. That e-mail is the alarm working.cd ~/lbv-cloud/04-vpc && terraform output
public_subnet_ids lists two subnets. cloud-05 is applied: aws ecr
describe-repositories --repository-names lbv-api answers, and v1 is
in it.exercises/py-10-fastapi-service,
uv run pytest -q prints 9 passed (or 8 passed, 1 skipped, plus 3 deselected).
That folder must contain integration/test_live.py. It landed together with this
page, so pull the repo if you cloned it earlier. Docker Desktop is running.terraform version prints 1.11 or later, and AWS_PROFILE=lbv-admin
is an assumed role, not :root (the same check as cloud-03 to cloud-05).Nothing is created today. terraform plan only reads, and costs $0.00. The module
is deliberately flat: no for_each, no count, one resource per thing in
the scene, so the plan reads top to bottom as the list of what would bill. Layout, next to
cloud-04's module:
~/lbv-cloud/ modules/vpc/ # cloud-04 modules/fargate-service/ # new: variables.tf, main.tf, outputs.tf 04-vpc/ 05-ecr/ 06-fargate/ # new: versions.tf, main.tf, outputs.tf
1. The module's inputs — modules/fargate-service/variables.tf. The defaults
are the sizes the meter starts from:
# modules/fargate-service/variables.tf variable "name" { type = string description = "Prefix for every resource name" } variable "region" { type = string } variable "vpc_id" { type = string } variable "subnet_ids" { type = list(string) description = "Public subnets in two AZs: the ALB needs two, and the tasks run in them with a public IP" } variable "image" { type = string description = "Full image reference, <repository-url>:<tag>" } variable "ecr_repository_arn" { type = string } variable "container_port" { type = number default = 8000 # py-10's Dockerfile: uvicorn --port 8000 } variable "cpu" { type = number default = 256 # CPU units: 1024 = 1 vCPU, so 0.25 vCPU } variable "memory" { type = number default = 512 # MiB: 0.5 GB } variable "desired_count" { type = number default = 1 } variable "health_path" { type = string default = "/readyz" # 503 until the model has loaded, 200 after } variable "log_retention_days" { type = number default = 7 }
2. The resources — modules/fargate-service/main.tf. Read each comment
before you paste. The two security groups are the part people get wrong. Terraform removes
the allow-all egress rule that AWS puts on a new group, so every way out is written down here.
The task needs HTTPS out to ECR, to the S3 bucket that holds image layers, and to CloudWatch
Logs. Without that rule, the image pull times out.
# modules/fargate-service/main.tf # ---- logs: where the container's stdout lands --------------------------------------- resource "aws_cloudwatch_log_group" "this" { name = "/ecs/${var.name}" retention_in_days = var.log_retention_days # 7: old lines are deleted, so storage cannot pile up } # ---- the task execution role: ECS itself uses it to pull the image and ship the logs -- resource "aws_iam_role" "execution" { name = "${var.name}-task-execution" assume_role_policy = jsonencode({ Version = "2012-10-17" Statement = [{ Effect = "Allow" Principal = { Service = "ecs-tasks.amazonaws.com" } Action = "sts:AssumeRole" }] }) } resource "aws_iam_role_policy" "execution" { name = "ecr-pull-and-logs-only" role = aws_iam_role.execution.id policy = jsonencode({ Version = "2012-10-17" Statement = [ { # the login token is account-wide: this one action cannot be scoped to a repository Effect = "Allow" Action = "ecr:GetAuthorizationToken" Resource = "*" }, { # pull from lbv-api, and nothing else Effect = "Allow" Action = ["ecr:BatchCheckLayerAvailability", "ecr:GetDownloadUrlForLayer", "ecr:BatchGetImage"] Resource = var.ecr_repository_arn }, { # write streams into this one log group Effect = "Allow" Action = ["logs:CreateLogStream", "logs:PutLogEvents"] Resource = "${aws_cloudwatch_log_group.this.arn}:*" }, ] }) } # ---- security groups: internet -> ALB :80 -> task :8000, and the task's HTTPS out ----- resource "aws_security_group" "alb" { name = "${var.name}-alb" description = "HTTP from the internet to the ALB" vpc_id = var.vpc_id } resource "aws_vpc_security_group_ingress_rule" "alb_http" { security_group_id = aws_security_group.alb.id cidr_ipv4 = "0.0.0.0/0" ip_protocol = "tcp" from_port = 80 to_port = 80 } resource "aws_vpc_security_group_egress_rule" "alb_to_task" { security_group_id = aws_security_group.alb.id referenced_security_group_id = aws_security_group.task.id ip_protocol = "tcp" from_port = var.container_port to_port = var.container_port } resource "aws_security_group" "task" { name = "${var.name}-task" description = "The container port, from the ALB only" vpc_id = var.vpc_id } resource "aws_vpc_security_group_ingress_rule" "task_from_alb" { security_group_id = aws_security_group.task.id referenced_security_group_id = aws_security_group.alb.id # the task's public IP answers nobody else ip_protocol = "tcp" from_port = var.container_port to_port = var.container_port } resource "aws_vpc_security_group_egress_rule" "task_https_out" { security_group_id = aws_security_group.task.id cidr_ipv4 = "0.0.0.0/0" ip_protocol = "tcp" from_port = 443 to_port = 443 description = "ECR, image layers in S3, CloudWatch Logs: all HTTPS" } # ---- the load balancer: a stable DNS name in front of tasks whose IPs change ---------- resource "aws_lb" "this" { name = var.name load_balancer_type = "application" internal = false subnets = var.subnet_ids # two AZs: one ALB node, and one public IPv4, in each security_groups = [aws_security_group.alb.id] } resource "aws_lb_target_group" "this" { name = var.name port = var.container_port protocol = "HTTP" target_type = "ip" # a Fargate task registers by its ENI's IP; there is no instance vpc_id = var.vpc_id deregistration_delay = 5 # default 300 s: destroy would wait five minutes draining each task health_check { path = var.health_path matcher = "200" interval = 10 timeout = 5 healthy_threshold = 2 # two passing checks, 10 s apart, before any traffic unhealthy_threshold = 2 } } resource "aws_lb_listener" "http" { load_balancer_arn = aws_lb.this.arn port = 80 protocol = "HTTP" default_action { type = "forward" target_group_arn = aws_lb_target_group.this.arn } } # ---- ECS: a cluster (a namespace), a task definition (a document), a service (the loop) - resource "aws_ecs_cluster" "this" { name = var.name } resource "aws_ecs_task_definition" "this" { family = var.name requires_compatibilities = ["FARGATE"] network_mode = "awsvpc" cpu = var.cpu memory = var.memory execution_role_arn = aws_iam_role.execution.arn runtime_platform { operating_system_family = "LINUX" cpu_architecture = "X86_64" # the linux/amd64 image you cross-build on the Mac } container_definitions = jsonencode([{ name = "api" image = var.image essential = true portMappings = [{ containerPort = var.container_port, protocol = "tcp" }] logConfiguration = { logDriver = "awslogs" options = { awslogs-group = aws_cloudwatch_log_group.this.name awslogs-region = var.region awslogs-stream-prefix = "api" } } healthCheck = { # python:3.12-slim has no curl; urlopen raises on a 503, so its exit code is the verdict command = ["CMD-SHELL", "python -c \"import urllib.request; urllib.request.urlopen('http://127.0.0.1:${var.container_port}${var.health_path}', timeout=2)\" || exit 1"] interval = 10 timeout = 5 retries = 3 startPeriod = 10 } }]) } resource "aws_ecs_service" "this" { name = var.name cluster = aws_ecs_cluster.this.id task_definition = aws_ecs_task_definition.this.arn desired_count = var.desired_count launch_type = "FARGATE" network_configuration { subnets = var.subnet_ids security_groups = [aws_security_group.task.id] assign_public_ip = true # no NAT: this address is the task's only road to ECR and CloudWatch } load_balancer { target_group_arn = aws_lb_target_group.this.arn container_name = "api" container_port = var.container_port } health_check_grace_period_seconds = 30 # ignore the ALB's verdict while the model loads deployment_circuit_breaker { # a deployment whose tasks keep failing is stopped, not retried forever enable = true rollback = true } depends_on = [aws_lb_listener.http] # the target group must be attached to the ALB first }
3. What the root reads back — modules/fargate-service/outputs.tf:
# modules/fargate-service/outputs.tf
output "alb_dns_name" {
value = aws_lb.this.dns_name
}
output "target_group_arn" {
value = aws_lb_target_group.this.arn
}
output "cluster_name" {
value = aws_ecs_cluster.this.name
}
output "service_name" {
value = aws_ecs_service.this.name
}
4. The root module's pins and backend — 06-fargate/versions.tf. This is the
same bucket as cloud-03 to cloud-05, with a fourth key. Change only the bucket name:
# 06-fargate/versions.tf
terraform {
required_version = "~> 1.11"
required_providers {
aws = {
source = "hashicorp/aws"
version = "~> 6.0"
}
}
backend "s3" {
bucket = "lbv-tfstate-yourname-x7q2"
key = "lbv-track/06-fargate/terraform.tfstate"
region = "ap-south-1"
use_lockfile = true
encrypt = true
}
}
provider "aws" {
region = "ap-south-1"
default_tags {
tags = {
Project = "lbv-track"
Item = "cloud-06"
}
}
}
5. Wire it — 06-fargate/main.tf. The VPC's IDs come from cloud-04's
state, so nothing is copied by hand. The repository is looked up by name, because
cloud-05 already made it. The image tag is a variable. v2 is py-10's service,
which Day 16's first step pushes. (v1 is cloud-01's app. It has
/health but no /readyz, so the health check would never pass on
it.)
# 06-fargate/main.tf
variable "image_tag" {
type = string
default = "v2"
}
data "terraform_remote_state" "vpc" {
backend = "s3"
config = {
bucket = "lbv-tfstate-yourname-x7q2"
key = "lbv-track/04-vpc/terraform.tfstate"
region = "ap-south-1"
}
}
data "aws_ecr_repository" "api" {
name = "lbv-api"
}
module "api" {
source = "../modules/fargate-service"
name = "lbv-api"
region = "ap-south-1"
vpc_id = data.terraform_remote_state.vpc.outputs.vpc_id
subnet_ids = data.terraform_remote_state.vpc.outputs.public_subnet_ids
image = "${data.aws_ecr_repository.api.repository_url}:${var.image_tag}"
ecr_repository_arn = data.aws_ecr_repository.api.arn
}
# 06-fargate/outputs.tf
output "alb_dns_name" {
value = module.api.alb_dns_name
}
output "target_group_arn" {
value = module.api.target_group_arn
}
output "cluster_name" {
value = module.api.cluster_name
}
output "service_name" {
value = module.api.service_name
}
6. Init, check, plan to a file. -out tfplan saves exactly this plan, so
Day 16's apply does what you read today and nothing else:
cd ~/lbv-cloud/06-fargate export AWS_PROFILE=lbv-admin terraform init terraform fmt ../modules/fargate-service && terraform fmt && terraform validate # Success! The configuration is valid. terraform plan -out tfplan # ... # Plan: 15 to add, 0 to change, 0 to destroy. # # Saved the plan to: tfplan
7. Read the plan's cost, line by line. Take each + create block in the plan
output and find it in this table. Twelve of the fifteen resources cost nothing on their own.
Three lines carry the whole bill: the ALB, the tasks the service starts, and the public IPv4
addresses both of them hold. That is the answer cloud-06 asks for ("the hourly cost of ALB +
task + public IP").
| # | planned resource (module.api.…) | hourly, approx. | why it exists |
|---|---|---|---|
| 1 | aws_cloudwatch_log_group.this | $0.00 — $0.50/GB ingested, first 5 GB free; today's traffic is kilobytes | where stdout lands; 7-day retention so it cannot grow forever |
| 2 | aws_iam_role.execution | $0.00 | the identity ECS assumes to pull the image and write logs |
| 3 | aws_iam_role_policy.execution | $0.00 | least privilege: pull from lbv-api, write to this log group, nothing else |
| 4 | aws_security_group.alb | $0.00 | the ALB's firewall |
| 5 | aws_vpc_security_group_ingress_rule.alb_http | $0.00 | the only door from the internet: :80 |
| 6 | aws_vpc_security_group_egress_rule.alb_to_task | $0.00 | the ALB may only talk to the tasks, on :8000 |
| 7 | aws_security_group.task | $0.00 | the task's firewall |
| 8 | aws_vpc_security_group_ingress_rule.task_from_alb | $0.00 | :8000 from the ALB's group only; the task's public IP answers nobody else |
| 9 | aws_vpc_security_group_egress_rule.task_https_out | $0.00 | HTTPS out to ECR, S3 and CloudWatch; without it the pull times out |
| 10 | aws_lb.this | $0.0225/h + $0.008 per LCU-hour, + 2 public IPv4 × $0.005/h | one stable DNS name, health-checked routing; bills every hour it exists, traffic or not |
| 11 | aws_lb_target_group.this | $0.00 | the list of task IPs the ALB may send to, and the /readyz check that admits them |
| 12 | aws_lb_listener.http | $0.00 | :80 → forward to the target group |
| 13 | aws_ecs_cluster.this | $0.00 | a namespace; ECS itself carries no charge, Fargate bills the tasks |
| 14 | aws_ecs_task_definition.this | $0.00 | a document: image, size, port, logs, health check |
| 15 | aws_ecs_service.this | per task: 0.25 × $0.04048 + 0.5 × $0.004446 = $0.0123/h, + 1 public IPv4 × $0.005/h | keeps desired_count tasks running and registers them with the target group |
While applied, all 15 cost ~$0.05/h: $0.0123 for the task, $0.0225 for the ALB and $0.015 for three public IPv4 addresses. Left on for a 720-hour month, that is ~$35.89, which the meter's step 2 shows. These are us-east-1 rates, approx.: verify on the pricing page today.
terraform show -no-color tfplan | grep '^Plan:' prints
Plan: 15 to add, 0 to change, 0 to destroy. These are the fifteen rows above, and
no more.terraform show -json tfplan | jq -r '.resource_changes[] | select(.mode == "managed") | .type' | sort | uniq -c
# 1 aws_cloudwatch_log_group
# 1 aws_ecs_cluster
# 1 aws_ecs_service
# 1 aws_ecs_task_definition
# 1 aws_iam_role
# 1 aws_iam_role_policy
# 1 aws_lb
# 1 aws_lb_listener
# 1 aws_lb_target_group
# 2 aws_security_group
# 2 aws_vpc_security_group_egress_rule
# 2 aws_vpc_security_group_ingress_ruleterraform show -json tfplan | jq '[.resource_changes[] | select(.type == "aws_nat_gateway" or .type == "aws_eip" or .type == "aws_db_instance")] | length'
# 0There is no teardown for cloud-06: a plan creates nothing, and
tfplan is a file on your Mac. Press the button once the three Verify lines above
are true and you can say, from the table, what the ALB, the task and a public IPv4 address
each cost per hour.
This is self-attestation. The site cannot see your terminal, so pressing the button is you telling The Path that the plan reads 15 to add and that the jq check printed 0.
The meter runs from step 2 to step 7. At the rates above, the whole session costs about five cents. The one expensive mistake is to stop before step 7.
1. Push py-10's service as lbv-api:v2. This is the same login and the same
cross-build as cloud-05, but run from py-10's folder. Its main.py loads the model
about 3 s after start (until then /readyz answers 503) and writes JSON log lines.
The tag is immutable, so v2 can never be re-pointed. Like v1, it
matches no lifecycle rule:
cd ~/path/to/ai-ml-roadmap/exercises/py-10-fastapi-service # your finished py-10 export AWS_PROFILE=lbv-admin ACCT=$(aws sts get-caller-identity --query Account --output text) REPO=$ACCT.dkr.ecr.ap-south-1.amazonaws.com/lbv-api aws ecr get-login-password --region ap-south-1 \ | docker login --username AWS --password-stdin "$ACCT.dkr.ecr.ap-south-1.amazonaws.com" docker buildx build --platform linux/amd64 --provenance=false --push -t "$REPO:v2" . aws ecr describe-images --repository-name lbv-api --image-ids imageTag=v2 \ --query 'imageDetails[0].imageSizeInBytes' # the compressed size Fargate will pull; the meter guesses ~55 MB
2. Apply the saved plan, and start the clock. If Terraform says the saved plan is stale
(the state changed since Day 15), run terraform plan -out tfplan again and re-read
the count first:
cd ~/lbv-cloud/06-fargate
START=$(date +%s)
terraform apply tfplan
# ...
# Apply complete! Resources: 15 added, 0 changed, 0 destroyed.
# Outputs:
# alb_dns_name = "lbv-api-<id>.ap-south-1.elb.amazonaws.com"
3. Wait for the target to read healthy — number one, the cold start. This covers
everything from apply to the first target that passes two /readyz
checks: creating the ALB, placing the task, pulling the image, loading the model. The meter's
~55 s estimate covers only the part after the task is placed. Creating the ALB comes first and
usually takes longer, so your number will be bigger:
TG=$(terraform output -raw target_group_arn)
until [ "$(aws elbv2 describe-target-health --target-group-arn "$TG" \
--query 'length(TargetHealthDescriptions[?TargetHealth.State==`healthy`])' --output text)" -ge 1 ]; do
sleep 5
done
echo "cold start: $(( $(date +%s) - START )) s from apply to first healthy target"
aws elbv2 describe-target-health --target-group-arn "$TG" \
--query 'TargetHealthDescriptions[].[Target.Id,TargetHealth.State]' --output text
# 10.0.x.x healthy
4. Run py-10's live test against the ALB — number two, the p50. The three tests in
integration/ check /healthz and /readyz, make 20
/ask calls that must echo the X-Request-Id you send, and time the
/stream first chunk. They report their numbers at the end of the run:
ALB=$(terraform output -raw alb_dns_name) curl -s "http://$ALB/healthz" # {"status":"ok"} cd ~/path/to/ai-ml-roadmap/exercises/py-10-fastapi-service BASE_URL="http://$ALB" uv run pytest -q -m integration # ... [100%] # ============== live numbers (record them for day 20) ============== # /ask p50 over 20 calls: <n> ms (min <n>, max <n>) # last request id (grep the logs for it): live-<8 hex> # /stream first chunk: <n> ms (whole stream <n> ms, <n> reads) # 3 passed, 9 deselected
5. Watch the logs. Each /ask writes one JSON line with its request id.
Find the id the test printed:
aws logs tail /ecs/lbv-api --since 10m --follow
# ... api/api/<task-id> {"ts": "...", "level": "INFO", "logger": "askapi", "msg": "ask served", "request_id": "live-<8 hex>"}
# Ctrl-C to stop following
6. Break one deployment on purpose (cloud-07's "a deliberate bad image tag"). Point the service at a tag that does not exist. That replaces the task definition, and the service tries to start tasks from the new one. The old task keeps serving while the new ones fail, and the circuit breaker stops the deployment after a few failures and rolls back. Diagnose it from ECS, not from the ALB:
cd ~/lbv-cloud/06-fargate terraform apply -var image_tag=does-not-exist # Plan: 1 to add, 1 to change, 1 to destroy. (task definition replaced, service updated) curl -s "http://$ALB/healthz" # still {"status":"ok"}: the old task is still serving aws ecs describe-services --cluster lbv-api --services lbv-api \ --query 'services[0].events[:5].[createdAt,message]' --output text # recent events: tasks started, then stopped; after a few, the deployment is marked failed and rolled back aws ecs describe-tasks --cluster lbv-api --query 'tasks[].stoppedReason' --output text --tasks \ $(aws ecs list-tasks --cluster lbv-api --desired-status STOPPED --query 'taskArns[]' --output text) # CannotPullContainerError: ... lbv-api:does-not-exist ... not found
The exact wording of the events changes between ECS releases. The stopped
task's stoppedReason is the line that names the cause. No need to re-apply
v2: step 7 destroys everything anyway.
7. Destroy — the same session.
cd ~/lbv-cloud/06-fargate
terraform destroy
# Plan: 0 to add, 0 to change, 15 to destroy. ... Destroy complete! Resources: 15 destroyed.
8. Number three, the bill: tomorrow morning. Open Billing → Bills (or Cost Explorer, grouped by service, daily). You should find one day with lines under Elastic Container Service, Elastic Load Balancing and the VPC's public IPv4 addresses, together about five cents. Write the three numbers down for the day-20 write-up: the cold start, the p50 and the bill line.
describe-target-health read healthy, and the
integration run ended 3 passed, 9 deselected.aws logs tail showed an "ask served" line carrying
the request id the test printed.CannotPullContainerError, while /healthz on the ALB still answered.Step 7 is the teardown. Check that it really left nothing (the plan item's teardown: ALB, ENI, log group, EIP all gone; check Cost Explorer next day):
aws ecs list-services --cluster lbv-api --output text # prints nothing, or ClusterNotFoundException: either way no service is left aws elbv2 describe-load-balancers --query 'LoadBalancers[].LoadBalancerName' # [] VPC=$(cd ~/lbv-cloud/04-vpc && terraform output -raw vpc_id) aws ec2 describe-network-interfaces --filters Name=vpc-id,Values=$VPC \ --query 'NetworkInterfaces[].[NetworkInterfaceId,Description]' --output text # nothing (an "ELB app/lbv-api/..." ENI can linger a few minutes after the ALB is gone; re-run) aws logs describe-log-groups --log-group-name-prefix /ecs/lbv-api --query 'logGroups[].logGroupName' # [] aws ec2 describe-addresses --query 'Addresses[].PublicIp' # []: this HCL never made an Elastic IP. The task's and the ALB's addresses were AWS-owned and went with them
What remains and what it costs, approx., verify on the pricing page today:
cloud-04's VPC ($0.00: nothing runs in it), cloud-05's repository with v1 and
v2 in it (~$0.10 per GB-month on the compressed sizes, a few cents), and this
item's state file in the cloud-03 bucket. Nothing in cloud-06/07 bills after the destroy.
This is self-attestation. The site cannot see your AWS account, so checking the box and pressing the button is you telling The Path that the live test passed and that the service is destroyed.
destroy stops that. Health and readiness matter here
too: the target group and the container check both probe /readyz, so no request
reaches a task before its model is loaded. The cold start you measure is mostly waiting for
things that are not your code: the ALB being created, the task being placed, the image being
pulled, two health checks ten seconds apart.