Cloud 06/07 — py-10's service on ECS Fargate behind a load balancer: read the plan, run it for an hour, destroy it

In Java you drop a WAR into a Tomcat that is already running and already paid for: the server is a sunk cost, and undeploying frees nothing. Fargate has no server for the cost to sink into. Every priced box below starts billing by the hour when terraform apply returns, whether or not a request ever arrives. The load balancer bills the most, and it is also the box you are most likely to forget. So this is two days of one habit. Day 15 (cloud-06) writes the module and reads the plan's cost line by line before anything exists. Day 16 (cloud-07) applies it, runs py-10's live test against the public URL, records three numbers, and destroys it the same session.

~90 min + ~90 min~$0.05 for the hour (approx.) 15 resourcesdestroy the same day cloud-06 · cloud-07

What this creates

Fifteen Terraform resources in ap-south-1, built into cloud-04's VPC, using only its two public subnets. No NAT gateway is involved. An internet-facing Application Load Balancer (aws_lb) sits in both subnets and has a listener on :80 that forwards to a target group of IP targets on :8000. The target group health-checks /readyz. An ECS cluster runs a service with desired_count = 1. The service starts tasks from a task definition: 0.25 vCPU and 0.5 GB, running py-10's image from cloud-05's repository, with /readyz as the container health check as well. The container's stdout goes to a CloudWatch log group with 7-day retention. A task execution role may pull from that one repository and write to that one log group, and nothing else. Two security groups let the internet reach the ALB on :80, and let the ALB (and only the ALB) reach the task on :8000.

Where the money goes: Fargate bills vCPU-hours and GB-hours per second, with a one-minute minimum. The ALB bills every hour or partial hour it exists, plus LCUs for traffic. And every public IPv4 address bills $0.005/h: the task has one, because in a public subnet with no NAT that address is its only route to ECR and CloudWatch. The ALB has one per subnet. AWS's ELB pricing page says: "You will incur standard public IPv4 address charges for all the addresses you consume with load balancers." The meter uses AWS's published us-east-1 rates, which I read on 2026-09-11: Fargate Linux/x86 at $0.000011244 per vCPU-second and $0.000001235 per GB-second (so $0.04048 per vCPU-hour and $0.004446 per GB-hour), an ALB at $0.0225 per hour plus $0.008 per LCU-hour, public IPv4 at $0.005 per hour, and CloudWatch Logs at $0.50 per GB ingested (5 GB free). A month is 720 hours. Prices in ap-south-1 differ, so they are approx.: verify on the pricing page today. The cold-start line is an estimate. Day 16 measures your real number.

Preconditions

Day 15 · cloud-06 — write the module, plan it, read the cost

Nothing is created today. terraform plan only reads, and costs $0.00. The module is deliberately flat: no for_each, no count, one resource per thing in the scene, so the plan reads top to bottom as the list of what would bill. Layout, next to cloud-04's module:

~/lbv-cloud/
  modules/vpc/                 # cloud-04
  modules/fargate-service/     # new: variables.tf, main.tf, outputs.tf
  04-vpc/  05-ecr/
  06-fargate/                  # new: versions.tf, main.tf, outputs.tf

1. The module's inputs — modules/fargate-service/variables.tf. The defaults are the sizes the meter starts from:

# modules/fargate-service/variables.tf
variable "name" {
  type        = string
  description = "Prefix for every resource name"
}

variable "region" {
  type = string
}

variable "vpc_id" {
  type = string
}

variable "subnet_ids" {
  type        = list(string)
  description = "Public subnets in two AZs: the ALB needs two, and the tasks run in them with a public IP"
}

variable "image" {
  type        = string
  description = "Full image reference, <repository-url>:<tag>"
}

variable "ecr_repository_arn" {
  type = string
}

variable "container_port" {
  type    = number
  default = 8000 # py-10's Dockerfile: uvicorn --port 8000
}

variable "cpu" {
  type    = number
  default = 256 # CPU units: 1024 = 1 vCPU, so 0.25 vCPU
}

variable "memory" {
  type    = number
  default = 512 # MiB: 0.5 GB
}

variable "desired_count" {
  type    = number
  default = 1
}

variable "health_path" {
  type    = string
  default = "/readyz" # 503 until the model has loaded, 200 after
}

variable "log_retention_days" {
  type    = number
  default = 7
}

2. The resources — modules/fargate-service/main.tf. Read each comment before you paste. The two security groups are the part people get wrong. Terraform removes the allow-all egress rule that AWS puts on a new group, so every way out is written down here. The task needs HTTPS out to ECR, to the S3 bucket that holds image layers, and to CloudWatch Logs. Without that rule, the image pull times out.

# modules/fargate-service/main.tf

# ---- logs: where the container's stdout lands ---------------------------------------
resource "aws_cloudwatch_log_group" "this" {
  name              = "/ecs/${var.name}"
  retention_in_days = var.log_retention_days # 7: old lines are deleted, so storage cannot pile up
}

# ---- the task execution role: ECS itself uses it to pull the image and ship the logs --
resource "aws_iam_role" "execution" {
  name = "${var.name}-task-execution"
  assume_role_policy = jsonencode({
    Version = "2012-10-17"
    Statement = [{
      Effect    = "Allow"
      Principal = { Service = "ecs-tasks.amazonaws.com" }
      Action    = "sts:AssumeRole"
    }]
  })
}

resource "aws_iam_role_policy" "execution" {
  name = "ecr-pull-and-logs-only"
  role = aws_iam_role.execution.id
  policy = jsonencode({
    Version = "2012-10-17"
    Statement = [
      { # the login token is account-wide: this one action cannot be scoped to a repository
        Effect   = "Allow"
        Action   = "ecr:GetAuthorizationToken"
        Resource = "*"
      },
      { # pull from lbv-api, and nothing else
        Effect   = "Allow"
        Action   = ["ecr:BatchCheckLayerAvailability", "ecr:GetDownloadUrlForLayer", "ecr:BatchGetImage"]
        Resource = var.ecr_repository_arn
      },
      { # write streams into this one log group
        Effect   = "Allow"
        Action   = ["logs:CreateLogStream", "logs:PutLogEvents"]
        Resource = "${aws_cloudwatch_log_group.this.arn}:*"
      },
    ]
  })
}

# ---- security groups: internet -> ALB :80 -> task :8000, and the task's HTTPS out -----
resource "aws_security_group" "alb" {
  name        = "${var.name}-alb"
  description = "HTTP from the internet to the ALB"
  vpc_id      = var.vpc_id
}

resource "aws_vpc_security_group_ingress_rule" "alb_http" {
  security_group_id = aws_security_group.alb.id
  cidr_ipv4         = "0.0.0.0/0"
  ip_protocol       = "tcp"
  from_port         = 80
  to_port           = 80
}

resource "aws_vpc_security_group_egress_rule" "alb_to_task" {
  security_group_id            = aws_security_group.alb.id
  referenced_security_group_id = aws_security_group.task.id
  ip_protocol                  = "tcp"
  from_port                    = var.container_port
  to_port                      = var.container_port
}

resource "aws_security_group" "task" {
  name        = "${var.name}-task"
  description = "The container port, from the ALB only"
  vpc_id      = var.vpc_id
}

resource "aws_vpc_security_group_ingress_rule" "task_from_alb" {
  security_group_id            = aws_security_group.task.id
  referenced_security_group_id = aws_security_group.alb.id # the task's public IP answers nobody else
  ip_protocol                  = "tcp"
  from_port                    = var.container_port
  to_port                      = var.container_port
}

resource "aws_vpc_security_group_egress_rule" "task_https_out" {
  security_group_id = aws_security_group.task.id
  cidr_ipv4         = "0.0.0.0/0"
  ip_protocol       = "tcp"
  from_port         = 443
  to_port           = 443
  description       = "ECR, image layers in S3, CloudWatch Logs: all HTTPS"
}

# ---- the load balancer: a stable DNS name in front of tasks whose IPs change ----------
resource "aws_lb" "this" {
  name               = var.name
  load_balancer_type = "application"
  internal           = false
  subnets            = var.subnet_ids # two AZs: one ALB node, and one public IPv4, in each
  security_groups    = [aws_security_group.alb.id]
}

resource "aws_lb_target_group" "this" {
  name                 = var.name
  port                 = var.container_port
  protocol             = "HTTP"
  target_type          = "ip" # a Fargate task registers by its ENI's IP; there is no instance
  vpc_id               = var.vpc_id
  deregistration_delay = 5 # default 300 s: destroy would wait five minutes draining each task

  health_check {
    path                = var.health_path
    matcher             = "200"
    interval            = 10
    timeout             = 5
    healthy_threshold   = 2 # two passing checks, 10 s apart, before any traffic
    unhealthy_threshold = 2
  }
}

resource "aws_lb_listener" "http" {
  load_balancer_arn = aws_lb.this.arn
  port              = 80
  protocol          = "HTTP"

  default_action {
    type             = "forward"
    target_group_arn = aws_lb_target_group.this.arn
  }
}

# ---- ECS: a cluster (a namespace), a task definition (a document), a service (the loop) -
resource "aws_ecs_cluster" "this" {
  name = var.name
}

resource "aws_ecs_task_definition" "this" {
  family                   = var.name
  requires_compatibilities = ["FARGATE"]
  network_mode             = "awsvpc"
  cpu                      = var.cpu
  memory                   = var.memory
  execution_role_arn       = aws_iam_role.execution.arn

  runtime_platform {
    operating_system_family = "LINUX"
    cpu_architecture        = "X86_64" # the linux/amd64 image you cross-build on the Mac
  }

  container_definitions = jsonencode([{
    name         = "api"
    image        = var.image
    essential    = true
    portMappings = [{ containerPort = var.container_port, protocol = "tcp" }]
    logConfiguration = {
      logDriver = "awslogs"
      options = {
        awslogs-group         = aws_cloudwatch_log_group.this.name
        awslogs-region        = var.region
        awslogs-stream-prefix = "api"
      }
    }
    healthCheck = {
      # python:3.12-slim has no curl; urlopen raises on a 503, so its exit code is the verdict
      command     = ["CMD-SHELL", "python -c \"import urllib.request; urllib.request.urlopen('http://127.0.0.1:${var.container_port}${var.health_path}', timeout=2)\" || exit 1"]
      interval    = 10
      timeout     = 5
      retries     = 3
      startPeriod = 10
    }
  }])
}

resource "aws_ecs_service" "this" {
  name            = var.name
  cluster         = aws_ecs_cluster.this.id
  task_definition = aws_ecs_task_definition.this.arn
  desired_count   = var.desired_count
  launch_type     = "FARGATE"

  network_configuration {
    subnets          = var.subnet_ids
    security_groups  = [aws_security_group.task.id]
    assign_public_ip = true # no NAT: this address is the task's only road to ECR and CloudWatch
  }

  load_balancer {
    target_group_arn = aws_lb_target_group.this.arn
    container_name   = "api"
    container_port   = var.container_port
  }

  health_check_grace_period_seconds = 30 # ignore the ALB's verdict while the model loads

  deployment_circuit_breaker { # a deployment whose tasks keep failing is stopped, not retried forever
    enable   = true
    rollback = true
  }

  depends_on = [aws_lb_listener.http] # the target group must be attached to the ALB first
}

3. What the root reads back — modules/fargate-service/outputs.tf:

# modules/fargate-service/outputs.tf
output "alb_dns_name" {
  value = aws_lb.this.dns_name
}

output "target_group_arn" {
  value = aws_lb_target_group.this.arn
}

output "cluster_name" {
  value = aws_ecs_cluster.this.name
}

output "service_name" {
  value = aws_ecs_service.this.name
}

4. The root module's pins and backend — 06-fargate/versions.tf. This is the same bucket as cloud-03 to cloud-05, with a fourth key. Change only the bucket name:

# 06-fargate/versions.tf
terraform {
  required_version = "~> 1.11"

  required_providers {
    aws = {
      source  = "hashicorp/aws"
      version = "~> 6.0"
    }
  }

  backend "s3" {
    bucket       = "lbv-tfstate-yourname-x7q2"
    key          = "lbv-track/06-fargate/terraform.tfstate"
    region       = "ap-south-1"
    use_lockfile = true
    encrypt      = true
  }
}

provider "aws" {
  region = "ap-south-1"

  default_tags {
    tags = {
      Project = "lbv-track"
      Item    = "cloud-06"
    }
  }
}

5. Wire it — 06-fargate/main.tf. The VPC's IDs come from cloud-04's state, so nothing is copied by hand. The repository is looked up by name, because cloud-05 already made it. The image tag is a variable. v2 is py-10's service, which Day 16's first step pushes. (v1 is cloud-01's app. It has /health but no /readyz, so the health check would never pass on it.)

# 06-fargate/main.tf
variable "image_tag" {
  type    = string
  default = "v2"
}

data "terraform_remote_state" "vpc" {
  backend = "s3"
  config = {
    bucket = "lbv-tfstate-yourname-x7q2"
    key    = "lbv-track/04-vpc/terraform.tfstate"
    region = "ap-south-1"
  }
}

data "aws_ecr_repository" "api" {
  name = "lbv-api"
}

module "api" {
  source = "../modules/fargate-service"

  name               = "lbv-api"
  region             = "ap-south-1"
  vpc_id             = data.terraform_remote_state.vpc.outputs.vpc_id
  subnet_ids         = data.terraform_remote_state.vpc.outputs.public_subnet_ids
  image              = "${data.aws_ecr_repository.api.repository_url}:${var.image_tag}"
  ecr_repository_arn = data.aws_ecr_repository.api.arn
}
# 06-fargate/outputs.tf
output "alb_dns_name" {
  value = module.api.alb_dns_name
}

output "target_group_arn" {
  value = module.api.target_group_arn
}

output "cluster_name" {
  value = module.api.cluster_name
}

output "service_name" {
  value = module.api.service_name
}

6. Init, check, plan to a file. -out tfplan saves exactly this plan, so Day 16's apply does what you read today and nothing else:

cd ~/lbv-cloud/06-fargate
export AWS_PROFILE=lbv-admin
terraform init
terraform fmt ../modules/fargate-service && terraform fmt && terraform validate
# Success! The configuration is valid.
terraform plan -out tfplan
# ...
# Plan: 15 to add, 0 to change, 0 to destroy.
#
# Saved the plan to: tfplan

7. Read the plan's cost, line by line. Take each + create block in the plan output and find it in this table. Twelve of the fifteen resources cost nothing on their own. Three lines carry the whole bill: the ALB, the tasks the service starts, and the public IPv4 addresses both of them hold. That is the answer cloud-06 asks for ("the hourly cost of ALB + task + public IP").

#planned resource (module.api.…)hourly, approx.why it exists
1aws_cloudwatch_log_group.this$0.00 — $0.50/GB ingested, first 5 GB free; today's traffic is kilobyteswhere stdout lands; 7-day retention so it cannot grow forever
2aws_iam_role.execution$0.00the identity ECS assumes to pull the image and write logs
3aws_iam_role_policy.execution$0.00least privilege: pull from lbv-api, write to this log group, nothing else
4aws_security_group.alb$0.00the ALB's firewall
5aws_vpc_security_group_ingress_rule.alb_http$0.00the only door from the internet: :80
6aws_vpc_security_group_egress_rule.alb_to_task$0.00the ALB may only talk to the tasks, on :8000
7aws_security_group.task$0.00the task's firewall
8aws_vpc_security_group_ingress_rule.task_from_alb$0.00:8000 from the ALB's group only; the task's public IP answers nobody else
9aws_vpc_security_group_egress_rule.task_https_out$0.00HTTPS out to ECR, S3 and CloudWatch; without it the pull times out
10aws_lb.this$0.0225/h + $0.008 per LCU-hour, + 2 public IPv4 × $0.005/hone stable DNS name, health-checked routing; bills every hour it exists, traffic or not
11aws_lb_target_group.this$0.00the list of task IPs the ALB may send to, and the /readyz check that admits them
12aws_lb_listener.http$0.00:80 → forward to the target group
13aws_ecs_cluster.this$0.00a namespace; ECS itself carries no charge, Fargate bills the tasks
14aws_ecs_task_definition.this$0.00a document: image, size, port, logs, health check
15aws_ecs_service.thisper task: 0.25 × $0.04048 + 0.5 × $0.004446 = $0.0123/h, + 1 public IPv4 × $0.005/hkeeps desired_count tasks running and registers them with the target group

While applied, all 15 cost ~$0.05/h: $0.0123 for the task, $0.0225 for the ALB and $0.015 for three public IPv4 addresses. Left on for a 720-hour month, that is ~$35.89, which the meter's step 2 shows. These are us-east-1 rates, approx.: verify on the pricing page today.

Verify — Day 15

There is no teardown for cloud-06: a plan creates nothing, and tfplan is a file on your Mac. Press the button once the three Verify lines above are true and you can say, from the table, what the ALB, the task and a public IPv4 address each cost per hour.

This is self-attestation. The site cannot see your terminal, so pressing the button is you telling The Path that the plan reads 15 to add and that the jq check printed 0.

Day 16 · cloud-07 — apply, test it live, record three numbers, destroy

The meter runs from step 2 to step 7. At the rates above, the whole session costs about five cents. The one expensive mistake is to stop before step 7.

1. Push py-10's service as lbv-api:v2. This is the same login and the same cross-build as cloud-05, but run from py-10's folder. Its main.py loads the model about 3 s after start (until then /readyz answers 503) and writes JSON log lines. The tag is immutable, so v2 can never be re-pointed. Like v1, it matches no lifecycle rule:

cd ~/path/to/ai-ml-roadmap/exercises/py-10-fastapi-service   # your finished py-10
export AWS_PROFILE=lbv-admin
ACCT=$(aws sts get-caller-identity --query Account --output text)
REPO=$ACCT.dkr.ecr.ap-south-1.amazonaws.com/lbv-api
aws ecr get-login-password --region ap-south-1 \
  | docker login --username AWS --password-stdin "$ACCT.dkr.ecr.ap-south-1.amazonaws.com"
docker buildx build --platform linux/amd64 --provenance=false --push -t "$REPO:v2" .
aws ecr describe-images --repository-name lbv-api --image-ids imageTag=v2 \
  --query 'imageDetails[0].imageSizeInBytes'
# the compressed size Fargate will pull; the meter guesses ~55 MB

2. Apply the saved plan, and start the clock. If Terraform says the saved plan is stale (the state changed since Day 15), run terraform plan -out tfplan again and re-read the count first:

cd ~/lbv-cloud/06-fargate
START=$(date +%s)
terraform apply tfplan
# ...
# Apply complete! Resources: 15 added, 0 changed, 0 destroyed.
# Outputs:
# alb_dns_name = "lbv-api-<id>.ap-south-1.elb.amazonaws.com"

3. Wait for the target to read healthy — number one, the cold start. This covers everything from apply to the first target that passes two /readyz checks: creating the ALB, placing the task, pulling the image, loading the model. The meter's ~55 s estimate covers only the part after the task is placed. Creating the ALB comes first and usually takes longer, so your number will be bigger:

TG=$(terraform output -raw target_group_arn)
until [ "$(aws elbv2 describe-target-health --target-group-arn "$TG" \
    --query 'length(TargetHealthDescriptions[?TargetHealth.State==`healthy`])' --output text)" -ge 1 ]; do
  sleep 5
done
echo "cold start: $(( $(date +%s) - START )) s from apply to first healthy target"
aws elbv2 describe-target-health --target-group-arn "$TG" \
  --query 'TargetHealthDescriptions[].[Target.Id,TargetHealth.State]' --output text
# 10.0.x.x    healthy

4. Run py-10's live test against the ALB — number two, the p50. The three tests in integration/ check /healthz and /readyz, make 20 /ask calls that must echo the X-Request-Id you send, and time the /stream first chunk. They report their numbers at the end of the run:

ALB=$(terraform output -raw alb_dns_name)
curl -s "http://$ALB/healthz"      # {"status":"ok"}
cd ~/path/to/ai-ml-roadmap/exercises/py-10-fastapi-service
BASE_URL="http://$ALB" uv run pytest -q -m integration
# ...                                                    [100%]
# ============== live numbers (record them for day 20) ==============
# /ask p50 over 20 calls: <n> ms (min <n>, max <n>)
# last request id (grep the logs for it): live-<8 hex>
# /stream first chunk: <n> ms (whole stream <n> ms, <n> reads)
# 3 passed, 9 deselected

5. Watch the logs. Each /ask writes one JSON line with its request id. Find the id the test printed:

aws logs tail /ecs/lbv-api --since 10m --follow
# ... api/api/<task-id> {"ts": "...", "level": "INFO", "logger": "askapi", "msg": "ask served", "request_id": "live-<8 hex>"}
# Ctrl-C to stop following

6. Break one deployment on purpose (cloud-07's "a deliberate bad image tag"). Point the service at a tag that does not exist. That replaces the task definition, and the service tries to start tasks from the new one. The old task keeps serving while the new ones fail, and the circuit breaker stops the deployment after a few failures and rolls back. Diagnose it from ECS, not from the ALB:

cd ~/lbv-cloud/06-fargate
terraform apply -var image_tag=does-not-exist
# Plan: 1 to add, 1 to change, 1 to destroy.   (task definition replaced, service updated)
curl -s "http://$ALB/healthz"      # still {"status":"ok"}: the old task is still serving
aws ecs describe-services --cluster lbv-api --services lbv-api \
  --query 'services[0].events[:5].[createdAt,message]' --output text
# recent events: tasks started, then stopped; after a few, the deployment is marked failed and rolled back
aws ecs describe-tasks --cluster lbv-api --query 'tasks[].stoppedReason' --output text --tasks \
  $(aws ecs list-tasks --cluster lbv-api --desired-status STOPPED --query 'taskArns[]' --output text)
# CannotPullContainerError: ... lbv-api:does-not-exist ... not found

The exact wording of the events changes between ECS releases. The stopped task's stoppedReason is the line that names the cause. No need to re-apply v2: step 7 destroys everything anyway.

7. Destroy — the same session.

cd ~/lbv-cloud/06-fargate
terraform destroy
# Plan: 0 to add, 0 to change, 15 to destroy. ... Destroy complete! Resources: 15 destroyed.

8. Number three, the bill: tomorrow morning. Open Billing → Bills (or Cost Explorer, grouped by service, daily). You should find one day with lines under Elastic Container Service, Elastic Load Balancing and the VPC's public IPv4 addresses, together about five cents. Write the three numbers down for the day-20 write-up: the cold start, the p50 and the bill line.

Verify — Day 16

Teardown — cloud-07

Step 7 is the teardown. Check that it really left nothing (the plan item's teardown: ALB, ENI, log group, EIP all gone; check Cost Explorer next day):

aws ecs list-services --cluster lbv-api --output text
# prints nothing, or ClusterNotFoundException: either way no service is left
aws elbv2 describe-load-balancers --query 'LoadBalancers[].LoadBalancerName'
# []
VPC=$(cd ~/lbv-cloud/04-vpc && terraform output -raw vpc_id)
aws ec2 describe-network-interfaces --filters Name=vpc-id,Values=$VPC \
  --query 'NetworkInterfaces[].[NetworkInterfaceId,Description]' --output text
# nothing (an "ELB app/lbv-api/..." ENI can linger a few minutes after the ALB is gone; re-run)
aws logs describe-log-groups --log-group-name-prefix /ecs/lbv-api --query 'logGroups[].logGroupName'
# []
aws ec2 describe-addresses --query 'Addresses[].PublicIp'
# []: this HCL never made an Elastic IP. The task's and the ALB's addresses were AWS-owned and went with them

What remains and what it costs, approx., verify on the pricing page today: cloud-04's VPC ($0.00: nothing runs in it), cloud-05's repository with v1 and v2 in it (~$0.10 per GB-month on the compressed sizes, a few cents), and this item's state file in the cloud-03 bucket. Nothing in cloud-06/07 bills after the destroy.

This is self-attestation. The site cannot see your AWS account, so checking the box and pressing the button is you telling The Path that the live test passed and that the service is destroyed.

Takeaways: read a plan the way you read a bill. Of fifteen resources, twelve are free: rules, roles, documents, namespaces. Three lines carry the money, and one of them, the load balancer, bills by the hour whether or not a request ever arrives, along with the public addresses it holds. Scaling the service to zero stops the task line but not the ALB line. Only destroy stops that. Health and readiness matter here too: the target group and the container check both probe /readyz, so no request reaches a task before its model is loaded. The cold start you measure is mostly waiting for things that are not your code: the ALB being created, the task being placed, the image being pulled, two health checks ten seconds apart.