ECS Fargate Autoscaling Architecture
This write-up covers how I set up ECS Fargate with target tracking autoscaling to handle dynamic workloads across 100+ microservices while maintaining 99.99% uptime.
Architecture Overview
Task Definition
The key to efficient Fargate usage is right-sizing your task definitions. I use a tiered approach:
resource "aws_ecs_task_definition" "service" {
family = "my-service"
network_mode = "awsvpc"
requires_compatibilities = ["FARGATE"]
cpu = var.task_cpu
memory = var.task_memory
execution_role_arn = aws_iam_role.ecs_execution.arn
task_role_arn = aws_iam_role.ecs_task.arn
container_definitions = jsonencode([
{
name = "app"
image = "${aws_ecr_repository.app.repository_url}:${var.image_tag}"
portMappings = [
{
containerPort = 8080
protocol = "tcp"
}
]
logConfiguration = {
logDriver = "awslogs"
options = {
"awslogs-group" = aws_cloudwatch_log_group.service.name
"awslogs-region" = var.region
"awslogs-stream-prefix" = "ecs"
}
}
healthCheck = {
command = ["CMD-SHELL", "curl -f http://localhost:8080/health || exit 1"]
interval = 30
timeout = 5
retries = 3
startPeriod = 60
}
secrets = [
{
name = "DB_PASSWORD"
valueFrom = aws_secretsmanager_secret.db_password.arn
}
]
}
])
}
Autoscaling Configuration
Target tracking is the cleanest approach — you set a target metric value and AWS handles the rest:
resource "aws_appautoscaling_target" "ecs" {
max_capacity = var.max_tasks
min_capacity = var.min_tasks
resource_id = "service/${aws_ecs_cluster.main.name}/${aws_ecs_service.app.name}"
scalable_dimension = "ecs:service:DesiredCount"
service_namespace = "ecs"
}
# Scale based on CPU utilization
resource "aws_appautoscaling_policy" "cpu" {
name = "${var.service_name}-cpu-scaling"
policy_type = "TargetTrackingScaling"
resource_id = aws_appautoscaling_target.ecs.resource_id
scalable_dimension = aws_appautoscaling_target.ecs.scalable_dimension
service_namespace = aws_appautoscaling_target.ecs.service_namespace
target_tracking_scaling_policy_configuration {
target_value = 60.0
scale_in_cooldown = 300
scale_out_cooldown = 60
predefined_metric_specification {
predefined_metric_type = "ECSServiceAverageCPUUtilization"
}
}
}
# Scale based on Memory utilization
resource "aws_appautoscaling_policy" "memory" {
name = "${var.service_name}-memory-scaling"
policy_type = "TargetTrackingScaling"
resource_id = aws_appautoscaling_target.ecs.resource_id
scalable_dimension = aws_appautoscaling_target.ecs.scalable_dimension
service_namespace = aws_appautoscaling_target.ecs.service_namespace
target_tracking_scaling_policy_configuration {
target_value = 70.0
scale_in_cooldown = 300
scale_out_cooldown = 60
predefined_metric_specification {
predefined_metric_type = "ECSServiceAverageMemoryUtilization"
}
}
}
Service Configuration with Deployment Circuit Breaker
The circuit breaker is critical — it prevents bad deployments from taking down your entire service:
resource "aws_ecs_service" "app" {
name = var.service_name
cluster = aws_ecs_cluster.main.id
task_definition = aws_ecs_task_definition.service.arn
desired_count = var.desired_count
launch_type = "FARGATE"
network_configuration {
subnets = var.private_subnet_ids
security_groups = [aws_security_group.ecs_tasks.id]
}
load_balancer {
target_group_arn = aws_lb_target_group.app.arn
container_name = "app"
container_port = 8080
}
deployment_circuit_breaker {
enable = true
rollback = true
}
deployment_controller {
type = "ECS"
}
lifecycle {
ignore_changes = [desired_count]
}
}
Key Design Decisions
- Scale-out cooldown at 60s — Aggressive scale-out ensures we handle traffic spikes quickly
- Scale-in cooldown at 300s — Conservative scale-in prevents thrashing during fluctuating load
- Circuit breaker with rollback — Bad deploys auto-revert without manual intervention
ignore_changeson desired_count — Terraform won't fight autoscaling decisions- Health check with startPeriod — Gives containers 60s to warm up before health checks kick in
Results
- 99.99% uptime maintained across 100+ microservices
- 60% better resource utilization vs fixed-capacity provisioning
- Zero manual scaling interventions — autoscaling handles everything from 3 AM traffic dips to Black Friday spikes