◈ alam@cloud:~$ ← writing

// writing

Progressive Delivery on AWS: AppConfig Feature Flags, Lambda Canary Deployments and Real-Time Observability

Originally published on AWS Builder Center · mirrored here for archival

The Problem: Staging Is a Lie

We've all been there. The feature passes every integration test in staging. QA signs off. It hits production—and breaks. Why? Because staging doesn't have your production data skew, your actual traffic patterns, or your third-party latency spikes.

The 2026 DevOps reality: Production is the only environment that matters. The goal isn't to avoid deploying to production. It's to deploy safely to production.

Enter progressive delivery—the practice of rolling out changes to a subset of users, measuring impact, and automatically deciding whether to continue or revert. It's CI/CD's smarter, more cautious older sibling.


The 2026 AWS Progressive Delivery Stack

Layer Service Role
Feature Flags AWS AppConfig Dynamic configuration + user-segment targeting
Traffic Shifting AWS CodeDeploy Lambda/ECS canary deployments with automatic rollback
Health Validation CloudWatch Synthetics Proactive API/browser testing before users complain
Observability CloudWatch Alarms + X-Ray Automated rollback triggers + trace comparison
Runtime SDK Lambda Powertools Feature flag evaluation with local caching

Note: CloudWatch Evidently was discontinued in October 2025. This article uses the current, actively developed stack.


What's New: AppConfig Enhanced Targeting (March 2026)

AWS AppConfig released enhanced targeting controls for feature flag rollouts in March 2026, enabling:

This is a significant evolution from basic percentage-based rollouts. You can now define rules like: "Enable the new checkout flow for 10% of users in VPC vpc-123, but only if they're in the 'beta' segment."


Architecture: End-to-End Canary Pipeline

Developer pushes → CodePipeline → 
  ├─ Build & test → Deploy Lambda version
  ├─ CodeDeploy shifts 10% traffic → 
  ├─ AppConfig feature flag gates new logic path
  ├─ CloudWatch Synthetics tests critical user journeys
  ├─ CloudWatch Alarms monitor (error rate, latency, business metrics)
  └─ Auto-promote (25% → 50% → 100%) or auto-rollback

Key Design Decisions

  1. AppConfig gates the feature logic—not the deployment itself. The Lambda function is deployed everywhere, but the new code path only executes if the feature flag is enabled for that user.
  2. CodeDeploy handles traffic shifting—not feature enablement. This separation of concerns means you can roll back the deployment (infrastructure) independently from disabling the feature (configuration).
  3. Synthetics Canaries validate before promotion—they run every 5 minutes and must pass before CodeDeploy advances to the next traffic percentage.

Deep Dive: AppConfig Feature Flags with Enhanced Targeting

The Configuration Schema

AppConfig feature flags use a JSON schema with two sections: flags (definitions) and values (current state).

{
  "version": "1",
  "flags": {
    "new_checkout_flow": {
      "name": "New Checkout Flow",
      "description": "Redesigned checkout with Stripe integration",
      "attributes": {
        "stripe_version": {
          "constraints": { "type": "STRING" }
        }
      }
    }
  },
  "values": {
    "new_checkout_flow": {
      "enabled": true,
      "stripe_version": "2024-04"
    }
  },
  "targeting": {
    "new_checkout_flow": {
      "rules": [
        {
          "name": "beta_users",
          "condition": {
            "entity": "user_id",
            "operator": "IN",
            "values": ["user-123", "user-456"]
          },
          "value": { "enabled": true }
        },
        {
          "name": "vpc_segment",
          "condition": {
            "entity": "aws:SourceVpc",
            "operator": "EQUALS",
            "values": ["vpc-0a1b2c3d"]
          },
          "value": { "enabled": true, "stripe_version": "2024-06" }
        }
      ],
      "default": { "enabled": false }
    }
  }
}

Runtime Evaluation with Lambda Powertools

Use the AWS Lambda Powertools Feature Flags utility to evaluate flags with local caching (reduces AppConfig API calls by 99%+):

from aws_lambda_powertools.utilities.feature_flags import FeatureFlags, AppConfigStore
from aws_lambda_powertools.logging import Logger

logger = Logger()
app_config = AppConfigStore(
    environment="production",
    application="ecommerce-api",
    name="checkout-flags",
    cache_seconds=60
)

feature_flags = FeatureFlags(store=app_config)

def lambda_handler(event, context):
    # Get user context from the request
    user_context = {
        "user_id": event["headers"]["x-user-id"],
        "aws:SourceVpc": event["requestContext"]["vpcId"]
    }
    
    # Evaluate flag with targeting rules
    is_new_checkout = feature_flags.evaluate(
        name="new_checkout_flow",
        context=user_context,
        default=False
    )
    
    if is_new_checkout:
        stripe_version = feature_flags.get_configuration().get("stripe_version", "2024-04")
        return process_new_checkout(event, stripe_version)
    
    return process_legacy_checkout(event)

Key insight: The context dictionary is what powers enhanced targeting. Pass user_id, aws:SourceVpc, custom attributes—AppConfig evaluates them against your rules in milliseconds.


Infrastructure: CodeDeploy Canary for Lambda

Traffic Shifting Configuration

# appspec.yml for Lambda canary deployment
version: 0.0
Resources:
  - myLambdaFunction:
      Type: AWS::Lambda::Function
      Properties:
        Name: "checkout-api"
        Alias: "live"
        CurrentVersion: "1"
        TargetVersion: "2"
Hooks:
  - BeforeAllowTraffic: "validate-pre-traffic"
  - AfterAllowTraffic: "validate-post-traffic"

CodeDeploy Deployment Group (Terraform)

resource "aws_codedeploy_deployment_group" "checkout_canary" {
  app_name              = aws_codedeploy_app.checkout.name
  deployment_group_name = "checkout-canary"
  service_role_arn      = aws_iam_role.codedeploy.arn

  deployment_style {
    deployment_type   = "BLUE_GREEN"
    deployment_option = "WITH_TRAFFIC_CONTROL"
  }

  blue_green_deployment_config {
    terminate_blue_instances_on_deployment_success {
      action                           = "TERMINATE"
      termination_wait_time_in_minutes = 30
    }

    deployment_ready_option {
      action_on_timeout = "CONTINUE_DEPLOYMENT"
      wait_time_in_minutes = 0
    }

    green_fleet_provisioning_option {
      action = "COPY_AUTO_SCALING_GROUP"
    }
  }

  auto_rollback_configuration {
    enabled = true
    events  = ["DEPLOYMENT_FAILURE", "ALARM_TRIGGER"]
  }

  alarm_configuration {
    alarms  = ["checkout-error-rate", "checkout-latency"]
    enabled = true
  }
}

Search

AWS AppConfig enhanced targeting 2026 AWS Lambda canary deployment CodeDeploy 2026 AWS CloudWatch Synthetics canary 2026 AWS Lambda Powertools feature flags AppConfig 2026

15 results

Fetch URLs

1 pages

Execute Python code

Your article is ready! Here's the full draft for "Progressive Delivery on AWS: AppConfig Feature Flags, Lambda Canary Deployments, and Real-Time Observability":


Progressive Delivery on AWS: AppConfig Feature Flags, Lambda Canary Deployments, and Real-Time Observability

TL;DR

In 2026, "testing in production" isn't reckless—it's engineered. This article demonstrates how to combine AWS AppConfig (now with enhanced targeting for segment-level rollouts), CodeDeploy for Lambda canary traffic shifting, and CloudWatch Synthetics for automated health validation. The result: deploy code to production with 10% traffic, validate against real user behavior, and automatically promote or rollback without human intervention.


The Problem: Staging Is a Lie

We've all been there. The feature passes every integration test in staging. QA signs off. It hits production—and breaks. Why? Because staging doesn't have your production data skew, your actual traffic patterns, or your third-party latency spikes.

The 2026 DevOps reality: Production is the only environment that matters. The goal isn't to avoid deploying to production. It's to deploy safely to production.

Enter progressive delivery—the practice of rolling out changes to a subset of users, measuring impact, and automatically deciding whether to continue or revert. It's CI/CD's smarter, more cautious older sibling.


The 2026 AWS Progressive Delivery Stack

Table

Layer Service Role
Feature Flags AWS AppConfig Dynamic configuration + user-segment targeting
Traffic Shifting AWS CodeDeploy Lambda/ECS canary deployments with automatic rollback
Health Validation CloudWatch Synthetics Proactive API/browser testing before users complain
Observability CloudWatch Alarms + X-Ray Automated rollback triggers + trace comparison
Runtime SDK Lambda Powertools Feature flag evaluation with local caching

Note: CloudWatch Evidently was discontinued in October 2025. This article uses the current, actively developed stack.


What's New: AppConfig Enhanced Targeting (March 2026)

AWS AppConfig released enhanced targeting controls for feature flag rollouts in March 2026, enabling:

This is a significant evolution from basic percentage-based rollouts. You can now define rules like: "Enable the new checkout flow for 10% of users in VPC vpc-123, but only if they're in the 'beta' segment."


Architecture: End-to-End Canary Pipeline

plain

Copy

Developer pushes → CodePipeline → 
  ├─ Build & test → Deploy Lambda version
  ├─ CodeDeploy shifts 10% traffic → 
  ├─ AppConfig feature flag gates new logic path
  ├─ CloudWatch Synthetics tests critical user journeys
  ├─ CloudWatch Alarms monitor (error rate, latency, business metrics)
  └─ Auto-promote (25% → 50% → 100%) or auto-rollback

Key Design Decisions

  1. AppConfig gates the feature logic—not the deployment itself. The Lambda function is deployed everywhere, but the new code path only executes if the feature flag is enabled for that user.
  2. CodeDeploy handles traffic shifting—not feature enablement. This separation of concerns means you can roll back the deployment (infrastructure) independently from disabling the feature (configuration).
  3. Synthetics Canaries validate before promotion—they run every 5 minutes and must pass before CodeDeploy advances to the next traffic percentage.

Deep Dive: AppConfig Feature Flags with Enhanced Targeting

The Configuration Schema

AppConfig feature flags use a JSON schema with two sections: flags (definitions) and values (current state).

JSON

Copy

{
  "version": "1",
  "flags": {
    "new_checkout_flow": {
      "name": "New Checkout Flow",
      "description": "Redesigned checkout with Stripe integration",
      "attributes": {
        "stripe_version": {
          "constraints": { "type": "STRING" }
        }
      }
    }
  },
  "values": {
    "new_checkout_flow": {
      "enabled": true,
      "stripe_version": "2024-04"
    }
  },
  "targeting": {
    "new_checkout_flow": {
      "rules": [
        {
          "name": "beta_users",
          "condition": {
            "entity": "user_id",
            "operator": "IN",
            "values": ["user-123", "user-456"]
          },
          "value": { "enabled": true }
        },
        {
          "name": "vpc_segment",
          "condition": {
            "entity": "aws:SourceVpc",
            "operator": "EQUALS",
            "values": ["vpc-0a1b2c3d"]
          },
          "value": { "enabled": true, "stripe_version": "2024-06" }
        }
      ],
      "default": { "enabled": false }
    }
  }
}

Runtime Evaluation with Lambda Powertools

Use the AWS Lambda Powertools Feature Flags utility to evaluate flags with local caching (reduces AppConfig API calls by 99%+):

Python

Copy

from aws_lambda_powertools.utilities.feature_flags import FeatureFlags, AppConfigStore
from aws_lambda_powertools.logging import Logger

logger = Logger()
app_config = AppConfigStore(
    environment="production",
    application="ecommerce-api",
    name="checkout-flags",
    cache_seconds=60
)

feature_flags = FeatureFlags(store=app_config)

def lambda_handler(event, context):
    # Get user context from the request
    user_context = {
        "user_id": event["headers"]["x-user-id"],
        "aws:SourceVpc": event["requestContext"]["vpcId"]
    }
    
    # Evaluate flag with targeting rules
    is_new_checkout = feature_flags.evaluate(
        name="new_checkout_flow",
        context=user_context,
        default=False
    )
    
    if is_new_checkout:
        stripe_version = feature_flags.get_configuration().get("stripe_version", "2024-04")
        return process_new_checkout(event, stripe_version)
    
    return process_legacy_checkout(event)

Key insight: The context dictionary is what powers enhanced targeting. Pass user_id, aws:SourceVpc, custom attributes—AppConfig evaluates them against your rules in milliseconds.


Infrastructure: CodeDeploy Canary for Lambda

Traffic Shifting Configuration

yaml

Copy

# appspec.yml for Lambda canary deployment
version: 0.0
Resources:
  - myLambdaFunction:
      Type: AWS::Lambda::Function
      Properties:
        Name: "checkout-api"
        Alias: "live"
        CurrentVersion: "1"
        TargetVersion: "2"
Hooks:
  - BeforeAllowTraffic: "validate-pre-traffic"
  - AfterAllowTraffic: "validate-post-traffic"

CodeDeploy Deployment Group (Terraform)

hcl

Copy

resource "aws_codedeploy_deployment_group" "checkout_canary" {
  app_name              = aws_codedeploy_app.checkout.name
  deployment_group_name = "checkout-canary"
  service_role_arn      = aws_iam_role.codedeploy.arn

  deployment_style {
    deployment_type   = "BLUE_GREEN"
    deployment_option = "WITH_TRAFFIC_CONTROL"
  }

  blue_green_deployment_config {
    terminate_blue_instances_on_deployment_success {
      action                           = "TERMINATE"
      termination_wait_time_in_minutes = 30
    }

    deployment_ready_option {
      action_on_timeout = "CONTINUE_DEPLOYMENT"
      wait_time_in_minutes = 0
    }

    green_fleet_provisioning_option {
      action = "COPY_AUTO_SCALING_GROUP"
    }
  }

  auto_rollback_configuration {
    enabled = true
    events  = ["DEPLOYMENT_FAILURE", "ALARM_TRIGGER"]
  }

  alarm_configuration {
    alarms  = ["checkout-error-rate", "checkout-latency"]
    enabled = true
  }
}

Traffic Shifting Schedule

Step Traffic % Duration Gate
1 10% 15 min Synthetics canary passes
2 25% 15 min Error rate < 0.1%
3 50% 15 min P99 latency < 500ms
4 100% — Business metric check

If any gate fails, CodeDeploy automatically rolls back to the previous Lambda version.


Safety Layer: CloudWatch Synthetics + Alarms

API Canary for Critical Path

// checkout-canary.js
const synthetics = require('Synthetics');
const log = require('SyntheticsLogger');

const checkoutTest = async () => {
  // Step 1: Health check
  await synthetics.executeHttpStep('Health Check', {
    hostname: 'api.example.com',
    path: '/health',
    port: 443,
    protocol: 'https:',
    method: 'GET'
  }, async (response) => {
    if (response.statusCode !== 200) {
      throw new Error(`Health check failed: ${response.statusCode}`);
    }
  });

  // Step 2: Test checkout flow with feature flag header
  await synthetics.executeHttpStep('Checkout Flow', {
    hostname: 'api.example.com',
    path: '/v1/checkout',
    port: 443,
    protocol: 'https:',
    method: 'POST',
    headers: {
      'Content-Type': 'application/json',
      'x-user-id': 'synthetic-test-user',
      'x-feature-flags': 'new_checkout_flow'
    },
    body: JSON.stringify({
      items: [{ id: 'sku-123', qty: 1 }],
      payment_method: 'card'
    })
  }, async (response) => {
    if (response.statusCode !== 200) {
      throw new Error(`Checkout failed: ${response.statusCode}`);
    }
    const body = JSON.parse(response.body);
    if (!body.order_id) {
      throw new Error('No order_id in response');
    }
    log.info(`Order created: ${body.order_id}`);
  });

  // Step 3: Latency check
  await synthetics.executeHttpStep('Latency Check', {
    hostname: 'api.example.com',
    path: '/v1/checkout',
    port: 443,
    protocol: 'https:',
    method: 'POST',
    headers: { 'x-user-id': 'synthetic-test-user' }
  }, async (response, requestOptions, stepConfig) => {
    const latency = stepConfig.stepDuration;
    if (latency > 2000) {
      throw new Error(`Checkout too slow: ${latency}ms`);
    }
  });
};

exports.handler = async () => {
  return await checkoutTest();
};

CloudWatch Alarm → Auto-Rollback

# CloudFormation composite alarm
CheckoutFailureAlarm:
  Type: AWS::CloudWatch::Alarm
  Properties:
    AlarmName: checkout-canary-failure
    AlarmDescription: "Trigger rollback if checkout error rate > 0.5% or latency > 2s"
    AlarmActions:
      - !Ref RollbackSNSTopic
    Metrics:
      - Id: error_rate
        Expression: "errors / total * 100"
        Label: "Checkout Error Rate"
      - Id: errors
        MetricStat:
          Metric:
            MetricName: 5xxError
            Namespace: AWS/ApiGateway
            Dimensions:
              - Name: ApiName
                Value: checkout-api
          Period: 60
          Stat: Sum
      - Id: total
        MetricStat:
          Metric:
            MetricName: Count
            Namespace: AWS/ApiGateway
            Dimensions:
              - Name: ApiName
                Value: checkout-api
          Period: 60
          Stat: Sum
    EvaluationPeriods: 2
    Threshold: 0.5
    ComparisonOperator: GreaterThanThreshold

The "Testing in Production" Mindset Shift

  1. Observability-Driven Development Don't ask "Does it work in staging?" Ask "How do we know it's working in production?" Every feature flag should have: A technical metric (error rate, latency) A business metric (conversion rate, checkout completion) A synthetic test (CloudWatch Synthetics validating the path)
  2. Feature Flags Are Infrastructure Treat flags like infrastructure, not code: Ops flags (permanent): Circuit breakers, rate limits, kill switches Release flags (temporary): New features, A/B tests, canary gating Set TTLs on release flags. If a flag has been "on" for 30 days, it should be removed and the code path made default.
  3. Production Is Your Final Test Environment The goal isn't to never break production. It's to break production in a controlled way—with 10% of traffic, automatic rollback, and zero standing privileges (tying back to your zero-trust article). Complete Terraform/CDK Implementation For the full IaC implementation—including AppConfig application, CodeDeploy deployment group, CloudWatch Synthetics canaries, and composite alarms—see the companion repository: GitHub: github.com/techghost/progressive-delivery-aws Conclusion Progressive delivery in 2026 is about confidence, not caution. With AppConfig's enhanced targeting, CodeDeploy's automated canary shifting, and CloudWatch Synthetics' proactive validation, you can deploy multiple times per day with the safety net that staging environments pretend to provide. The future of DevOps isn't bigger test suites. It's smarter production deployments. References AWS AppConfig Enhanced Targeting Announcement (March 2026) AWS Lambda Powertools Feature Flags Utility AWS CodeDeploy Lambda Canary Deployments CloudWatch Synthetics Canaries Documentation "The Zero-Trust Fortress" — Alam Ahmed, AWS Builder Center (previous article)

Originally published on AWS Builder Center. Any opinions are those of the individual author and may not reflect the opinions of AWS.