The Problem: Staging Is a Lie
We've all been there. The feature passes every integration test in staging. QA signs off. It hits production—and breaks. Why? Because staging doesn't have your production data skew, your actual traffic patterns, or your third-party latency spikes.
The 2026 DevOps reality: Production is the only environment that matters. The goal isn't to avoid deploying to production. It's to deploy safely to production.
Enter progressive delivery—the practice of rolling out changes to a subset of users, measuring impact, and automatically deciding whether to continue or revert. It's CI/CD's smarter, more cautious older sibling.
The 2026 AWS Progressive Delivery Stack
| Layer | Service | Role |
|---|---|---|
| Feature Flags | AWS AppConfig | Dynamic configuration + user-segment targeting |
| Traffic Shifting | AWS CodeDeploy | Lambda/ECS canary deployments with automatic rollback |
| Health Validation | CloudWatch Synthetics | Proactive API/browser testing before users complain |
| Observability | CloudWatch Alarms + X-Ray | Automated rollback triggers + trace comparison |
| Runtime SDK | Lambda Powertools | Feature flag evaluation with local caching |
Note: CloudWatch Evidently was discontinued in October 2025. This article uses the current, actively developed stack.
What's New: AppConfig Enhanced Targeting (March 2026)
AWS AppConfig released enhanced targeting controls for feature flag rollouts in March 2026, enabling:
- Segment-level precision: Target flags to specific user segments (e.g.,
aws:SourceVpc, custom attributes, percentage + rule combinations) - Entity stickiness: Use entity identifiers to ensure users consistently see the same flag variant across sessions
- Individual user targeting: Fine-grained control via AppConfig Agent with individual user IDs
- Gradual rollouts: Deploy changes over minutes or hours, not seconds
This is a significant evolution from basic percentage-based rollouts. You can now define rules like: "Enable the new checkout flow for 10% of users in VPC vpc-123, but only if they're in the 'beta' segment."
Architecture: End-to-End Canary Pipeline
Developer pushes → CodePipeline →
├─ Build & test → Deploy Lambda version
├─ CodeDeploy shifts 10% traffic →
├─ AppConfig feature flag gates new logic path
├─ CloudWatch Synthetics tests critical user journeys
├─ CloudWatch Alarms monitor (error rate, latency, business metrics)
└─ Auto-promote (25% → 50% → 100%) or auto-rollback
Key Design Decisions
- AppConfig gates the feature logic—not the deployment itself. The Lambda function is deployed everywhere, but the new code path only executes if the feature flag is enabled for that user.
- CodeDeploy handles traffic shifting—not feature enablement. This separation of concerns means you can roll back the deployment (infrastructure) independently from disabling the feature (configuration).
- Synthetics Canaries validate before promotion—they run every 5 minutes and must pass before CodeDeploy advances to the next traffic percentage.
Deep Dive: AppConfig Feature Flags with Enhanced Targeting
The Configuration Schema
AppConfig feature flags use a JSON schema with two sections: flags (definitions) and values (current state).
{
"version": "1",
"flags": {
"new_checkout_flow": {
"name": "New Checkout Flow",
"description": "Redesigned checkout with Stripe integration",
"attributes": {
"stripe_version": {
"constraints": { "type": "STRING" }
}
}
}
},
"values": {
"new_checkout_flow": {
"enabled": true,
"stripe_version": "2024-04"
}
},
"targeting": {
"new_checkout_flow": {
"rules": [
{
"name": "beta_users",
"condition": {
"entity": "user_id",
"operator": "IN",
"values": ["user-123", "user-456"]
},
"value": { "enabled": true }
},
{
"name": "vpc_segment",
"condition": {
"entity": "aws:SourceVpc",
"operator": "EQUALS",
"values": ["vpc-0a1b2c3d"]
},
"value": { "enabled": true, "stripe_version": "2024-06" }
}
],
"default": { "enabled": false }
}
}
}Runtime Evaluation with Lambda Powertools
Use the AWS Lambda Powertools Feature Flags utility to evaluate flags with local caching (reduces AppConfig API calls by 99%+):
from aws_lambda_powertools.utilities.feature_flags import FeatureFlags, AppConfigStore
from aws_lambda_powertools.logging import Logger
logger = Logger()
app_config = AppConfigStore(
environment="production",
application="ecommerce-api",
name="checkout-flags",
cache_seconds=60
)
feature_flags = FeatureFlags(store=app_config)
def lambda_handler(event, context):
# Get user context from the request
user_context = {
"user_id": event["headers"]["x-user-id"],
"aws:SourceVpc": event["requestContext"]["vpcId"]
}
# Evaluate flag with targeting rules
is_new_checkout = feature_flags.evaluate(
name="new_checkout_flow",
context=user_context,
default=False
)
if is_new_checkout:
stripe_version = feature_flags.get_configuration().get("stripe_version", "2024-04")
return process_new_checkout(event, stripe_version)
return process_legacy_checkout(event)Key insight: The
contextdictionary is what powers enhanced targeting. Passuser_id,aws:SourceVpc, custom attributes—AppConfig evaluates them against your rules in milliseconds.
Infrastructure: CodeDeploy Canary for Lambda
Traffic Shifting Configuration
# appspec.yml for Lambda canary deployment
version: 0.0
Resources:
- myLambdaFunction:
Type: AWS::Lambda::Function
Properties:
Name: "checkout-api"
Alias: "live"
CurrentVersion: "1"
TargetVersion: "2"
Hooks:
- BeforeAllowTraffic: "validate-pre-traffic"
- AfterAllowTraffic: "validate-post-traffic"CodeDeploy Deployment Group (Terraform)
resource "aws_codedeploy_deployment_group" "checkout_canary" {
app_name = aws_codedeploy_app.checkout.name
deployment_group_name = "checkout-canary"
service_role_arn = aws_iam_role.codedeploy.arn
deployment_style {
deployment_type = "BLUE_GREEN"
deployment_option = "WITH_TRAFFIC_CONTROL"
}
blue_green_deployment_config {
terminate_blue_instances_on_deployment_success {
action = "TERMINATE"
termination_wait_time_in_minutes = 30
}
deployment_ready_option {
action_on_timeout = "CONTINUE_DEPLOYMENT"
wait_time_in_minutes = 0
}
green_fleet_provisioning_option {
action = "COPY_AUTO_SCALING_GROUP"
}
}
auto_rollback_configuration {
enabled = true
events = ["DEPLOYMENT_FAILURE", "ALARM_TRIGGER"]
}
alarm_configuration {
alarms = ["checkout-error-rate", "checkout-latency"]
enabled = true
}
}
Search
AWS AppConfig enhanced targeting 2026 AWS Lambda canary deployment CodeDeploy 2026 AWS CloudWatch Synthetics canary 2026 AWS Lambda Powertools feature flags AppConfig 2026
15 results
Fetch URLs
1 pages
Execute Python code
Your article is ready! Here's the full draft for "Progressive Delivery on AWS: AppConfig Feature Flags, Lambda Canary Deployments, and Real-Time Observability":
Progressive Delivery on AWS: AppConfig Feature Flags, Lambda Canary Deployments, and Real-Time Observability
TL;DR
In 2026, "testing in production" isn't reckless—it's engineered. This article demonstrates how to combine AWS AppConfig (now with enhanced targeting for segment-level rollouts), CodeDeploy for Lambda canary traffic shifting, and CloudWatch Synthetics for automated health validation. The result: deploy code to production with 10% traffic, validate against real user behavior, and automatically promote or rollback without human intervention.
The Problem: Staging Is a Lie
We've all been there. The feature passes every integration test in staging. QA signs off. It hits production—and breaks. Why? Because staging doesn't have your production data skew, your actual traffic patterns, or your third-party latency spikes.
The 2026 DevOps reality: Production is the only environment that matters. The goal isn't to avoid deploying to production. It's to deploy safely to production.
Enter progressive delivery—the practice of rolling out changes to a subset of users, measuring impact, and automatically deciding whether to continue or revert. It's CI/CD's smarter, more cautious older sibling.
The 2026 AWS Progressive Delivery Stack
Table
| Layer | Service | Role |
|---|---|---|
| Feature Flags | AWS AppConfig | Dynamic configuration + user-segment targeting |
| Traffic Shifting | AWS CodeDeploy | Lambda/ECS canary deployments with automatic rollback |
| Health Validation | CloudWatch Synthetics | Proactive API/browser testing before users complain |
| Observability | CloudWatch Alarms + X-Ray | Automated rollback triggers + trace comparison |
| Runtime SDK | Lambda Powertools | Feature flag evaluation with local caching |
Note: CloudWatch Evidently was discontinued in October 2025. This article uses the current, actively developed stack.
What's New: AppConfig Enhanced Targeting (March 2026)
AWS AppConfig released enhanced targeting controls for feature flag rollouts in March 2026, enabling:
- Segment-level precision: Target flags to specific user segments (e.g.,
aws:SourceVpc, custom attributes, percentage + rule combinations) - Entity stickiness: Use entity identifiers to ensure users consistently see the same flag variant across sessions
- Individual user targeting: Fine-grained control via AppConfig Agent with individual user IDs
- Gradual rollouts: Deploy changes over minutes or hours, not seconds
This is a significant evolution from basic percentage-based rollouts. You can now define rules like: "Enable the new checkout flow for 10% of users in VPC vpc-123, but only if they're in the 'beta' segment."
Architecture: End-to-End Canary Pipeline
plain
Copy
Developer pushes → CodePipeline →
├─ Build & test → Deploy Lambda version
├─ CodeDeploy shifts 10% traffic →
├─ AppConfig feature flag gates new logic path
├─ CloudWatch Synthetics tests critical user journeys
├─ CloudWatch Alarms monitor (error rate, latency, business metrics)
└─ Auto-promote (25% → 50% → 100%) or auto-rollback
Key Design Decisions
- AppConfig gates the feature logic—not the deployment itself. The Lambda function is deployed everywhere, but the new code path only executes if the feature flag is enabled for that user.
- CodeDeploy handles traffic shifting—not feature enablement. This separation of concerns means you can roll back the deployment (infrastructure) independently from disabling the feature (configuration).
- Synthetics Canaries validate before promotion—they run every 5 minutes and must pass before CodeDeploy advances to the next traffic percentage.
Deep Dive: AppConfig Feature Flags with Enhanced Targeting
The Configuration Schema
AppConfig feature flags use a JSON schema with two sections: flags (definitions) and values (current state).
JSON
Copy
{
"version": "1",
"flags": {
"new_checkout_flow": {
"name": "New Checkout Flow",
"description": "Redesigned checkout with Stripe integration",
"attributes": {
"stripe_version": {
"constraints": { "type": "STRING" }
}
}
}
},
"values": {
"new_checkout_flow": {
"enabled": true,
"stripe_version": "2024-04"
}
},
"targeting": {
"new_checkout_flow": {
"rules": [
{
"name": "beta_users",
"condition": {
"entity": "user_id",
"operator": "IN",
"values": ["user-123", "user-456"]
},
"value": { "enabled": true }
},
{
"name": "vpc_segment",
"condition": {
"entity": "aws:SourceVpc",
"operator": "EQUALS",
"values": ["vpc-0a1b2c3d"]
},
"value": { "enabled": true, "stripe_version": "2024-06" }
}
],
"default": { "enabled": false }
}
}
}Runtime Evaluation with Lambda Powertools
Use the AWS Lambda Powertools Feature Flags utility to evaluate flags with local caching (reduces AppConfig API calls by 99%+):
Python
Copy
from aws_lambda_powertools.utilities.feature_flags import FeatureFlags, AppConfigStore
from aws_lambda_powertools.logging import Logger
logger = Logger()
app_config = AppConfigStore(
environment="production",
application="ecommerce-api",
name="checkout-flags",
cache_seconds=60
)
feature_flags = FeatureFlags(store=app_config)
def lambda_handler(event, context):
# Get user context from the request
user_context = {
"user_id": event["headers"]["x-user-id"],
"aws:SourceVpc": event["requestContext"]["vpcId"]
}
# Evaluate flag with targeting rules
is_new_checkout = feature_flags.evaluate(
name="new_checkout_flow",
context=user_context,
default=False
)
if is_new_checkout:
stripe_version = feature_flags.get_configuration().get("stripe_version", "2024-04")
return process_new_checkout(event, stripe_version)
return process_legacy_checkout(event)Key insight: The
contextdictionary is what powers enhanced targeting. Passuser_id,aws:SourceVpc, custom attributes—AppConfig evaluates them against your rules in milliseconds.
Infrastructure: CodeDeploy Canary for Lambda
Traffic Shifting Configuration
yaml
Copy
# appspec.yml for Lambda canary deployment
version: 0.0
Resources:
- myLambdaFunction:
Type: AWS::Lambda::Function
Properties:
Name: "checkout-api"
Alias: "live"
CurrentVersion: "1"
TargetVersion: "2"
Hooks:
- BeforeAllowTraffic: "validate-pre-traffic"
- AfterAllowTraffic: "validate-post-traffic"CodeDeploy Deployment Group (Terraform)
hcl
Copy
resource "aws_codedeploy_deployment_group" "checkout_canary" {
app_name = aws_codedeploy_app.checkout.name
deployment_group_name = "checkout-canary"
service_role_arn = aws_iam_role.codedeploy.arn
deployment_style {
deployment_type = "BLUE_GREEN"
deployment_option = "WITH_TRAFFIC_CONTROL"
}
blue_green_deployment_config {
terminate_blue_instances_on_deployment_success {
action = "TERMINATE"
termination_wait_time_in_minutes = 30
}
deployment_ready_option {
action_on_timeout = "CONTINUE_DEPLOYMENT"
wait_time_in_minutes = 0
}
green_fleet_provisioning_option {
action = "COPY_AUTO_SCALING_GROUP"
}
}
auto_rollback_configuration {
enabled = true
events = ["DEPLOYMENT_FAILURE", "ALARM_TRIGGER"]
}
alarm_configuration {
alarms = ["checkout-error-rate", "checkout-latency"]
enabled = true
}
}
Traffic Shifting Schedule
| Step | Traffic % | Duration | Gate |
|---|---|---|---|
| 1 | 10% | 15 min | Synthetics canary passes |
| 2 | 25% | 15 min | Error rate < 0.1% |
| 3 | 50% | 15 min | P99 latency < 500ms |
| 4 | 100% | — | Business metric check |
If any gate fails, CodeDeploy automatically rolls back to the previous Lambda version.
Safety Layer: CloudWatch Synthetics + Alarms
API Canary for Critical Path
// checkout-canary.js
const synthetics = require('Synthetics');
const log = require('SyntheticsLogger');
const checkoutTest = async () => {
// Step 1: Health check
await synthetics.executeHttpStep('Health Check', {
hostname: 'api.example.com',
path: '/health',
port: 443,
protocol: 'https:',
method: 'GET'
}, async (response) => {
if (response.statusCode !== 200) {
throw new Error(`Health check failed: ${response.statusCode}`);
}
});
// Step 2: Test checkout flow with feature flag header
await synthetics.executeHttpStep('Checkout Flow', {
hostname: 'api.example.com',
path: '/v1/checkout',
port: 443,
protocol: 'https:',
method: 'POST',
headers: {
'Content-Type': 'application/json',
'x-user-id': 'synthetic-test-user',
'x-feature-flags': 'new_checkout_flow'
},
body: JSON.stringify({
items: [{ id: 'sku-123', qty: 1 }],
payment_method: 'card'
})
}, async (response) => {
if (response.statusCode !== 200) {
throw new Error(`Checkout failed: ${response.statusCode}`);
}
const body = JSON.parse(response.body);
if (!body.order_id) {
throw new Error('No order_id in response');
}
log.info(`Order created: ${body.order_id}`);
});
// Step 3: Latency check
await synthetics.executeHttpStep('Latency Check', {
hostname: 'api.example.com',
path: '/v1/checkout',
port: 443,
protocol: 'https:',
method: 'POST',
headers: { 'x-user-id': 'synthetic-test-user' }
}, async (response, requestOptions, stepConfig) => {
const latency = stepConfig.stepDuration;
if (latency > 2000) {
throw new Error(`Checkout too slow: ${latency}ms`);
}
});
};
exports.handler = async () => {
return await checkoutTest();
};CloudWatch Alarm → Auto-Rollback
# CloudFormation composite alarm
CheckoutFailureAlarm:
Type: AWS::CloudWatch::Alarm
Properties:
AlarmName: checkout-canary-failure
AlarmDescription: "Trigger rollback if checkout error rate > 0.5% or latency > 2s"
AlarmActions:
- !Ref RollbackSNSTopic
Metrics:
- Id: error_rate
Expression: "errors / total * 100"
Label: "Checkout Error Rate"
- Id: errors
MetricStat:
Metric:
MetricName: 5xxError
Namespace: AWS/ApiGateway
Dimensions:
- Name: ApiName
Value: checkout-api
Period: 60
Stat: Sum
- Id: total
MetricStat:
Metric:
MetricName: Count
Namespace: AWS/ApiGateway
Dimensions:
- Name: ApiName
Value: checkout-api
Period: 60
Stat: Sum
EvaluationPeriods: 2
Threshold: 0.5
ComparisonOperator: GreaterThanThresholdThe "Testing in Production" Mindset Shift
- Observability-Driven Development Don't ask "Does it work in staging?" Ask "How do we know it's working in production?" Every feature flag should have: A technical metric (error rate, latency) A business metric (conversion rate, checkout completion) A synthetic test (CloudWatch Synthetics validating the path)
- Feature Flags Are Infrastructure Treat flags like infrastructure, not code: Ops flags (permanent): Circuit breakers, rate limits, kill switches Release flags (temporary): New features, A/B tests, canary gating Set TTLs on release flags. If a flag has been "on" for 30 days, it should be removed and the code path made default.
- Production Is Your Final Test Environment The goal isn't to never break production. It's to break production in a controlled way—with 10% of traffic, automatic rollback, and zero standing privileges (tying back to your zero-trust article). Complete Terraform/CDK Implementation For the full IaC implementation—including AppConfig application, CodeDeploy deployment group, CloudWatch Synthetics canaries, and composite alarms—see the companion repository: GitHub: github.com/techghost/progressive-delivery-aws Conclusion Progressive delivery in 2026 is about confidence, not caution. With AppConfig's enhanced targeting, CodeDeploy's automated canary shifting, and CloudWatch Synthetics' proactive validation, you can deploy multiple times per day with the safety net that staging environments pretend to provide. The future of DevOps isn't bigger test suites. It's smarter production deployments. References AWS AppConfig Enhanced Targeting Announcement (March 2026) AWS Lambda Powertools Feature Flags Utility AWS CodeDeploy Lambda Canary Deployments CloudWatch Synthetics Canaries Documentation "The Zero-Trust Fortress" — Alam Ahmed, AWS Builder Center (previous article)
Originally published on AWS Builder Center. Any opinions are those of the individual author and may not reflect the opinions of AWS.