A few months ago I pasted a Kiro-generated Terraform module into our staging environment. It was clean HCL, well-structured, followed our steering file conventions. Then it tried to create an S3 bucket with public-read ACLs and a Lambda execution role sporting AdministratorAccess.
The SCP caught the S3 violation immediately. The IAM policy slipped through during a window when our RCP propagation was still syncing. Nothing broke — we had canary deployments enabled — but it was a wake-up call.
When AI writes your infrastructure, "shift-left" isn't enough. You need guardrails that understand intent, not just syntax. And you need delivery mechanisms that treat every AI-generated change as potentially hostile until proven otherwise.
This is how we combined four ideas — spec-driven AI development, runtime condition contracts, progressive delivery with policy guardrails, and Lambda MicroVMs — into something that actually works at scale.
The Problem: AI Agents Don't Understand Context
Kiro is impressive. Give it a requirements document in EARS notation, and it'll generate a design doc, task breakdown, and implementation that mostly compiles. For application code, this is transformative. For infrastructure, it's dangerous — because the agent doesn't know your VPC topology, your compliance boundaries, or which data stores are allowed to talk to which compute.
The traditional answer is post-generation linting: run checkov, tfsec, or tflint against the output. But that's like checking a car's brakes after it's already on the motorway. By the time the scan fails, the developer has mentally committed to the solution. The friction is in the wrong place.
What we needed was a way to tell the agent upfront what the runtime environment looks like — not as comments in a steering file, but as a machine-readable contract the agent must satisfy before it generates a single line of HCL.
Runtime Conditions as the Contract Layer A friend of mine Colin Lacy, has been working on something called Runtime Conditions Profiles — a way for workloads to declare their external dependencies abstractly. The workload says "I need a key-value cache, preferably Redis," not "here's the connection string to redis-prod-01.internal." A platform adapter fulfills that contract based on environment, policy, and availability.
We borrowed that idea and applied it to Kiro's spec phase.
Instead of letting Kiro guess what resources a workload needs, we feed it a Runtime Conditions Profile as part of the requirements input. The profile declares:
- What integrations are needed (DynamoDB table, SNS topic, S3 bucket with object lock)
- What interfaces they must satisfy (encryption at rest, VPC-scoped access, specific IAM action boundaries)
- What configuration shape the workload expects (env var names, not values)
Kiro then generates Terraform that references these conditions rather than hardcoding resources. The agent doesn't create a DynamoDB table — it creates a data source lookup against a table that the platform adapter has already validated and provisioned.
Here's what that looks like in practice. Our platform adapter emits a profile snippet like this:
conditions:
- name: audit-log-store
kind: database
interface:
type: document
engine: dynamodb
spec:
encryption: aws_managed
point_in_time_recovery: true
deletion_protection: true
configuration:
env:
- property: table_name
name: AUDIT_LOG_TABLE
iam:
actions:
- dynamodb:PutItem
- dynamodb:Query
resources:
- "arn:aws:dynamodb:*:*:table/audit-log-*"Kiro sees this and generates Terraform that:
- Looks up the existing table via
aws_dynamodb_tabledata source - Creates an IAM policy with only the listed actions
- Injects
AUDIT_LOG_TABLEas an environment variable - Never attempts to create or modify the table itself
The agent can't over-provision because the contract doesn't give it permission to. It can't under-secure because the interface spec mandates encryption and recovery. And it can't hardcode ARNs because the profile only provides env var names, not values.
But Contracts Can Be Broken
Here's the thing about AI agents: they're non-deterministic. Feed the same requirements to Kiro twice, and you might get two different implementations. One follows the contract. One "optimises" by inlining a resource creation it thinks is more efficient. Both pass terraform validate. Only one passes policy.
We learned this the hard way when an agent-generated module created a secondary IAM role with sts:AssumeRole to a principal outside our organisation. The SCP blocked it — eventually. But SCP propagation isn't instant, and during that window, the role existed in a partial state that confused our monitoring.
That incident taught us that Runtime Conditions Profiles handle intent, but we still need enforcement at multiple layers. And enforcement needs to be as progressive as the delivery itself.
Enter Lambda MicroVMs: The Isolated Canary Cage
In June 2026, AWS launched Lambda MicroVMs in preview — a new serverless compute primitive that gives you VM-level isolated sandboxes with no shared kernel or resources between sessions. Rapid launch and resume, full lifecycle control, state preservation up to 8 hours, and zero infrastructure to manage.
At first glance, this looks like a tool for interactive notebooks or sandboxed code execution. For us, it solved a completely different problem: how do you safely run a live canary of AI-generated infrastructure when you don't trust the code enough to put it in your VPC?
Our old live canary deployed to a single AZ in production, scoped to synthetic traffic. That worked for application changes, but infrastructure changes are different. A misconfigured security group or IAM policy doesn't just fail — it opens holes. Running that in a production-adjacent environment still felt too close to the edge.
Lambda MicroVMs changed the calculus. Each canary now runs inside its own MicroVM — kernel-isolated from everything else, with no shared state between sessions. If the agent-generated Terraform creates a rogue IAM policy or exposes an unintended endpoint, the blast radius is contained to that single MicroVM. We can pause it, inspect its state, and decide whether to resume or destroy — all through the lifecycle API.
The 8-hour state preservation is the feature we didn't know we needed. Our old canaries ran for 30 minutes. That's enough for synthetic traffic but misses slower-burn issues — IAM policy eventual consistency drift, CloudWatch metric lag, or cost anomalies that only show up after an hour. With MicroVMs, we run canaries for 4-6 hours, preserving full state throughout. If something looks off at hour three, we snapshot the MicroVM state, attach it to our incident response pipeline, and investigate without losing context.
Progressive Policy: The Four-Stage Pipeline
You can't "canary" a VPC. You can, however, canary the deployment of infrastructure changes — and now you can do it inside an isolated sandbox.
Our updated pipeline looks like this:
Stage 1: Spec Validation (Pre-Generation)
Before Kiro generates anything, we validate the Runtime Conditions Profile against our platform registry. The registry knows:
- Which DynamoDB tables exist in which accounts
- Which IAM action boundaries apply to which workload types
- Which VPC endpoints are available for private access
If the profile requests a condition the registry can't satisfy — say, a Redis cluster in an account that only has ElastiCache Serverless — the pipeline fails before generation starts. No wasted compute, no false hope.
Stage 2: Static Guardrailing (Post-Generation, Pre-Plan)
Kiro outputs Terraform. We run a custom policy engine that does three things:
- Resource boundary check: Does the plan create anything outside the profile's declared conditions? An unexpected
aws_iam_roleis an automatic fail. - Permission scope check: Do IAM policies exceed the actions listed in the profile? We parse the JSON policy and compare against the condition's
iam.actionslist. - Network topology check: Does any resource expose a public endpoint when the profile specifies VPC-scoped access?
These aren't generic rules. They're generated from the profile itself, so they're tailored to each workload. A Lambda that only needs DynamoDB read access gets a different guardrail set than one that needs S3 write + KMS decrypt.
Stage 3: Shadow Apply (Pre-Production)
Terraform plan is executed against a mirror environment that replicates production's SCP and RCP boundaries but uses stub resources. This catches policy violations that static analysis misses — like the SCP propagation delay I mentioned earlier. The shadow environment has artificially delayed SCP sync (we inject a 5-minute lag) to simulate worst-case propagation.
If shadow apply fails, the pipeline stops. The developer gets a Kiro-readable error trace that the agent can use to regenerate the fix.
Stage 4: MicroVM Live Canary (Isolated Production Validation)
Here's where Lambda MicroVMs come in. Instead of deploying to a single AZ in our production VPC, we spin up a dedicated MicroVM with:
- A copy of the agent-generated Terraform
- Read-only access to a sandbox AWS account with production-equivalent SCPs
- Synthetic traffic generators that mimic real workload patterns
- CloudWatch agent, GuardDuty detector, and IAM Access Analyzer enabled
The MicroVM runs for up to 4 hours (we've found that's the sweet spot). During that window:
- CloudWatch Synthetics canaries validate infrastructure health
- GuardDuty watches for anomalous API calls
- IAM Access Analyzer alerts on unexpected cross-account access
- Cost anomaly detection flags oversized resources (AI agents love creating xlarge instances)
- Our custom policy engine continuously evaluates the running state against the Runtime Conditions Profile
If any signal turns red, we pause the MicroVM — preserving its full state — and trigger automatic rollback. The incident data, including the MicroVM snapshot, feeds back into Kiro's steering files as a negative example.
If all signals stay green for the full window, we destroy the MicroVM and promote the Terraform to production deployment via CodeDeploy.
The SCP of AI: When Guardrails Become the Architecture
In my previous articles I wrote about the gap between "attached" and "enforced." With AI-generated infrastructure, that gap is existential. An agent can generate and apply a non-compliant resource faster than SCPs can propagate to deny it.
Lambda MicroVMs don't fix SCP propagation. What they do is give us a contained environment where propagation delay doesn't matter — because the MicroVM is isolated from production resources by design. Even if an SCP hasn't propagated yet, the MicroVM can't touch production. It's a sandbox by architecture, not just by policy.
But — and this is critical — isolation is not governance. A MicroVM isolates the execution environment. It does not govern what the code inside does. An AI agent running inside a MicroVM can still create resources in the sandbox account, exfiltrate data to an external endpoint, or spin up expensive instances that burn budget. The isolation contains the blast radius, but the guardrails still need to enforce the rules.
That's why our five-layer stack now looks like this:
Layer Mechanism What It Blocks
Spec Runtime Conditions Profile Invalid intent before generation
Static Policy engine (profile-derived) Non-compliant HCL before plan
Shadow Mirrored SCP/RCP environment Propagat ion-delayed violations
Live Lambda MicroVM with synthetic traffic Runtime misbehaviour in isolation
Organisational SCPs + RCPs + IAM Access Analyzer Persistent policy violations
The insight is that SCPs are your last line of defence, not your first. If an AI agent is hitting SCP denials regularly, your earlier layers are broken. We track "SCP block rate" as a platform health metric. A spike means our spec validation, static guardrailing, or MicroVM canary has a gap.
We also use Resource Control Policies (RCPs) — the newer sibling to SCPs — to enforce data perimeter rules. Where SCPs control who can do what, RCPs control which resources can be accessed from where. This is critical for AI-generated code because agents don't understand your network topology. An RCP that denies S3 access from outside a specific VPC endpoint doesn't care whether the access came from a human or an agent. It just works.
What We Learned (So Far)
AI-generated infrastructure is not "infrastructure with AI assistance." It's a different paradigm that requires different controls. You wouldn't let a junior engineer deploy to production without code review. Don't let an AI agent deploy without equivalent — or stronger — guardrails.
Runtime Conditions Profiles changed how we think about platform contracts. By making dependencies explicit and machine-readable, we shifted the conversation from "did the agent write good Terraform?" to "did the agent satisfy the platform contract?" The second question is much easier to answer automatically.
Progressive delivery applies to infrastructure, not just apps. The shadow apply + MicroVM canary pattern has caught issues that static analysis never would — particularly around IAM policy eventual consistency, SCP propagation windows, and runtime cost anomalies. If you're running AI agents in production, you need this.
Lambda MicroVMs are the missing piece for untrusted code validation. The VM-level isolation means we can run canaries longer, with more realistic traffic, without risking production. The 8-hour state preservation lets us investigate issues that only surface after hours, not minutes. And the lifecycle API means we can pause, inspect, and resume — treating each canary as a forensic artifact, not just a pass/fail gate.
Kiro's steering files are necessary but not sufficient. They guide the agent's style and conventions, but they don't enforce runtime constraints. Combining steering files with Runtime Conditions Profiles and MicroVM isolation gives you how the code looks, what it can touch, and where it can run.
A Practical Starting Point
If you're looking to implement something similar, here's the minimal viable setup I'd recommend:
- Define one Runtime Conditions Profile for a standard workload type (e.g., "internal API with DynamoDB backend"). Keep it narrow.
- Integrate it into Kiro's requirements phase as a structured input, not a comment.
- Build a static policy engine that parses the generated Terraform and validates against the profile's
conditionslist. Start with resource boundary checks — they're the highest ROI. - Add a shadow apply stage in your CI pipeline using a mirrored account with production-equivalent SCPs. Inject artificial delay if your org's propagation is fast.
- Spin up a Lambda MicroVM for live canary validation. Start with 1-hour runs and scale up as you build confidence. Use the lifecycle API to pause and inspect any failures.
- Track SCP block rate as a metric. If it's above zero, your earlier layers need work.
This isn't theoretical. We've been running this in production for three months across a dozen services. The agent-generated code isn't perfect — it still needs human review for architectural decisions — but it no longer scares us.
And honestly? That's the bar. AI-generated infrastructure shouldn't be exciting. It should be boring, predictable, and heavily guarded. The canary sings, but only after the guards have checked its credentials — and only from inside its cage.
Originally published on AWS Builder Center. Any opinions are those of the individual author and may not reflect the opinions of AWS.