What Is Amazon Bedrock AgentCore? Runtime, Pricing, V2
Amazon Bedrock AgentCore is AWS's platform for running AI agents in production: a serverless runtime where every user session gets its own isolated microVM, plus a set of building blocks the agent can use, including memory, a tool gateway that turns your APIs into MCP tools, identity, a sandboxed code interpreter, a cloud browser, observability, evaluations and a policy engine. You can use any framework (LangGraph, CrewAI, Strands, LlamaIndex, the OpenAI Agents SDK, Google ADK) and any model, inside Bedrock or outside it, and you pay per second for what runs. On September 18, 2026 AWS shipped a second-generation runtime, V2, that starts agents from a snapshot in about two seconds whatever the image size. This guide explains what each AgentCore piece does in plain English, how the runtime and its sessions work, what it all costs, and the pricing twist in V2 that nobody puts in the headline: the new runtime charges 43% more per vCPU-hour, and it can still be the cheaper one.
Jake had built a small agent for Ethan's shop that looked up a customer's repair ticket, drafted a reply and, if asked, checked a supplier's website for a part's price. On his laptop it was twenty lines of Python and a model key. Then a client asked to run it for their own help desk, with their own customers, and the questions changed overnight. Where does it run when Jake's laptop is closed? What stops customer A's conversation leaking into customer B's? Where does the supplier password live? Who sees what it did when something goes wrong? "That's four different AWS services," Ethan said, reading over his shoulder. "Or it's one," Jake said, "and it's called AgentCore." They were both half right, and working out which half is what this page is for.
If Bedrock itself is new to you, start with our plain-English guide to Amazon Bedrock, the model marketplace underneath all of this. And if the word "agent" still sounds like marketing, our honest answer to what an AI agent is takes five minutes: a model that does not just answer, but takes actions with tools, in a loop, until a job is done. AgentCore is the place those loops run when real users depend on them.
What Bedrock AgentCore actually is: one platform, twelve pieces
The cleanest way to understand AgentCore is to stop thinking of it as one product. It is a set of modular services that share a name, a console and a billing page, and that work together or entirely on their own. You can host an agent on AgentCore Runtime and never touch Memory. You can put your company's APIs behind AgentCore Gateway and run the agent somewhere else entirely. That flexibility is the point, and it is also why the product is confusing on first contact. Here is the whole map in one table.
| Piece | What it does, plainly | Jake's shop would use it to |
|---|---|---|
| Runtime | Hosts your agent's container; one isolated microVM per user session | Run the ticket agent for the client's customers |
| Harness | A managed agent loop: give it a model, a system prompt and tools in one API call, and it runs the loop for you | Skip writing the loop at all for a simple agent |
| Memory | Short-term memory for the conversation, long-term memory that survives across sessions | Remember that this customer prefers text over calls |
| Gateway | Turns your APIs, Lambda functions and OpenAPI specs into MCP tools; connects existing MCP servers | Expose the ticket system as a "look up ticket" tool |
| Identity | Who may call the agent, and the credentials the agent uses to call other services | Hold the supplier-site password so the code never sees it |
| Code Interpreter | A sandbox where the agent can run Python, JavaScript or TypeScript | Total a month of repair costs from a spreadsheet |
| Browser | A managed cloud browser the agent can drive, compatible with Playwright and BrowserUse | Check the supplier's price page |
| Observability | Traces of every step the agent took, in OpenTelemetry format | See why the agent quoted the wrong part |
| Evaluations | Scores the quality of the agent's answers, built-in or custom | Catch the week its replies got worse after a prompt change |
| Policy | Deterministic rules on every tool call through Gateway, written in plain language or a Cedar-compatible policy language | "Never issue a refund over $50 without a human" |
| Registry | A catalog of approved agents, MCP servers, tools and skills | Not yet; this is for companies with many teams |
| Payments, Optimization | Micro-payments for paid APIs; A/B-tested prompt and tool-description improvements | Later, if ever |
Two naming traps before we go deeper. AgentCore is not Bedrock Agents. Bedrock Agents is the older feature where you build an agent by filling in forms in the console; AgentCore is for teams who write their own agent code in a framework and need somewhere serious to run it. And AgentCore is not tied to Bedrock's models. The runtime works with models inside Bedrock and outside it, including OpenAI, Google Gemini, Mistral and Meta Llama models. The Bedrock in the name tells you which part of AWS sells it, not which model you must use.
What stitches the pieces together is open protocols. Tools speak MCP, the Model Context Protocol. Agents can talk to other agents over A2A, the agent-to-agent protocol. Traces are OpenTelemetry. That is why the framework list is long: AgentCore does not care how your agent thinks, only how it connects.
How AgentCore Runtime works: one microVM per session
Runtime is the piece most people come for, so it gets the most space. You package your agent as a container image, push it to Amazon ECR, and create an agent runtime that points to it. AgentCore gives that runtime a version (version 1), a DEFAULT endpoint that points to the latest version, and an ARN you invoke. Every time you update the image or the settings, a new immutable version appears and the default endpoint moves to it; you can create named endpoints such as dev, test and prod that stay pinned to older versions, which is how you roll out and roll back without downtime.
The part that makes AgentCore different from running your agent on Lambda or Fargate is the session. Each conversation carries a runtimeSessionId, and each session runs in its own microVM: a tiny virtual machine with its own CPU, memory and filesystem, not shared with any other session. If you have read our piece on containers versus virtual machines, this is the best of both: you ship a container, but each user gets the hardware-level isolation of a VM.
- A session lives up to 8 hours of total runtime, which is what makes long, asynchronous agent jobs possible.
- A session ends after 15 minutes of inactivity, or at the 8-hour limit, or if it becomes unhealthy.
- When it ends, the whole microVM is terminated and its memory is sanitized. A later request with the same session ID gets a brand-new environment.
- Session state is not storage. Anything that must survive belongs in AgentCore Memory or your own database, never on the session's disk.
That last point is the one Jake's customer-isolation worry comes down to. Customer A's session and customer B's session never share a process, a filesystem or a slice of memory, because they never share a machine. When A's conversation goes quiet for fifteen minutes, A's machine is destroyed and scrubbed. There is no cache to leak, because there is no shared cache. For a regulated client, that single sentence is often worth the whole price of the platform.
Runtime speaks three protocols. HTTP for ordinary request and response, with streaming for partial answers and a WebSocket API for two-way, real-time conversations. MCP, so you can host an MCP server on Runtime and let other agents use it as a tool. And A2A, so agents can find and call each other. Your container must answer a health check at /ping; the runtime uses it to know when the agent is ready and, for long jobs, whether it is still busy.
Who can call it, and what it can call
Runtime has two directions of security, and both come from AgentCore Identity. Inbound decides who may invoke the agent: either AWS IAM with SigV4 signatures, or OAuth 2.0 tokens from an identity provider such as Amazon Cognito, Okta or Microsoft Entra ID, validated against a discovery URL, allowed audiences and allowed clients before your code runs. Outbound decides how the agent reaches other services: OAuth or API keys held by AgentCore Identity, used either on behalf of the signed-in user (user-delegated) or with the agent's own service credentials (autonomous). The credentials never appear in your agent's code or logs. Jake's supplier password goes into Identity once, and the agent asks for it by name at the moment it needs it.
Runtime V2: the snapshot runtime from September 18, 2026
On September 18, 2026, AWS released the next generation of the runtime, called platform version V2. You choose it per agent runtime with a single field, platformVersion, set to V1 or V2. V1 is still the default. Two things changed, and both are about what happens underneath your code.
Cold starts. V1 boots and initializes your environment every time a new session needs a microVM, so the bigger your container image, the longer the first reply takes. V2 prepares your environment once, takes a snapshot of it, and restores that snapshot for every new session. AWS's own test with an empty echo agent measured the 75th-percentile cold start like this:
| Container image size | V1 cold start (P75) | V2 cold start (P75) |
|---|---|---|
| 200 MB | about 5.4 seconds | about 2 seconds |
| 2 GB | about 30 seconds | about 2 seconds |
Thirty seconds is the difference between an agent that feels broken and one that feels instant, and it is the number that matters if your agent carries a heavy image full of libraries, or a local model. If you have fought Lambda cold starts, the snapshot idea will be familiar: pay the startup cost once, then restore.
Billing. V1 holds your session's peak memory allocation for the life of the session. V2 pages memory in on demand, reclaims memory that your agent releases or stops touching, and bills the actual amount over time; AWS also says you are not charged for idle CPU waiting on I/O, which for an agent that spends most of its life waiting on a model's reply is most of its life. That is the mechanism behind the pricing twist in the next section.
What V2 changes about your code and your deployments
V2 is not a free switch, and the documentation is honest about why. Read these before you flip it.
- Create and update take minutes, not seconds. Preparing the snapshot is a one-time cost per version, so a V2 runtime sits in
CREATINGorUPDATINGfor several minutes. Callupdateordeletebefore it reaches a terminal state and you get aConflictException; pollget-agent-runtimeuntil it isREADYor ends inFAILED. - Your container must report healthy within 120 seconds. The snapshot is taken at the first healthy
/ping. Report healthy only after initialization is complete, so the snapshot captures a ready agent; miss the 120-second window and creation fails with a health-check error. - Code that runs "once at startup" now runs once, ever. Anything you compute at startup, a random seed, a timestamp, a token fetched on boot, is frozen into the snapshot and restored into every session. AWS publishes a guide to optimizing agents for V2 for exactly this reason; read it before migrating.
- Smaller environment-variable budget. V2 currently caps total environment variables at 1.5 KB for direct code deployments and 2.5 KB for container agents, against 4 KB on V1, and fails with
ValidationExceptionif you exceed it. AWS says it will raise the limit to match V1. - No CloudFormation or CDK support yet. Neither can set
platformVersiontoday. If your deployments are infrastructure as code, V2 means a console or CLI step until that lands. - Five Regions. V2 is available in US East (N. Virginia), US East (Ohio), US West (Oregon), Europe (Ireland) and Asia Pacific (Tokyo).
Moving an existing runtime is one update call with --platform-version V2; leaving the field out on an update keeps whatever version the runtime already had. Old snapshots are cleaned up automatically once no endpoint points to them, which can take up to eight hours because sessions already running on a snapshot are allowed to finish.
aws bedrock-agentcore-control create-agent-runtime \
--agent-runtime-name "ticket-agent" \
--role-arn "arn:aws:iam::111122223333:role/AgentExecutionRole" \
--agent-runtime-artifact '{"containerConfiguration":{"containerUri":"111122223333.dkr.ecr.us-west-2.amazonaws.com/ticket-agent:latest"}}' \
--network-configuration '{"networkMode":"PUBLIC"}' \
--platform-version V2
aws bedrock-agentcore-control get-agent-runtime \
--agent-runtime-id ticket-agent-ABCDE12345 --query platformVersion
One small oddity worth knowing so it does not waste your afternoon: the create call's response does not include platformVersion. Confirm it with get-agent-runtime, as above.
AgentCore pricing, and why the pricier runtime can cost less
AgentCore has no upfront fee and no minimum. Each piece bills on its own meter, per second where time is involved. These are the published rates on the pricing page as of September 29, 2026.
| Piece | Price |
|---|---|
| Runtime V1 | $0.0895 per vCPU-hour, $0.00945 per GB-hour |
| Runtime V2, consumption | $0.1276 per vCPU-hour, $0.0169 per GB-hour |
| Runtime V2, committed baseline (launching October 2026) | $0.0997 per vCPU-hour, $0.0132 per GB-hour |
| Browser, Code Interpreter | $0.0895 per vCPU-hour, $0.00945 per GB-hour |
| Web search | $7.00 per 1,000 queries |
| Gateway | $0.005 per 1,000 invocations; $0.025 per 1,000 searches; $0.02 per 100 tools indexed per month |
| Identity | $0.010 per 1,000 token or API-key requests; no charge when used through Runtime or Gateway |
| Memory | $0.25 per 1,000 short-term events; $0.75 per 1,000 long-term records per month (built-in); $0.50 per 1,000 retrievals |
| Evaluations | Built-in: $0.0024 per 1,000 input and $0.012 per 1,000 output tokens; custom: $1.50 per 1,000 evaluations; batch 25% cheaper |
| Policy | $0.000025 per authorization request; $0.13 per 1,000 tokens of natural-language authoring |
| Registry | First 5,000 records, 1 million searches and 2 million list or get calls free each month |
The model itself is not on this bill. AgentCore charges for hosting and plumbing; the tokens your agent sends to a model are billed by Bedrock or by whichever provider you call. Our Bedrock pricing guide covers that side.
The V2 pricing twist, worked out
Look at the two runtime rows again. V2 costs 43% more per vCPU-hour than V1 ($0.1276 against $0.0895) and 79% more per GB-hour ($0.0169 against $0.00945). Read that way, V2 is simply the expensive one. AWS's own description of V2 is "a higher rate but on far fewer GB-hours," and the second half of that sentence is the whole story.
Here is why, with one ordinary agent session. Say Jake's agent has 1 vCPU and 2 GB of memory, a customer asks one question, the agent spends 30 seconds actually computing and waiting on the model, and then the customer goes quiet. The session stays alive until the 15-minute idle timeout ends it, so it lives about 15.5 minutes.
- On V1, the session holds its peak 2 GB for all 15.5 minutes. Memory: 2 GB × 0.258 hours × $0.00945 = about $0.0049. CPU, assuming only the 30 busy seconds are billed: about $0.0007. Memory is roughly 87% of the session's cost, and most of that memory was paid for while nobody was talking.
- On V2, the CPU line is at most about $0.0011 for the same 30 seconds, and less if part of that time was spent waiting on the model, which V2 does not bill as CPU. The memory line depends on how much memory V2 reclaims after the answer. Because V2's memory rate is 1.79 times V1's, V2 comes out cheaper on memory whenever your agent's average memory over the session is below about 56% of its peak.
For a chat agent that answers and then waits, the average is usually far below that line, because the peak lasts seconds and the waiting lasts minutes. For an agent that loads a big model into memory and keeps it hot the whole time, the average stays near the peak, and V1 is cheaper. That is the honest rule: V2 saves money on bursty, mostly-idle sessions and costs more on busy, memory-heavy ones. Run a week of real traffic on each and compare in Cost Explorer before you commit a fleet.
For scale, AWS's own published example of a customer-support agent on V2 handling a million sessions a month comes to about $6,703. A browser-driving travel-booking agent at 100,000 sessions is about $1,227; a data-analysis agent making 30,000 Code Interpreter runs is about $109; and an HR assistant behind Gateway with 200 tools and 50 million interactions is about $2,250. None of those include model tokens.
Gateway, Identity and Policy: how the agent reaches your systems safely
If Runtime is where the agent lives, Gateway is the only door it should use to reach your systems. Its targets come in three kinds: MCP targets, where the gateway aggregates OpenAPI specs, Smithy models, Lambda functions and existing MCP servers into one virtual MCP server; HTTP targets, passed straight through without translation, which is how it fronts other agents and A2A traffic; and inference targets, which route model calls across providers behind one endpoint. Each target can carry its own credential provider: no auth, the gateway's IAM service role, OAuth 2.0 client credentials, or an API key. One-click integrations exist for Salesforce, Slack, Jira, Asana and Zendesk. You point Gateway at what you already have, a Lambda function, an OpenAPI spec, a SaaS integration such as Salesforce, Zoom, Jira or Slack, or an existing MCP server, and Gateway presents it to the agent as a set of MCP tools behind one endpoint. The agent sees "look up ticket" and "check part price"; it never sees your internal URLs or keys. With hundreds of tools, the Gateway search API lets the agent find the right one by meaning rather than loading every tool description into its prompt, which saves tokens as well as confusion. And since September 28, 2026, Gateway can return the complete MCP tools list in a single response with pagination turned off, for clients that expect one list.
Identity sits in two places, as described above: checking who calls the agent, and holding the credentials the agent uses to call out. Its practical gift is that secrets stop living in environment variables. Policy is the newest guardrail and the most interesting one. It intercepts every tool call that goes through Gateway, before it executes, and checks it against rules you write either in plain language or in Dogwood, a policy language compatible with AWS's open-source Cedar. A rule such as "refunds over $50 require a human" is enforced by the gateway, not by asking the model nicely in a system prompt. That difference, deterministic rules instead of instructions, is the same idea behind the sandbox approach in our OpenShell guide, applied to tool calls in the cloud rather than processes on a laptop.
Memory, Code Interpreter and Browser: what each can hold and how long
Memory has two halves. Short-term memory stores the raw turns of a session as events, each tied to a session ID, so the agent can reload a whole conversation even after a restart or when the customer comes back to continue it; events can carry metadata such as a product category or case type so the agent can find the right old conversation without scanning everything. Long-term memory is generated asynchronously in the background: after events land, an extraction pass pulls out durable insights (facts, preferences, session summaries) and a consolidation pass merges them with what is already stored, and the agent retrieves them later by semantic search. Remember why this exists: the session's microVM is destroyed after fifteen idle minutes, so anything the agent should know tomorrow must live in Memory, not on the session's disk. Memory works with LangGraph, LangChain, Strands and LlamaIndex out of the box, and the built-in extraction strategies can be used as-is, with your own overrides, or replaced by your own.
Code Interpreter is a sandbox where the agent writes and runs Python, JavaScript or TypeScript, isolated from everything else, with state kept between executions inside a session, file upload and download, and shell and AWS CLI commands. Sessions default to 15 minutes and can be extended to eight hours; inline file uploads go up to 100 MB, and files staged through S3 with terminal commands up to 5 GB, which is how an agent chews through gigabytes of CSV without hitting API limits. When you create your own interpreter you choose the network mode, Sandbox (no outside access) or Public, and the execution role that decides which AWS resources it can touch. Stop sessions when you are done; they bill until they end.
Browser is a managed cloud browser the agent drives over WebSocket automation endpoints, through Playwright, Nova Act or Strands, to fill forms, click through sites and read pages that have no API. The AWS-managed browser (aws.browser.v1) is the quick start; a custom browser adds session recording (DOM changes, actions, console and network events, replayable from S3 in the console), custom network settings and a specific execution role. Sessions default to 15 minutes with an 8-hour maximum, and a Live View endpoint lets a human watch, and take over, a session in real time, which is the humane answer to an agent stuck on a CAPTCHA. Both tools bill at the same per-second vCPU and GB rates as Runtime V1.
Observability and Evaluations: seeing what the agent did, and how well
Every step your agent takes on AgentCore can be traced: the model calls, the tool calls, the retrievals, the time each took and what went in and out. The traces are standard OpenTelemetry, so they flow into Amazon CloudWatch and any other stack that reads that format. On top of the traces sits Evaluations, which scores the quality of the answers with built-in checks such as correctness, helpfulness and tool-selection quality, third-party metrics, or your own rubric judged by a model or a Lambda function. The newest place to read all of this is CloudWatch Omni, released September 23, 2026, which puts agent traces and evaluation scores beside ordinary application health and runs its evaluators through AgentCore. If you only take one production habit from this page, make it this one: an agent can return a fast, error-free, well-formatted answer that is simply wrong, and only an evaluation catches that.
Getting started with the AgentCore CLI, step by step
You have three ways in, and which one you pick depends on how much code you want to write: the managed Harness (a config file, no loop to write), a code-based agent in a framework you already know, or the raw AWS SDKs for full control. The first two share one command-line tool, and the whole path from empty folder to a deployed, invocable agent is short enough to show in full.
Prerequisites: Node.js 20 or later (the CLI ships as an npm package), Python 3.10 or later for the agent code, AWS credentials configured, and Docker only if you choose the container build. Bedrock enables access to Amazon, Anthropic, Meta and Mistral foundation models by default; models from other providers need the model-access step first.
npm install -g @aws/agentcore # the CLI
agentcore --version
agentcore create # wizard: Harness or Agent; framework; model provider; memory; build type
cd MyProject
agentcore dev # local server with hot reload + an agent inspector in the browser
agentcore deploy # zips (or builds a container), provisions with CDK, creates the endpoint
agentcore status # shows the runtime ARN
agentcore invoke --prompt "Hello, what can you do?"
agentcore add memory # then: gateway, credential, evaluator, agent
agentcore logs --since 30m --level error
agentcore traces list
Five details from that flow save the most time. The wizard's choices: framework (Strands Agents, LangChain or LangGraph, Google ADK, OpenAI Agents SDK), model provider (Bedrock, Anthropic, OpenAI, Gemini), memory (none, short-term, or long- and short-term), and build type. Build types: CodeZip, the default, packages your code as a zip to S3 and needs no Docker; Container builds an image and needs a running Docker daemon. Deploy runs the AWS CDK under the hood, so the first deploy bootstraps your account and takes a few minutes; agentcore deploy --dry-run previews the change and --yes approves the bootstrap without a prompt. Cleanup is two commands: agentcore remove all then agentcore deploy, which detects the empty configuration and tears the resources down. And one gotcha: if agentcore --version errors instead of printing a version, an older Python package of the same name is shadowing the new CLI; run pip uninstall bedrock-agentcore-starter-toolkit and open a new terminal.
From your own code, invoking a deployed agent is one SDK call, with a fresh session ID per conversation and the same ID reused for follow-ups. Your identity needs the bedrock-agentcore:InvokeAgentRuntime permission; agents protected by OAuth are called over HTTPS with a bearer token instead of the SDK.
import boto3, json, uuid
client = boto3.client("bedrock-agentcore")
resp = client.invoke_agent_runtime(
agentRuntimeArn="arn:aws:bedrock-agentcore:us-west-2:111122223333:runtime/ticket-agent-ABCDE12345",
runtimeSessionId=str(uuid.uuid4()),
payload=json.dumps({"prompt": "Look up ticket 4471"}).encode(),
qualifier="DEFAULT",
)
print(json.loads(b"".join(resp["response"]).decode()))
Before you start, confirm three things in your account: the Region you picked supports what you need (Runtime microVMs run in 22 Regions including GovCloud US-West, but V2 in five, Memory in 15 and the Registry in 5), your execution role can reach the model you chose, and model access is enabled. If the model call fails with an access error, our Bedrock AccessDeniedException fix walks through every cause.
Where AgentCore runs: 22 Regions, unevenly
The Region map is the most uneven part of the product, and it decides your architecture more than any feature does. Runtime microVMs, Gateway, Identity, the built-in tools, Observability, Policy and Evaluations run in all 22 supported Regions, from N. Virginia to Hyderabad, Thailand and GovCloud (US-West). The rest is patchier, and this is the table to check before you promise a customer a Region.
| Piece | Regions |
|---|---|
| Runtime microVMs, Gateway, Identity, Code Interpreter and Browser, Observability, Policy, Evaluations | All 22, including GovCloud (US-West) |
| Runtime V2 (snapshot platform) | N. Virginia, Ohio, Oregon, Ireland, Tokyo |
| Memory, Harness | 15 Regions; not N. California, Milan, Spain, Malaysia, Hyderabad or Thailand |
| Runtime Instances (EC2-backed) | N. Virginia, Ohio, Oregon, Frankfurt, Ireland, Mumbai, Singapore, Sydney, Tokyo |
| Registry | N. Virginia, Oregon, Ireland, Sydney, Tokyo |
| Web search tool | N. Virginia, Ireland, Tokyo |
The practical rule: if you need Memory, Harness or V2, plan around Virginia, Oregon, Ireland or Tokyo, where everything overlaps. An EU customer who wants data to stay in Frankfurt gets Runtime, Gateway and Memory there but not V2 cold starts or the Registry today. Check the Region table in the AgentCore guide before every new deployment; it changed several times in 2026 and will again.
AgentCore vs Lambda, Fargate and doing it yourself
You can absolutely run an agent without AgentCore, so here is when it earns its place. Lambda is cheap and familiar, but a Lambda invocation tops out at 15 minutes for normal synchronous calls, keeps no per-user state, and shares execution environments across invocations of the same function; an agent session that needs an hour, isolation per user and memory between turns is working against the grain. Fargate or ECS gives you long-running containers, but isolation per user is yours to build, as are scaling to zero, identity, tool access and tracing. AgentCore is the option where per-session microVM isolation, the 8-hour session, scale to zero, and the security plumbing come built in, and you pay per second for it. For a hobby agent, Lambda is fine. For an agent that talks to other people's customers, the isolation alone usually settles it.
And against running an agent on your own laptop in a sandbox, the comparison is not really a competition: OpenShell-style sandboxes are for letting an agent loose on your machine safely; AgentCore is for serving an agent to other people. Many teams will use both, one for development and one for production.
For teams and IT: rolling AgentCore out without regrets
- Decide the platform version per agent, on data. Start new chat-style agents on V2 in a supported Region; keep memory-heavy agents and anything deployed by CloudFormation or CDK on V1 until the numbers and the tooling say otherwise.
- Put every tool behind Gateway, and every sensitive tool behind a Policy rule. Deterministic rules on tool calls are the control auditors understand.
- Move secrets into Identity on day one. An agent's environment variables are small on V2 anyway; use that as the push to get keys out of them entirely.
- Pin production to a named endpoint. Let
DEFAULTfollow the latest version in dev; point prod at a specific version and move it deliberately. - Turn on tracing and at least one custom evaluation before launch. The first incident with an agent is usually a wrong answer, not an error.
- Tag and watch the bill by agent. Runtime, Browser, Code Interpreter, Memory and Evaluations are separate meters, and the model tokens are on a different bill again.
What is Amazon Bedrock AgentCore?
A set of managed AWS services for building, deploying and operating AI agents securely at scale with any framework and any model. It includes Runtime (serverless hosting with one isolated microVM per session), Harness (a managed agent loop), Memory, Gateway (APIs as MCP tools), Identity, Code Interpreter, Browser, Observability, Evaluations, Policy and Registry. The pieces work together or independently.
What is the difference between Bedrock Agents and Bedrock AgentCore?
Bedrock Agents is the older feature for building an agent through the console with configuration. AgentCore is for running your own agent code, written in a framework such as LangGraph, CrewAI or Strands, on managed infrastructure with isolation, memory, tools, identity and observability built in. AgentCore also works with models outside Bedrock.
How much does Bedrock AgentCore cost?
There is no upfront fee. Runtime V1 costs $0.0895 per vCPU-hour and $0.00945 per GB-hour; Runtime V2 costs $0.1276 and $0.0169 on consumption pricing, billed per second on actual use. Gateway is $0.005 per 1,000 invocations, Memory $0.25 per 1,000 short-term events, Evaluations from $0.0024 per 1,000 input tokens, and Policy $0.000025 per authorization. Model tokens are billed separately.
What is AgentCore Runtime V2?
The second-generation runtime released September 18, 2026. It starts each session from a snapshot, giving cold starts of about 2 seconds at the 75th percentile whether the image is 200 MB or 2 GB, against 5.4 to 30 seconds on V1. It reclaims unused memory during a session and bills actual use. You enable it per runtime with platformVersion set to V2.
Is AgentCore Runtime V2 cheaper than V1?
It depends on the workload. V2's rates are 43% higher per vCPU-hour and 79% higher per GB-hour, but it bills memory by actual use rather than the session's peak, and does not charge for idle CPU waiting on I/O. For bursty, mostly-idle chat agents it is usually cheaper; for sessions that hold most of their memory the whole time, V1 is usually cheaper.
How long can an AgentCore Runtime session last?
Up to 8 hours of total runtime. A session ends after 15 minutes of inactivity, at the 8-hour limit, or if it becomes unhealthy. When it ends, its microVM is terminated and memory is sanitized, and a later request with the same session ID starts a fresh environment. Use AgentCore Memory for anything that must persist.
Which frameworks and models work with AgentCore?
Runtime works with custom code and open-source frameworks including LangGraph, CrewAI, LlamaIndex, Google ADK, the OpenAI Agents SDK and Strands Agents, and with models inside or outside Bedrock, including OpenAI, Google Gemini, Amazon Nova, Meta Llama and Mistral. It supports the HTTP, MCP and A2A protocols.
What is AgentCore Gateway?
A service that turns your APIs, Lambda functions, OpenAPI specs and SaaS integrations into Model Context Protocol tools behind one endpoint, and connects existing MCP servers. Agents discover and call tools through it without seeing internal URLs or credentials. It costs $0.005 per 1,000 invocations, $0.025 per 1,000 searches and $0.02 per 100 tools indexed per month.
What is AgentCore Harness?
A managed agent loop: you specify a model, a system prompt and tools in a single API call, and Harness handles orchestration, tool execution, memory and response generation. Each session runs in an isolated microVM with filesystem and shell access, and you can bring your own container image. It works with Bedrock, OpenAI, Google Gemini and OpenAI-compatible providers.
What is AgentCore Policy?
A capability that checks every tool call passing through AgentCore Gateway against deterministic rules before it executes. You write rules in natural language or in Dogwood, a policy language compatible with Cedar. It costs $0.000025 per authorization request and $0.13 per 1,000 tokens of natural-language authoring.
Which Regions support AgentCore Runtime V2?
At launch, US East (N. Virginia), US East (Ohio), US West (Oregon), Europe (Ireland) and Asia Pacific (Tokyo). V1 remains the default in every Region where AgentCore runs.
Can I deploy AgentCore Runtime V2 with CloudFormation or CDK?
Not yet. AWS CloudFormation and the AWS CDK do not currently support setting platformVersion. Create or update V2 runtimes through the console, the AWS CLI or the SDKs until support lands.
Why does my AgentCore V2 runtime fail to create?
The most common causes are a container that does not report healthy on /ping within 120 seconds of startup, environment variables above V2's current limit of 1.5 KB for direct code or 2.5 KB for containers, and calling update or delete while the runtime is still CREATING or UPDATING, which returns ConflictException. V2 creation also takes several minutes, so poll until READY.
How does AgentCore keep one customer's session from seeing another's?
Each session runs in its own microVM with dedicated CPU, memory and filesystem, never shared with another session. When a session ends, the microVM is terminated and its memory is sanitized. Nothing persists between sessions except what you deliberately store in AgentCore Memory or your own systems.
Does AgentCore include the cost of the AI model?
No. AgentCore bills for hosting, tools, memory, identity, evaluations and policy. The tokens your agent sends to a model are billed separately by Amazon Bedrock or by the model provider you call.
Where are AgentCore observability traces shown?
Traces are emitted in OpenTelemetry format and delivered to Amazon CloudWatch, where CloudWatch's generative AI observability views and CloudWatch Omni can show them, alongside any other OpenTelemetry-compatible monitoring stack you use.
Jake's client said yes, in the end, and the deciding line was not a feature, it was the microVM. "Your customers never share a machine with anyone else's, and when they stop talking, their machine is destroyed." Ethan's line was about the bill: "Don't pick the runtime by its rate card. Pick it by how your agent waits." Both of those are true of almost every agent you will build. Start small with one runtime, one Gateway tool and tracing on, measure a week of real sessions, and let the numbers choose V1 or V2 for you.
📌 If you keep one line from this page
V2 costs more per hour and less per session when your agent mostly waits.
One microVM per session, wiped after 15 idle minutes; choose V1 or V2 by how your agent spends its time, not by the rate card.
Revision note. Written September 29, 2026, eleven days after AgentCore Runtime V2 became available, from the AgentCore developer guide, the V2 launch announcement and technical post, and the AgentCore pricing page as published that day. The single-session cost comparison uses AWS's published rates and an illustrative session; your agent's real memory curve decides the answer. Prices, Regions and the CloudFormation and CDK gap will change; this page gets a dated update when they do.