Skip to main content

Amazon Bedrock Guide (2026) | Complete Guide to AWS Generative AI Platform

Introduction​

Amazon Bedrock is AWS’s fully managed generative AI service that provides secure, enterprise-grade access to foundation models (FMs) from leading AI providers through a unified API. Bedrock is AWS’s control plane for enterprise AIβ€”bundling the model catalog, higher-level building blocks (Agents, Knowledge Bases, Guardrails, Prompt Management), and tight integration with the rest of AWS (IAM, KMS, VPC endpoints, CloudWatch, CloudTrail).

By 2026, Bedrock has evolved from a model gateway into a comprehensive enterprise AI platform. It offers roughly 100 serverless models from Amazon, Anthropic, Meta, Mistral, Cohere, and othersβ€”including Claude Sonnet 4.6 and Opus 4.6, the Amazon Nova family, and open-weight options. The platform now includes Bedrock Agents for multi-step task automation, Knowledge Bases for managed RAG pipelines, Guardrails for responsible AI controls, Custom Model Import for bringing your own models, and AgentCore as a managed agent runtime.

This guide covers everything from Bedrock’s architecture and model catalog to its agentic capabilities, RAG pipelines, safety controls, pricing, and enterprise use cases.

What Is Amazon Bedrock?​

Amazon Bedrock is a fully managed service that offers a choice of high-performing foundation models from leading AI companies through a single serverless API. It abstracts away the complexity of managing inference infrastructure, allowing developers to build and scale generative AI applications without provisioning GPUs or managing model deployments.

Key Characteristics​

CharacteristicDescription
Fully managedNo infrastructure to provision or manage
ServerlessPay only for what you use, no minimum commitment
Unified APIOne API for all supported models
Enterprise securityIAM, KMS encryption, VPC endpoints, CloudTrail logging
AWS-nativeDeep integration with AWS services
Model diversity~100 models from multiple providers

The Bedrock Architecture​

β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚ Application Layer β”‚
β”‚ (Lambda, ECS, EKS, EC2) β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
↓
β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚ Amazon Bedrock β”‚
β”‚ β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β” β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β” β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β” β”‚
β”‚ β”‚ Agents β”‚ β”‚ Knowledge β”‚ β”‚ Guardrails β”‚ β”‚
β”‚ β”‚ β”‚ β”‚ Bases β”‚ β”‚ β”‚ β”‚
β”‚ β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜ β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜ β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜ β”‚
β”‚ β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β” β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β” β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β” β”‚
β”‚ β”‚ Model β”‚ β”‚ Prompt β”‚ β”‚ Evaluation β”‚ β”‚
β”‚ β”‚ Evaluation β”‚ β”‚ Management β”‚ β”‚ β”‚ β”‚
β”‚ β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜ β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜ β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜ β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
↓
β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚ Foundation Models β”‚
β”‚ Anthropic β”‚ Amazon Nova β”‚ Meta Llama β”‚ Mistral β”‚ Cohere β”‚
β”‚ AI21 β”‚ Stability β”‚ DeepSeek β”‚ MiniMax β”‚ GLM β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜

Supported Foundation Models​

As of 2026, Bedrock offers roughly 100 serverless models across multiple providers.

Model Providers​

Anthropic Claude – The Claude family (Sonnet, Opus, Haiku) is the most common choice for agents due to strong tool use, long context handling, and document reasoning. Claude Sonnet 4.6 is available at $3.30/$16.50 per MTok.

Amazon Nova – Amazon’s own foundation model family covering multiple tiers:

  • Nova Premier ($2.50/$12.50 per MTok) – Most capable
  • Nova Pro ($0.80/$3.20 per MTok) – Balanced performance
  • Nova Lite ($0.06/$0.24 per MTok) – Cost-effective
  • Nova Micro ($0.035/$0.14 per MTok) – Lowest latency
  • Nova Canvas – Image generation
  • Nova Reel – Video generation
  • Nova Sonic – Speech understanding

Meta Llama – Open-weight options for cost-sensitive workloads or fine-tuning. Llama 3.1 405B: $5.32/$16.00 per MTok.

Mistral AI – Strong European option for multilingual workloads. Mixtral 8Γ—7B: $0.59/$0.91 per MTok.

Cohere – Command R+ tuned for RAG and tool use. Command Text v14: $1.50/$2.00 per MTok.

AI21 Labs – Jamba hybrid Mamba-Transformer architecture with long context.

Stability AI – Image generation models.

Additional models – DeepSeek V3.2, MiniMax M2.1, GLM 4.7, Kimi K2.5, Qwen3 Coder Next.

Model Selection Strategy​

Bedrock’s strength is model diversity in one integrationβ€”route classification to Nova Micro, extraction to Nova Lite, and complex reasoning to Claude Sonnet 4.6 from the same codebase. Switching models is a parameter change, not a re-architecture.

Core Features​

1. Bedrock Agents​

Amazon Bedrock Agents uses the reasoning of foundation models, APIs, and data to break down user requests, gather relevant information, and efficiently complete tasksβ€”freeing teams to focus on high-value work.

How agents work: You define a base foundation model, an instructions prompt defining the agent’s role and constraints, one or more action groups (backed by Lambda functions with OpenAPI schemas), optional Knowledge Base attachments for retrieval, and optional session memory. At runtime, the agent receives a user prompt, plans steps, calls tools through action groups, and composes a final response.

Key capabilities:

  • Multi-agent collaboration: Multiple specialized agents work together under a supervisor agent
  • Memory retention: Agents remember historical interactions for personalized experiences
  • Code interpretation: Dynamically generate and execute code in a secure environment
  • RAG integration: Securely connect to company data sources

2. Amazon Bedrock AgentCore​

AgentCore is a managed runtime platform for building, connecting, and optimizing agents at scale. It runs the orchestration loop, executes tools, manages the context window, persists state, and isolates each session.

Key capabilities introduced in 2026:

  • Three knowledge layers: Organizational (Managed Knowledge Base), web (Web Search), and paid knowledge
  • Managed harness: Production-grade agent runtime with dynamic scaling
  • Model flexibility: Choose any model and switch providers mid-session
  • Web Search: Ground agents in current, accurate web knowledge
  • Guardrails integration: Enforce controls that scale as agents grow more capable

3. Knowledge Bases (Managed RAG)​

Amazon Bedrock Knowledge Bases provide managed RAG pipelines that give foundation models and agents contextual information from private data sources for more relevant, accurate, and customized responses.

Managed Knowledge Base (GA June 2026) abstracts away the complexity of building and managing RAG pipelines, allowing developers to focus on business outcomes rather than infrastructure management.

Key features:

  • Native data connectors: Six pre-built ingestion connectors for Amazon S3, SharePoint, Confluence, Web Crawler, Google Drive, and OneDrive
  • Smart Parsing: Automatically selects the right parsing strategy for each data type
  • Agentic Retriever: Multi-turn, multi-hop retrieval across one or multiple knowledge bases
  • Multimodal support: RAG for images, audio, and video

S3 Vectors integration: Knowledge Bases on S3 Vectors collapses retrieval-layer economics by an order of magnitude for storage-bound workloads. One user reported paying β€œcents per day instead of $700/month minimum”.

4. Guardrails (Responsible AI)​

Amazon Bedrock Guardrails provides a policy layer that sits between your application and any model in the catalog. A single guardrail applies the same rules across Claude, Llama, Nova, and Mistral.

Key capabilities:

  • Content filters: Detect and filter harmful content across hate, violence, sexual, insults, and misconduct categories
  • Prompt attack detection: Identify jailbreak, prompt injection, and prompt leakage
  • Sensitive information filters: Detect supported PII entity types
  • Denied topics: Define topics the model should not discuss

Automated Reasoning checks (June 2026) use formal verification techniques to validate AI model outputs with mathematical rigor, delivering up to 99% accuracy in detecting correct responses. This makes AWS the first major cloud provider to integrate automated reasoning in generative AI offerings.

InvokeGuardrailChecks API (June 2026) provides granular, per-request control over which safeguards to run at each step of your agent loop.

5. Custom Model Import​

Custom Model Import enables the import and use of customized models alongside existing foundation models through a single serverless, unified API. You can leverage native Bedrock toolingβ€”Knowledge Bases, Guardrails, and Agentsβ€”with imported custom models.

Supported variants include DeepSeek-R1-Distill-Llama-8B and 70B.

6. Model Evaluation​

Bedrock provides tools for comparing model performance through human evaluation and automated evaluation, helping you select the right model for your use case.

Enterprise Use Cases​

Enterprise Chatbots & Knowledge Assistants​

Build secure, RAG-powered chatbots that answer questions from internal documents. Bedrock Agents with Knowledge Bases retrieve and synthesize information from enterprise data sources.

Customer Support Automation​

Agents can classify tickets, retrieve knowledge base articles, and draft responses. With Guardrails and session memory, agents provide consistent, personalized support.

Software Development Assistance​

Bedrock supports OpenAI GPT-5.5, GPT-5.4, and Codex models, enabling code generation, review, and debugging within the AWS ecosystem.

Document Intelligence​

Process and analyze documents at scale using RAG pipelines with Smart Parsing for different content types.

Business Process Automation​

Multi-agent collaboration enables multiple specialized agents to work together on complex workflows under a supervisor agent.

Regulated Workloads​

Guardrails, CloudTrail logging, VPC endpoints, and IAM-scoped model access support compliance requirements.

Amazon Bedrock Pricing​

Bedrock pricing has five or six moving parts that are not obvious until you have lived through a billing cycle. The service charges through four core modes:

On-Demand Pricing​

Pay per 1,000 input tokens, per 1,000 output tokens, per image, or per second of generated video. No commitment, no minimum.

Example rates (June 2026):

ModelInput (per MTok)Output (per MTok)
Claude Sonnet 4.6$3.30$16.50
Nova Premier$2.50$12.50
Nova Pro$0.80$3.20
Nova Lite$0.06$0.24
Nova Micro$0.035$0.14
Llama 3.1 405B$5.32$16.00
Cohere Command v14$1.50$2.00

Rates vary by regionβ€”verify current pricing at the official Amazon Bedrock Pricing page.

Provisioned Throughput​

Reserve dedicated capacity for a specific model, billed hourly whether used or not. Requires 1-month or 6-month commitments. Rates range from ~$21/hour (Meta Llama) to ~$50/hour (Stability AI).

Batch Inference​

Run asynchronous jobs at 50% off the On-Demand rate.

Prompt Caching​

Cache repeated input context (system prompts, large knowledge snippets) and pay up to 90% off the input-token portion.

Additional Costs​

  • Knowledge Bases vector storage: OpenSearch Serverless carries a ~$350/month minimum
  • Bedrock Guardrails: $0.15 per 1,000 text units
  • Bedrock Flows: $0.035 per 1,000 visual node transitions

Cost Optimization​

The single biggest lever is model routing: send simple requests to Nova Lite ($0.06/MTok) or Nova Micro ($0.035/MTok), and reserve Claude Sonnet 4.6 or Nova Premier only for queries that need it.

Amazon Bedrock vs Other Enterprise AI Platforms​

Amazon Bedrock vs Azure AI Foundry​

DimensionAWS BedrockAzure AI Foundry
Model ecosystem~100 models, widest open-source coverageOpenAI partnership, HuggingFace models
Best forAWS-native teamsMicrosoft 365 teams
IdentityIAMEntra ID
Agent runtimeAgentCoreAgent Service

Amazon Bedrock vs Google Vertex AI​

DimensionAWS BedrockGoogle Vertex AI
Best forModel breadth and EU optionsGemini and long context
Agent runtimeAgentCoreVertex AI Agent Engine

Amazon Bedrock vs Amazon SageMaker​

DimensionBedrockSageMaker
PurposeManaged foundation modelsFull ML lifecycle
CustomizationFine-tuning, Custom Model ImportFull model training
Target usersApplication developersML engineers, data scientists

Best Practices​

1. Choose the Right Model​

Start with Nova Lite or Nova Micro for simple tasks. Use Claude Sonnet 4.6 or Nova Premier for complex reasoning. Switch models via parameter change.

2. Use Knowledge Bases for Enterprise RAG​

Let Managed Knowledge Base handle ingestion, chunking, embedding, and retrieval. Use S3 Vectors for cost-sensitive workloads.

3. Apply Guardrails in Production​

Deploy Guardrails with content filters, denied topics, and PII detection. Use the InvokeGuardrailChecks API for granular control.

4. Monitor Inference Costs​

Track token usage by model. Route simple requests to smaller models to cut spend significantly.

5. Implement Least-Privilege IAM​

Each tool call runs under an IAM role you control, so the agent never sees credentials or services it is not authorized for.

6. Use AgentCore for Production Agents​

Let AgentCore handle orchestration, state management, and scaling. Use Web Search and Managed Knowledge Base for broader knowledge access.

7. Optimize Prompts Before Fine-Tuning​

Test prompt engineering before investing in model customization.

Frequently Asked Questions​

What is Amazon Bedrock?​

Amazon Bedrock is AWS’s fully managed service providing access to foundation models from Amazon, Anthropic, Meta, Mistral, and others for building generative AI applications.

Which models does Bedrock support?​

Bedrock offers roughly 100 serverless models including Anthropic Claude, Amazon Nova, Meta Llama, Mistral, Cohere, AI21, Stability AI, DeepSeek, and more.

What are Bedrock Agents?​

Bedrock Agents wrap a model with planning, tool use, and optional memory to automate multi-step tasks by connecting with company systems, APIs, and data sources.

What are Knowledge Bases?​

Knowledge Bases provide managed RAG pipelines, giving agents contextual information from private data sources for more accurate responses. Managed Knowledge Base became GA in June 2026.

What are Guardrails?​

Guardrails provide a policy layer for content filtering, prompt attack detection, PII protection, and automated reasoning checks.

How does Bedrock differ from SageMaker?​

Bedrock is for managed foundation models with pay-per-token pricing. SageMaker is for full ML lifecycle including custom model training and deployment.

Is Bedrock suitable for enterprise AI?​

Yes. Bedrock offers enterprise security (IAM, KMS, VPC), compliance features (Guardrails, CloudTrail), and deep AWS integration.

Can Bedrock build RAG applications?​

Yes. Bedrock Knowledge Bases provide managed RAG with native connectors, Smart Parsing, and Agentic Retriever.

  • [Amazon Bedrock Tutorial]
  • [Amazon Bedrock Pricing Guide]
  • [Amazon Bedrock Agents Guide]
  • [Amazon Bedrock Knowledge Bases Guide]
  • [Amazon Bedrock Guardrails Guide]
  • [Amazon Bedrock API Guide]

Conclusion​

Amazon Bedrock has evolved from a model gateway into a comprehensive enterprise AI platform. By 2026, it offers roughly 100 serverless models, managed Agents for multi-step automation, Knowledge Bases for RAG, Guardrails for responsible AI, and AgentCore as a production runtime for agents.

The platform’s key strengthsβ€”model diversity, AWS-native security, managed infrastructure, and enterprise governanceβ€”make it the natural choice for organizations already on AWS. The addition of Managed Knowledge Base (June 2026), AgentCore with Web Search and three knowledge layers, and Automated Reasoning checks in Guardrails positions Bedrock as one of the most comprehensive enterprise AI platforms available.

When to choose Bedrock:

  • Your team already runs on AWS
  • You need multiple foundation models behind one API
  • You require enterprise security, compliance, and governance
  • You want managed RAG, agents, and guardrails without building from scratch

When to consider alternatives:

  • Azure AI Foundry: If your team lives in Microsoft 365
  • Google Vertex AI: If you need Gemini deep integration
  • Self-hosting: If you have steady, high-volume workloads and want lower per-token costs

For AWS-native teams building production generative AI applications, Bedrock provides the most complete, secure, and scalable foundation available in 2026.