Skip to main content

AWS Bedrock

Nastech Agent supports Amazon Bedrock as a native provider. This gives you full access to the Bedrock ecosystem: IAM authentication, Guardrails, cross-region inference profiles, and all foundation models.

Nastech routes each model family through the API that serves it best:

Model familyAPI routeWhy
Anthropic ClaudeAnthropic SDK (AnthropicBedrock)Prompt caching, thinking budgets, adaptive thinking — features not exposed via Converse
OpenAI GPT-5.5 / GPT-5.6 (Sol, Terra, Luna)Bedrock Mantle OpenAI Responses endpoint (bedrock-mantle.<region>.api.aws/openai/v1)These models are Mantle-only — their model cards list bedrock-runtime/Converse as unsupported
Everything else (Nova, DeepSeek, Llama, GPT-OSS, …)Native Converse API (bedrock-runtime)Full Bedrock feature set: Guardrails, inference profiles, streaming

All three routes share the same AWS credential chain and region resolution — no separate configuration is needed. Requests to the Mantle endpoint are authenticated with AWS_BEARER_TOKEN_BEDROCK when set, or SigV4-signed via the standard boto3 credential chain otherwise.

Prerequisites​

  • AWS credentials — any source supported by the boto3 credential chain:
    • IAM instance role (EC2, ECS, Lambda — zero config)
    • AWS_ACCESS_KEY_ID + AWS_SECRET_ACCESS_KEY environment variables
    • AWS_PROFILE for SSO or named profiles
    • aws configure for local development
  • boto3 — install with cd ~/.nastech/nastech-agent && uv pip install -e ".[bedrock]"
  • IAM permissions — at minimum:
    • bedrock:InvokeModel and bedrock:InvokeModelWithResponseStream (for inference)
    • bedrock:ListFoundationModels and bedrock:ListInferenceProfiles (for model discovery)
    • bedrock:GetInferenceProfile (only if model.default is an application inference profile ARN — used to size the context window from the wrapped model)
EC2 / ECS / Lambda

On AWS compute, attach an IAM role with AmazonBedrockFullAccess and you're done. No API keys, no .env configuration — Nastech detects the instance role automatically.

Quick Start​

# Install with Bedrock support
cd ~/.nastech/nastech-agent && uv pip install -e ".[bedrock]"

# Select Bedrock as your provider
nastech model
# → Choose "More providers..." → "AWS Bedrock"
# → Select your region and model

# Start chatting
nastech chat

Configuration​

After running nastech model, your ~/.nastech/config.yaml will contain:

model:
default: us.anthropic.claude-sonnet-4-6
provider: bedrock
base_url: https://bedrock-runtime.us-east-2.amazonaws.com

bedrock:
region: us-east-2

Region​

Set the AWS region in any of these ways (highest priority first):

  1. bedrock.region in config.yaml
  2. AWS_REGION environment variable
  3. AWS_DEFAULT_REGION environment variable
  4. Default: us-east-1

Guardrails​

To apply Amazon Bedrock Guardrails to all model invocations:

bedrock:
region: us-east-2
guardrail:
guardrail_identifier: "abc123def456" # From the Bedrock console
guardrail_version: "1" # Version number or "DRAFT"
stream_processing_mode: "async" # "sync" or "async"
trace: "disabled" # "enabled", "disabled", or "enabled_full"

The guardrail is attached on the Converse route (guardrailConfig) and on the Claude route (InvokeModel headers via the Anthropic Bedrock SDK, so prompt caching and thinking are kept). A blocked request surfaces as a content-filter refusal rather than as model text. stream_processing_mode only applies to Converse.

AWS does not apply Guardrails to the Mantle Responses endpoint used by openai.gpt-5.x models (AWS docs); use a Converse-served model when a guardrail is required.

Model Discovery​

Nastech auto-discovers available models via the Bedrock control plane. You can customize discovery:

bedrock:
discovery:
enabled: true
provider_filter: ["anthropic", "amazon"] # Only show these providers
refresh_interval: 3600 # Cache for 1 hour

Prompt caching (cachePoint)​

Nastech automatically applies prompt caching on the Bedrock Converse API path by inserting cachePoint markers after the system prompt, tool definitions, and the latest message. Because sending a cachePoint block to a model that doesn't support it raises a ValidationException, markers are only added for models on a known-good allowlist (Anthropic Claude and Amazon Nova model IDs); unknown models default to no cache markers. Claude models normally use the AnthropicBedrock SDK path, which has its own prompt caching — the Converse cachePoint path covers Nova and the bearer-token Claude fallback. No configuration needed; cache reads/writes show up in usage accounting.

Context-window probing​

For models whose context window isn't in Nastech' static table, Nastech can probe the real limit by sending oversized requests at fixed tiers (~1.3M and ~2.2M tokens) and parsing the maximum reported in Bedrock's length-validation error. Probed values feed the same metadata cache as the static table; stale cached entries that under-report a model's window (e.g. entries seeded before a model's 1M window went GA) are dropped automatically in favor of the larger known value.

Application inference profiles. An ARN such as arn:aws:bedrock:us-west-2:123456789012:application-inference-profile/abcdef123456 names no model, so neither the probe nor the static table can size it. Nastech calls bedrock:GetInferenceProfile in the ARN's region and sizes the window from the model the profile wraps (1M for a profile wrapping Claude Sonnet 4.6). Without that permission the 128,000-token default applies and a WARNING names the profile; set model.context_length explicitly to override either way.

Available Models​

Bedrock models use inference profile IDs for on-demand invocation. The nastech model picker shows these automatically, with recommended models at the top:

ModelIDNotes
Claude Sonnet 4.6us.anthropic.claude-sonnet-4-6Recommended — best balance of speed and capability
Claude Opus 4.6us.anthropic.claude-opus-4-6-v1Most capable
Claude Haiku 4.5us.anthropic.claude-haiku-4-5-20251001-v1:0Fastest Claude
OpenAI GPT-5.6 Solopenai.gpt-5.6-solOpenAI frontier model (via Bedrock Mantle)
OpenAI GPT-5.6 Terraopenai.gpt-5.6-terraBalanced (via Bedrock Mantle)
OpenAI GPT-5.6 Lunaopenai.gpt-5.6-lunaFast, affordable (via Bedrock Mantle)
OpenAI GPT-5.5openai.gpt-5.5Previous OpenAI flagship (via Bedrock Mantle)
Amazon Nova Prous.amazon.nova-pro-v1:0Amazon's flagship
Amazon Nova Microus.amazon.nova-micro-v1:0Fastest, cheapest
DeepSeek V3.2deepseek.v3.2Strong open model
Llama 4 Scout 17Bus.meta.llama4-scout-17b-instruct-v1:0Meta's latest
Cross-Region Inference

Models prefixed with us. use cross-region inference profiles, which provide better capacity and automatic failover across AWS regions. Models prefixed with global. route across all available regions worldwide. OpenAI openai.* model IDs are served by Bedrock Mantle in the configured region and don't use inference-profile prefixes.

Switching Models Mid-Session​

Use the /model command during a conversation:

/model us.amazon.nova-pro-v1:0
/model deepseek.v3.2
/model us.anthropic.claude-opus-4-6-v1

Diagnostics​

nastech doctor

The doctor checks:

  • Whether AWS credentials are available (env vars, IAM role, SSO)
  • Whether boto3 is installed
  • Whether the Bedrock API is reachable (ListFoundationModels)
  • Number of available models in your region

Gateway (Messaging Platforms)​

Bedrock works with all Nastech gateway platforms (Telegram, Discord, Slack, Feishu, etc.). Configure Bedrock as your provider, then start the gateway normally:

nastech gateway setup
nastech gateway start

The gateway reads config.yaml and uses the same Bedrock provider configuration.

Troubleshooting​

"No API key found" / "No AWS credentials"​

Nastech checks for credentials in this order:

  1. AWS_BEARER_TOKEN_BEDROCK
  2. AWS_ACCESS_KEY_ID + AWS_SECRET_ACCESS_KEY
  3. AWS_PROFILE
  4. EC2 instance metadata (IMDS)
  5. ECS container credentials
  6. Lambda execution role

If none are found, run aws configure or attach an IAM role to your compute instance.

"Invocation of model ID ... with on-demand throughput isn't supported"​

Use an inference profile ID (prefixed with us. or global.) instead of the bare foundation model ID. For example:

  • ❌ anthropic.claude-sonnet-4-6
  • ✅ us.anthropic.claude-sonnet-4-6

"ThrottlingException"​

You've hit the Bedrock per-model rate limit. Nastech automatically retries with backoff. To increase limits, request a quota increase in the AWS Service Quotas console.

One-Click AWS Deployment​

For a fully automated deployment on EC2 with CloudFormation:

sample-nastech-agent-on-aws-with-bedrock — creates VPC, IAM role, EC2 instance, and configures Bedrock automatically. Deploy in any region with one click.