SaaS Architecture & Economics | AI in the EU series

    Controlling the Cost of Claude on Amazon Bedrock

    Understand regional pricing, budgets, cost attribution and quota planning for Claude workloads on Amazon Bedrock.

    Siva SadhuBy Siva SadhuFounder and Principal ConsultantLast checked 29 September 2026
    Measured token flows and cost controls across European cloud regions

    One of the worries in the conversation that started this series was billing. People assumed Claude on Amazon Bedrock would be hard to predict. It isn’t, but the controls don’t set themselves up.

    Claude on Amazon Bedrock is usage-based rather than seat-based. For Claude Sonnet 4.5, Haiku 4.5, Opus 4.5 and later models that use Anthropic’s Global-versus-Regional pricing structure, Regional endpoints are priced at a 10% premium to Global. That is not a universal rule for every historical Claude model, so check the current model price before forecasting.

    How you’re billed

    • Through AWS Marketplace, on your AWS bill. Claude is a third-party model on Bedrock, so charges show up on your AWS bill and in Cost Explorer under the model provider, not under Amazon Bedrock (AWS: Claude Sonnet 5 model card). Finance teams looking under “Bedrock” won’t find it.
    • Input and output are priced separately. On the models in the current table, output tokens cost five times as much as input tokens, so long answers can dominate the bill.
    • By source Region, with no extra charge for routing. Geographic cross-Region inference has no routing fee. You pay the price of the Region you call from (AWS Alps blog).
    • By service tier. Bedrock has Standard, Priority, Flex and Reserved tiers, but not every model supports every tier. Claude Sonnet 5, for example, currently supports Standard only (AWS: Claude Sonnet 5 model card).

    What EU routing costs

    For Claude Sonnet 4.5, Haiku 4.5, Opus 4.5 and later models, Anthropic documents a 10% premium for Regional and multi-region endpoints compared with Global. The table below uses that rule only for models that currently have an EU-routable option on Bedrock.

    ModelList input, $/1M tokensList output, $/1M tokensEU input, $/1M tokensEU output, $/1M tokens
    Claude Opus 5.54.0020.004.4022.00
    Claude Opus 55.0025.005.5027.50
    Claude Sonnet 4.5 / 4.63.0015.003.3016.50
    Claude Sonnet 52.0010.002.2011.00
    Claude Haiku 4.51.005.001.105.50

    Prices move quickly. The figures above were checked on 29 September 2026. AWS sets the Bedrock prices you are actually billed, so verify the Amazon Bedrock pricing page before you quote them externally.

    To put the premium in perspective: 1,000 requests a day on Sonnet 5, each with 2,000 input tokens and 500 output tokens, is about $9.90 a day at the EU Regional rate, or roughly $300 a month. The same token volume at the Global rate is about $270 a month. Sonnet 5.5 is priced at the same $2/$10 list rate, but as of 29 September it does not yet have an EU Geo or EU In-Region Bedrock option, so it is not included in the EU table.

    What I would put in place before the first team gets access

    1. Set budgets and alerts first. Put an AWS Budget on the account that runs Claude, with alerts at, say, 50%, 80% and 100% of the monthly figure. Remember that the cost appears under the model provider in Cost Explorer.
    2. Give each team its own account or profile. Anthropic recommends a dedicated AWS account for Claude Code to make cost tracking simpler (Claude Code docs). For applications, application inference profiles let each team or product call Claude through its own profile, so usage can be tracked separately (AWS: Inference profiles).
    3. Choose models on purpose. Opus is materially more expensive than Sonnet for many workloads. Claude Code can also change defaults as new models launch, so pin the model you intend to fund instead of inheriting a moving default.
    4. Use prompt caching. Repeated system prompts and documents can be cached, and for most models a cache hit costs 10% of the normal input price (Anthropic: Pricing). Check caching support for your model and Region on Bedrock.
    5. Use cheaper tiers for batch work where the model supports them. Flex is designed for jobs that aren’t urgent.

    Quotas can be a bigger surprise than the bill

    The first problem teams hit in the EU is often a quota, not the invoice. Your AWS account has default Bedrock quotas, and AWS notes they can vary with regional factors, payment history and approval of increase requests (AWS: Claude Sonnet 5 model card). An increase is something you ask for, not something you’re guaranteed.

    • Ask for increases early, weeks before launch, with a realistic estimate of your volume.
    • Use EU cross-Region profiles where you don’t need single-Region processing, so the load is spread across EU Regions.
    • Plan a fallback that stays in the EU. If one Region runs out of quota, fall back to another EU Region, never to Global. For single-Region workloads on bedrock-mantle, Anthropic lists a default quota of 2 million input tokens per minute (Anthropic: Claude in Amazon Bedrock).

    Want the cost model before the rollout?

    Wolkn Minds can estimate spend from representative workloads, set up budgets and attribution, and help you request the quotas you actually need.

    Sources

    Open-weight model cores deployed across European infrastructureNext in the AI in the EU seriesOpen-Weight Coding Models in the EU: Qwen, Gemma, Mistral and Where to Run ThemCompare open-weight coding models and the European cloud, sovereign and self-hosted environments where they can run.Read next

    Related reading

    Want a second view on your SaaS architecture or costs?

    A SaaS Architecture Review or Cloud Cost and Unit Economics Review gives you findings, options and a prioritised roadmap in one to three weeks.