Agent web search APIs in 2026: $5 to $35 per 1,000 queries compared

Six web grounding APIs priced from vendor pages, August 2026: a 5.4x spread on the same 200,000 queries.

Read time
18 min
Word count
2.8K
Sections
13
FAQs
8
Share
Cost comparison of six agent web search grounding APIs in 2026
Published rates for six agent web grounding APIs, checked 3 August 2026.
On this page · 13 sections
  1. The headline rates, from the vendors' own pricing pages
  2. What 200,000 searches a month actually costs
  3. The variable that actually decides enterprise deals: where the query goes
  4. The citation display rule nobody budgets for
  5. Billing units: where the estimate breaks
  6. Why Microsoft customers have no easy exit
  7. Cost per task: a worked model
  8. India-specific considerations
  9. What the launch customers actually said
  10. How to choose
  11. FAQ
  12. How eCorpIT can help
  13. References

Summary. Six production web grounding APIs are priced between $5 and $35 per 1,000 calls as of 3 August 2026. Brave Search sits at $5.00 per 1,000 requests, Amazon Bedrock AgentCore Web Search and Exa Search at $7 per 1,000, Grounding with Bing Search at $14 per 1,000 transactions, and Gemini's Grounding with Google Search at $14 per 1,000 search queries on Gemini 3 models but $35 per 1,000 grounded prompts on Gemini 2.5. AWS published its own worked example on the AgentCore pricing page: 200,000 research queries per month costs $1,403.00. The same 200,000 queries cost $2,800 on Grounding with Bing and $5,425 on Gemini 2.5 after the free allowance. That is a 5.4x spread on an identical workload, and price is not even the variable that decides most enterprise deals. Data egress and citation display rules are.

Amazon made Web Search on Bedrock AgentCore generally available in June 2026 and added an explicit pricing statement to the AWS News Blog launch post on 18 June 2026. Microsoft retired the standalone Bing Search APIs on 11 August 2025 and pushed every remaining customer into Grounding with Bing inside Azure AI Foundry. Google bills grounding per individual search query, not per prompt, so a single user turn can bill three or four times. Those three facts, not the headline rate card, are what move a real bill.

The headline rates, from the vendors' own pricing pages

Every figure below comes from the provider's published pricing page, checked on 3 August 2026. Rates are for the standard, non-enterprise tier.

Provider Published rate Free allowance Rate limit stated
Brave Search API $5.00 per 1,000 requests $5 in credits every month 50 requests per second
Amazon Bedrock AgentCore Web Search $7.00 per 1,000 queries Up to $200 Free Tier credits for new AWS customers Not published on the pricing page
Exa Search $7 per 1,000 requests, up to 10 results $20 on signup, then $10 in credits monthly Custom QPS on Enterprise only
Tavily Search $0.008 per credit, 1 credit per basic search 1,000 credits per month Not published on the credits page
Grounding with Bing Search $14 per 1,000 transactions None published 150 transactions per second, 1 million per day
Grounding with Google Search $14 per 1,000 search queries on Gemini 3; $35 per 1,000 grounded prompts on Gemini 2.5 5,000 prompts per month on Gemini 3; 1,500 requests per day on Gemini 2.5 paid tier Model rate limits apply

Brave's Answers product is priced separately at $4.00 per 1,000 queries plus $5.00 per million input tokens and $5.00 per million output tokens, and it is capped at 2 requests per second against the Search product's 50. Exa's Answer endpoint is $5 per 1,000 requests, its Deep Search tiers are $12 to $15 per 1,000, and its Monitors endpoint is $15 per 1,000. Tavily's advanced search depth costs 2 credits instead of 1, which doubles the effective rate to $0.016 per search on pay-as-you-go.

Brave also ships an LLM Context endpoint, launched on 6 February 2026, that bills at the same $5.00 per 1,000 requests but returns pre-extracted content with explicit token budgets: maximum_number_of_tokens defaults to 8,192 and accepts 1,024 to 32,768, and maximum_number_of_tokens_per_url defaults to 4,096. That control matters more than the search price for anyone paying per input token downstream, because it caps what the grounding step can push into the model's context window.

What 200,000 searches a month actually costs

Rate cards compare badly because the billable unit differs. Amazon and Brave bill a request. Microsoft bills a "transaction", counted by the number of tool calls per run. Google bills each individual search query the model decides to issue. Tavily bills credits, and one Research call can consume between 15 and 250 credits on model=pro.

The table below prices the same 200,000 monthly searches on each provider. AWS's own figure is used where AWS published one.

Provider Unit price 200,000 searches per month Notes on the arithmetic
Brave Search $5.00 per 1,000 $1,000.00 Straight multiplication; $5 monthly credit ignored
Tavily, Growth plan $0.005 per credit $1,000.00 Basic depth only; advanced depth doubles it to $2,000
Exa Search $7 per 1,000 $1,400.00 Assumes 10 results or fewer; each result above 10 adds $1 per 1,000
AgentCore Web Search $7 per 1,000 $1,403.00 AWS's published example: $1,400 search plus $3.00 of Gateway InvokeTool calls
Tavily, pay-as-you-go $0.008 per credit $1,600.00 Basic depth; no monthly commitment
Grounding with Bing Search $14 per 1,000 $2,800.00 Transactions counted per tool call, so retries bill again
Gemini 3, Grounding with Google Search $14 per 1,000 $2,730.00 First 5,000 prompts free, then 195,000 queries billed
Gemini 2.5, Grounding with Google Search $35 per 1,000 $5,425.00 1,500 requests per day free on the paid tier, roughly 45,000 a month

Two numbers in that table deserve a second look. The Gemini 2.5 line is 5.4x the Brave line for work that produces broadly comparable snippets, and it is the default anyone on gemini-2.5-flash inherits without changing a line of code. Moving the same agent to a Gemini 3 model drops grounding from $35 to $14 per 1,000, a 60% cut that has nothing to do with model quality. Teams running an LLM hybrid routing and API spend decision framework should treat the grounding rate as a routing input, not a footnote.

The second is AgentCore's $1,403.00. AWS reached it by adding 600,000 Gateway InvokeTool calls at $5 per million to the $1,400 of search, because Web Search on AgentCore is a connector target on AgentCore Gateway rather than a standalone endpoint. Gateway's own meter is cheap ($0.005 per 1,000 API invocations, $0.025 per 1,000 Search API calls, and $0.02 per 100 tools indexed per month), but it is a second meter, and data egress from Gateway to a customer-owned VPC carries a $0.006 per GB data processing charge in commercial AWS Regions.

The variable that actually decides enterprise deals: where the query goes

Amazon's launch post is unusually direct about the pitch. Web Search runs "with zero data egress from customer's secured AWS environment", and the post argues you can "meet enterprise governance policies without sending user prompts and retrieval queries to external search API providers outside of AWS".

Microsoft documents the opposite posture in plain language. The Microsoft Learn page for the Grounding with Bing tools, last updated 3 April 2026, states: "The Microsoft Data Protection Addendum doesn't apply to data sent to Grounding with Bing Search or Grounding with Bing Custom Search. When you use these services, your data flows outside the Azure compliance and Geo boundary. This also means use of these services waives all elevated Government Community Cloud security and compliance commitments, including data sovereignty and screened/citizenship-based support, as applicable."

The same page narrows what actually leaves: "the only information sent to Bing is the Bing search query, tool parameters, and your resource key. The service doesn't send any end user-specific information." So the exposure is the generated query string, not the user's prompt. For a healthcare or BFSI agent whose queries are themselves derived from patient or account context, that distinction is thinner than it looks, and it is the kind of detail that belongs in an enterprise AI agent governance layer rather than in a procurement spreadsheet.

Brave publishes "full-funnel Zero Data Retention" only on its Enterprise plan, and Exa lists Zero Data Retention under Enterprise too, alongside SLAs and MSAs. Neither is available on the self-serve rate the comparison table quotes. If your data protection impact assessment requires ZDR, the $5 and $7 rates are not the rates you will pay.

Provider Where the query goes Zero data retention Region control
AgentCore Web Search Stays inside the customer's AWS environment; Amazon's own index Not published as a separate term US East (N. Virginia) only at GA
Grounding with Bing Outside the Azure compliance and Geo boundary; DPA does not apply Not published Not stated on the tool page
Grounding with Google Search Google Search infrastructure Not published on the pricing page Not stated on the pricing page
Brave Search API Brave's independent index Enterprise plan only Not published
Exa Exa's index Enterprise plan only Custom indexes on Enterprise
Tavily Tavily infrastructure Enterprise plan, described as enterprise-grade security and privacy Not published

The citation display rule nobody budgets for

Grounding with Bing carries a contractual UI obligation. The Microsoft Learn page states that developers and end users "don't have access to raw content returned from Grounding with Bing Search", that the model response includes citations plus "a link to the Bing query used for the search", and that "these two references must be retained and displayed in the exact form provided by Microsoft, as per Grounding with Bing Search's Use and Display Requirements". The page's display section is more specific still: "you need to display both website URLs and Bing search query URLs in your custom interface."

That is front-end work with a deadline attached to a contract, not an optional nicety. If your agent surfaces answers inside a chat widget, a Slack app and a mobile screen, the requirement lands three times. Amazon's Web Search returns "snippets, source URLs, titles, and publication dates" without an equivalent published display mandate, and Brave and Exa return results you can render as you choose.

Billing units: where the estimate breaks

Every one of these providers bills something other than "a user question", and the gap between the two is where budgets fail.

Google is explicit in a footnote on its pricing page: "A customer-submitted request to Gemini may result in one or more queries to Google Search. You will be charged for each individual search query performed." The free allowance, meanwhile, is expressed in prompts per month, not queries. A research agent that decomposes one question into four searches burns four billable units against an allowance counted in ones. Google does offer a hedge: with dynamic retrieval, "only requests that contain at least one grounding support URL from the web in their response are charged".

Microsoft counts the same way. Transactions "are counted by the number of tool calls per run", and Microsoft's documentation notes the model "might decide to invoke the tool again for more information and context". A verbose planner is a billing decision.

Tavily's spread is the widest of the six. A basic search is 1 credit and an advanced search 2, but a single Research call ranges from 15 to 250 credits on model=pro and 4 to 110 on model=mini. At the pay-as-you-go rate of $0.008 per credit, one Research call costs between $0.12 and $2.00. Exa's Agent endpoint behaves similarly: fixed effort levels cost $0.012 to $1.00 per run, while the default effort: "auto" meters Agent Compute Units at $0.10 each and bills up to $5 per run.

Model the searches per task, not the tasks. An agent averaging 3.5 searches per user question at 100,000 questions a month is a 350,000-query workload, and on Gemini 2.5 that is $10,675 before a single token of inference. The same discipline that governs every enterprise AI agent production use case applies here: the unit you are billed on is rarely the unit you designed around.

Why Microsoft customers have no easy exit

Microsoft retired the Bing Search APIs on 11 August 2025. The lifecycle announcement is blunt: "Any existing instances of Bing Search APIs will be decommissioned completely, and the product will no longer be available to be used or new customer signup." Customers were directed to Grounding with Bing Search inside Azure AI Agents.

That migration is not an endpoint swap. Grounding with Bing exists only as a resource or tool inside Azure AI Foundry Agent Service, or as a Web Knowledge Source in Azure AI Search and Foundry Knowledge. Microsoft's pricing page says so twice. A team that used to make a plain REST call now needs an Azure subscription, the Contributor or Owner role to create Bing resources, and the Foundry Project Manager role to create project connections. The capacity ceiling is generous at 150 transactions per second and 1 million transactions per day, but the architectural commitment is the point.

Anyone weighing that against a portable alternative is really weighing platform lock-in against a $7 to $14 per 1,000 difference, which at 200,000 queries a month is $1,400. That gap is small enough that a short migration can consume a year of the saving, so the lock-in argument has to carry the decision on its own.

Cost per task: a worked model

Take a competitive intelligence agent that answers 50,000 questions a month, issues 3 searches per question, and needs 20 results per search rather than 10.

  • 150,000 searches a month.
  • On Exa, the base is 150,000 x $7 per 1,000 = $1,050, plus 10 extra results per search at $1 per 1,000 results: 1,500,000 extra results x $0.001 = $1,500. Total $2,550, because the results overage is larger than the base.
  • On Brave, 150,000 x $5 per 1,000 = $750, with result depth included in the request price.
  • On AgentCore, 150,000 x $7 per 1,000 = $1,050, plus Gateway invocations at $0.005 per 1,000.
  • On Grounding with Bing, 150,000 x $14 per 1,000 = $2,100 before any retries.

Exa's per-result meter is the trap. It is clearly documented, "every result above 10 adds $1 / 1k results", and it is easy to miss when a developer sets numResults: 25 to improve recall. Recall tuning is a pricing decision on Exa and is not on Brave. The general lesson is the one that shows up in every agentic RAG versus classic RAG cost analysis: the knob that improves quality is usually the knob that moves the bill.

India-specific considerations

None of the six providers publishes an India region for their grounding endpoint. AgentCore Web Search was generally available in US East (N. Virginia) only at launch, and the other five publish no regional endpoint list at all on the pages cited here. For an Indian enterprise, that means the search query, and whatever business context the model encoded into it, crosses a border on every call.

Under the Digital Personal Data Protection Act 2023, the practical question is whether the query string carries personal data. A query like "refund policy for order 4471 placed by A. Sharma" plainly does; "GST rate for imported industrial pumps" plainly does not. The engineering answer is a query sanitisation step between the planner and the search tool, so that the model's search string is rewritten to remove identifiers before it leaves your VPC. Building it is a day of work and it converts an unbounded exposure into a bounded one. Teams already running PII redaction before LLM calls can extend the same filter to the tool boundary.

The arithmetic matters for Indian budget holders too. At 200,000 searches a month, the gap between Brave at $1,000 and Gemini 2.5 at $5,425 is $4,425 a month, or $53,100 a year, on a line item most teams never priced before the prototype shipped. Convert that at your own booked rate before the annual plan is signed. It is a real trade, and it should be made deliberately rather than inherited from whichever SDK the prototype used.

What the launch customers actually said

Two named customers went on record in Amazon's launch post, and both talked about governance rather than result quality.

Nicholas Larus-Stone, Head of AI Agents at Benchling, said: "Scientists using Benchling AI can now ask about a target they're actively working on and get answers grounded in both their institutional data in Benchling and published literature. The result is more complete science, and hypothesis generation done right. Because we're using the Web Search tool on Amazon Bedrock AgentCore, customers have a secure, governed environment to bring that high quality published data into their workflows without compromising how they manage their data."

Iskander Sanchez-Rola, Senior Director of AI & Innovation at Gen Digital, was more specific about the reason: "What we value most is that AWS uses its own search index and keep queries within our trusted AWS environment."

Neither claimed a relevance advantage. In a market where six vendors return broadly similar snippets, the differentiator being marketed is custody of the query.

How to choose

A short decision path that matches the published facts:

  1. If you are already on AgentCore Gateway and your workload can live in US East (N. Virginia), Web Search at $7 per 1,000 removes a vendor and a data-egress argument. Verify current Regional availability before you commit, because the launch post pointed to a roadmap.
  1. If cost per query dominates and you can render citations freely, Brave at $5.00 per 1,000 with 50 requests per second is the cheapest published rate of the six.
  1. If you need multi-step research with structured output rather than raw snippets, Exa's Deep Search at $12 to $15 per 1,000 is priced against the work it replaces. Watch the per-result overage.
  1. If you are standardised on Azure AI Foundry, Grounding with Bing at $14 per 1,000 is effectively the only supported path since the standalone APIs were retired. Budget the display-requirement front-end work.
  1. If you are calling Gemini anyway, move to a Gemini 3 model before you enable grounding. The rate difference is $14 against $35 per 1,000.
  1. Whichever you pick, put it behind a gateway you control so switching is a config change. The same reasoning drives AI gateway comparisons across LiteLLM, Cloudflare, Kong and Bifrost.

The real cost is usually the migration, not the per-query rate. Design for a swap on day one and the rate card stops being a lock-in decision.

FAQ

How eCorpIT can help

eCorpIT builds and operates production AI agents for enterprises across India and global markets, including the grounding, gateway and governance layers that sit between a model and the open web. Our senior engineering teams model the per-query economics before a line of integration code is written, then build the query sanitisation and provider abstraction that let you change vendor without changing the agent. As a CMMI Level 5, ISO 27001:2022 certified and MSME certified organisation, we design these systems aligned with DPDP Act requirements for Indian data handling. Talk to us via /contact-us/ about pricing your agent's web grounding before it reaches production.

References

  1. Announcing Web Search on Amazon Bedrock AgentCore: Ground your AI agents in current, accurate web knowledge. AWS News Blog, 17 June 2026, pricing statement added 18 June 2026.
  1. Amazon Bedrock AgentCore Pricing. AWS, rates for Web Search, Gateway, Runtime and the published Web Search cost example.
  1. Grounding with Bing pricing. Microsoft, plan table showing $14 per 1,000 transactions and capacity limits.
  1. Use Grounding with Bing Search tools with the agents API. Microsoft Learn, updated 3 April 2026, on data boundaries, transaction counting and display requirements.
  1. Bing Search APIs retiring on August 11, 2025. Microsoft Lifecycle announcement.
  1. Gemini Developer API pricing. Google, grounding rates by model generation and the per-search-query billing footnote.
  1. Credits and pricing. Tavily documentation, plan table and per-endpoint credit costs.
  1. Pricing. Exa documentation, per-endpoint rates, result overage and Agent effort pricing.
  1. Brave Search API pricing. Brave, Search and Answers plans, capacity and Enterprise ZDR.
  1. Amazon Bedrock AgentCore Gateway documentation. AWS, on Web Search as an MCP connector target.
  1. Announcing Web Search on Amazon Bedrock AgentCore for Agentic Web Retrieval. AWS What's New, on GA Region and zero data egress.
  1. LLM Context: web search for agents and chatbots. Brave Search API documentation, launched 6 February 2026, on token-budget controls for agent grounding.

Last updated: 3 August 2026.

Frequently asked

Quick answers.

01 What does Amazon Bedrock AgentCore Web Search cost?
Web Search on Amazon Bedrock AgentCore is priced at $7.00 per 1,000 queries, with no upfront commitment, according to the AgentCore pricing page checked on 3 August 2026. New AWS customers receive up to $200 in Free Tier credits. It became generally available in June 2026 in the US East (N. Virginia) Region.
02 Which agent web search API is cheapest in August 2026?
Brave Search publishes the lowest standard rate of the six compared here, at $5.00 per 1,000 requests, including web results, LLM Context, news, videos and images, with 50 requests per second of capacity and $5 in free credits every month. Tavily's Growth plan reaches the same $0.005 per credit for basic-depth searches.
03 Why is Grounding with Google Search $14 in some places and $35 in others?
Google prices grounding differently by model generation. Gemini 3 models bill $14 per 1,000 search queries after 5,000 free prompts a month, while Gemini 2.5 and 2.0 models bill $35 per 1,000 grounded prompts after the free daily allowance. Moving an agent to a Gemini 3 model cuts the grounding rate by 60%.
04 Does Grounding with Bing keep my data inside Azure?
No. Microsoft's documentation states that data sent to Grounding with Bing flows outside the Azure compliance and Geo boundary, and that the Microsoft Data Protection Addendum does not apply to it. Only the search query, tool parameters and your resource key are sent, not end-user-specific information, but that query still crosses the boundary.
05 What happened to the standalone Bing Search APIs?
Microsoft retired the Bing Search APIs on 11 August 2025. The lifecycle announcement says existing instances were decommissioned completely and the product is no longer available for use or new signups. Microsoft directed customers to Grounding with Bing Search inside Azure AI Foundry Agent Service, which requires an Azure subscription and specific role assignments.
06 How do I estimate my monthly grounding bill?
Count searches, not user questions. Google bills each individual search query a model issues, and Microsoft counts transactions by tool calls per run, so one user turn can bill several times. Multiply expected questions by average searches per question, then apply the per-1,000 rate for your chosen provider.
07 Are these rates available with zero data retention?
Not at the self-serve prices quoted here. Brave lists full-funnel Zero Data Retention on its Enterprise plan only, and Exa lists Zero Data Retention under Enterprise alongside SLAs and MSAs. If a data protection impact assessment requires ZDR, budget for enterprise terms rather than the published per-1,000 rate.
08 Does the display of citations carry contractual obligations?
For Grounding with Bing, yes. Microsoft requires that the citation links and the Bing search query link be retained and displayed in the exact form provided, in every interface where results appear. That is front-end work in each surface your agent serves. The other providers cited here publish no equivalent display mandate.

About the author

Manu Shukla

Founder & Director

Founder of eCorpIT. Hands-on engineer leading senior-only delivery for AI apps, custom software, and cloud systems for global clients.

Subscribe

One engineering note a week. No fluff, no spam.

Senior-architect playbooks on AI agents, mobile apps, cloud, security, data, and marketing — delivered every Wednesday.

Past the reading

Read enough. Let's build something.

A senior architect responds in 24 working hours with scope, indicative cost, and a timeline. NDA before any technical conversation.