July 29, 2026

Public Benchmarks Won't Tell You If a Search API Is Right for Your Agent

Nimble leads the DRACO benchmark. Here's why that shouldn't be the number you evaluate us on.

clock
4
min read
Copied!

Charlie Klein

linkedin
Director of Product Marketing
No items found.
Public Benchmarks Won't Tell You If a Search API Is Right for Your Agent
July 29, 2026

Public Benchmarks Won't Tell You If a Search API Is Right for Your Agent

Nimble leads the DRACO benchmark. Here's why that shouldn't be the number you evaluate us on.

clock
4
min read
Copied!

Charlie Klein

linkedin
Director of Product Marketing
No items found.
Public Benchmarks Won't Tell You If a Search API Is Right for Your Agent

Every web search and research vendor points to a public benchmark to prove they're the best. We can play that game too, and right now we're winning it.

But we don't think that's how you should evaluate a search system for your agent.

We lead DRACO

DRACO is a research benchmark that scores 100 questions across domains on two axes: factual accuracy and the breadth and depth of the research behind the answer.

Nimble's Web Search Agents lead the field on both axes. At x-high effort, Nimble sits ahead of Parallel's ultra8x, Exa's xhigh, and OpenAI's deep research on accuracy while covering more ground per question. Even at high effort, Nimble matches the top competitor score on accuracy with better coverage, at a lower cost per query.

We're obviously happy with the result, but it's also a limited one.

Why generic benchmarks fall short

A public benchmark measures performance on a fixed set of questions written by someone who has never seen your product.

Your agent doesn't ask those questions. A lead enrichment agent runs the same twelve field lookups against thousands of companies. A social listening agent pulls sentiment from forums and review pages that a general benchmark never touches. A compliance agent needs primary filings, not the top blog post about them.

Scoring well on general research questions tells you a system is competent. It doesn't tell you how it will behave on the queries you actually send it in production, against the sources you actually care about.

That gap is where most agent teams get burned. The search API looked great in the eval and returned noise on day one.

Vertical benchmarks

So we built benchmarks that look more like production.

Each vertical benchmark is 100 questions written against a specific domain, with the sources, vocabulary, and output expectations that domain requires.

The results shift when the questions get specific. Nimble leads GTM at 88.8% against 86.1% and 84.4%. Company research at 96.1%. Market analysis at 84.0%. Social media monitoring at 83.5%, where the gap over the next system widens to a point and a half and over five points against the third.

The point isn't the individual numbers. It's that the ranking and the margins change depending on what you're asking. A system that wins on general research questions doesn't automatically win on review scraping or supplier monitoring.

Use the benchmark that matches your agent

If you're building lead enrichment, look at GTM. If you're building a brand monitoring tool, look at social media. If you're building due diligence workflows, look at company research.

That's a more honest signal than a single leaderboard number, because it's closer to the work your agent does every day.

Why Web Search Agents win on vertical tasks

Vertical benchmarks are also where the architecture matters most.

Generic search runs the same retrieval logic for every caller. A supply chain risk agent and a consumer sentiment agent hit the same index and get the same treatment.

Web Search Agents specialize. Every run stores its outputs, sources, and search history, and memory records which retrieval paths produced the right data. That knowledge steers adaptive crawling: the agent learns which sources hold your information and goes deeper on them, while skipping the ones that produce noise.

Two things follow from that:

  • Accuracy climbs. The agent stops guessing where your data lives and starts knowing.
  • Cost drops. Fewer wasted searches, fewer tokens spent parsing pages that were never going to help.

There's nothing to configure. You start sending research requests to the API and Nimble learns the domain from your traffic.

Run your own benchmark

The best benchmark is the one you write yourself, using the questions your agent asks in production.

Take a hundred of your real queries, run them against Nimble and against whatever you're using today, and score the outputs against what your product actually needs.

Learn more about Web Search Agents

Start a free trial

FAQ

Answers to frequently asked questions

No items found.