Parallel Quality Benchmarks
Give your AI the highest-quality web search tools available
When building applications that rely on web data to make decisions or answer questions, nothing matters more than accuracy. These benchmarks help to measure different web search offerings on their ability to answer prompts accurately. By obsessing over accuracy, we consistently lead the market with state-of-the-art quality. In addition to leading in accuracy, Parallel often leads in pricing.
Fast Web Search
Agentic Web Research
Accuracy (%)
P50 SEARCH LATENCY (ms)
ACCURACY (%)
Latency: p50 client-side wall clock per search API request, in ms, shown on a log scale (best across runs). OpenAI Web Search is omitted (single search-call latency not available).
**Dataset**
BrowseComp[BrowseComp], created by OpenAI, contains 1,266 questions that require persistent browsing to locate hard-to-find, entangled information on the web.
**Evaluation methodology**
Multi-hop evaluation: a GPT-5.4 agent runs with up to 20 tool calls (search_web, plus web_fetch for engines with an extract API: Parallel, Exa, and Tavily; Brave and SerpAPI are search-only). Answers are graded by an LLM judge (GPT-5.4, per-suite grader prompts). Benchmarks were run across multiple sessions, with the best observed scores selected for each provider.
Latency: search-call latency is the client-side wall clock measured around a single provider search API request, from a client in us-central; we report the p50 across all questions (best across runs). OpenAI Web Search scored 57.7% accuracy on this suite but is not plotted because its single search-call latency is not available.
**Testing dates**
Evals were run between July 10 and 12, 2026.
## Parallel Quality Benchmarks
Give your AI the highest-quality web search tools available
When building applications that rely on web data to make decisions or answer questions, nothing matters more than accuracy. These benchmarks help to measure different web search offerings on their ability to answer prompts accurately. By obsessing over accuracy, we consistently lead the market with state-of-the-art quality. In addition to leading in accuracy, Parallel often leads in pricing.
### Fast Web Search
#### BrowseComp
| Series | Model | p50 Search Latency (ms) | Accuracy (%) | | -------- | ----------------- | ----------------------- | ------------ | | Parallel | Parallel Turbo | 216 | 51 | | Others | Exa Instant | 361 | 33.7 | | Others | Tavily Ultra Fast | 357 | 19.3 | | Others | Brave Search | 430 | 38.3 | | Others | SerpAPI | 999 | 23.3 |
Latency: p50 client-side wall clock per search API request, in ms, shown on a log scale (best across runs). OpenAI Web Search is omitted (single search-call latency not available).
**Dataset**
BrowseComp[BrowseComp], created by OpenAI, contains 1,266 questions that require persistent browsing to locate hard-to-find, entangled information on the web.
**Evaluation methodology**
Multi-hop evaluation: a GPT-5.4 agent runs with up to 20 tool calls (search_web, plus web_fetch for engines with an extract API: Parallel, Exa, and Tavily; Brave and SerpAPI are search-only). Answers are graded by an LLM judge (GPT-5.4, per-suite grader prompts). Benchmarks were run across multiple sessions, with the best observed scores selected for each provider.
Latency: search-call latency is the client-side wall clock measured around a single provider search API request, from a client in us-central; we report the p50 across all questions (best across runs). OpenAI Web Search scored 57.7% accuracy on this suite but is not plotted because its single search-call latency is not available.
**Testing dates**
Evals were run between July 10 and 12, 2026.
#### HLE
| Series | Model | p50 Search Latency (ms) | Accuracy (%) | | -------- | ----------------- | ----------------------- | ------------ | | Parallel | Parallel Turbo | 220 | 52.7 | | Others | Exa Instant | 358 | 49.3 | | Others | Tavily Ultra Fast | 243 | 42 | | Others | Brave Search | 563 | 47.7 | | Others | SerpAPI | 865 | 40 |
Latency: p50 client-side wall clock per search API request, in ms, shown on a log scale (best across runs). OpenAI Web Search is omitted (single search-call latency not available).
**Dataset**
Humanity's Last Exam (HLE)[Humanity's Last Exam (HLE)], created by CAIS and Scale AI, is a benchmark of expert-written questions at the frontier of human knowledge across dozens of subjects.
**Evaluation methodology**
Multi-hop evaluation: a GPT-5.4 agent runs with up to 20 tool calls (search_web, plus web_fetch for engines with an extract API: Parallel, Exa, and Tavily; Brave and SerpAPI are search-only). Answers are graded by an LLM judge (GPT-5.4, per-suite grader prompts).
Latency: search-call latency is the client-side wall clock measured around a single provider search API request, from a client in us-central; we report the p50 across all questions (best across runs). OpenAI Web Search scored 66% accuracy on this suite but is not plotted because its single search-call latency is not available.
**Testing dates**
Evals were run between July 10 and 12, 2026.
#### WebWalker
| Series | Model | p50 Search Latency (ms) | Accuracy (%) | | -------- | ----------------- | ----------------------- | ------------ | | Parallel | Parallel Turbo | 217 | 75.7 | | Others | Exa Instant | 336 | 65 | | Others | Tavily Ultra Fast | 240 | 63.7 | | Others | Brave Search | 503 | 65.7 | | Others | SerpAPI | 761 | 50.7 |
Latency: p50 client-side wall clock per search API request, in ms, shown on a log scale (best across runs). OpenAI Web Search is omitted (single search-call latency not available).
**Dataset**
WebWalkerQA[WebWalkerQA] evaluates an agent's ability to traverse the web — navigating through linked pages to find information that a single search does not surface.
**Evaluation methodology**
Multi-hop evaluation: a GPT-5.4 agent runs with up to 20 tool calls (search_web, plus web_fetch for engines with an extract API: Parallel, Exa, and Tavily; Brave and SerpAPI are search-only). Answers are graded by an LLM judge (GPT-5.4, per-suite grader prompts).
Latency: search-call latency is the client-side wall clock measured around a single provider search API request, from a client in us-central; we report the p50 across all questions (best across runs). OpenAI Web Search scored 80.7% accuracy on this suite but is not plotted because its single search-call latency is not available.
**Testing dates**
Evals were run between July 10 and 12, 2026.
#### Coding
| Series | Model | p50 Search Latency (ms) | Accuracy (%) | | -------- | ----------------- | ----------------------- | ------------ | | Parallel | Parallel Turbo | 216 | 79.7 | | Others | Exa Instant | 341 | 76.7 | | Others | Tavily Ultra Fast | 208 | 71.9 | | Others | Brave Search | 514 | 64.3 | | Others | SerpAPI | 683 | 54 |
Latency: p50 client-side wall clock per search API request, in ms, shown on a log scale (best across runs). OpenAI Web Search is omitted (single search-call latency not available).
**Dataset**
A proprietary coding dataset derived from production queries to Parallel's search API.
**Evaluation methodology**
Multi-hop evaluation: a GPT-5.4 agent runs with up to 20 tool calls (search_web, plus web_fetch for engines with an extract API: Parallel, Exa, and Tavily; Brave and SerpAPI are search-only). Answers are graded by an LLM judge (GPT-5.4, per-suite grader prompts).
Latency: search-call latency is the client-side wall clock measured around a single provider search API request, from a client in us-central; we report the p50 across all questions (best across runs). OpenAI Web Search scored 76.7% accuracy on this suite but is not plotted because its single search-call latency is not available.
**Testing dates**
Evals were run between July 10 and 12, 2026.
#### SimpleQA
| Series | Model | p50 Search Latency (ms) | Accuracy (%) | | -------- | ----------------------- | ----------------------- | ------------ | | Parallel | Parallel Search (Turbo) | 240 | 91 | | Others | Exa Instant | 335 | 89.3 | | Others | Tavily Ultra Fast | 150 | 72 | | Others | Brave Search | 475 | 87 | | Others | SerpAPI | 652 | 76.7 |
Latency: p50 client-side wall clock per search API request, in ms, shown on a log scale (best across runs).
**Dataset**
SimpleQA[SimpleQA], created by OpenAI, contains 4,326 short, fact-seeking questions across a variety of domains.
**Evaluation methodology**
Single-step evaluation: the raw question is sent as the search query (num_results=10, with an equal ~1,000 character-per-result content budget for every engine) and GPT-5.4 (reasoning: high) synthesizes an answer from the search results only. Answers are graded by an LLM judge (GPT-5.4, per-suite grader prompts).
Latency: search-call latency is the client-side wall clock measured around a single provider search API request, from a client in us-central; we report the p50 across all questions (best across runs).
**Testing dates**
Evals were run between July 10 and 12, 2026.
### Agentic Web Research
#### DeepSearchQA
| Series | Model | Cost (CPM) | Accuracy (%) | | -------- | ---------------------------------- | ---------- | ------------ | | Parallel | Parallel Task Ultra | 300 | 70 | | Parallel | Parallel Task Ultra2x | 600 | 77 | | Parallel | Parallel Task Ultra4x | 1200 | 81 | | Parallel | Parallel Task Ultra8x | 2400 | 82 | | Others | GPT 5.4 with code execution | 701 | 63 | | Others | Gemini 3.1 Pro with code execution | 707 | 62 | | Others | Opus 4-6 with PTC | 36231 | 58 | | Others | Perplexity Sonar Deep Research | 883 | 28 |
CPM: USD per 1000 requests. Cost is shown on a Log scale.
### Methodology
**Evaluation criteria**
Accuracy refers to answers that are "fully correct": a response is fully correct if and only if the submitted set is semantically identical to the ground-truth set. The agent must identify all correct answers while including zero incorrect answers.
**Evaluation sample**
We ran all benchmarks on a random 100-question[ random 100-question] subset of the original dataset. This subset was held constant across experiments with our own agents and with competitors.
**Experiment Setup: **We evaluate all systems using their highest-quality configurations with no budget constraints. For Gemini 3.1 Pro, GPT-5.4, and Opus 4.6, we use their respective agent harnesses along with web browsing and code execution tools. For Exa, we initially attempted to use exa-deep-max, but encountered persistent 524 API errors. As a result, we use exa-deep-reasoning for benchmarking. For Perplexity, we benchmarked them using their Sonar Pro API.
**Benchmark dates**
All testing was conducted between April 1 and April 6, 2026.