Linkup

Benchmarks

State-of-the-art across public & verticalized benchmarks.

Public benchmark

SimpleQA

Factual accuracy

Verified SimpleQA measures factual accuracy for search-augmented systems on web-grounded, fact-seeking questions. It is the baseline for evaluating how precisely a search API retrieves the right answer.

SimpleQA SimpleQA chart

Details

Benchmark details

Linkup leads on public research suites and on vertical evaluations built around real GTM, legal, and finance workflows, the same accuracy standard measured across both.

Methodology and source repositories are linked below.

Public benchmark

SimpleQA

Factual accuracy

About

Verified SimpleQA measures factual accuracy for search-augmented systems on web-grounded, fact-seeking questions. It is the baseline for evaluating how precisely a search API retrieves the right answer.

Evaluation

Evaluated on OpenAI's SimpleQA suite with an F1-score across sub-second web search APIs.

Result

Linkup Fast leads at 92%, ahead of Perplexity Sonar, Gemini 3 Flash, Serper, Tavily, and Exa.

Results

ProviderAccuracy
  • Linkup Fast92%
  • Perplexity Sonar87%
  • Gemini 3 Flash85%
  • Serper Fast83%
  • Tavily Fast79%
  • Exa Fast63%
View methodology

Public benchmark

SealQA-0

Hard factual retrieval

About

SealQA-0 is a challenge benchmark designed for frontier models to fail on fact-seeking questions where standard web search breaks down. It stresses deep research systems on hard, ambiguous retrieval.

Evaluation

Accuracy is measured on SealQA-0 across research-oriented products.

Result

Linkup Research leads at 61%, ahead of Parallel Ultra 8X, GPT-5, Exa Research Pro, Perplexity DR, and Tavily Research.

Results

ProviderAccuracy
  • Linkup Research61%
  • Parallel Ultra 8X56%
  • GPT-548%
  • Exa Research Pro45%
  • Perplexity DR38%
  • Tavily Research24%
View methodology

Use case

GTM

People & company intelligence

About

Public GTM benchmarks of how web-search APIs handle production go-to-market work: people enrichment, pre-meeting briefs, and job-change freshness, plus company answer quality and funding accuracy.

Evaluation

Linkup is compared to Exa, Perplexity, and Parallel on independent ground truth (profile DB, Crunchbase).

Result

Linkup leads every GTM signal: Freshness 74%, Enrichment 94%, Richness 64.8, Answer quality 80.7%, and Funding 83%.

Results

People · Freshness1 / 5

ProviderCaught
  • Linkup74%
  • Exa14%
  • Parallel11%
  • Perplexity9%
Access benchmark repository

Use case

Legal

Contract redlining with search

About

RedlineBench × Linkup measures whether live web search makes a frontier model a better contract negotiator. A Linkup-backed search skill is layered on Crosby's RedlineBench so the agent can look up market norms while it redlines.

Evaluation

Same model (GPT-5.5), same prompts, and Crosby's published 3-judge grading across 140 tasks; the only change is that the agent can search. Score is turn-weighted against the published closed-book baseline.

Result

GPT-5.5 + Linkup search scores 58.6%, +8.1 points over the published closed-book baseline of 50.5%, ahead of other frontier models on the board.

Results

ProviderScore
  • GPT-5.5 + Linkup58.6%
  • GPT-5.550.5%
  • Claude Fable 547.3%
  • Gemini 3.5 Flash45.1%
  • Claude Opus 4.844.4%
Access benchmark repository

Use case

Finance

FinSearchComp retrieval

About

FinSearchComp is a financial-search benchmark of 635 questions written by 70+ professional analysts, each with a single gold answer verified against primary sources (SEC filings, exchange data) and graded strictly right/wrong.

Evaluation

We run the two historical Global tiers identically across every system, taking each provider's best research/deep config.

Result

Linkup research is 1st on both tiers (T2 82.4% and T3 58.3%), ahead of Perplexity, Parallel, Tavily, and Exa.

Results

T21 / 2

ProviderAccuracy
  • Linkup research82.4%
  • Parallel ultra73.1%
  • Perplexity72.3%
  • Exa research_pro42.0%
  • Tavily research_pro40.3%
Access benchmark repository