Press Esc to close

OpenAI Web Search debuts fifth on Artificial Analysis's search benchmark

OpenAI Web Search debuts fifth on Artificial Analysis's search benchmark
OpenAI wordmark in black at the center of a white banner surrounded by colorful doodles, shapes and ASCII art

Image: OpenAI

OpenAI's built-in web search tool now has a spot on an independent leaderboard, and it's the first of its kind there. Artificial Analysis added OpenAI Web Search to its Search Index, which measures how much a search tool helps an AI model answer hard questions. It scored 74, the fifth-best provider on the board behind Perplexity, Octen, Parallel and Brave.

The twist is how it was tested. Every other entry is a standalone Search API plugged into Artificial Analysis's open-source Stirrup agent harness, where the model searches, opens pages and decides when it's done, within 25 turns. OpenAI's entry skips that. The model makes one call to OpenAI's Responses API with the built-in web_search tool, and OpenAI runs the whole search loop on its side, according to the methodology.

Every setup uses the same model, GPT-5.6 Luna at medium reasoning. The index is a straight average of three tests: DeepSearchQA (broad research questions with list answers), BrowseComp (hard facts that take several hops to find) and AA-Omniscience, a private set of 600 factual questions.

With no search at all, the model scores 33, so OpenAI's tool more than doubles it. Perplexity Search (medium) leads at 80, followed by Perplexity (high) at 79, Octen at 77, and Parallel (advanced) and Brave (LLM context) at 75. Counting every product variant, OpenAI ties for seventh of 26 with You.com, Nimble and Exa.

Where it lands depends on the job. On AA-Omniscience's fact questions, its 72% accuracy ranks third of 26, one point behind leader Firecrawl. On BrowseComp's multi-hop puzzles it manages about 74%, well short of Perplexity and Octen at 86% to 87%.

Price is the stronger story. Artificial Analysis puts OpenAI Web Search at about $0.05 per task: roughly $0.04 in search fees at OpenAI's $10 per 1,000 calls, plus under a cent for the model. That undercuts Parallel (advanced) at about $0.06 and Perplexity (medium) at about $0.07, and it's cheaper than 17 of the 25 Search API products. Octen is the outlier, at about $0.024 per task while scoring three points higher.

It's also lean: OpenAI billed about 40,000 input tokens per task, search results included, versus 125,000 for the leanest Search API. Speed is middling at about 30 seconds per task, and that may run a little high because OpenAI doesn't report search time, so the measurement includes OpenAI's own processing. Octen is fastest at about 16 seconds.

For anyone building an AI agent and picking a search provider, built-in search is now a credible default. If your agent already runs on OpenAI's models, you get near top-tier results with one tool setting, one bill and no second vendor, plus domain filters and a full list of sources consulted.

But you're buying a package, not just a search engine. OpenAI decides when and how often to search, so you see less and can't tune the loop like you can in your own harness. The tool also lives inside OpenAI's API, while standalone Search APIs work with any model. And OpenAI's docs require apps to show its inline citations as visible, clickable links.

The token savings matter more than they look. Search results count as model input either way, and a cheap model like Luna keeps that gap small. With a pricier model, 40,000 versus 125,000-plus tokens per task becomes real money.

For deep, multi-step research, Perplexity and Octen still have a clear edge, and Octen pairs a higher score with half the price and the fastest times. Just remember this is one model, one setting and one test mix, so run some of your own real queries before you commit.

Comments