Ontology 1: Benchmarks
By Aditri Bhagirath, Alex Gunnarson
We've built Ontology 1, a neurosymbolic model that answers complex, conversational, multimodal product queries more accurately than Amazon or Google Shopping, despite having indexed only 1%1 of their catalog.
Best in-class precision with smaller datasets
In the benchmark below, Ontology 1 wins 52 of 90 searches outright. (Google wins 19, and Amazon 16.) Ontology achieves 63.0% accuracy in its top 10 results, compared to Google's 54.3% and Amazon's 46.9%. We also compare multimodal queries. Amazon lacks multimodal search support, and Google, while it returns results, routinely misses the query (Fig. 8). As far as we're aware, Onton is the only platform that handles complex, intent-heavy multimodal queries like those in Fig. 8.
Methodology
We introduce Subtext-Decor-90, in which three independent multimodal judges — Claude Opus 4.8, Gemini 3.1 Pro, and GPT-5.52 — evaluate 90 text queries against Onton, Amazon, and Google Shopping3, scoring the top 10 results from each engine. (Complete code and data may be found here.)
Subtext-Decor-90 was curated to include intent-heavy queries with rich aesthetic descriptors, negation, cultural references, and emotional framing. These are cases that resist traditional keyword and vector retrieval. They also reflect where product search is heading. In 2025, Bloomreach found that 57% of shoppers have used AI to help them shop, and 41% now search using natural language rather than keywords. A 2026 study by Klaviyo similarly found that consumers are abandoning keywords in favor of full phrases and questions, with daily AI users routinely searching with queries of eight or more words — sometimes entire paragraphs.
For each query in Subtext-Decor-90, the judges saw all three result screenshots in a single call. Each judge scored the first 10 visible ranked result cards in each screenshot (left-to-right, top-to-bottom). P@104 was aggregated as the mean across judges, with 10,000-resample bootstrap CIs5 over queries. We report precision rather than recall6, which is unanswerable here without full access to Amazon and Google’s catalogs.7
Image and multimodal searches are excluded from Subtext-Decor-90 because a 1:1 comparison with Amazon and Google isn't possible. Amazon Lens is built around finding live products with a mobile camera, and doesn’t support multimodal queries at all. Google Lens supports multimodal, but unlike Google Shopping, doesn’t exclusively return products. We provide a separate Onton-vs.-Google comparison on 10 image and multimodal queries below (Fig. 8).
| Query ID | Search query | Onton | Amazon | |
|---|---|---|---|---|
| 8 | Lighting that makes my apartment feel like a Tokyo cocktail bar at 11pm | 0.93 | 0.57 | 0.37 |
| 14 | Gross looking art | 1.00 | 0.17 | 0.50 |
| 51 | laundry hamper I won't hate looking at for 10 years | 0.97 | 0.30 | 0.43 |
| 19 | A chair that's comfortable for crying in | 0.73 | 0.60 | 0.70 |
| 6 | A sofa my husband won't call feminine and I won't call a man cave | 0.60 | 0.00 | 0.03 |
| 18 | Bedroom furniture similar to Call Me By Your Name | 0.53 | 0.43 | 0.33 |
| 28 | Rug that hides cat puke but isn't beige | 0.93 | 0.47 | 0.10 |
Google returned beige rugs for “isn’t beige”:
Rug that hides cat puke but isn't beige
and rather utilitarian and standard-looking laundry hampers for an intent-heavy query:
laundry hamper I won't hate looking at for 10 years
Amazon returned zero results for this sofa query:
A sofa my husband won't call feminine and I won't call a man cave
and surfaced LED ice cubes for this lighting query:
Lighting that makes my apartment feel like a Tokyo cocktail bar at 11pm
Analysis
The judges grade on different scales: Gemini 3.1 Pro is harshest (mean P@10 ~0.43 across engines), Opus 4.8 the most generous (~0.70), and GPT-5.5 sits in the middle (~0.51).
| Judge | Onton | Amazon | |
|---|---|---|---|
| Claude Opus 4.8 | 0.759 | 0.641 | 0.696 |
| Gemini 3.1 Pro | 0.528 | 0.342 | 0.426 |
| GPT 5.5 | 0.603 | 0.422 | 0.507 |
While Krippendorff's alpha9 across the three judges is 0.465 (interval, on P@10), all three judges put the engines in the same order: Onton, then Google, then Amazon. This is to say the exact P@10 values are noisy and judge-dependent, but the directional finding is robust across judges: Onton outperforms the next-best engine by a clear margin under every judge.
Pearson correlation tells the complementary half of the story. Unlike Krippendorff's alpha, Pearson is invariant to how harsh or lenient a judge is, so it captures agreement on which results are good rather than on absolute scores. Pairwise P@10:
- Gemini 3.1 Pro vs. GPT-5.5: r = 0.730
- Opus 4.8 vs. GPT-5.5: r = 0.569
- Opus 4.8 vs. Gemini 3.1 Pro: r = 0.498
The two harsher graders agree with each other most. Claude's scores bunch up at the high end because it grades easy, which drags down its correlation with the other two, but it still produces the same ordering.
Failure Cases
When we take a look at judge evaluations for Amazon/Google, the recurring qualitative note is “misinterprets the query”. This is more of a semantic ceiling than a ranking gap. On the other hand, Ontology’s failure cases are concentrated in functional spec-related queries where Amazon’s category metadata dominates.
A few examples of Ontology 1 failure cases:
- “lamp that won't wake my partner if I read at 3am” (Onton 0.4, Amazon 0.9)
- “something to put on a weirdly deep windowsill” (Onton 0.07, Amazon 0.67)
We address these through two primary strategies:
- Expanding our product catalog: Even accounting for our current single-vertical focus of home decor and furniture, our catalog is is significantly smaller than Amazon’s and Google’s due to our focus on higher-signal, non-sponsored products. However, we’re actively working on catalog breadth.
- Self-learning: Ontology 1 improves itself10 through a scientific-method-shaped learning loop. The meaning of "won't wake my partner" is not just captured by a string or a vector, but it becomes increasingly precise and interconnected with other concepts in the graph. (For instance, to not wake someone up, a lamp might need to be low-light or red-light). The learnings then get leveraged on subsequent user searches.
Performance by search engine
Amazon does well on functional-spec queries where there's a clean keyword-to-category match (dimensions, materials, fixed attributes), and falls apart on aesthetic modifiers and stacked negation.
Google Shopping does well when it can pull up a curated section for a structured constraint, and misses on subtle intent and explicit negation.
Onton does well on the queries where you have to model what the person actually means instead of matching keywords.
Image and multimodal search
Below we compare 10 image and multimodal queries on Onton and Google. Google surfaces visually relevant results that often miss the query. (See row 6 where instead of surfacing mirrors, the primary results surface articles related to farmhouse decor).
| Query Image | Query Text | Onton | |
|---|---|---|---|
| chairs fitting this vibe | |||
| chairs fitting this vibe | |||
| bed like this but black | |||
| crib like this but pink | |||
| mirrors fitting this vibe | |||
| desk in this vibe | |||
| i'd like chair like this but budget version, and maybe red | |||
| lamps in this aesthetic, but tall |
Looking forward
Next, we want to harden the benchmark and broaden its scope.
Hardening
Richer ranking metrics. P@10 weights all ten slots equally, so an engine earns no credit for ranking its best result first. Adding NDCG@1011 on graded (rather than binary) relevance judgments would reward both correct ordering and degree of quality.
Judge normalization. Because the three judges grade on different severity scales, normalizing each judge's scores before aggregating would tighten our confidence intervals.
Live, longitudinal evaluation. A one-shot screenshot benchmark is a snapshot, and search engines change. Because Subtext-Decor-90 is automated, we can re-run it on a regular cadence — tracking both how the baselines shift and how far Ontology's own learning loop moves its numbers over time.
Broadening
Larger, broader query sets. Subtext-Decor-90 deliberately targets the difficult queries Onton was built for. Scaling up the number of queries (e.g. to a hypothetical Subtext-Decor-10K or beyond) would increase benchmark confidence. Adding in “easier” queries would test how Ontology 1 holds up across the full distribution of real shopping behavior — and let us report complementary metrics like result diversity, seller reputation, and freshness that matter more once basic relevance is met.
More engines. We benchmarked against Amazon and Google because that's where the large majority of product searches start, but extending Subtext-Decor-90 to platforms like Pinterest, Wayfair, and Etsy would give a broader view of the search landscape.
Beyond decor. Subtext-Decor-90 is scoped to home decor and furniture because that's the vertical Onton indexes today. But the methodology isn't specific to decor. Ontology 1 is a general model, and the benchmark can likewise generalize — to Subtext-Shopping-N across verticals, and eventually Subtext-N beyond shopping.
Coverage estimates. Pooled relative recall (see footnote 7) is an imperfect proxy for recall, but it would give a first, bounded read on coverage alongside our precision numbers.
Catalog and vertical expansion. Our single-vertical, higher-signal catalog is a primary source of Ontology's remaining failure cases. Growing catalog breadth — and moving beyond home decor — is the clearest path to closing the functional-spec gap with Amazon and Google.
Accelerating the learning loop. Ontology’s failure cases shrink as it learns. Tightening that loop is the lever we expect to move our numbers most between benchmark runs.
Outreach
If you're a researcher excited by neurosymbolic systems, your organization has a search or discovery problem that conventional tools handle poorly, or you've found a use for Ontology we haven't thought of, we’d love to talk ([email protected]).
Footnotes
1 Amazon has ~600M unique products (ASINs) as of 2025 (~100x Onton), and ~2.6B listings worldwide as of 3/2026. Google Shopping has 50B+ listings as of 1/2026, with no published number on uniques; if Amazon’s ratio holds, one could extrapolate to 11.5B (~1000x Onton).
2 At the time of this benchmark, Claude Opus 4.8, Gemini 3.1 Pro, and GPT-5.5 were arguably the most capable publicly available frontier models. Drawing them from three different labs reduces the risk of correlated judgments.
3 A large majority of product searches start on Amazon or Google, making them the natural baseline. We plan to benchmark against Wayfair, Pinterest, and others over time.
4 P@10, or Precision at 10, measures the proportion of relevant items in the top 10 results. For each query, we average the three judges' P@10 into a single per-query score. An engine's reported P@10 is the mean of those scores across all 90 queries. To get the 95% confidence interval, we use the bootstrap: we draw 90 queries at random (with replacement) from our set, recompute the engine’s mean, and repeat this 10,000 times. The interval is the middle 95% of those means. Per-query P@10 is discrete and skewed, not normally distributed, so bootstrapping — which assumes nothing about the distribution's shape — is the natural choice here.
6 Precision asks what fraction of results were relevant. Recall asks what fraction of all relevant products the engine managed to return.
7 The usual proxy, pooled relative recall, asks what fraction an engine captured from the relevant results that all engines returned. It may or may not be a reasonable estimate of coverage, but coverage is outside the scope of this benchmark.
8 See Footnote 1.
9 Krippendorff's alpha is a measure of inter-rater reliability.
10 See Ontology 1 release announcement.
11 NDCG (Normalized Discounted Cumulative Gain, introduced in Järvelin & Kekäläinen 2002 rewards placing relevant results near the top: each result's relevance is discounted by how far down the list it appears, then normalized against the best possible ordering so scores fall in [0, 1].


