Onton

Onton is a next-generation search and discovery engine. We’re a tight-knit San Francisco-based team building the future of shopping.

Onton

Ontology 1: Benchmarks

Research

We've built Ontology 1, a neurosymbolic model that answers complex, conversational, multimodal product queries more accurately than Amazon or Google Shopping, despite having indexed only 1%1 of their catalog.

Best in-class precision with smaller datasets

In the benchmark below, Ontology 1 wins 52 of 90 searches outright. (Google wins 19, and Amazon 16.) Ontology achieves 63.0% accuracy in its top 10 results, compared to Google's 54.3% and Amazon's 46.9%. We also compare multimodal queries. Amazon lacks multimodal search support, and Google, while it returns results, routinely misses the query (Fig. 8). As far as we're aware, Onton is the only platform that handles complex, intent-heavy multimodal queries like those in Fig. 8.

Methodology

We introduce Subtext-Decor-90, in which three independent multimodal judges — Claude Opus 4.8, Gemini 3.1 Pro, and GPT-5.52 — evaluate 90 text queries against Onton, Amazon, and Google Shopping3, scoring the top 10 results from each engine. (Complete code and data may be found here.)

Subtext-Decor-90 was curated to include intent-heavy queries with rich aesthetic descriptors, negation, cultural references, and emotional framing. These are cases that resist traditional keyword and vector retrieval. They also reflect where product search is heading. In 2025, Bloomreach found that 57% of shoppers have used AI to help them shop, and 41% now search using natural language rather than keywords. A 2026 study by Klaviyo similarly found that consumers are abandoning keywords in favor of full phrases and questions, with daily AI users routinely searching with queries of eight or more words — sometimes entire paragraphs.

For each query in Subtext-Decor-90, the judges saw all three result screenshots in a single call. Each judge scored the first 10 visible ranked result cards in each screenshot (left-to-right, top-to-bottom). P@104 was aggregated as the mean across judges, with 10,000-resample bootstrap CIs5 over queries. We report precision rather than recall6, which is unanswerable here without full access to Amazon and Google’s catalogs.7

Image and multimodal searches are excluded from Subtext-Decor-90 because a 1:1 comparison with Amazon and Google isn't possible. Amazon Lens is built around finding live products with a mobile camera, and doesn’t support multimodal queries at all. Google Lens supports multimodal, but unlike Google Shopping, doesn’t exclusively return products. We provide a separate Onton-vs.-Google comparison on 10 image and multimodal queries below (Fig. 8).

Average P@1095% confidence interval
Onton
Google Shopping
Amazon
0.630
[0.571, 0.688]
0.543
[0.490, 0.596]
0.469
[0.417, 0.521]
0.40.50.60.7
Query wins
Outright winsTotal wins
Onton
Google Shopping
Amazon
52
19
16
582519
0204060
Fig. 1: P@10 Note that in three cases, Ontology returned fewer than 10 results due to insufficient catalog size. (This is why outright wins sum to 87 instead of 90.) Amazon and Google have product catalogs that are currently 100-1000 times larger than Onton’s.8 The Avg. P@10 below counts the missing slots in those three queries as non-relevant. If we instead excluded the empty slots from scoring, the values would be Onton 0.665, Google 0.549, and Amazon 0.459.
Query IDSearch queryOntonAmazonGoogle
8Lighting that makes my apartment feel like a Tokyo cocktail bar at 11pm0.930.570.37
14Gross looking art1.000.170.50
51laundry hamper I won't hate looking at for 10 years0.970.300.43
19A chair that's comfortable for crying in0.730.600.70
6A sofa my husband won't call feminine and I won't call a man cave0.600.000.03
18Bedroom furniture similar to Call Me By Your Name0.530.430.33
28Rug that hides cat puke but isn't beige0.930.470.10
Fig. 2: Ranking Sample A sample of representative queries where Onton outperforms Google and Amazon.

Google returned beige rugs for “isn’t beige”:

Rug that hides cat puke but isn't beige

Google
Onton
Fig. 3: Google vs. Onton results for Query 28 (“rug that hides cat puke but isn’t beige”).

and rather utilitarian and standard-looking laundry hampers for an intent-heavy query:

laundry hamper I won't hate looking at for 10 years

Google
Onton
Fig. 4: Google vs. Onton results for Query 51 (“laundry hamper I won't hate looking at for 10 years”).

Amazon returned zero results for this sofa query:

A sofa my husband won't call feminine and I won't call a man cave

Amazon
Onton
Fig. 5: Amazon vs. Onton results for Query 6 (“a sofa my husband won’t call feminine and I won’t call a man cave”)

and surfaced LED ice cubes for this lighting query:

Lighting that makes my apartment feel like a Tokyo cocktail bar at 11pm

Amazon
Onton
Fig. 6: Amazon vs. Onton results for Query 8 (“lighting that makes my apartment feel like a Tokyo cocktail bar at 11pm”)

Analysis

The judges grade on different scales: Gemini 3.1 Pro is harshest (mean P@10 ~0.43 across engines), Opus 4.8 the most generous (~0.70), and GPT-5.5 sits in the middle (~0.51).

JudgeOntonAmazonGoogle
Claude Opus 4.80.7590.6410.696
Gemini 3.1 Pro0.5280.3420.426
GPT 5.50.6030.4220.507
Fig. 7: Per-judge avg. P@10 by source

While Krippendorff's alpha9 across the three judges is 0.465 (interval, on P@10), all three judges put the engines in the same order: Onton, then Google, then Amazon. This is to say the exact P@10 values are noisy and judge-dependent, but the directional finding is robust across judges: Onton outperforms the next-best engine by a clear margin under every judge.

Pearson correlation tells the complementary half of the story. Unlike Krippendorff's alpha, Pearson is invariant to how harsh or lenient a judge is, so it captures agreement on which results are good rather than on absolute scores. Pairwise P@10:

  • Gemini 3.1 Pro vs. GPT-5.5: r = 0.730
  • Opus 4.8 vs. GPT-5.5: r = 0.569
  • Opus 4.8 vs. Gemini 3.1 Pro: r = 0.498

The two harsher graders agree with each other most. Claude's scores bunch up at the high end because it grades easy, which drags down its correlation with the other two, but it still produces the same ordering.

Failure Cases

When we take a look at judge evaluations for Amazon/Google, the recurring qualitative note is “misinterprets the query”. This is more of a semantic ceiling than a ranking gap. On the other hand, Ontology’s failure cases are concentrated in functional spec-related queries where Amazon’s category metadata dominates.

A few examples of Ontology 1 failure cases:

  • “lamp that won't wake my partner if I read at 3am” (Onton 0.4, Amazon 0.9)
  • “something to put on a weirdly deep windowsill” (Onton 0.07, Amazon 0.67)

We address these through two primary strategies:

  1. Expanding our product catalog: Even accounting for our current single-vertical focus of home decor and furniture, our catalog is is significantly smaller than Amazon’s and Google’s due to our focus on higher-signal, non-sponsored products. However, we’re actively working on catalog breadth.
  2. Self-learning: Ontology 1 improves itself10 through a scientific-method-shaped learning loop. The meaning of "won't wake my partner" is not just captured by a string or a vector, but it becomes increasingly precise and interconnected with other concepts in the graph. (For instance, to not wake someone up, a lamp might need to be low-light or red-light). The learnings then get leveraged on subsequent user searches.

Performance by search engine

Amazon does well on functional-spec queries where there's a clean keyword-to-category match (dimensions, materials, fixed attributes), and falls apart on aesthetic modifiers and stacked negation.

Google Shopping does well when it can pull up a curated section for a structured constraint, and misses on subtle intent and explicit negation.

Onton does well on the queries where you have to model what the person actually means instead of matching keywords.

Image and multimodal search

Below we compare 10 image and multimodal queries on Onton and Google. Google surfaces visually relevant results that often miss the query. (See row 6 where instead of surfacing mirrors, the primary results surface articles related to farmhouse decor).

Query ImageQuery TextOntonGoogle
chairs fitting this vibe
chairs fitting this vibe
bed like this but black
crib like this but pink
mirrors fitting this vibe
desk in this vibe
i'd like chair like this but budget version, and maybe red
lamps in this aesthetic, but tall
Fig. 8: Onton vs. Google, image and multimodal search

Looking forward

Next, we want to harden the benchmark and broaden its scope.

Hardening

Richer ranking metrics. P@10 weights all ten slots equally, so an engine earns no credit for ranking its best result first. Adding NDCG@1011 on graded (rather than binary) relevance judgments would reward both correct ordering and degree of quality.

Judge normalization. Because the three judges grade on different severity scales, normalizing each judge's scores before aggregating would tighten our confidence intervals.

Live, longitudinal evaluation. A one-shot screenshot benchmark is a snapshot, and search engines change. Because Subtext-Decor-90 is automated, we can re-run it on a regular cadence — tracking both how the baselines shift and how far Ontology's own learning loop moves its numbers over time.

Broadening

Larger, broader query sets. Subtext-Decor-90 deliberately targets the difficult queries Onton was built for. Scaling up the number of queries (e.g. to a hypothetical Subtext-Decor-10K or beyond) would increase benchmark confidence. Adding in “easier” queries would test how Ontology 1 holds up across the full distribution of real shopping behavior — and let us report complementary metrics like result diversity, seller reputation, and freshness that matter more once basic relevance is met.

More engines. We benchmarked against Amazon and Google because that's where the large majority of product searches start, but extending Subtext-Decor-90 to platforms like Pinterest, Wayfair, and Etsy would give a broader view of the search landscape.

Beyond decor. Subtext-Decor-90 is scoped to home decor and furniture because that's the vertical Onton indexes today. But the methodology isn't specific to decor. Ontology 1 is a general model, and the benchmark can likewise generalize — to Subtext-Shopping-N across verticals, and eventually Subtext-N beyond shopping.

Coverage estimates. Pooled relative recall (see footnote 7) is an imperfect proxy for recall, but it would give a first, bounded read on coverage alongside our precision numbers.

Catalog and vertical expansion. Our single-vertical, higher-signal catalog is a primary source of Ontology's remaining failure cases. Growing catalog breadth — and moving beyond home decor — is the clearest path to closing the functional-spec gap with Amazon and Google.

Accelerating the learning loop. Ontology’s failure cases shrink as it learns. Tightening that loop is the lever we expect to move our numbers most between benchmark runs.

Outreach

If you're a researcher excited by neurosymbolic systems, your organization has a search or discovery problem that conventional tools handle poorly, or you've found a use for Ontology we haven't thought of, we’d love to talk ([email protected]).

Footnotes

1 Amazon has ~600M unique products (ASINs) as of 2025 (~100x Onton), and ~2.6B listings worldwide as of 3/2026. Google Shopping has 50B+ listings as of 1/2026, with no published number on uniques; if Amazon’s ratio holds, one could extrapolate to 11.5B (~1000x Onton).

2 At the time of this benchmark, Claude Opus 4.8, Gemini 3.1 Pro, and GPT-5.5 were arguably the most capable publicly available frontier models. Drawing them from three different labs reduces the risk of correlated judgments.

3 A large majority of product searches start on Amazon or Google, making them the natural baseline. We plan to benchmark against Wayfair, Pinterest, and others over time.

4 P@10, or Precision at 10, measures the proportion of relevant items in the top 10 results. For each query, we average the three judges' P@10 into a single per-query score. An engine's reported P@10 is the mean of those scores across all 90 queries. To get the 95% confidence interval, we use the bootstrap: we draw 90 queries at random (with replacement) from our set, recompute the engine’s mean, and repeat this 10,000 times. The interval is the middle 95% of those means. Per-query P@10 is discrete and skewed, not normally distributed, so bootstrapping — which assumes nothing about the distribution's shape — is the natural choice here.

6 Precision asks what fraction of results were relevant. Recall asks what fraction of all relevant products the engine managed to return.

7 The usual proxy, pooled relative recall, asks what fraction an engine captured from the relevant results that all engines returned. It may or may not be a reasonable estimate of coverage, but coverage is outside the scope of this benchmark.

8 See Footnote 1.

9 Krippendorff's alpha is a measure of inter-rater reliability.

10 See Ontology 1 release announcement.

11 NDCG (Normalized Discounted Cumulative Gain, introduced in Järvelin & Kekäläinen 2002 rewards placing relevant results near the top: each result's relevance is discounted by how far down the list it appears, then normalized against the best possible ordering so scores fall in [0, 1].

More in research