Imagine a homeware retailer trading in Dubai and Riyadh. The Arabic storefront looks finished. Product descriptions have been translated, banners approved and the language switch tested. Yet a shopper searching in Arabic for a coffee machine sees accessories first. Another mixes an Arabic category with an English brand and finds nothing. Arabic ecommerce search is failing at the point where the customer has already told the business what they want.
The executive question is specific: can we prove that the search experience understands the buying intent of our customers before paying for a new engine or an AI assistant? I would begin with a small set of queries and human judgments about the products that should appear. Without that evidence, a search project becomes a debate about features, with no agreed definition of a useful result.
Arabic ecommerce search needs its own acceptance set
Take a sample of search terms from the existing store, subject to your data-handling rules. Remove personal details and restrict access to raw logs. Add phrases collected from merchandising and customer service, clearly labelling those as proposed tests rather than observed demand. If there is no search history, start with that labelled test set and replace assumptions as evidence arrives.
Cover exact product names, broad categories, use cases, model numbers, spelling variation and queries that mix scripts. Keep UAE and Saudi market context attached to each case. Language and country are different dimensions: choosing Arabic should not silently decide currency, assortment or delivery eligibility.
Algolia's multilingual search documentation explicitly recognises searches combining Arabic words with Roman-alphabet brand names. It describes trade-offs between separate language indices and a shared multilingual index. That is an architecture choice to test against your queries, not a universal rule that one approach always wins.
Use five checks before changing the engine
1. Agree what each query means
Create a worksheet with the query, intended market, likely meaning, acceptable products and unacceptable results. Ask an Arabic-speaking merchandiser and someone responsible for the buying journey to judge it together. Where they disagree, record the ambiguity. Do not manufacture one correct answer for a phrase that genuinely has several meanings.
For an exact model query, the requested model should be easy to find when offered. For a category query, several relevant products may be sensible. For an unavailable item, the interface should explain the absence and distinguish alternatives. A replacement is a commercial suggestion; it should not be presented as the product requested.
Include negative cases. A compatible accessory is not the main appliance. A similar model number is not an exact match. A larger pack is not necessarily a substitute for a single unit. Those distinctions give the search team a useful boundary when broader matching appears to improve coverage.
2. Inspect how language becomes searchable text
Ask the technical team to show how representative queries and product fields are processed. Normalisation makes selected text variations comparable. Stemming reduces related word forms. These operations can help matching, but they need testing against product names, brands and identifiers that should retain their exact meaning.
Elastic's Arabic analyser documentation shows a pipeline including digit handling, stop words, Arabic normalisation and stemming, with provision for terms excluded from stemming. The useful purchasing question is whether your implementation exposes and tests these choices. An Arabic language checkbox says little about the result for your catalogue.
Run the same intent through Arabic script, English brand spelling and the mixed forms customers actually use. Include numeral variants where relevant. If transliterated Arabic appears in your logs, test it explicitly rather than assuming every language feature covers it. Keep failures grouped by cause so a catalogue omission is not mistaken for a language-processing defect.
3. Govern synonyms as merchandising decisions
A synonym rule expands what a search may retrieve. Give each rule an owner, examples that should improve and examples that must remain unchanged. Category vocabulary, common spelling and brand aliases deserve separate consideration. Avoid a large unreviewed synonym upload generated from translations or an AI model.
Algolia distinguishes ordinary and one-way synonyms: the latter expand a term in one direction without automatically applying the reverse relationship. That distinction matters commercially. A broad category may reasonably retrieve a particular product family, while a specific family query should not automatically retrieve every product in the category.
Review the first results after every change, including exact model searches. Keep an expiry or review date for seasonal and campaign vocabulary. A temporary promotion should not leave a permanent rule that distorts normal shopping months later. Store the previous version so the team can reverse a bad change.
4. Separate finding a product from promoting it
Ask the team to show results before and after merchandising boosts. Stock, popularity, margin and paid placement may affect the order, but leadership should see when they displace a closer match. Define which exact queries must preserve the requested product's visibility and how unavailable products should be handled.
Carry the selected market into availability and the destination page. The adjacent product feed quality audit checks the offer advertised outside the store. This test concerns discovery inside it: a relevant result must still open the intended variant and a truthful buying option. Fix shared catalogue defects once, then verify both surfaces.
Keep zero-result searches separate from searches that return irrelevant products. Reducing the first by filling every page with vaguely related merchandise can make the second worse. Provide a useful route to refine a query, browse the relevant category or see clearly labelled alternatives.
5. Measure relevance before claiming revenue
Elastic's ranking evaluation guidance uses representative queries and manually rated results to measure search quality. Apply that discipline even if your platform is different. Score whether the first results satisfy the agreed intent, and report separately for exact products, categories and mixed-language searches.
Keep some queries outside the tuning set for a later check. Otherwise, the team can optimise the demonstration while missing the rest of the store. Retest after catalogue, synonym, ranking or model changes. An aggregate improvement should not conceal a regression in Arabic model searches that matter to a major category.
Then observe the customer journey: refinements, product visits, add-to-basket actions, purchases and time to results. Compare controlled groups where traffic supports a useful test. Search users and browsing users begin with different intent, so their conversion-rate difference alone does not prove the engine caused additional sales.
Buy against the customer's words
Bring this evidence into an AI investment decision before approving semantic retrieval or a conversational layer. A stronger model may help a demonstrated gap. It should earn that role against the same catalogue, queries and buying constraints as the current system.
Arabic ecommerce search is ready to expand when the business can explain what improved, which queries still fail and who owns the next correction. Put the customer's words beside the first products shown. That comparison is more useful than another impressive search demo.