What is semantic search?

Learning center5 min read

Semantic search is a retrieval method that matches a query to documents by meaning instead of by shared words. It converts both the query and the documents into lists of numbers called embeddings, then returns the documents whose embeddings sit closest to the query's. A page about "cheap flights to Lisbon" can match the question "low cost air travel to Portugal" even though the two share almost no words.

Most AI search products use semantic search somewhere in their pipeline, usually alongside older keyword methods. Knowing how it works explains why some passages get retrieved and quoted while others on the same page are ignored.

How does semantic search work?

A semantic search system has an indexing side, which runs ahead of time, and a query side, which runs when someone asks a question.

  1. Split the content. Each page is cut into chunks, usually a few sentences to a few paragraphs long. A chunk is the unit that gets matched and returned.
  2. Create embeddings. An embedding model reads each chunk and outputs a vector, a fixed-length list of numbers (often several hundred or a few thousand). The model is trained so that texts with similar meaning produce vectors that point in similar directions.
  3. Store the vectors. The vectors go into a vector index, a database built to find nearby vectors quickly among millions or billions of entries.
  4. Embed the query. When a question arrives, the same embedding model turns it into a vector.
  5. Find the nearest neighbors. The index compares the query vector with stored vectors, usually with a measure called cosine similarity, and returns the chunks with the highest scores.

The embedding model is a neural network related to a large language model, trained on large amounts of text so that it learns which words and phrases tend to mean the same thing. When it embeds a query, it places the text at a point in a space where related ideas cluster together. No lookup happens at this step.

Keyword search, also called lexical search, scores documents by the words they share with the query. The most common scoring formula, BM25, rewards documents that contain the query terms often and gives more weight to rare terms than to common ones. Semantic search ignores exact wording and compares meaning.

Keyword searchSemantic search
What it comparesShared wordsEmbedding vectors
Synonyms and paraphrasesMissed unless listedUsually matched
Exact strings (product codes, names, error messages)Matched reliablySometimes missed or blurred
Long, conversational questionsWeak, many words carry little signalStrong
Explaining why a result matchedEasy, the words are visibleHard, the match is numerical

The two methods fail on different kinds of query. A question like "why is my laptop loud when idle" suits semantic search, because good answers may talk about fan noise and background processes. A query for a model number like "XPS 9530" suits keyword search, because an embedding model may treat it as similar to other model numbers and return the wrong product.

Why do search systems combine semantic and keyword retrieval?

Because the two methods have opposite weaknesses, most production systems run both and merge the results. This is called hybrid search. A common approach takes the top results from each method and combines their rankings into one list, so a passage that scores well on either side has a chance to survive.

The merged list is then usually passed to a reranker, a separate model that reads the query and each candidate passage together and gives a more careful relevance score. Reranking is slower than vector lookup, so it only runs on a short list of candidates. The order it produces decides which passages move on to the next stage, a process covered in how AI search chooses which sources to cite.

AI search is built on retrieval-augmented generation: the system retrieves text first and the language model writes its answer from that text. Semantic search can appear at several points in that process.

  • Retrieving pages. Some search indexes use embeddings to find candidate pages for each query the model writes.
  • Selecting passages. After pages are fetched and split into passages, embeddings help score which passages inside a page match the question.
  • Matching follow-up questions. In a conversation, the query the system searches with is often a rewritten version of what the user typed, and embeddings help match that rewritten intent.

This is one reason AI search behaves differently from a classic results page, as described in how AI search differs from traditional search. A traditional engine ranks whole pages for a short query. An AI system compares individual passages against long, natural-language questions, and semantic search handles that kind of matching well.

What does semantic search mean for a website?

Since matching happens at the level of chunks, each passage on a page competes on its own. A passage that answers one question clearly, using the terms a reader would use, produces an embedding that sits close to that question. If a passage mixes several topics, its embedding lands somewhere between them and may be close to none.

A few practical consequences follow from the mechanism:

  • Write sections that make sense without the text around them. A chunk that starts with "As mentioned above" loses the context the embedding needed.
  • Name things explicitly. Product names, brand names, locations and specifications should appear in the passage that discusses them, because keyword retrieval still handles exact strings better than embeddings.
  • Cover the different ways people phrase a question. Semantic search tolerates paraphrase, but a page that addresses the actual problem a person describes will match more of their questions than a page that only uses internal jargon.
  • Keep pages fetchable. None of this applies if the AI crawlers that collect content for a given product are blocked from reading the page.

Hybrid retrieval and reranking change with each query and each run, so the same page can be retrieved for one phrasing of a question and missed for another. Testing a range of realistic prompts over time shows which phrasings actually surface a page. That approach is described in measuring AI visibility.

Start a 14-day free trial
and get your AI visibility report