Full-text search vs. semantic search: A full comparison

Compare full-text search and semantic search, including their differences, use cases, advantages, limitations, and other factors.

Maya Shin

Maya Shin

Head of Marketing @ Meilisearch·@mayya_shin·LinkedIn

·17 min read
Full-text search vs. semantic search: A full comparison

Share the article

Full-text search matches the exact keywords in your query with the exact words in a document. It is one of the most established approaches to information retrieval.

It is most commonly used for documentation and knowledge bases, e-commerce search, code and log search, and similar use cases; wherever users know or can closely approximate the exact terminology they’re looking for.

The main limitation of full-text search is that the user query and the relevant document need enough lexical overlap for the search engine to connect them.

Semantic search overcomes this limitation.

It represents both queries and documents as embeddings and retrieves results based on similar meaning rather than shared terms.

However, that doesn't automatically make semantic search better than full-text search.

In this guide, we will explore the definitions, inner workings, and applications of both search types, so you can choose which suits your needs better.

Perhaps you might even combine both approaches. Yes, it’s possible!

Full-text search relies heavily on how many words in the user query appear in a document. If a document contains the exact words from the user's query, it’s considered a match and is returned in the search results.

For example, a user can search software documentation using the term ‘API authentication error.’ This would typically retrieve pages containing terms such as ‘API,’ ‘authentication,’ and ‘error.’ The documents that contain more of these terms – or contain them in more important locations – will rank higher.

The advantage of a full-text search is that it’s precise. More common words are disregarded (like ‘the’ or ‘a’) and more specific words are matched against a document’s content.

BM25 is a widely used full-text ranking algorithm that scores documents using signals such as term frequency and document length.

However, other algorithms are used as well, such as in the case of Meilisearch. Meilisearch uses a multi-criteria, recursive bucket-ranking system in which ranking rules are applied sequentially to refine groups of matching documents. More on that after the jump.

Semantic search focuses on interpreting the meaning behind the words in the query, rather than requiring the query and the document to use the same phrasing.

For example, an e-commerce user can search for ‘something to keep coffee hot on my commute.’ Semantic search will retrieve products described as ‘insulated travel mug,’ ‘thermal tumbler,’ or similar even though the product descriptions don’t use the exact same words used in the user query.

This helps with natural-language queries, paraphrasing, concept discovery, and vocabulary mismatches between users and indexed content. Modern semantic search often relies on natural language processing (NLP) models to represent and compare text meaning.

However, semantic retrieval depends heavily on factors such as the embedding model and how documents are represented before embedding.

How does full-text search work?

A full-text search finds keywords in a document, but first it takes several steps to parse the user query and normalize words.

Here is a breakdown of a full-text search process:

How full-text search finds a document

  1. Document analysis and tokenization: At index time, searchable document text is split into tokens. Depending on the configuration, the system can normalize case, punctuation, word forms, and other textual variations.
  2. Index creation: The processed tokens are stored in an inverted index. This is a data structure that maps each searchable term to the documents in which it appears.
  3. Query analysis: When the user submits a query, the search engine processes it into searchable terms using compatible tokenization and normalization rules. This means that the query terms can be looked up against the same representations created during indexing.
  4. Candidate retrieval: The engine searches the inverted index for the query terms and identifies the documents that satisfy its matching rules. Some engines may also consider prefix matching, spelling variations, or phrase matching.
  5. Ranking and return: The engine orders the matching documents using its relevance algorithm and returns them to the user.

How does semantic search work?

Semantic search works with vectors and the meaning behind words in the user query.

Here is a breakdown of a semantic search process:

How semantic search finds related results

  1. Document preparation: The system determines what information from each document best represents it semantically. This might mean embedding the complete text, only selected fields, or smaller chunks of longer documents.
  2. Embedding generation: An embedding model converts each document representation (from the previous step) into a dense numerical vector. Text with similar meaning occupies nearby regions of the embedding space.
  3. Vector index creation: The document embeddings are stored in a vector index. This is the foundation of vector search: systems commonly use Approximate Nearest Neighbor (ANN) techniques so they don’t need to compare a query against every vector in the database.
  4. User query embedding: When a user types in a query, it passes through a compatible embedding model. The result of this is a query vector in the same vector space as the indexed documents.
  5. Ranking and return: The vector index performs a similarity search to find the document vector nearest to the query vector using similarity or distance functions. Metrics such as cosine similarity are common for text embeddings, though the specific ones will depend on implementation. The best-matching documents are then returned as semantic search results.

Full-text and semantic search each have advantages, but semantic search is necessary for conversational results.

If you want users to talk to AI agents and ask broad questions that return results, such as in an AI assistant or chatbot, you need semantic search.

Here is a table that compares the two search approaches:

Full-text searchSemantic search
Main functionMatch exact wordsMatch word meanings
Key componentsText analysis, inverted indexing, and a lexical relevance-ranking systemEmbedding model, vector embeddings, vector index, and nearest-neighbor retrieval
Document storageInverted indexVector index
Relevance measurementLexical signals such as matched query terms, exactness, term proximity, field importance, typo distanceSimilarity or distance between the query embedding and document embeddings
Best forExact term matchesMatches including synonyms or similar phrases
CostDoes not require embedding generation; typically has lower compute and storage overhead for the retrieval layerAdds embedding-generation and vector-indexing costs; query-time cost and latency vary significantly depending on whether the embedder runs locally or through an external API
ExampleSearch for ‘return policy,’ so a search would return documents with ‘return’ and ‘policy’ in the text.Search for ‘return policy’ returns documents that include similar meanings such as ‘send back’ or ‘return purchased product’

What problems do full-text search and semantic search solve?

Full-text search is great when you need a simple, affordable document search. For example, suppose you want to create a product search based on product name and price. You can build full-text search in-house without external LLMs, saving on costs.

For contextual matches across a broad range of documents, you need semantic search.

Let’s say you need an AI assistant to answer questions from a sales manager about the performance of revenue for a certain product. A basic search cannot handle the nuances of a conversational question, but a semantic search can find documents that broadly match the query’s intent.

Here is a comparison table for both types of search:

Full-text searchSemantic search
Find known products, records, or documents using names, SKUs, serial numbers, error codes, case numbers, or other precise termsFind products or content from a natural-language description even when the query does not contain the terminology used in the indexed data
Prioritize precise lexical matching for domain-specific terminology, proper nouns, abbreviations, and identifiersBridge vocabulary differences such as paraphrases, related concepts, and cross-lingual semantic similarity
Provide low-latency retrieval without generating query embeddingsImprove recall for ambiguous, descriptive, or conversational queries where lexical overlap is weak
Power site search, documentation search, e-commerce search, code search, and other keyword-driven interfacesPower natural-language knowledge retrieval, recommendation/discovery experiences, and retrieval components for AI assistants or RAG systems

Your business use case will determine if full-text search or semantic search is best for you.

Full-text retrieval can operate entirely on text indexes. This avoids the cost of generating, storing, and searching dense embeddings.

Semantic search, on the other hand, requires an embedding model, a vector index, and additional indexing work whenever searchable content changes.

If you want to work with AI agents and build personalized assistants for business queries, you need the power of semantic search.

Here is a table of advantages that will help you decide:

Full-text searchSemantic search
• High precision for exact terminology, identifiers, names, and phrases
• No embedding-generation or vector-storage layer required
• Low-latency retrieval suitable for interactive and search-as-you-type experiences
• Deterministic relevance behavior that is comparatively straightforward to inspect, tune, and test
• Mature controls for ranking, phrase matching, field weighting, filters, synonyms, and domain terminology
• Retrieval based on conceptual similarity rather than lexical overlap alone
• Can retrieve paraphrases and semantically related wording without manually defining every synonym
• Handles descriptive and natural-language queries more effectively
• Can bridge vocabulary differences between the user's query and indexed content
• Can support multilingual or cross-lingual retrieval when the selected embedding model is designed for it

Keep in mind that if you’re worried about users making spelling mistakes or you need synonyms, you don’t necessarily need to implement semantic search.

In Meilisearch, full-text search includes built-in features like typo tolerance and configurable synonyms, which can be the hybrid solution you’re looking for.

Full-text search is limited mainly by lexical matching. Relevance drops when users and documents use different terminology to describe the same concept, unless the search engine can bridge that gap.

Semantic search introduces model-dependent relevance behavior, more complex infrastructure, and additional indexing and query-time computation.

For engineering teams, the practical question is which failure modes are acceptable for their search workload.

Here is a breakdown of each search type’s limitations:

Full-text searchSemantic search
• Lexical mismatch can reduce recall when relevant documents use substantially different wording from the query
• Understanding of broader intent or conceptual similarity is limited
• Aggressive stemming, stop-word handling, tokenization, or fuzzy matching can introduce false positives if configured poorly
• Can underperform on exact identifiers, rare proper nouns, product codes, and highly specific terminology if semantic similarity overwhelms lexical precision
• Relevance depends heavily on the embedding model and how well it represents the application's domain, languages, and query patterns
• Query embedding can add latency and external API dependency when embeddings are generated by a remote provider

What is the difference between a full-text index and a semantic index?

The main difference between a full-text index and a semantic index is how they store their indexes.

A full-text index commonly uses an inverted index. Instead of storing documents as continuous blocks of text, it maps searchable terms to posting lists identifying the documents in which each term shows up.

At query time, the search engine doesn’t need to scan every document. It simply looks up query terms directly in the index.

A semantic index stores vector embeddings generated from documents or document representations.

At query time, an embedding model converts the query into a vector and the system performs nearest-neighbor search to identify document vectors that are closest to it in the embedding space.

So the key distinction is that an inverted index is optimized for looking up lexical signals, while a vector index is optimized for finding nearby points in an embedding space.

Full-text search remains the right retrieval approach for many applications, particularly when users search using known terminology, identifiers, names, or short keyword queries.

Here are a few examples of real-world use cases for a full-text search:

  • E-commerce searches: Find products using codes, names, SKUs, titles, ISBNs, serial numbers, model numbers, or any other data that must be an exact match.
  • Legal searches: Find cases, filings, or other legal documents using case numbers, party names, statutory terminology, or exact phrases.
  • Code search: Locate functions, classes, variables, error messages, or known code fragments across a codebase.
  • Log errors: Find application or infrastructure events using timestamps, service names, exception messages, request IDs, or error codes.
  • API documentation: Retrieve endpoint names, methods, configuration options, parameters, or specific technical concepts from documentation.
  • Email search: Find messages using sender and recipient names, subject-line terms, attachment names, or words within the message body.
  • Medical coding: Retrieve records using billing codes, procedure codes, record identifiers, medication names, or other standardized terminology.

Semantic search is most valuable when the searcher's intent can be expressed in different ways, and the indexed content may not contain the same vocabulary as the query.

Here are a few real-world examples of semantic search:

  • Customer service: Retrieve troubleshooting articles or help-center content for questions phrased differently from the documentation.
  • Recruiting: Retrieve candidate profiles based on conceptually related skills, roles, and experience rather than relying exclusively on exact job-title or keyword matches.
  • Enterprise knowledge search: Find relevant internal documents when employees describe a problem or concept without knowing the exact terminology used in the source material.
  • Content recommendation: Surface articles, videos, products, or other content that is conceptually related to a user's query even when there is little direct keyword overlap.

Full-text relevance depends on more than whether the correct words exist somewhere in the index.

The way text is tokenized and normalized, which fields are searchable, and how the engine handles spelling variations, synonyms, proximity, and ranking, all affect which documents qualify as matches and where they appear in the result set.

Here is a breakdown of the factors impacting full-text search accuracy:

FactorHow it affects accuracy
TokenizationPoor tokenization can break domain-specific names, compounds, punctuation-heavy identifiers, or language-specific text into tokens that no longer match how users search
Normalization and language processingOverly aggressive normalization can collapse distinct terms, while under-normalization can miss valid variations
Stop wordsRemoving a word that carries meaning in a specific domain can change the query or prevent an intended match
SynonymsIncomplete synonym mapping leaves vocabulary gaps while overly broad synonym sets can introduce irrelevant results
Searchable fieldsField order or weighting also affects whether, for example, an exact title match outranks the same term appearing only in a long description
Typo tolerance and fuzzy matchingPermissive matching improves recall for misspellings, but too much tolerance can increase false positives

For Meilisearch, several of these factors are directly configurable.

searchableAttributes controls which fields participate in full-text search and their relative ranking importance. Typo tolerance, stop words, synonyms, and ranking rules provide additional control over recall and relevance.

Choosing the right embedding model is key for semantic search. However, document representation, embedding consistency, vector retrieval, and evaluation methodology can also affect the final results.

Here is a list of the most significant factors that affect semantic search accuracy:

FactorHow it affects accuracy
Embedding-model fitA model that performs strongly on general English benchmarks may not represent specialized legal, medical, technical, or multilingual terminology equally well
Document representationThe text sent to the embedder determines what semantic information the vector can represent. Including irrelevant fields adds noise, while omitting important fields removes signals needed to retrieve the document
Chunking strategyFor long content, chunks that are too large can combine unrelated concepts into one representation, while chunks that are too small can remove the context needed to interpret a passage
Embedding consistencyChanging embedding models, dimensions, preprocessing, or model-specific query/document instructions without re-embedding the relevant data can make vectors incomparable
Data qualitySemantic retrieval cannot recover information that is absent, duplicated, outdated, or incorrectly represented in the indexed source data

How does Meilisearch combine full-text and semantic searches?

Meilisearch supports full-text, semantic, and hybrid searches within the same index. This means you don't need to run a separate keyword search engine and vector database, then build your own logic to combine their results.

To enable semantic search, you configure an embedder. Meilisearch can generate embeddings automatically using providers such as OpenAI, Hugging Face, or Ollama, or work with vectors you generate yourself.

You can also use documentTemplate to control what information gets embedded. For example, an e-commerce index might embed the product title, description, and category while leaving out fields that don't add useful semantic context.

At query time, the main control is semanticRatio:

  • 0.0 = full-text search only
  • 1.0 = semantic search only
  • Anything in between = hybrid search

So you can choose how much weight each retrieval method gets.

For example, exact searches such as a SKU, error code, or product name could use more full-text weight. A query such as ‘waterproof shoes for long hikes’ may benefit from more semantic weight because relevant products don't necessarily contain those exact words.

In the case of hybrid search, Meilisearch merges lexical and semantic result sets based on their relevance scores.

Meilisearch can also use the embedder's distribution setting to normalize semantic scores before combining them with full-text scores. This prevents a particular embedding model's numerical behavior from skewing hybrid ranking.

The result is that you can tune both sides independently:

  • Full-text relevance: Ranking rules, typo tolerance, synonyms, searchable attributes, and filters.
  • Semantic relevance: Embedding model, documentTemplate, and score distribution.
  • The balance between them: semanticRatio.

Full-text and semantic search solve different retrieval problems.

If users search with known product names, identifiers, or technical terms, full-text search is usually the better fit.

If they describe what they want in natural language that is different from your indexed content, semantic search can recover results that keyword matching would miss.

Most real search experiences get both types of queries.

That's where hybrid search becomes useful: you don't have to force every query through the same retrieval strategy.

Meilisearch lets you run full-text and semantic retrieval in the same search engine. And because the two sides remain configurable, you can tune full-text ranking separately from embeddings and semantic relevance instead of treating hybrid search as a black box.

To find out how Meilisearch can help build your search functions, try Meilisearch today.

Maya Shin

Maya Shin

Head of Marketing @ Meilisearch

Related articles