When semantic search helps
A customer asks about returning an item while the relevant document is called an order refund policy. Literal matching may miss the connection. Semantic search represents text as numeric vectors and ranks their similarity. pgvector brings that capability into PostgreSQL, allowing a bounded pilot to keep document metadata and embeddings together. It is worth evaluating when the team already operates PostgreSQL and the workload fits its capacity, latency and maintenance constraints.
An embedding does not replace the source document or its authorization policy. Keep the original text, tenant, document version and source address. Evaluate the embedding model on your language and vocabulary: a model that works on general English can struggle with Persian product terms. Queries and documents need the same model version and dimensions. Our three-dimensional vectors are deliberately synthetic so the SQL is reproducible; they are not a demonstration of real language understanding.
Start with data quality and permissions
Remove duplicates and expired versions before embedding. Split longer guides into traceable sections with document IDs and meaningful headings. Very short chunks lose context; very long chunks mix unrelated topics. Our editorial recommendation is to compare two chunk sizes on actual questions rather than adopt a popular number without evidence. Record the source for every retrieved chunk so a user can verify the answer and report a wrong or outdated document.
An exact query computes distance against relevant vectors. HNSW and IVFFlat can reduce search time on larger datasets, but approximate retrieval requires evaluation. Match the index operator class to the distance operator; the example uses cosine distance throughout. Apply tenant permissions before showing results. Approximate search with a filter can return fewer qualifying results than requested. Increasing the limit is not a substitute for testing index behavior, selectivity and the actual query plan.
Build a test set with known answers: exact wording, synonyms, typos, product names and questions with no answer. Measure whether the expected document appears in the first five results and record p95 latency. Keep lexical matching for identifiers and exact product codes. Compare hybrid and vector-only retrieval on the same questions. A few impressive assistant responses are insufficient evidence. Changing the embedding model requires a controlled re-embedding process with a separate version.
Code example and verification
This educational example demonstrates the implementation path. Check the stated runtime and prerequisites in a test environment; the notes explain what remains before production use.
CREATE EXTENSION IF NOT EXISTS vector;
CREATE TABLE demo_documents (
id bigint GENERATED ALWAYS AS IDENTITY PRIMARY KEY,
tenant_id integer NOT NULL,
title text NOT NULL,
embedding vector(3) NOT NULL
);
INSERT INTO demo_documents (tenant_id,title,embedding) VALUES
(1,'Return policy','[1,0,0]'),
(1,'Shipping guide','[0,1,0]'),
(2,'Private policy','[0.99,0.01,0]');
CREATE INDEX ON demo_documents
USING hnsw (embedding vector_cosine_ops);
SELECT id,title,1-(embedding <=> '[1,0,0]') AS similarity
FROM demo_documents
WHERE tenant_id=1
ORDER BY embedding <=> '[1,0,0]'
LIMIT 5;With the extension installed, run this in a disposable database. It should exclude tenant 2 and rank Return policy first. The three-dimensional embeddings are test fixtures. Bind query parameters and obtain the tenant from verified identity in an application. This WHERE clause is not a complete multitenant policy; evaluate database roles and RLS separately.
Measure before expanding
Start with read-only documents from one department and a defined user group. Record the model, vector dimensions, dataset size, index settings, permission boundaries and failed queries. Review refresh costs and recovery before expanding. The acceptance criterion is a reliable route to a relevant, current and permitted document. Include user feedback and source verification in that criterion rather than optimizing distance scores alone.
Implementation checklist
- Install vector in a test database and match the column dimensions to the selected model.
- Version query and document embeddings together; avoid logging private embeddings.
- Compare exact, approximate and hybrid retrieval on the same test set.
- Test tenant isolation with a user who must not see the document.
Practical explanations and recommendations are Liyan Knowledge editorial analysis.Sources: pgvector — official documentation
This Liyan Knowledge article is an editorial synthesis based on the original source.View original source





