OneRuby.devAN ENGINEERING NOTEBOOK

AI · 5 min read

FAISS, Pinecone and Weaviate: Compare the Search Contract First

Compare FAISS, Pinecone and Weaviate search contracts through a runnable FAISS cosine, filtering and persistence experiment.

A vector query can return the right nearest neighbors and still produce the wrong application result. In the local experiment here, the global top two records belong to the red tenant. Filtering those two results for the blue tenant returns nothing. Searching within the blue tenant returns two valid records.

That is a more useful starting point for comparing FAISS, Pinecone and Weaviate than a latency table copied from incompatible workloads. Before choosing an index or service, define the metric, ID mapping and filter boundary that a correct result must satisfy.

Establish an exact local baseline

The runnable example uses Python 3.11.5, NumPy 2.2.6 and FAISS CPU 1.11.0. It does not call Pinecone or Weaviate, create a cloud account, or claim a hosted-service benchmark. Four two-dimensional vectors make every ranking inspectable.

Both stored vectors and queries are normalized, then searched with IndexFlatIP inside IndexIDMap2. The first component performs exhaustive inner-product search; the second keeps the supplied integer IDs rather than exposing insertion positions as application identities.

FAISS documents the metric distinction directly: inner product is maximized, while its L2 index returns squared Euclidean distance. With normalized vectors, squared L2 equals 2 - 2 * inner_product, so the ranking agrees with cosine similarity. The official metric notes explain these conventions. A test checks both rankings and the numerical relationship on the same fixture.

Normalization also defines an input boundary. A zero vector has no cosine direction. Nonfinite coordinates, empty matrices and mismatched query dimensions are rejected. Silently indexing a zero vector and calling its score “cosine similarity” would give the application a result without a defined metric.

Filtering belongs before the result limit

The IDs are 101 and 102 for red, 201 and 202 for blue. A query along the first coordinate produces global IDs [101, 102]. Removing records outside blue then gives an empty list. This is not evidence that blue has no documents; the earlier global limit discarded its candidates.

The small implementation constructs an exact index containing the allowed rows before searching:

Python
def search(self, query, k=2, tenant=None):
if type(k) is not int or k < 1:
raise ValueError("positive integer k required")
q=normalized([query])
if q.shape[1] != self.vectors.shape[1]:
raise ValueError("query dimension mismatch")
# Rebuild a small exact allowlisted index for this demonstration.
# A FAISS index does not supply an application's authorization policy.
mask=np.array([tenant is None or t == tenant for t in self.tenants])
if not mask.any(): return []
idx=self.index
if tenant is not None:
idx=faiss.IndexIDMap2(faiss.IndexFlatIP(q.shape[1]))
idx.add_with_ids(self.vectors[mask],self.ids[mask])
scores, ids=idx.search(q,min(k,int(mask.sum())))
return [(int(i),float(s)) for i,s in zip(ids[0],scores[0]) if i != -1]

Rebuilding an index for each query is appropriate only for this tiny demonstration. It makes the intended semantics explicit without pretending to implement an efficient multi-tenant database. The caller must derive the allowed tenant from authenticated application state; accepting an arbitrary tenant string from an untrusted request would not enforce access control.

The recorded output is global_top2: [101, 102], postfilter_blue: [], and prefilter_blue: [201, 202]. Increasing the global candidate count can hide the problem on this fixture, but a fixed oversampling multiplier does not guarantee enough authorized matches for every distribution.

Download the implementation, tests, requirements and reproduction instructions. After installing the dependencies:

Terminal
python3 -B -m unittest -v test_example.py

Five tests pass. They cover metric equivalence, filtering, empty scopes, invalid vectors and metadata, oversized k, and persistence. Results never expose FAISS's missing-neighbor sentinel as a real record ID.

What changes between the three choices

ChoiceWhere this application contract livesWhat this note verifies
FAISSYour code owns metadata, allowed IDs, persistence and service boundaries around an indexExact CPU search, explicit IDs, prefilter semantics and local round trip
PineconeQueries target an index/namespace and can include metadata filtersDocumentation contract only; no remote query or consistency test
WeaviateCollections and query filters combine stored objects with vector searchDocumentation contract only; no server or client integration test

Pinecone's filter documentation describes metadata conditions on searches. Its namespace documentation matters too: writing and querying must target the intended namespace. A correct embedding in the wrong namespace is not retrieved by searching somewhere else.

Weaviate's filtering design explains its allow-list approach to prefiltered search. Its vector-search documentation distinguishes distance metrics and vector query inputs. Do not carry a cosine-similarity threshold unchanged into a distance field where smaller values are better.

Those documented interfaces do not prove identical behavior for every dataset or index setting. A migration test should submit the same IDs, vectors, allowed scope and query to both systems, then compare the returned records under the chosen metric. Service-specific client calls belong in a separately pinned integration test.

Persistence needs the metadata too

The example writes a FAISS index plus JSON containing IDs and tenants. Reloading reconstructs vectors by their IDs, restores the metadata and reproduces the local search result. Saving only the index would lose the application's tenant association.

This is a trusted local round trip, not a defensive loader for arbitrary uploaded indexes. FAISS explicitly warns that index loading does not validate hostile files in its index I/O documentation. The example also makes no crash-atomic update promise: a real service must publish matching index and metadata versions together.

Once the exact baseline is correct, an approximate index can be evaluated against it on representative queries and filters. Measure recall, update behavior, latency and memory under that workload; include sparse tenant scopes and deleted records. Choosing a managed service then becomes a question about an evidenced search contract and operating requirements, rather than which product wins an unsupported universal speed claim.

Found a mistake or tried a different approach?

Send Alex a note ↗