AI · 5 min read
FAISS, Pinecone and Weaviate: Compare the Search Contract First
Compare FAISS, Pinecone and Weaviate search contracts through a runnable FAISS cosine, filtering and persistence experiment.
A vector query can return the right nearest neighbors and still produce the wrong application result. In the local experiment here, the global top two records belong to the red tenant. Filtering those two results for the blue tenant returns nothing. Searching within the blue tenant returns two valid records.
That is a more useful starting point for comparing FAISS, Pinecone and Weaviate than a latency table copied from incompatible workloads. Before choosing an index or service, define the metric, ID mapping and filter boundary that a correct result must satisfy.
Establish an exact local baseline
The runnable example uses Python 3.11.5, NumPy 2.2.6 and FAISS CPU 1.11.0. It does not call Pinecone or Weaviate, create a cloud account, or claim a hosted-service benchmark. Four two-dimensional vectors make every ranking inspectable.
Both stored vectors and queries are normalized, then searched with IndexFlatIP inside IndexIDMap2. The first component performs exhaustive inner-product search; the second keeps the supplied integer IDs rather than exposing insertion positions as application identities.
FAISS documents the metric distinction directly: inner product is maximized, while its L2 index returns squared Euclidean distance. With normalized vectors, squared L2 equals 2 - 2 * inner_product, so the ranking agrees with cosine similarity. The official metric notes explain these conventions. A test checks both rankings and the numerical relationship on the same fixture.
Normalization also defines an input boundary. A zero vector has no cosine direction. Nonfinite coordinates, empty matrices and mismatched query dimensions are rejected. Silently indexing a zero vector and calling its score “cosine similarity” would give the application a result without a defined metric.
Filtering belongs before the result limit
The IDs are 101 and 102 for red, 201 and 202 for blue. A query along the first coordinate produces global IDs [101, 102]. Removing records outside blue then gives an empty list. This is not evidence that blue has no documents; the earlier global limit discarded its candidates.
The small implementation constructs an exact index containing the allowed rows before searching:
def search(self, query, k=2, tenant=None): if type(k) is not int or k < 1: raise ValueError("positive integer k required") q=normalized([query]) if q.shape[1] != self.vectors.shape[1]: raise ValueError("query dimension mismatch") # Rebuild a small exact allowlisted index for this demonstration. # A FAISS index does not supply an application's authorization policy. mask=np.array([tenant is None or t == tenant for t in self.tenants]) if not mask.any(): return [] idx=self.index if tenant is not None: idx=faiss.IndexIDMap2(faiss.IndexFlatIP(q.shape[1])) idx.add_with_ids(self.vectors[mask],self.ids[mask]) scores, ids=idx.search(q,min(k,int(mask.sum()))) return [(int(i),float(s)) for i,s in zip(ids[0],scores[0]) if i != -1]Rebuilding an index for each query is appropriate only for this tiny demonstration. It makes the intended semantics explicit without pretending to implement an efficient multi-tenant database. The caller must derive the allowed tenant from authenticated application state; accepting an arbitrary tenant string from an untrusted request would not enforce access control.
The recorded output is global_top2: [101, 102], postfilter_blue: [], and prefilter_blue: [201, 202]. Increasing the global candidate count can hide the problem on this fixture, but a fixed oversampling multiplier does not guarantee enough authorized matches for every distribution.
Download the implementation, tests, requirements and reproduction instructions. After installing the dependencies:
python3 -B -m unittest -v test_example.pyFive tests pass. They cover metric equivalence, filtering, empty scopes, invalid vectors and metadata, oversized k, and persistence. Results never expose FAISS's missing-neighbor sentinel as a real record ID.
What changes between the three choices
| Choice | Where this application contract lives | What this note verifies |
|---|---|---|
| FAISS | Your code owns metadata, allowed IDs, persistence and service boundaries around an index | Exact CPU search, explicit IDs, prefilter semantics and local round trip |
| Pinecone | Queries target an index/namespace and can include metadata filters | Documentation contract only; no remote query or consistency test |
| Weaviate | Collections and query filters combine stored objects with vector search | Documentation contract only; no server or client integration test |
Pinecone's filter documentation describes metadata conditions on searches. Its namespace documentation matters too: writing and querying must target the intended namespace. A correct embedding in the wrong namespace is not retrieved by searching somewhere else.
Weaviate's filtering design explains its allow-list approach to prefiltered search. Its vector-search documentation distinguishes distance metrics and vector query inputs. Do not carry a cosine-similarity threshold unchanged into a distance field where smaller values are better.
Those documented interfaces do not prove identical behavior for every dataset or index setting. A migration test should submit the same IDs, vectors, allowed scope and query to both systems, then compare the returned records under the chosen metric. Service-specific client calls belong in a separately pinned integration test.
Persistence needs the metadata too
The example writes a FAISS index plus JSON containing IDs and tenants. Reloading reconstructs vectors by their IDs, restores the metadata and reproduces the local search result. Saving only the index would lose the application's tenant association.
This is a trusted local round trip, not a defensive loader for arbitrary uploaded indexes. FAISS explicitly warns that index loading does not validate hostile files in its index I/O documentation. The example also makes no crash-atomic update promise: a real service must publish matching index and metadata versions together.
Once the exact baseline is correct, an approximate index can be evaluated against it on representative queries and filters. Measure recall, update behavior, latency and memory under that workload; include sparse tenant scopes and deleted records. Choosing a managed service then becomes a question about an evidenced search contract and operating requirements, rather than which product wins an unsupported universal speed claim.
Found a mistake or tried a different approach?
Send Alex a note ↗