OneRuby.devAN ENGINEERING NOTEBOOK

AI · 5 min read

Inside Elasticsearch: Terms, Positions and the Search You Asked For

Trace terms, positions and BM25 with a tested Python model, then separate those mechanics from Elasticsearch analysis, refresh and shard behavior.

Three documents contain the same two words: “green tea”, “tea green”, and “green herbal tea”. A search for both words can return all three. A phrase search should return only the first. An index that stores nothing but document IDs cannot explain the difference; it also needs positions.

That small example is a useful way into Elasticsearch. The interesting questions are what got indexed, how the query was interpreted, and which information the engine retained. Start there before changing boosts or adding shards.

The downloadable Python experiment implements a tiny positional index and a BM25 teaching calculation. It runs locally without Elasticsearch. It is not a replacement for Lucene, a performance model, or a reproduction of Elasticsearch's exact scores.

The index runs in the other direction

A document-oriented view maps an ID to its text. An inverted index maps a term to the documents containing it. In the experiment, a posting also records each occurrence's position:

Python
for doc_id, text in documents.items():
words = tokens(text); self.lengths[doc_id] = len(words)
for position, word in enumerate(words):
self.postings[word].setdefault(doc_id, []).append(position)

For green, the three documents have positions zero, one and zero. For tea, they have positions one, zero and two. Intersecting the document-ID sets finds all three. Requiring tea's position to be exactly one greater than green's leaves only document one.

Repeated terms expose another mistake that a set-only representation would hide. “tea tea” requires two adjacent occurrences. The fixture tests that a document containing just one “tea” does not match. Positions are evidence about sequence, not simply another way to count words.

Real Lucene indexes use specialized storage and execution strategies rather than these Python dictionaries. The dictionary is useful precisely because you can inspect every posting and disprove a mistaken match by hand.

Analysis changes what “the same word” means

Before a term reaches the index, analysis can tokenize text and transform the tokens. Elasticsearch's standard analyzer lowercases tokens and uses Unicode-aware segmentation. It does not enable stemming, and its default stop-word list is empty. Thus “foxes” does not automatically become “fox” just because the analyzer is called standard.

Our tokenizer is deliberately smaller: lowercase ASCII letters and digits, with everything else treated as a separator. It retains “the” and does not stem. It should not be used to predict how a Unicode name, an email address or a hyphenated identifier is analyzed by Elasticsearch.

A match query analyzes its text. A term query asks for an exact indexed term, which is why Elastic warns against casually using term queries on text fields. In the model, the literal term Tea misses the lowercase posting. That is a model demonstration of the distinction, not a promise that every field mapping lowercases input.

For an unexpected real result, inspect the field mapping and the tokens from the _analyze API. Then inspect the actual query. Index-time and search-time analyzers can differ intentionally; saying that every query always uses the same analyzer is too broad.

BM25 rewards evidence without counting forever

Matching identifies candidates. Ranking orders them. Elasticsearch's default text similarity is BM25, with documented defaults k1=1.2 and b=0.75. The similarity reference explains term-frequency saturation and field-length normalization.

The experiment uses the familiar single-term BM25 form: inverse document frequency, multiplied by a frequency factor whose denominator grows with frequency and normalized length. Its tests compare relationships rather than treating its floating-point values as Elasticsearch outputs.

With length normalization disabled for one controlled test, adding a second occurrence increases the score, and adding a third increases it again by a smaller amount. Repetition helps, but the reward saturates. In a second test, two documents contain the term once; the shorter document scores higher with b=0.75. Setting b=0 removes that length difference from this model.

Actual scores also depend on field statistics, boosts, query structure and Lucene's implementation details. A shard can see different term statistics from the whole index. A tutorial formula therefore cannot tell you the exact order of a distributed production query. Use the explain API to investigate a particular document and query.

A shard contains a changing collection of segments

A shard is a Lucene index; several shards may live on the same node. Five primary shards do not imply five machines or equal disk usage. Replicas, allocation rules, routing and document distribution all matter.

Within a shard, new indexed data becomes searchable when a refresh opens a search view that includes it. A refresh is not the same event as a durability-oriented flush. Elastic's near-real-time search explanation connects visibility to Lucene segments and the filesystem cache. If a test writes and immediately searches, its refresh assumptions belong in the test.

Segments are immutable. Updates introduce a new document version and mark the older one deleted; merges can later reclaim that space. More segments mean more structures to manage, but “force merge every busy index at night” is not a sensible consequence. Elastic recommends force merge for read-only indexes, and documents substantial temporary disk requirements. It is an operational decision, not a routine fix for a slow query.

Reproduce one distinction at a time

Download the model and tests. Run python3 test_search_model.py, then python3 search_model.py. Ten tests passed on Python 3.11.5. The demo prints all three IDs for the term intersection and only ID one for the phrase.

There is no Elasticsearch server, refresh simulation, segment merger or distributed query execution in these checks. Those mechanisms are explained from primary documentation. To extend the experiment, index the same three texts into a disposable Elasticsearch index with an explicit mapping; compare match, match_phrase, _analyze and _explain responses.

That gives debugging a concrete order: inspect the stored terms, check candidate selection, examine score contributions, then investigate visibility and distributed execution. Each step asks a smaller question than “why is search wrong?”

Found a mistake or tried a different approach?

Send Alex a note ↗