Core Concepts
Whoosh-NG is a pure-Python search engine library. This guide explains the main concepts you need to understand to use it effectively.
Architecture
Whoosh-NG follows a layered architecture:
Key Components
Index
An Index is the top-level container for your searchable documents. It manages one or more segments on disk.
from whoosh.index import create_in, open_dir
# Create a new index
ix = create_in("indexdir", schema)
# Open an existing index
ix = open_dir("indexdir")
Schema
The Schema defines the fields that documents in your index can have. Each field has a type that determines how it is indexed and stored.
from whoosh.fields import Schema, TEXT, ID, NUMERIC
schema = Schema(
title=TEXT(stored=True),
path=ID(stored=True, unique=True),
content=TEXT,
rating=NUMERIC(float, stored=True)
)
Writer
An IndexWriter lets you add, update, and delete documents in the index.
writer = ix.writer()
writer.add_document(title="Hello", content="World")
writer.commit()
Searcher
A Searcher lets you query the index and retrieve results.
with ix.searcher() as s:
results = s.search("hello")
Query Parser
The QueryParser converts a query string into a query object that the searcher can execute.
from whoosh.qparser import QueryParser
qp = QueryParser("content", schema)
query = qp.parse("hello world")
Modern Features
Plugin System
Plugins extend Whoosh-NG without modifying the core. Plugins can:
- Register new vector providers
- Add FastAPI endpoints
- Provide custom analyzers
- Hook into the middleware pipeline
from whoosh.plugins.manager import PluginManager
# Load plugins from entry points
PluginManager.load_plugins()
# Or register manually
PluginManager.register("my_plugin", MyPlugin())
Middleware Pipeline
Middleware intercepts indexing and search operations:
from whoosh.middleware import Middleware, MiddlewareContext
class LoggingMiddleware(Middleware):
def before_search(self, context: MiddlewareContext):
print(f"Searching: {context.query}")
return context
def after_search(self, context: MiddlewareContext):
print(f"Found: {len(context.results) if context.results else 0} results")
return context
Vector Search
Vector fields enable semantic search using embeddings:
from whoosh.fields import Schema, TEXT, VectorField
schema = Schema(
content=TEXT,
embedding=VectorField(dimensions=384)
)
Event Bus
The event system allows loose coupling between components:
from whoosh.event_bus import EventBus, DocumentIndexed
bus = EventBus()
@bus.subscribe
def on_document_indexed(event: DocumentIndexed):
print(f"Document indexed: {event.docnum}")
Data Flow
Indexing Flow
- Application calls
writer.add_document() - Schema validates and analyzes fields
- Middleware
before_indexhooks run - Document is written to segment
- Middleware
after_indexhooks run DocumentIndexedevent is publishedcommit()merges segments and writes TOC
Search Flow
- Application calls
searcher.search(query) - Query is parsed into query tree
- Middleware
before_searchhooks run - Searcher executes query against segments
- Results are scored and sorted
- Middleware
after_searchhooks run SearchExecutedevent is published- Results are returned to application
Design Principles
- Composability: Components combine via
|and+operators - Zero-cost abstractions: No middleware = no overhead
- Sync-first: Core is synchronous; async is opt-in
- Plugin isolation: Plugins cannot break the core
- Type safety: Comprehensive type hints throughout