Middleware
The middleware pipeline allows you to intercept and modify indexing and search operations. It is the primary extension mechanism for cross-cutting concerns like logging, caching, metrics, and security.
Core Concepts
A middleware is a class that implements hooks into the indexing and search lifecycle:
from whoosh.middleware.base import Middleware
from whoosh.middleware.context import MiddlewareContext
class MyMiddleware(Middleware):
def before_search(self, context: MiddlewareContext) -> MiddlewareContext:
# Modify context.query or context.metadata
return context
def after_search(self, context: MiddlewareContext) -> MiddlewareContext:
# Access context.results
return context
Available Hooks
| Hook | When | Common Uses |
|---|---|---|
startup(context) | Middleware initialized | Open connections, warm caches |
shutdown(context) | Middleware torn down | Close connections, flush buffers |
before_index(context) | Before document added | Validation, enrichment, compression flags |
after_index(context) | After document added | Metrics, events, cache invalidation |
before_delete(context) | Before document deleted | Audit logging, access control |
after_delete(context) | After document deleted | Metrics, cache invalidation |
before_search(context) | Before query executes | Query rewriting, caching, auth |
after_search(context) | After results returned | Logging, metrics, result modification |
on_error(context, exc) | On exception | Error handling, fallbacks |
on_commit(context) | After commit | Metrics, notifications |
Built-in Middlewares
MetricsMiddleware
Tracks basic statistics:
from whoosh.middleware import MetricsMiddleware
metrics = MetricsMiddleware()
# After operations:
stats = metrics.get_metrics()
# Returns: {"documents_indexed": N, "searches_executed": N}
CacheMiddleware
Caches search results in memory:
from whoosh.middleware import CacheMiddleware
cache = CacheMiddleware()
# Check cache
cached = cache.get_cached("user query string")
# Store manually
cache.set_cached("user query string", results)
CompressionMiddleware
Marks documents for compression at the backend level:
from whoosh.middleware import CompressionMiddleware
compression = CompressionMiddleware()
# Sets document["_compressed"] = True
EncryptionMiddleware
Marks documents for encryption at the backend level:
from whoosh.middleware import EncryptionMiddleware
encryption = EncryptionMiddleware()
# Sets document["_encrypted"] = True
MiddlewareChain
Orchestrates middleware execution:
from whoosh.middleware import MiddlewareChain
chain = MiddlewareChain([
MetricsMiddleware(),
CacheMiddleware()
])
# Execute before hook
context = MiddlewareContext("search")
context.query = "test"
context = chain.run_before("before_search", context)
# ... core operation ...
# Execute after hook
context = chain.run_after("after_search", context)
Execution Order
before_*hooks run in registration orderafter_*hooks run in reverse order- If a hook raises
StopOperation, the pipeline aborts - If
fail_open=False, exceptions propagate immediately
Integration
With Writer
from whoosh.middleware.integration import apply_middleware_to_writer
writer = apply_middleware_to_writer(ix.writer(), chain.middlewares)
with writer:
writer.add_document(title="Hello", content="World")
With Searcher
from whoosh.middleware.integration import apply_middleware_to_searcher
searcher = apply_middleware_to_searcher(ix.searcher(), chain.middlewares)
results = searcher.search("query")
With PluginManager
from whoosh.plugins.manager import PluginManager
# Plugins can provide middleware
PluginManager.load_plugins()
chain = PluginManager.get_middleware_chain()
Custom Middleware Example
class RequestLoggingMiddleware(Middleware):
"""Log all search requests."""
def before_search(self, context: MiddlewareContext) -> MiddlewareContext:
context.metadata["request_id"] = generate_request_id()
logger.info(f"Search: {context.query}")
return context
def after_search(self, context: MiddlewareContext) -> MiddlewareContext:
logger.info(f"Found {len(context.results)} results")
return context
class RateLimitMiddleware(Middleware):
"""Abort searches exceeding rate limit."""
def before_search(self, context: MiddlewareContext) -> MiddlewareContext:
if not rate_limiter.allow(context):
raise StopOperation("Rate limit exceeded")
return context
class QueryEnrichmentMiddleware(Middleware):
"""Add synonyms to the query."""
def before_search(self, context: MiddlewareContext) -> MiddlewareContext:
if context.query:
context.query += " " + get_synonyms(context.query)
return context
Error Handling
class ResilientMiddleware(Middleware):
"""Continue on non-critical errors."""
def after_search(self, context: MiddlewareContext) -> MiddlewareContext:
try:
send_to_analytics(context.results)
except Exception:
# Log but don't fail the search
logger.warning("Analytics failed", exc_info=True)
return context
Best Practices
- Stateless: Use
context.metadatafor per-request data - Fail fast: Only use
fail_open=Truefor non-critical middleware - Order matters: Place caching before metrics, auth before routing
- Performance: Keep hooks lightweight; use async for I/O
- Testing: Mock the context object to test middleware in isolation
Provider Integration via Middleware
Many Whoosh-NG providers integrate into the indexing and search pipeline through middleware hooks. This is the standard pattern for cross-cutting concerns that need to transform documents, queries, or results.
Provider-to-Middleware Mapping
| Provider | Middleware | Hooks Used | Purpose |
|---|---|---|---|
StorageProvider | StorageMiddleware | before_index, on_commit | Tags context with storage backend; writes commit checkpoints |
StemmerProvider | StemmingMiddleware | before_index, before_search | Stems document fields and query text |
SynonymProvider | SynonymExpansionMiddleware | before_index, before_search | Expands documents and queries with synonyms |
VectorProvider | (built into Whoosh core) | N/A (segment format) | Registered in VectorRegistry; resolved at search time from segment metadata |
AutocompleteProvider | (standalone or registry) | N/A | Used directly via .search() or via AutocompleteRegistry |
How providers flow through the pipeline
┌──────────────────┐
│ PluginManager │
│ .load_plugins() │
└────────┬─────────┘
│
┌──────────────┼──────────────┐
│ │ │
┌────────▼──────┐ ┌────▼─────┐ ┌──────▼────────┐
│ VectorPlugin │ │AutoPlugin│ │ Other Plugins │
│ │ │ │ │ │
│ VectorRegistry│ │AutoReg. │ │ ProviderReg. │
│ .register() │ │.register│ │ .register() │
└───────────────┘ └─────────┘ └───────────────┘
┌──────────────────┐
│ MiddlewareChain │
│ │
│ ┌──────────────┐ │
│ │ before_index │ │
│ │ hooks │ │
│ │ (ordered) │ │
│ └──────┬───────┘ │
│ │ │
│ ┌──────▼───────┐ │
│ │ Writer │ │
│ │ .add_doc() │ │
│ └──────┬───────┘ │
│ │ │
│ ┌──────▼───────┐ │
│ │ on_commit │ │
│ │ hooks │ │
│ └──────────────┘ │
└──────────────────┘
┌──────────────────┐
│ Searcher │
│ │
│ ┌──────────────┐ │
│ │before_search │ │
│ │ hooks │ │
│ └──────┬───────┘ │
│ │ │
│ ┌──────▼───────┐ │
│ │ Query │ │
│ │ execution │ │
│ └──────┬───────┘ │
│ │ │
│ ┌──────▼───────┐ │
│ │ after_search │ │
│ │ hooks │ │
│ └──────────────┘ │
└──────────────────┘
Example: Full provider pipeline
from whoosh.middleware.chain import MiddlewareChain
from whoosh.middleware.wrappers import MiddlewareWriter, MiddlewareSearcher
from whoosh_modern.middleware import (
StorageMiddleware,
StemmingMiddleware,
FileStorageProvider,
)
from whoosh_modern.linguistics.synonyms import (
SynonymExpansionMiddleware,
SynonymManager,
)
from whoosh_modern.analysis import get_stemmer
# 1. Build providers
storage = FileStorageProvider("/data/index")
stemmer = get_stemmer("auto", "english")
syn_manager = SynonymManager({"car": ["automobile", "vehicle"]})
# 2. Build middleware chain
chain = MiddlewareChain([
StorageMiddleware(storage, name="primary"),
StemmingMiddleware(stemmer=stemmer.stem, fields=["title", "content"]),
SynonymExpansionMiddleware(syn_manager),
])
# 3. Index with middleware
with MiddlewareWriter(ix.writer(), chain) as writer:
writer.add_document(title="Car for sale", content="House near beach")
# StorageMiddleware.before_index() tags context
# StemmingMiddleware stems "Car" → "car"
# SynonymExpansionMiddleware expands "Car" → "Car automobile vehicle"
writer.commit()
# StorageMiddleware.on_commit() writes checkpoint
# 4. Search with middleware
with MiddlewareSearcher(ix.searcher(), chain) as searcher:
# StemmingMiddleware stems query "cars" → "car"
# SynonymExpansionMiddleware expands "car" → "car automobile vehicle"
results = searcher.search("cars")
Key insight: middleware as provider adapter
Middleware acts as the adapter between providers and Whoosh's core pipeline:
Provider (domain logic)
│
▼
Middleware (pipeline integration)
│
▼
Whoosh core (generic engine)
- StorageProvider → StorageMiddleware → Whoosh writer/searcher
- StemmerProvider → StemmingMiddleware → Whoosh document/query
- SynonymProvider → SynonymExpansionMiddleware → Whoosh text fields
This separation keeps providers simple (single responsibility) while middleware handles the lifecycle integration (when and how to apply the provider).
Modern Middleware (Whoosh-NG 2.0)
Whoosh-NG 2.0 adds a modern middleware package (whoosh_modern.middleware) with wrap-style resilience middleware (retry, caching, logging) and hook-based middleware for storage, search, and analysis. For full details on the modern middleware architecture, plugin integration, and deployment, see the Middleware & Plugin Pipeline Guide.