Skip to main content

Middleware

The middleware pipeline allows you to intercept and modify indexing and search operations. It is the primary extension mechanism for cross-cutting concerns like logging, caching, metrics, and security.

Core Concepts​

A middleware is a class that implements hooks into the indexing and search lifecycle:

from whoosh.middleware.base import Middleware
from whoosh.middleware.context import MiddlewareContext

class MyMiddleware(Middleware):
def before_search(self, context: MiddlewareContext) -> MiddlewareContext:
# Modify context.query or context.metadata
return context

def after_search(self, context: MiddlewareContext) -> MiddlewareContext:
# Access context.results
return context

Available Hooks​

HookWhenCommon Uses
startup(context)Middleware initializedOpen connections, warm caches
shutdown(context)Middleware torn downClose connections, flush buffers
before_index(context)Before document addedValidation, enrichment, compression flags
after_index(context)After document addedMetrics, events, cache invalidation
before_delete(context)Before document deletedAudit logging, access control
after_delete(context)After document deletedMetrics, cache invalidation
before_search(context)Before query executesQuery rewriting, caching, auth
after_search(context)After results returnedLogging, metrics, result modification
on_error(context, exc)On exceptionError handling, fallbacks
on_commit(context)After commitMetrics, notifications

Built-in Middlewares​

MetricsMiddleware​

Tracks basic statistics:

from whoosh.middleware import MetricsMiddleware

metrics = MetricsMiddleware()
# After operations:
stats = metrics.get_metrics()
# Returns: {"documents_indexed": N, "searches_executed": N}

CacheMiddleware​

Caches search results in memory:

from whoosh.middleware import CacheMiddleware

cache = CacheMiddleware()

# Check cache
cached = cache.get_cached("user query string")

# Store manually
cache.set_cached("user query string", results)

CompressionMiddleware​

Marks documents for compression at the backend level:

from whoosh.middleware import CompressionMiddleware

compression = CompressionMiddleware()
# Sets document["_compressed"] = True

EncryptionMiddleware​

Marks documents for encryption at the backend level:

from whoosh.middleware import EncryptionMiddleware

encryption = EncryptionMiddleware()
# Sets document["_encrypted"] = True

MiddlewareChain​

Orchestrates middleware execution:

from whoosh.middleware import MiddlewareChain

chain = MiddlewareChain([
MetricsMiddleware(),
CacheMiddleware()
])

# Execute before hook
context = MiddlewareContext("search")
context.query = "test"
context = chain.run_before("before_search", context)

# ... core operation ...

# Execute after hook
context = chain.run_after("after_search", context)

Execution Order​

  • before_* hooks run in registration order
  • after_* hooks run in reverse order
  • If a hook raises StopOperation, the pipeline aborts
  • If fail_open=False, exceptions propagate immediately

Integration​

With Writer​

from whoosh.middleware.integration import apply_middleware_to_writer

writer = apply_middleware_to_writer(ix.writer(), chain.middlewares)

with writer:
writer.add_document(title="Hello", content="World")

With Searcher​

from whoosh.middleware.integration import apply_middleware_to_searcher

searcher = apply_middleware_to_searcher(ix.searcher(), chain.middlewares)
results = searcher.search("query")

With PluginManager​

from whoosh.plugins.manager import PluginManager

# Plugins can provide middleware
PluginManager.load_plugins()
chain = PluginManager.get_middleware_chain()

Custom Middleware Example​

class RequestLoggingMiddleware(Middleware):
"""Log all search requests."""

def before_search(self, context: MiddlewareContext) -> MiddlewareContext:
context.metadata["request_id"] = generate_request_id()
logger.info(f"Search: {context.query}")
return context

def after_search(self, context: MiddlewareContext) -> MiddlewareContext:
logger.info(f"Found {len(context.results)} results")
return context

class RateLimitMiddleware(Middleware):
"""Abort searches exceeding rate limit."""

def before_search(self, context: MiddlewareContext) -> MiddlewareContext:
if not rate_limiter.allow(context):
raise StopOperation("Rate limit exceeded")
return context

class QueryEnrichmentMiddleware(Middleware):
"""Add synonyms to the query."""

def before_search(self, context: MiddlewareContext) -> MiddlewareContext:
if context.query:
context.query += " " + get_synonyms(context.query)
return context

Error Handling​

class ResilientMiddleware(Middleware):
"""Continue on non-critical errors."""

def after_search(self, context: MiddlewareContext) -> MiddlewareContext:
try:
send_to_analytics(context.results)
except Exception:
# Log but don't fail the search
logger.warning("Analytics failed", exc_info=True)
return context

Best Practices​

  1. Stateless: Use context.metadata for per-request data
  2. Fail fast: Only use fail_open=True for non-critical middleware
  3. Order matters: Place caching before metrics, auth before routing
  4. Performance: Keep hooks lightweight; use async for I/O
  5. Testing: Mock the context object to test middleware in isolation

Provider Integration via Middleware​

Many Whoosh-NG providers integrate into the indexing and search pipeline through middleware hooks. This is the standard pattern for cross-cutting concerns that need to transform documents, queries, or results.

Provider-to-Middleware Mapping​

ProviderMiddlewareHooks UsedPurpose
StorageProviderStorageMiddlewarebefore_index, on_commitTags context with storage backend; writes commit checkpoints
StemmerProviderStemmingMiddlewarebefore_index, before_searchStems document fields and query text
SynonymProviderSynonymExpansionMiddlewarebefore_index, before_searchExpands documents and queries with synonyms
VectorProvider(built into Whoosh core)N/A (segment format)Registered in VectorRegistry; resolved at search time from segment metadata
AutocompleteProvider(standalone or registry)N/AUsed directly via .search() or via AutocompleteRegistry

How providers flow through the pipeline​

┌──────────────────┐
│ PluginManager │
│ .load_plugins() │
└────────┬─────────┘
│
┌──────────────┼──────────────┐
│ │ │
┌────────▼──────┐ ┌────▼─────┐ ┌──────▼────────┐
│ VectorPlugin │ │AutoPlugin│ │ Other Plugins │
│ │ │ │ │ │
│ VectorRegistry│ │AutoReg. │ │ ProviderReg. │
│ .register() │ │.register│ │ .register() │
└───────────────┘ └─────────┘ └───────────────┘

┌──────────────────┐
│ MiddlewareChain │
│ │
│ ┌──────────────┐ │
│ │ before_index │ │
│ │ hooks │ │
│ │ (ordered) │ │
│ └──────┬───────┘ │
│ │ │
│ ┌──────▼───────┐ │
│ │ Writer │ │
│ │ .add_doc() │ │
│ └──────┬───────┘ │
│ │ │
│ ┌──────▼───────┐ │
│ │ on_commit │ │
│ │ hooks │ │
│ └──────────────┘ │
└──────────────────┘

┌──────────────────┐
│ Searcher │
│ │
│ ┌──────────────┐ │
│ │before_search │ │
│ │ hooks │ │
│ └──────┬───────┘ │
│ │ │
│ ┌──────▼───────┐ │
│ │ Query │ │
│ │ execution │ │
│ └──────┬───────┘ │
│ │ │
│ ┌──────▼───────┐ │
│ │ after_search │ │
│ │ hooks │ │
│ └──────────────┘ │
└──────────────────┘

Example: Full provider pipeline​

from whoosh.middleware.chain import MiddlewareChain
from whoosh.middleware.wrappers import MiddlewareWriter, MiddlewareSearcher
from whoosh_modern.middleware import (
StorageMiddleware,
StemmingMiddleware,
FileStorageProvider,
)
from whoosh_modern.linguistics.synonyms import (
SynonymExpansionMiddleware,
SynonymManager,
)
from whoosh_modern.analysis import get_stemmer

# 1. Build providers
storage = FileStorageProvider("/data/index")
stemmer = get_stemmer("auto", "english")
syn_manager = SynonymManager({"car": ["automobile", "vehicle"]})

# 2. Build middleware chain
chain = MiddlewareChain([
StorageMiddleware(storage, name="primary"),
StemmingMiddleware(stemmer=stemmer.stem, fields=["title", "content"]),
SynonymExpansionMiddleware(syn_manager),
])

# 3. Index with middleware
with MiddlewareWriter(ix.writer(), chain) as writer:
writer.add_document(title="Car for sale", content="House near beach")
# StorageMiddleware.before_index() tags context
# StemmingMiddleware stems "Car" → "car"
# SynonymExpansionMiddleware expands "Car" → "Car automobile vehicle"
writer.commit()
# StorageMiddleware.on_commit() writes checkpoint

# 4. Search with middleware
with MiddlewareSearcher(ix.searcher(), chain) as searcher:
# StemmingMiddleware stems query "cars" → "car"
# SynonymExpansionMiddleware expands "car" → "car automobile vehicle"
results = searcher.search("cars")

Key insight: middleware as provider adapter​

Middleware acts as the adapter between providers and Whoosh's core pipeline:

Provider (domain logic)
│
▼
Middleware (pipeline integration)
│
▼
Whoosh core (generic engine)
  • StorageProvider → StorageMiddleware → Whoosh writer/searcher
  • StemmerProvider → StemmingMiddleware → Whoosh document/query
  • SynonymProvider → SynonymExpansionMiddleware → Whoosh text fields

This separation keeps providers simple (single responsibility) while middleware handles the lifecycle integration (when and how to apply the provider).

Modern Middleware (Whoosh-NG 2.0)​

Whoosh-NG 2.0 adds a modern middleware package (whoosh_modern.middleware) with wrap-style resilience middleware (retry, caching, logging) and hook-based middleware for storage, search, and analysis. For full details on the modern middleware architecture, plugin integration, and deployment, see the Middleware & Plugin Pipeline Guide.