Skip to main content

Fields API

Define the structure of your index with field types.

Schema​

class whoosh.fields.Schema

The Schema class defines the fields available in an index.

Constructor​

schema = Schema(
title=TEXT(stored=True),
content=TEXT,
path=ID(stored=True, unique=True),
tags=KEYWORD(lowercase=True),
rating=NUMERIC(float, stored=True),
published=DATETIME(stored=True),
active=BOOLEAN
)

Methods​

add()​

schema.add(
fieldname: str,
fieldtype,
glob: bool = False,
**kwargs
)

Add a field to the schema. If glob=True, the fieldname is treated as a glob pattern.

remove()​

schema.remove(fieldname: str, **kwargs)

Remove a field from the schema.

items()​

for name, field in schema.items():
print(name, field)

Return a list of (fieldname, field object) pairs.

names()​

names = schema.names()

Return a list of field names.

FieldType Base Class​

class whoosh.fields.FieldType

Base class for all field types.

Attributes​

AttributeTypeDescription
formatFormatDefines how the field is indexed
vectorFormat or NoneOptional per-document vector format
scorableboolWhether field length is stored (for BM25F)
storedboolWhether field value is stored in index
uniqueboolWhether field uniquely identifies documents

Methods​

index()​

Convert a value into indexed items.

indexable()​

Check if the value can be indexed.

spelling_fieldname()​

Return the field name used for spelling data.

spellable_words()​

Generate spellable words from a value.

Built-in Field Types​

TEXT​

whoosh.fields.TEXT(
stored: bool = False,
unique: bool = False,
phrase: bool = True,
analyzer: Analyzer = None,
field_boost: float = 1.0,
**kwargs
)

Full-text field with tokenization and optional phrase search.

Example:

title = TEXT(stored=True)
body = TEXT(analyzer=StemmingAnalyzer(), phrase=False)

ID​

whoosh.fields.ID(
stored: bool = False,
unique: bool = False,
field_boost: float = 1.0,
**kwargs
)

Untokenized identifier field. Stores the entire value as a single term.

Example:

path = ID(stored=True, unique=True)
slug = ID(stored=True)

KEYWORD​

whoosh.fields.KEYWORD(
stored: bool = False,
lowercase: bool = False,
commas: bool = False,
scorable: bool = False,
field_boost: float = 1.0,
**kwargs
)

Space or comma-separated keywords. Phrase search is not supported.

Example:

tags = KEYWORD(lowercase=True, commas=True, stored=True)

STORED​

whoosh.fields.STORED(
stored: bool = True,
unique: bool = False,
**kwargs
)

Stored-only field. Not indexed or searchable.

Example:

icon = STORED()
description = STORED()

NUMERIC​

whoosh.fields.NUMERIC(
numtype: type = int,
stored: bool = False,
unique: bool = False,
field_boost: float = 1.0,
**kwargs
)

Numeric field for integers or floats.

Example:

rating = NUMERIC(float, stored=True)
count = NUMERIC(int)
price = NUMERIC(float, stored=True, sortable=True)

DATETIME​

whoosh.fields.DATETIME(
stored: bool = False,
unique: bool = False,
field_boost: float = 1.0,
**kwargs
)

Date/time field. Stores datetime objects.

Example:

published = DATETIME(stored=True)
updated = DATETIME()

BOOLEAN​

whoosh.fields.BOOLEAN(
stored: bool = False,
unique: bool = False,
field_boost: float = 1.0,
**kwargs
)

Boolean field. Searchable with yes, no, true, false, 1, 0, t, f.

Example:

published = BOOLEAN(stored=True)

NGRAM​

whoosh.fields.NGRAM(
minsize: int = 2,
maxsize: int = 5,
stored: bool = False,
field_boost: float = 1.0,
**kwargs
)

Character n-gram field.


NGRAMWORDS​

whoosh.fields.NGRAMWORDS(
minsize: int = 2,
maxsize: int = 5,
stored: bool = False,
field_boost: float = 1.0,
**kwargs
)

Word-level n-gram field.


VectorField​

whoosh.fields.VectorField(
dimensions: int,
metric: str = "cosine",
provider: str = "numpy",
stored: bool = False,
**kwargs
)

Field for storing and searching vector embeddings.

Args:

  • dimensions (int): Embedding dimension (e.g., 384 for all-MiniLM-L6-v2).
  • metric (str): Similarity metric: "cosine", "euclidean", "dot".
  • provider (str): Vector provider name from registry.

Example:

embedding = VectorField(dimensions=384, metric="cosine", stored=True)

SchemaBuilder​

Fluent API for building schemas:

from whoosh.fields import SchemaBuilder

schema = (
SchemaBuilder()
.field("title", TEXT(stored=True))
.field("path", ID(stored=True, unique=True))
.field("content", TEXT)
.field("tags", KEYWORD(lowercase=True))
.field("published", DATETIME(stored=True))
.build()
)

Constants​

  • whoosh.fields.STORED: Stored-only field type
  • whoosh.fields.TEXT: Full-text field
  • whoosh.fields.ID: Identifier field
  • whoosh.fields.KEYWORD: Keyword field
  • whoosh.fields.NUMERIC: Numeric field
  • whoosh.fields.DATETIME: Date/time field
  • whoosh.fields.BOOLEAN: Boolean field