braindecode.models.TFMTokenizer#

class braindecode.models.TFMTokenizer(n_outputs=None, n_chans=None, chs_info=None, n_times=None, input_window_seconds=None, sfreq=200, embed_dim=64, codebook_size=8192, freq_patch_size=5, freq_encoder_depth=2, temporal_encoder_depth=2, decoder_depth=8, max_seq_len=1024, commitment_cost=1.0, activation=<class 'torch.nn.modules.activation.GELU'>, drop_prob=0.2)[source]#

Time-Frequency Motif (TFM) tokenizer from Pradeepkumar et al. (2026) [tfm2026].

Foundation Model Attention/Transformer

TFM-Tokenizer overview (Pradeepkumar et al., 2026, Fig. 2).

Each channel is tokenized independently. A frequency path (patch convolutions and a linear-attention transformer over the STFT magnitude of each frame) and a temporal path (strided convolution over the raw signal) are concatenated, contextualized by a temporal transformer and quantized by an EMA codebook. A decoder reconstructs the unmasked spectrogram from the quantized tokens. tokenize() returns every output; forward() returns the reconstruction only.

Parameters:
  • n_outputs (int | None) – Number of outputs of the model. This is the number of classes in the case of classification.

  • n_chans (int | None) – Number of EEG channels.

  • chs_info (list of dict) – Information about each individual EEG channel. This should be filled with info["chs"]. Refer to mne.Info for more details.

  • n_times (int | None) – Number of time samples of the input window.

  • input_window_seconds (float | None) – Length of the input window in seconds.

  • sfreq (float) – Sampling frequency of the EEG recordings.

  • embed_dim (int) – Embedding size; a multiple of 8 and of twice the number of frequency groups.

  • codebook_size (int) – Number of motifs in the codebook.

  • freq_patch_size (int) – Width and stride of the frequency patches; sfreq // 2 must be divisible by its square.

  • freq_encoder_depth (int) – Linear-attention blocks of the frequency encoder.

  • temporal_encoder_depth (int) – Linear-attention blocks of the temporal encoder.

  • decoder_depth (int) – Linear-attention blocks of the decoder.

  • max_seq_len (int) – Maximum number of frames.

  • commitment_cost (float) – Weight of the commitment loss.

  • activation (type[Module]) – Activation of the convolutional blocks.

  • drop_prob (float) – Dropout of the attention blocks.

Raises:

ValueError – If some input signal-related parameters are not specified: and can not be inferred.

Notes

The STFT uses a one-second Hann window and a half-second hop, so sfreq must be an even integer (200 Hz for the released weights) and a window of n_times samples gives 1 + (n_times - sfreq) // (sfreq // 2) tokens per channel.

Training differs from the reference code in one place; inference does not. The codebook term of the quantization loss carries no gradient (the codebook moves by EMA only), so at commitment_cost=1 the encoder still receives the commitment gradient (in the reference both terms use the straight-through tensor and their encoder gradients cancel).

Tokens from this class are identical to the reference tokenizer’s. With them, on CHB-MIT, the authors’ released fine-tuned classifier gives a balanced accuracy of 0.619 and our retraining with the authors’ code gives 0.611 ± 0.034 (10 seeds); the paper reports 0.647 ± 0.015.

Important

Pre-trained Weights Available

The authors’ tokenizer weights (MIT) are on the Hugging Face Hub at Jathurshan/TFM-Tokenizer under the reference module names; rename the keys to load them:

import torch
from huggingface_hub import hf_hub_download
from braindecode.models import TFMTokenizer

path = hf_hub_download(
    "Jathurshan/TFM-Tokenizer",
    "multiple_dataset_settings/Pretrained_tfm_tokenizer_2x2x8/"
    "tfm_tokenizer_last.pth",
)
renames = {
    "trans_freq_encoder.transformer.": "frequency_encoder.",
    "trans_temporal_encoder.transformer.": "temporal_encoder.",
    "trans_decoder.transformer.": "decoder.",
    "freq_patch_embedding_2_atten.": "frequency_attention.",
    "freq_patch_embedding_2.0.": "frequency_projection.",
    "freq_patch_embedding.": "frequency_patch_embedding.",
    "decoder.": "final_layer.",
    "quantizer.embedding.weight": "quantizer.embed",
    "quantizer.ema_w": "quantizer.embed_avg",
}
state = {
    next(
        (new + k[len(old) :] for old, new in renames.items() if k.startswith(old)), k
    ): v
    for k, v in torch.load(path, map_location="cpu").items()
}
state["quantizer.inited"] = torch.ones(1)
model = TFMTokenizer(sfreq=200)
model.load_state_dict(state)

Added in version 1.9.

Examples

>>> import torch
>>> from braindecode.models import TFMTokenizer
>>> model = TFMTokenizer(sfreq=200, codebook_size=256)
>>> x = torch.randn(2, 8, 1000)
>>> spec = model.compute_spectrogram(x)
>>> mask_a, mask_b = model.make_complementary_masks(spec)
>>> out = model.tokenize(x, spectrogram_mask=mask_a)
>>> out.token_ids.shape
torch.Size([2, 8, 9])
>>> loss = (out.reconstruction - spec).square().mean() + out.quantization_loss

References

[tfm2026]

Pradeepkumar, J., Piao, X., Chen, Z., Sun, J. (2026). Tokenizing Single-Channel EEG with Time-Frequency Motif Learning. ICLR 2026. https://openreview.net/forum?id=2sPmWHZ8Ir

Hugging Face Hub integration

When the optional huggingface_hub package is installed, all models automatically gain the ability to be pushed to and loaded from the Hugging Face Hub. Install with:

pip install braindecode[hub]

Pushing a model to the Hub:

from braindecode.models import TFMTokenizer

# Train your model
model = TFMTokenizer(n_chans=22, n_outputs=4, n_times=1000)
# ... training code ...

# Push to the Hub
model.push_to_hub(
    repo_id="username/my-tfmtokenizer-model",
    commit_message="Initial model upload",
)

Loading a model from the Hub:

from braindecode.models import TFMTokenizer

# Load pretrained model
model = TFMTokenizer.from_pretrained("username/my-tfmtokenizer-model")

# Load with a different number of outputs (head is rebuilt automatically)
model = TFMTokenizer.from_pretrained("username/my-tfmtokenizer-model", n_outputs=4)

Extracting features and replacing the head:

import torch

x = torch.randn(1, model.n_chans, model.n_times)
# Extract encoder features (consistent dict across all models)
out = model(x, return_features=True)
features = out["features"]

# Replace the classification head
model.reset_head(n_outputs=10)

Saving and restoring full configuration:

import json

config = model.get_config()            # all __init__ params
with open("config.json", "w") as f:
    json.dump(config, f)

model2 = TFMTokenizer.from_config(config)    # reconstruct (no weights)

All model parameters (both EEG-specific and model-specific such as dropout rates, activation functions, number of filters) are automatically saved to the Hub and restored when loading.

See Loading and Adapting Pretrained Foundation Models for a complete tutorial.

Methods

compute_spectrogram(x)[source]#

STFT magnitude of x (batch, channels, samples).

Returns a (batch, channels, n_freqs, n_frames) tensor.

Parameters:

x (Tensor) – The description is missing.

Return type:

Tensor

encode(x, spectrogram)[source]#

Embed x and its (batch * channels, n_freqs, n_frames) spectrogram.

Parameters:
  • x (Tensor) – The description is missing.

  • spectrogram (Tensor) – The description is missing.

Return type:

Tensor

forward(x, spectrogram_mask=None)[source]#

Return the reconstructed spectrogram (see tokenize()).

Parameters:
  • x (Tensor) – The description is missing.

  • spectrogram_mask (Tensor | None) – The description is missing.

Return type:

Tensor

make_complementary_masks(spectrogram, freq_mask_ratio=0.5, freq_bin_size=5, time_mask_ratio=0.5, time_bin_size=1)[source]#

Return a random time-frequency mask and its complement.

Groups of freq_bin_size bins and time_bin_size frames are hidden at the given ratios. One mask, shaped like spectrogram, is shared by every trial and channel, as in the reference.

Parameters:
  • spectrogram (Tensor) – The description is missing.

  • freq_mask_ratio (float) – The description is missing.

  • freq_bin_size (int) – The description is missing.

  • time_mask_ratio (float) – The description is missing.

  • time_bin_size (int) – The description is missing.

Return type:

tuple[Tensor, Tensor]

tokenize(x, spectrogram_mask=None)[source]#

Tokenize x and reconstruct its spectrogram.

Parameters:
  • x (Tensor) – EEG (batch, channels, samples) sampled at sfreq.

  • spectrogram_mask (Tensor | None) – Boolean mask shaped like compute_spectrogram() output; True bins are visible to the encoder. Default: all visible.

Return type:

TFMTokenizerOutput