BOLD
Multi-echo EPI at TR 2.0 s, in MNI152 volumes and on the fsaverage5 surface.
sub-0_ses-11_task-audiodec_run-0_bold.nii.gzOpen fMRI dataset · naturalistic speech · French podcasts
A deep fMRI dataset for decoding naturalistic speech in the human brain.
Replay
Whole-brain BOLD from one session, played back in sync with the podcast that produced it: the words, their translation, and the five seconds it takes blood to answer them.
From volume to film
Five steps take a stack of voxels to the flat, frame-by-frame view used in the player above. Scroll through them.
The scanner samples the head as a stack of slices made of 2 mm voxels. Each stack is one frame of BOLD: blood oxygenation, measured everywhere at once.
Speech is processed in the cortex, a folded sheet a few millimetres thick. The signal is sampled on that sheet at 20,484 points, the fsaverage5 surface.
About two thirds of the cortex is hidden inside the sulci. Inflating the surface brings it into view: light patches are gyri, dark ones are sulci.
Cut along the medial wall and a few relaxation lines, each hemisphere lies flat. The whole cortex now fits in one frame, like a film still.
In sequence, the frames become a film of listening. Sound reaches the auditory cortex first, then speech, words and meaning spread along the temporal lobe into frontal and parietal cortex.
The BOLD signal
fMRI does not see neurons. It sees the oxygenated blood that rushes in to refuel them, and that takes time: the response to a word peaks about five seconds after it is spoken, then settles over the next fifteen.
Every word in the dataset is time-stamped so that models can learn which words each frame is answering. Toggle the words, change the TR, and watch what the scanner gets to measure.
Design
Most naturalistic fMRI datasets scan many people for an hour or two. Cephalonauts One scans six people for about fifty hours each, so that a model can be trained on a single brain.
Fifty hours per brain is enough to fit an encoding or decoding model to one person, instead of averaging across many.
Contents
Each run ships with the audio that was played, every word and when it was heard, a translation, model features, and the questions the participant answered afterwards.
Multi-echo EPI at TR 2.0 s, in MNI152 volumes and on the fsaverage5 surface.
sub-0_ses-11_task-audiodec_run-0_bold.nii.gzThe exact waveform played in the scanner, aligned to the first trigger.
EP-0217_part0.wavEvery word with its onset and offset, in seconds from the first trigger.
EP-0217_part0_words.tsvAn English translation aligned sentence by sentence to the French.
EP-0217_part0_translation.tsvText and audio embeddings for every word, resampled to every TR.
EP-0217_part0_desc-text_embeddings.npyThree questions after each part check that the participant was listening.
EP-0217_part0_questions.jsonimport nibabel as nib import numpy as np import pandas as pd run = "sub-0_ses-11_task-audiodec_run-0" TR = 2.0 # BOLD on the left hemisphere: (n_TR, 10242) gii = nib.load(f"derivatives/…/{run}_space-fsaverage5_hemi-L_bold.func.gii") bold = np.stack([d.data for d in gii.darrays]) # words, with onsets in seconds from the first trigger words = pd.read_csv("stimuli/EP-0217_part0_words.tsv", sep="\t") # the frame where each word's response peaks (about 5 s later) words["peak_tr"] = ((words.onset + 5.0) // TR).astype(int)
Access
The dataset is shared in BIDS, with the preprocessed derivatives and every stimulus file. The links below are placeholders in this prototype.
Raw multi-echo BOLD, anatomy, events, preprocessed volumes and surfaces, and all stimuli.
Open the dataset Placeholder PreprintAcquisition protocol, preprocessing, data quality and baseline encoding and decoding models.
Read the paper Placeholder CodeReaders for every file in the layout above and the scripts behind the baseline models.
Browse the code@article{cephalonauts_one_2026,
title = {Cephalonauts One: a deep fMRI dataset for decoding naturalistic speech},
author = {Karavela},
year = {2026},
note = {Preprint}
}