Cephalonauts One

Open fMRI dataset  ·  naturalistic speech  ·  French podcasts

A deep fMRI dataset for decoding naturalistic speech in the human brain.

Now in the scanner sub-0 · ses-11 · run-0 TR 01 / 32
Au bout du chemin, il y a une grille verte.
6participants
50hfMRI per participant
110podcast episodes
470kwords, time-stamped
2.0sTR · multi-echo BOLD

Replay

Watch a brain listen

Whole-brain BOLD from one session, played back in sync with the podcast that produced it: the words, their translation, and the five seconds it takes blood to answer them.

sub-0 / ses-11 / run-0 task-audiodec · fsaverage5 Simulated preview
Drag to rotate · click the cortex to probe a point
01:17
00:00
02:21
Transcript · French · word-aligned0 / 0 words
−5 s ▸
BOLD monitor · z-scored · one dot per TRclick a row to light its region
Volume · MNI152 · 2 mmx −58 · y −24 · z 4
Axialz 4
Coronaly −24
Sagittalx −58

From volume to film

How a scan becomes a film

Five steps take a stack of voxels to the flat, frame-by-frame view used in the player above. Scroll through them.

  1. 1 / 5 · Volume

    Every two seconds, a whole brain

    The scanner samples the head as a stack of slices made of 2 mm voxels. Each stack is one frame of BOLD: blood oxygenation, measured everywhere at once.

    TR 2.0 s2 mm voxelsmulti-echo
  2. 2 / 5 · Surface

    Projected onto the cortex

    Speech is processed in the cortex, a folded sheet a few millimetres thick. The signal is sampled on that sheet at 20,484 points, the fsaverage5 surface.

    fsaverage510,242 vertices per hemisphere
  3. 3 / 5 · Inflated

    Inflated to show the folds

    About two thirds of the cortex is hidden inside the sulci. Inflating the surface brings it into view: light patches are gyri, dark ones are sulci.

  4. 4 / 5 · Flat

    Cut and laid flat

    Cut along the medial wall and a few relaxation lines, each hemisphere lies flat. The whole cortex now fits in one frame, like a film still.

  5. 5 / 5 · Film

    Play the frames

    In sequence, the frames become a film of listening. Sound reaches the auditory cortex first, then speech, words and meaning spread along the temporal lobe into frontal and parietal cortex.

    A1 → STS → MTG → IFG · AG

The BOLD signal

The brain answers five seconds late

fMRI does not see neurons. It sees the oxygenated blood that rushes in to refuel them, and that takes time: the response to a word peaks about five seconds after it is spoken, then settles over the next fifteen.

Every word in the dataset is time-stamped so that models can learn which words each frame is answering. Toggle the words, change the TR, and watch what the scanner gets to measure.

TR
last word ends2.9 s
BOLD peaks7.9 s
samples on the curve10

Every point, every frame

1,200 vertices × 32 TRs · rows grouped by region

Design

Deep, not wide

Most naturalistic fMRI datasets scan many people for an hour or two. Cephalonauts One scans six people for about fifty hours each, so that a model can be trained on a single brain.

Same scanner, two designs1 block = 1 h

Fifty hours per brain is enough to fit an encoding or decoding model to one person, instead of averaging across many.

Sessions per participant · one bar per run hover a run
3 / 3 questions 2 / 3 ≤ 1 / 3, flagged planned

Contents

Everything a decoder needs

Each run ships with the audio that was played, every word and when it was heard, a translation, model features, and the questions the participant answered afterwards.

func/

BOLD

Multi-echo EPI at TR 2.0 s, in MNI152 volumes and on the fsaverage5 surface.

sub-0_ses-11_task-audiodec_run-0_bold.nii.gz
stimuli/

Audio

The exact waveform played in the scanner, aligned to the first trigger.

EP-0217_part0.wav
stimuli/

Words

Every word with its onset and offset, in seconds from the first trigger.

EP-0217_part0_words.tsv
stimuli/

Translation

An English translation aligned sentence by sentence to the French.

EP-0217_part0_translation.tsv
derivatives/

Features

Text and audio embeddings for every word, resampled to every TR.

EP-0217_part0_desc-text_embeddings.npy
beh/

Comprehension

Three questions after each part check that the participant was listening.

EP-0217_part0_questions.json
BIDS layoutclick folders
  • cephalonauts-one/
    • dataset_description.json
    • participants.tsv · 6 rows
    • sub-0/
      • anat/
        • sub-0_T1w.nii.gz
      • ses-11/
        • func/
          • sub-0_ses-11_task-audiodec_run-0_echo-1_bold.nii.gz
          • sub-0_ses-11_task-audiodec_run-0_echo-2_bold.nii.gz
          • sub-0_ses-11_task-audiodec_run-0_bold.json
          • sub-0_ses-11_task-audiodec_run-0_events.tsv
          • run-1 … run-4
      • ses-01 … ses-50
    • sub-1 … sub-5
    • derivatives/
      • …_space-MNI152NLin2009cAsym_res-2_desc-preproc_bold.nii.gz
      • …_space-fsaverage5_hemi-L_bold.func.gii
      • …_space-fsaverage5_hemi-R_bold.func.gii
    • stimuli/
      • EP-0217_part0.wav
      • EP-0217_part0_words.tsv
      • EP-0217_part0_translation.tsv
      • EP-0217_part0_questions.json
Python · align words to frames
import nibabel as nib
import numpy as np
import pandas as pd

run = "sub-0_ses-11_task-audiodec_run-0"
TR = 2.0

# BOLD on the left hemisphere: (n_TR, 10242)
gii = nib.load(f"derivatives/…/{run}_space-fsaverage5_hemi-L_bold.func.gii")
bold = np.stack([d.data for d in gii.darrays])

# words, with onsets in seconds from the first trigger
words = pd.read_csv("stimuli/EP-0217_part0_words.tsv", sep="\t")

# the frame where each word's response peaks (about 5 s later)
words["peak_tr"] = ((words.onset + 5.0) // TR).astype(int)

Access

Get the data

The dataset is shared in BIDS, with the preprocessed derivatives and every stimulus file. The links below are placeholders in this prototype.

Cite · BibTeX (placeholder)
@article{cephalonauts_one_2026,
  title   = {Cephalonauts One: a deep fMRI dataset for decoding naturalistic speech},
  author  = {Karavela},
  year    = {2026},
  note    = {Preprint}
}