torch_em.data.datasets.medical.topbrain

The TopBrain dataset contains annotations for whole brain vessel anatomy segmentation in computed tomography angiography (CTA) and magnetic resonance angiography (MRA).

The data was curated for the TopBrain challenge (https://topbrain2025.grand-challenge.org, continued as https://topbrain2026.grand-challenge.org for the 2026 'version 2' edition), which extends the TopCoW challenge from the Circle of Willis to over 40 landmark brain vessel anatomies. The version is selected with the 'version' argument:

  • 'v1' (default): the original batch-1 training data release, which consists of 30 annotated angiographies (15 CTA and 15 MRA), reusing the same 15 patients that are also part of the TopCoW dataset (torch_em.data.datasets.medical.topcow). The multi-class segmentations label the brain vessel anatomy, see the label ids in the 'itksnap_labelmap_txt' folder of the downloaded data (40 classes for CTA, 42 classes for MRA, both include the 13 Circle of Willis classes from the TopCoW dataset).
  • 'v2': the 2026-08-17 data release, which adds the 'TA36' ground-truth: a unified, modality-agnostic 36-class vessel labeling (compatible with the TopAneu challenge, torch_em.data.datasets.medical.topaneu) for the combined batch-1 and batch-2 angiographies (25 CTA and 25 MRA in total), see the 'labelmap_jsons/labels_topbrain_v2_topaneu36class.json' file of the downloaded data. The modality is selected with the 'modality' argument ('ct' or 'mr').

The 'v1' data is located at https://doi.org/10.5281/zenodo.16623496. The 'v2' data is located at https://doi.org/10.5281/zenodo.21972006.

This dataset is the successor of the TopCoW challenge, published as https://doi.org/10.1056/aidbp2500994. Please cite it if you use this dataset in your research.

  1"""The TopBrain dataset contains annotations for whole brain vessel anatomy segmentation
  2in computed tomography angiography (CTA) and magnetic resonance angiography (MRA).
  3
  4The data was curated for the TopBrain challenge (https://topbrain2025.grand-challenge.org, continued as
  5https://topbrain2026.grand-challenge.org for the 2026 'version 2' edition), which extends the TopCoW
  6challenge from the Circle of Willis to over 40 landmark brain vessel anatomies. The version is selected
  7with the 'version' argument:
  8- 'v1' (default): the original batch-1 training data release, which consists of 30 annotated angiographies
  9  (15 CTA and 15 MRA), reusing the same 15 patients that are also part of the TopCoW dataset
 10  (`torch_em.data.datasets.medical.topcow`). The multi-class segmentations label the brain vessel anatomy,
 11  see the label ids in the 'itksnap_labelmap_txt' folder of the downloaded data (40 classes for CTA, 42
 12  classes for MRA, both include the 13 Circle of Willis classes from the TopCoW dataset).
 13- 'v2': the 2026-08-17 data release, which adds the 'TA36' ground-truth: a unified, modality-agnostic
 14  36-class vessel labeling (compatible with the TopAneu challenge, `torch_em.data.datasets.medical.topaneu`)
 15  for the combined batch-1 and batch-2 angiographies (25 CTA and 25 MRA in total), see the
 16  'labelmap_jsons/labels_topbrain_v2_topaneu36class.json' file of the downloaded data.
 17The modality is selected with the 'modality' argument ('ct' or 'mr').
 18
 19The 'v1' data is located at https://doi.org/10.5281/zenodo.16623496.
 20The 'v2' data is located at https://doi.org/10.5281/zenodo.21972006.
 21
 22This dataset is the successor of the TopCoW challenge, published as https://doi.org/10.1056/aidbp2500994.
 23Please cite it if you use this dataset in your research.
 24"""
 25
 26import os
 27from glob import glob
 28from natsort import natsorted
 29from typing import Union, Tuple, List, Optional, Literal
 30
 31from torch.utils.data import Dataset, DataLoader
 32
 33import torch_em
 34
 35from .. import util
 36
 37
 38URL = {
 39    "v1": "https://zenodo.org/records/16623496/files/TopBrain_Data_Release_Batch1_073025.zip",
 40    "v2": "https://zenodo.org/records/21972006/files/TopBrain_Data_Release_Batches1n2nTA36_081726.zip",
 41}
 42CHECKSUM = {
 43    "v1": "468af6ba9a3ff36c9a61c4e02dbb737d304219aec0e3b1f77df5d97e36304a35",
 44    "v2": "d70fb948ae830709396b36a06de807df40668751fc71be78eed98a54d55c1622",
 45}
 46DATA_DIR_NAME = {
 47    "v1": "TopBrain_Data_Release_Batch1_073025",
 48    "v2": "TopBrain_Data_Release_Batches1n2nTA36_081726",
 49}
 50
 51MODALITIES = ["ct", "mr"]
 52
 53VERSIONS = ["v1", "v2"]
 54
 55
 56def get_topbrain_data(
 57    path: Union[os.PathLike, str], download: bool = False, version: Literal["v1", "v2"] = "v1"
 58) -> str:
 59    """Download the TopBrain dataset.
 60
 61    Args:
 62        path: Filepath to a folder where the data is downloaded for further processing.
 63        download: Whether to download the data if it is not present.
 64        version: The version of the dataset. Either 'v1' (2025 batch-1 release) or 'v2' (2026 release with
 65            the 'TA36' ground-truth for the combined batch-1 and batch-2 angiographies).
 66
 67    Returns:
 68        Filepath where the data is downloaded.
 69    """
 70    if version not in VERSIONS:
 71        raise ValueError(f"'{version}' is not a valid version. Please choose one of {VERSIONS}.")
 72
 73    data_dir = os.path.join(path, DATA_DIR_NAME[version])
 74    if os.path.exists(data_dir):
 75        return data_dir
 76
 77    os.makedirs(path, exist_ok=True)
 78
 79    zip_path = os.path.join(path, f"{DATA_DIR_NAME[version]}.zip")
 80    util.download_source(path=zip_path, url=URL[version], download=download, checksum=CHECKSUM[version])
 81    util.unzip(zip_path=zip_path, dst=path)
 82
 83    return data_dir
 84
 85
 86def get_topbrain_paths(
 87    path: Union[os.PathLike, str],
 88    modality: Optional[Literal["ct", "mr"]] = None,
 89    download: bool = False,
 90    version: Literal["v1", "v2"] = "v1",
 91) -> Tuple[List[str], List[str]]:
 92    """Get paths to the TopBrain data.
 93
 94    Args:
 95        path: Filepath to a folder where the data is downloaded for further processing.
 96        modality: The angiography modality. Either 'ct' or 'mr'. If None, both modalities are returned.
 97        download: Whether to download the data if it is not present.
 98        version: The version of the dataset. Either 'v1' (2025 batch-1 release) or 'v2' (2026 release with
 99            the 'TA36' ground-truth for the combined batch-1 and batch-2 angiographies).
100
101    Returns:
102        List of filepaths for the image data.
103        List of filepaths for the label data.
104    """
105    data_dir = get_topbrain_data(path, download, version)
106
107    if modality is not None and modality not in MODALITIES:
108        raise ValueError(f"'{modality}' is not a valid modality. Please choose one of {MODALITIES}.")
109
110    if version == "v1":
111        modalities = MODALITIES if modality is None else [modality]
112
113        raw_paths, label_paths = [], []
114        for mod in modalities:
115            this_label_paths = natsorted(glob(os.path.join(data_dir, f"labelsTr_topbrain_{mod}", "*.nii.gz")))
116            # The images carry the channel suffix '_0000' of the nnU-Net format, the labels do not.
117            this_raw_paths = [
118                os.path.join(
119                    data_dir, f"imagesTr_topbrain_{mod}", os.path.basename(p).replace(".nii.gz", "_0000.nii.gz")
120                )
121                for p in this_label_paths
122            ]
123            assert len(this_raw_paths) > 0 and all(os.path.exists(p) for p in this_raw_paths)
124            raw_paths.extend(this_raw_paths)
125            label_paths.extend(this_label_paths)
126    else:
127        pattern = "topcow_*.nii.gz" if modality is None else f"topcow_{modality}_*.nii.gz"
128        label_paths = natsorted(glob(os.path.join(data_dir, "labelsTr_topbrain_v2_topaneu36class", pattern)))
129        # The images carry the channel suffix '_0000' of the nnU-Net format, the labels do not.
130        raw_paths = [
131            os.path.join(data_dir, "imagesTr_topbrain", os.path.basename(p).replace(".nii.gz", "_0000.nii.gz"))
132            for p in label_paths
133        ]
134        assert len(raw_paths) > 0 and all(os.path.exists(p) for p in raw_paths)
135
136    return raw_paths, label_paths
137
138
139def get_topbrain_dataset(
140    path: Union[os.PathLike, str],
141    patch_shape: Tuple[int, ...],
142    modality: Optional[Literal["ct", "mr"]] = None,
143    resize_inputs: bool = False,
144    download: bool = False,
145    version: Literal["v1", "v2"] = "v1",
146    **kwargs
147) -> Dataset:
148    """Get the TopBrain dataset for whole brain vessel anatomy segmentation.
149
150    Args:
151        path: Filepath to a folder where the data is downloaded for further processing.
152        patch_shape: The patch shape to use for training.
153        modality: The angiography modality. Either 'ct' or 'mr'. If None, both modalities are returned.
154        resize_inputs: Whether to resize inputs to the desired patch shape.
155        download: Whether to download the data if it is not present.
156        version: The version of the dataset. Either 'v1' (2025 batch-1 release) or 'v2' (2026 release with
157            the 'TA36' ground-truth for the combined batch-1 and batch-2 angiographies).
158        kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset`.
159
160    Returns:
161        The segmentation dataset.
162    """
163    raw_paths, label_paths = get_topbrain_paths(path, modality, download, version)
164
165    if resize_inputs:
166        resize_kwargs = {"patch_shape": patch_shape, "is_rgb": False}
167        kwargs, patch_shape = util.update_kwargs_for_resize_trafo(
168            kwargs=kwargs, patch_shape=patch_shape, resize_inputs=resize_inputs, resize_kwargs=resize_kwargs
169        )
170
171    return torch_em.default_segmentation_dataset(
172        raw_paths=raw_paths,
173        raw_key="data",
174        label_paths=label_paths,
175        label_key="data",
176        patch_shape=patch_shape,
177        is_seg_dataset=True,
178        **kwargs
179    )
180
181
182def get_topbrain_loader(
183    path: Union[os.PathLike, str],
184    batch_size: int,
185    patch_shape: Tuple[int, ...],
186    modality: Optional[Literal["ct", "mr"]] = None,
187    resize_inputs: bool = False,
188    download: bool = False,
189    version: Literal["v1", "v2"] = "v1",
190    **kwargs
191) -> DataLoader:
192    """Get the TopBrain dataloader for whole brain vessel anatomy segmentation.
193
194    Args:
195        path: Filepath to a folder where the data is downloaded for further processing.
196        batch_size: The batch size for training.
197        patch_shape: The patch shape to use for training.
198        modality: The angiography modality. Either 'ct' or 'mr'. If None, both modalities are returned.
199        resize_inputs: Whether to resize inputs to the desired patch shape.
200        download: Whether to download the data if it is not present.
201        version: The version of the dataset. Either 'v1' (2025 batch-1 release) or 'v2' (2026 release with
202            the 'TA36' ground-truth for the combined batch-1 and batch-2 angiographies).
203        kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset` or for the PyTorch DataLoader.
204
205    Returns:
206        The DataLoader.
207    """
208    ds_kwargs, loader_kwargs = util.split_kwargs(torch_em.default_segmentation_dataset, **kwargs)
209    dataset = get_topbrain_dataset(path, patch_shape, modality, resize_inputs, download, version, **ds_kwargs)
210    return torch_em.get_data_loader(dataset, batch_size, **loader_kwargs)
URL = {'v1': 'https://zenodo.org/records/16623496/files/TopBrain_Data_Release_Batch1_073025.zip', 'v2': 'https://zenodo.org/records/21972006/files/TopBrain_Data_Release_Batches1n2nTA36_081726.zip'}
CHECKSUM = {'v1': '468af6ba9a3ff36c9a61c4e02dbb737d304219aec0e3b1f77df5d97e36304a35', 'v2': 'd70fb948ae830709396b36a06de807df40668751fc71be78eed98a54d55c1622'}
DATA_DIR_NAME = {'v1': 'TopBrain_Data_Release_Batch1_073025', 'v2': 'TopBrain_Data_Release_Batches1n2nTA36_081726'}
MODALITIES = ['ct', 'mr']
VERSIONS = ['v1', 'v2']
def get_topbrain_data( path: Union[os.PathLike, str], download: bool = False, version: Literal['v1', 'v2'] = 'v1') -> str:
57def get_topbrain_data(
58    path: Union[os.PathLike, str], download: bool = False, version: Literal["v1", "v2"] = "v1"
59) -> str:
60    """Download the TopBrain dataset.
61
62    Args:
63        path: Filepath to a folder where the data is downloaded for further processing.
64        download: Whether to download the data if it is not present.
65        version: The version of the dataset. Either 'v1' (2025 batch-1 release) or 'v2' (2026 release with
66            the 'TA36' ground-truth for the combined batch-1 and batch-2 angiographies).
67
68    Returns:
69        Filepath where the data is downloaded.
70    """
71    if version not in VERSIONS:
72        raise ValueError(f"'{version}' is not a valid version. Please choose one of {VERSIONS}.")
73
74    data_dir = os.path.join(path, DATA_DIR_NAME[version])
75    if os.path.exists(data_dir):
76        return data_dir
77
78    os.makedirs(path, exist_ok=True)
79
80    zip_path = os.path.join(path, f"{DATA_DIR_NAME[version]}.zip")
81    util.download_source(path=zip_path, url=URL[version], download=download, checksum=CHECKSUM[version])
82    util.unzip(zip_path=zip_path, dst=path)
83
84    return data_dir

Download the TopBrain dataset.

Arguments:
  • path: Filepath to a folder where the data is downloaded for further processing.
  • download: Whether to download the data if it is not present.
  • version: The version of the dataset. Either 'v1' (2025 batch-1 release) or 'v2' (2026 release with the 'TA36' ground-truth for the combined batch-1 and batch-2 angiographies).
Returns:

Filepath where the data is downloaded.

def get_topbrain_paths( path: Union[os.PathLike, str], modality: Optional[Literal['ct', 'mr']] = None, download: bool = False, version: Literal['v1', 'v2'] = 'v1') -> Tuple[List[str], List[str]]:
 87def get_topbrain_paths(
 88    path: Union[os.PathLike, str],
 89    modality: Optional[Literal["ct", "mr"]] = None,
 90    download: bool = False,
 91    version: Literal["v1", "v2"] = "v1",
 92) -> Tuple[List[str], List[str]]:
 93    """Get paths to the TopBrain data.
 94
 95    Args:
 96        path: Filepath to a folder where the data is downloaded for further processing.
 97        modality: The angiography modality. Either 'ct' or 'mr'. If None, both modalities are returned.
 98        download: Whether to download the data if it is not present.
 99        version: The version of the dataset. Either 'v1' (2025 batch-1 release) or 'v2' (2026 release with
100            the 'TA36' ground-truth for the combined batch-1 and batch-2 angiographies).
101
102    Returns:
103        List of filepaths for the image data.
104        List of filepaths for the label data.
105    """
106    data_dir = get_topbrain_data(path, download, version)
107
108    if modality is not None and modality not in MODALITIES:
109        raise ValueError(f"'{modality}' is not a valid modality. Please choose one of {MODALITIES}.")
110
111    if version == "v1":
112        modalities = MODALITIES if modality is None else [modality]
113
114        raw_paths, label_paths = [], []
115        for mod in modalities:
116            this_label_paths = natsorted(glob(os.path.join(data_dir, f"labelsTr_topbrain_{mod}", "*.nii.gz")))
117            # The images carry the channel suffix '_0000' of the nnU-Net format, the labels do not.
118            this_raw_paths = [
119                os.path.join(
120                    data_dir, f"imagesTr_topbrain_{mod}", os.path.basename(p).replace(".nii.gz", "_0000.nii.gz")
121                )
122                for p in this_label_paths
123            ]
124            assert len(this_raw_paths) > 0 and all(os.path.exists(p) for p in this_raw_paths)
125            raw_paths.extend(this_raw_paths)
126            label_paths.extend(this_label_paths)
127    else:
128        pattern = "topcow_*.nii.gz" if modality is None else f"topcow_{modality}_*.nii.gz"
129        label_paths = natsorted(glob(os.path.join(data_dir, "labelsTr_topbrain_v2_topaneu36class", pattern)))
130        # The images carry the channel suffix '_0000' of the nnU-Net format, the labels do not.
131        raw_paths = [
132            os.path.join(data_dir, "imagesTr_topbrain", os.path.basename(p).replace(".nii.gz", "_0000.nii.gz"))
133            for p in label_paths
134        ]
135        assert len(raw_paths) > 0 and all(os.path.exists(p) for p in raw_paths)
136
137    return raw_paths, label_paths

Get paths to the TopBrain data.

Arguments:
  • path: Filepath to a folder where the data is downloaded for further processing.
  • modality: The angiography modality. Either 'ct' or 'mr'. If None, both modalities are returned.
  • download: Whether to download the data if it is not present.
  • version: The version of the dataset. Either 'v1' (2025 batch-1 release) or 'v2' (2026 release with the 'TA36' ground-truth for the combined batch-1 and batch-2 angiographies).
Returns:

List of filepaths for the image data. List of filepaths for the label data.

def get_topbrain_dataset( path: Union[os.PathLike, str], patch_shape: Tuple[int, ...], modality: Optional[Literal['ct', 'mr']] = None, resize_inputs: bool = False, download: bool = False, version: Literal['v1', 'v2'] = 'v1', **kwargs) -> torch.utils.data.dataset.Dataset:
140def get_topbrain_dataset(
141    path: Union[os.PathLike, str],
142    patch_shape: Tuple[int, ...],
143    modality: Optional[Literal["ct", "mr"]] = None,
144    resize_inputs: bool = False,
145    download: bool = False,
146    version: Literal["v1", "v2"] = "v1",
147    **kwargs
148) -> Dataset:
149    """Get the TopBrain dataset for whole brain vessel anatomy segmentation.
150
151    Args:
152        path: Filepath to a folder where the data is downloaded for further processing.
153        patch_shape: The patch shape to use for training.
154        modality: The angiography modality. Either 'ct' or 'mr'. If None, both modalities are returned.
155        resize_inputs: Whether to resize inputs to the desired patch shape.
156        download: Whether to download the data if it is not present.
157        version: The version of the dataset. Either 'v1' (2025 batch-1 release) or 'v2' (2026 release with
158            the 'TA36' ground-truth for the combined batch-1 and batch-2 angiographies).
159        kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset`.
160
161    Returns:
162        The segmentation dataset.
163    """
164    raw_paths, label_paths = get_topbrain_paths(path, modality, download, version)
165
166    if resize_inputs:
167        resize_kwargs = {"patch_shape": patch_shape, "is_rgb": False}
168        kwargs, patch_shape = util.update_kwargs_for_resize_trafo(
169            kwargs=kwargs, patch_shape=patch_shape, resize_inputs=resize_inputs, resize_kwargs=resize_kwargs
170        )
171
172    return torch_em.default_segmentation_dataset(
173        raw_paths=raw_paths,
174        raw_key="data",
175        label_paths=label_paths,
176        label_key="data",
177        patch_shape=patch_shape,
178        is_seg_dataset=True,
179        **kwargs
180    )

Get the TopBrain dataset for whole brain vessel anatomy segmentation.

Arguments:
  • path: Filepath to a folder where the data is downloaded for further processing.
  • patch_shape: The patch shape to use for training.
  • modality: The angiography modality. Either 'ct' or 'mr'. If None, both modalities are returned.
  • resize_inputs: Whether to resize inputs to the desired patch shape.
  • download: Whether to download the data if it is not present.
  • version: The version of the dataset. Either 'v1' (2025 batch-1 release) or 'v2' (2026 release with the 'TA36' ground-truth for the combined batch-1 and batch-2 angiographies).
  • kwargs: Additional keyword arguments for torch_em.default_segmentation_dataset.
Returns:

The segmentation dataset.

def get_topbrain_loader( path: Union[os.PathLike, str], batch_size: int, patch_shape: Tuple[int, ...], modality: Optional[Literal['ct', 'mr']] = None, resize_inputs: bool = False, download: bool = False, version: Literal['v1', 'v2'] = 'v1', **kwargs) -> torch.utils.data.dataloader.DataLoader:
183def get_topbrain_loader(
184    path: Union[os.PathLike, str],
185    batch_size: int,
186    patch_shape: Tuple[int, ...],
187    modality: Optional[Literal["ct", "mr"]] = None,
188    resize_inputs: bool = False,
189    download: bool = False,
190    version: Literal["v1", "v2"] = "v1",
191    **kwargs
192) -> DataLoader:
193    """Get the TopBrain dataloader for whole brain vessel anatomy segmentation.
194
195    Args:
196        path: Filepath to a folder where the data is downloaded for further processing.
197        batch_size: The batch size for training.
198        patch_shape: The patch shape to use for training.
199        modality: The angiography modality. Either 'ct' or 'mr'. If None, both modalities are returned.
200        resize_inputs: Whether to resize inputs to the desired patch shape.
201        download: Whether to download the data if it is not present.
202        version: The version of the dataset. Either 'v1' (2025 batch-1 release) or 'v2' (2026 release with
203            the 'TA36' ground-truth for the combined batch-1 and batch-2 angiographies).
204        kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset` or for the PyTorch DataLoader.
205
206    Returns:
207        The DataLoader.
208    """
209    ds_kwargs, loader_kwargs = util.split_kwargs(torch_em.default_segmentation_dataset, **kwargs)
210    dataset = get_topbrain_dataset(path, patch_shape, modality, resize_inputs, download, version, **ds_kwargs)
211    return torch_em.get_data_loader(dataset, batch_size, **loader_kwargs)

Get the TopBrain dataloader for whole brain vessel anatomy segmentation.

Arguments:
  • path: Filepath to a folder where the data is downloaded for further processing.
  • batch_size: The batch size for training.
  • patch_shape: The patch shape to use for training.
  • modality: The angiography modality. Either 'ct' or 'mr'. If None, both modalities are returned.
  • resize_inputs: Whether to resize inputs to the desired patch shape.
  • download: Whether to download the data if it is not present.
  • version: The version of the dataset. Either 'v1' (2025 batch-1 release) or 'v2' (2026 release with the 'TA36' ground-truth for the combined batch-1 and batch-2 angiographies).
  • kwargs: Additional keyword arguments for torch_em.default_segmentation_dataset or for the PyTorch DataLoader.
Returns:

The DataLoader.