torch_em.data.datasets.medical.topbrain
The TopBrain dataset contains annotations for whole brain vessel anatomy segmentation in computed tomography angiography (CTA) and magnetic resonance angiography (MRA).
The data was curated for the TopBrain challenge (https://topbrain2025.grand-challenge.org, continued as https://topbrain2026.grand-challenge.org for the 2026 'version 2' edition), which extends the TopCoW challenge from the Circle of Willis to over 40 landmark brain vessel anatomies. The version is selected with the 'version' argument:
- 'v1' (default): the original batch-1 training data release, which consists of 30 annotated angiographies
(15 CTA and 15 MRA), reusing the same 15 patients that are also part of the TopCoW dataset
(
torch_em.data.datasets.medical.topcow). The multi-class segmentations label the brain vessel anatomy, see the label ids in the 'itksnap_labelmap_txt' folder of the downloaded data (40 classes for CTA, 42 classes for MRA, both include the 13 Circle of Willis classes from the TopCoW dataset). - 'v2': the 2026-08-17 data release, which adds the 'TA36' ground-truth: a unified, modality-agnostic
36-class vessel labeling (compatible with the TopAneu challenge,
torch_em.data.datasets.medical.topaneu) for the combined batch-1 and batch-2 angiographies (25 CTA and 25 MRA in total), see the 'labelmap_jsons/labels_topbrain_v2_topaneu36class.json' file of the downloaded data. The modality is selected with the 'modality' argument ('ct' or 'mr').
The 'v1' data is located at https://doi.org/10.5281/zenodo.16623496. The 'v2' data is located at https://doi.org/10.5281/zenodo.21972006.
This dataset is the successor of the TopCoW challenge, published as https://doi.org/10.1056/aidbp2500994. Please cite it if you use this dataset in your research.
1"""The TopBrain dataset contains annotations for whole brain vessel anatomy segmentation 2in computed tomography angiography (CTA) and magnetic resonance angiography (MRA). 3 4The data was curated for the TopBrain challenge (https://topbrain2025.grand-challenge.org, continued as 5https://topbrain2026.grand-challenge.org for the 2026 'version 2' edition), which extends the TopCoW 6challenge from the Circle of Willis to over 40 landmark brain vessel anatomies. The version is selected 7with the 'version' argument: 8- 'v1' (default): the original batch-1 training data release, which consists of 30 annotated angiographies 9 (15 CTA and 15 MRA), reusing the same 15 patients that are also part of the TopCoW dataset 10 (`torch_em.data.datasets.medical.topcow`). The multi-class segmentations label the brain vessel anatomy, 11 see the label ids in the 'itksnap_labelmap_txt' folder of the downloaded data (40 classes for CTA, 42 12 classes for MRA, both include the 13 Circle of Willis classes from the TopCoW dataset). 13- 'v2': the 2026-08-17 data release, which adds the 'TA36' ground-truth: a unified, modality-agnostic 14 36-class vessel labeling (compatible with the TopAneu challenge, `torch_em.data.datasets.medical.topaneu`) 15 for the combined batch-1 and batch-2 angiographies (25 CTA and 25 MRA in total), see the 16 'labelmap_jsons/labels_topbrain_v2_topaneu36class.json' file of the downloaded data. 17The modality is selected with the 'modality' argument ('ct' or 'mr'). 18 19The 'v1' data is located at https://doi.org/10.5281/zenodo.16623496. 20The 'v2' data is located at https://doi.org/10.5281/zenodo.21972006. 21 22This dataset is the successor of the TopCoW challenge, published as https://doi.org/10.1056/aidbp2500994. 23Please cite it if you use this dataset in your research. 24""" 25 26import os 27from glob import glob 28from natsort import natsorted 29from typing import Union, Tuple, List, Optional, Literal 30 31from torch.utils.data import Dataset, DataLoader 32 33import torch_em 34 35from .. import util 36 37 38URL = { 39 "v1": "https://zenodo.org/records/16623496/files/TopBrain_Data_Release_Batch1_073025.zip", 40 "v2": "https://zenodo.org/records/21972006/files/TopBrain_Data_Release_Batches1n2nTA36_081726.zip", 41} 42CHECKSUM = { 43 "v1": "468af6ba9a3ff36c9a61c4e02dbb737d304219aec0e3b1f77df5d97e36304a35", 44 "v2": "d70fb948ae830709396b36a06de807df40668751fc71be78eed98a54d55c1622", 45} 46DATA_DIR_NAME = { 47 "v1": "TopBrain_Data_Release_Batch1_073025", 48 "v2": "TopBrain_Data_Release_Batches1n2nTA36_081726", 49} 50 51MODALITIES = ["ct", "mr"] 52 53VERSIONS = ["v1", "v2"] 54 55 56def get_topbrain_data( 57 path: Union[os.PathLike, str], download: bool = False, version: Literal["v1", "v2"] = "v1" 58) -> str: 59 """Download the TopBrain dataset. 60 61 Args: 62 path: Filepath to a folder where the data is downloaded for further processing. 63 download: Whether to download the data if it is not present. 64 version: The version of the dataset. Either 'v1' (2025 batch-1 release) or 'v2' (2026 release with 65 the 'TA36' ground-truth for the combined batch-1 and batch-2 angiographies). 66 67 Returns: 68 Filepath where the data is downloaded. 69 """ 70 if version not in VERSIONS: 71 raise ValueError(f"'{version}' is not a valid version. Please choose one of {VERSIONS}.") 72 73 data_dir = os.path.join(path, DATA_DIR_NAME[version]) 74 if os.path.exists(data_dir): 75 return data_dir 76 77 os.makedirs(path, exist_ok=True) 78 79 zip_path = os.path.join(path, f"{DATA_DIR_NAME[version]}.zip") 80 util.download_source(path=zip_path, url=URL[version], download=download, checksum=CHECKSUM[version]) 81 util.unzip(zip_path=zip_path, dst=path) 82 83 return data_dir 84 85 86def get_topbrain_paths( 87 path: Union[os.PathLike, str], 88 modality: Optional[Literal["ct", "mr"]] = None, 89 download: bool = False, 90 version: Literal["v1", "v2"] = "v1", 91) -> Tuple[List[str], List[str]]: 92 """Get paths to the TopBrain data. 93 94 Args: 95 path: Filepath to a folder where the data is downloaded for further processing. 96 modality: The angiography modality. Either 'ct' or 'mr'. If None, both modalities are returned. 97 download: Whether to download the data if it is not present. 98 version: The version of the dataset. Either 'v1' (2025 batch-1 release) or 'v2' (2026 release with 99 the 'TA36' ground-truth for the combined batch-1 and batch-2 angiographies). 100 101 Returns: 102 List of filepaths for the image data. 103 List of filepaths for the label data. 104 """ 105 data_dir = get_topbrain_data(path, download, version) 106 107 if modality is not None and modality not in MODALITIES: 108 raise ValueError(f"'{modality}' is not a valid modality. Please choose one of {MODALITIES}.") 109 110 if version == "v1": 111 modalities = MODALITIES if modality is None else [modality] 112 113 raw_paths, label_paths = [], [] 114 for mod in modalities: 115 this_label_paths = natsorted(glob(os.path.join(data_dir, f"labelsTr_topbrain_{mod}", "*.nii.gz"))) 116 # The images carry the channel suffix '_0000' of the nnU-Net format, the labels do not. 117 this_raw_paths = [ 118 os.path.join( 119 data_dir, f"imagesTr_topbrain_{mod}", os.path.basename(p).replace(".nii.gz", "_0000.nii.gz") 120 ) 121 for p in this_label_paths 122 ] 123 assert len(this_raw_paths) > 0 and all(os.path.exists(p) for p in this_raw_paths) 124 raw_paths.extend(this_raw_paths) 125 label_paths.extend(this_label_paths) 126 else: 127 pattern = "topcow_*.nii.gz" if modality is None else f"topcow_{modality}_*.nii.gz" 128 label_paths = natsorted(glob(os.path.join(data_dir, "labelsTr_topbrain_v2_topaneu36class", pattern))) 129 # The images carry the channel suffix '_0000' of the nnU-Net format, the labels do not. 130 raw_paths = [ 131 os.path.join(data_dir, "imagesTr_topbrain", os.path.basename(p).replace(".nii.gz", "_0000.nii.gz")) 132 for p in label_paths 133 ] 134 assert len(raw_paths) > 0 and all(os.path.exists(p) for p in raw_paths) 135 136 return raw_paths, label_paths 137 138 139def get_topbrain_dataset( 140 path: Union[os.PathLike, str], 141 patch_shape: Tuple[int, ...], 142 modality: Optional[Literal["ct", "mr"]] = None, 143 resize_inputs: bool = False, 144 download: bool = False, 145 version: Literal["v1", "v2"] = "v1", 146 **kwargs 147) -> Dataset: 148 """Get the TopBrain dataset for whole brain vessel anatomy segmentation. 149 150 Args: 151 path: Filepath to a folder where the data is downloaded for further processing. 152 patch_shape: The patch shape to use for training. 153 modality: The angiography modality. Either 'ct' or 'mr'. If None, both modalities are returned. 154 resize_inputs: Whether to resize inputs to the desired patch shape. 155 download: Whether to download the data if it is not present. 156 version: The version of the dataset. Either 'v1' (2025 batch-1 release) or 'v2' (2026 release with 157 the 'TA36' ground-truth for the combined batch-1 and batch-2 angiographies). 158 kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset`. 159 160 Returns: 161 The segmentation dataset. 162 """ 163 raw_paths, label_paths = get_topbrain_paths(path, modality, download, version) 164 165 if resize_inputs: 166 resize_kwargs = {"patch_shape": patch_shape, "is_rgb": False} 167 kwargs, patch_shape = util.update_kwargs_for_resize_trafo( 168 kwargs=kwargs, patch_shape=patch_shape, resize_inputs=resize_inputs, resize_kwargs=resize_kwargs 169 ) 170 171 return torch_em.default_segmentation_dataset( 172 raw_paths=raw_paths, 173 raw_key="data", 174 label_paths=label_paths, 175 label_key="data", 176 patch_shape=patch_shape, 177 is_seg_dataset=True, 178 **kwargs 179 ) 180 181 182def get_topbrain_loader( 183 path: Union[os.PathLike, str], 184 batch_size: int, 185 patch_shape: Tuple[int, ...], 186 modality: Optional[Literal["ct", "mr"]] = None, 187 resize_inputs: bool = False, 188 download: bool = False, 189 version: Literal["v1", "v2"] = "v1", 190 **kwargs 191) -> DataLoader: 192 """Get the TopBrain dataloader for whole brain vessel anatomy segmentation. 193 194 Args: 195 path: Filepath to a folder where the data is downloaded for further processing. 196 batch_size: The batch size for training. 197 patch_shape: The patch shape to use for training. 198 modality: The angiography modality. Either 'ct' or 'mr'. If None, both modalities are returned. 199 resize_inputs: Whether to resize inputs to the desired patch shape. 200 download: Whether to download the data if it is not present. 201 version: The version of the dataset. Either 'v1' (2025 batch-1 release) or 'v2' (2026 release with 202 the 'TA36' ground-truth for the combined batch-1 and batch-2 angiographies). 203 kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset` or for the PyTorch DataLoader. 204 205 Returns: 206 The DataLoader. 207 """ 208 ds_kwargs, loader_kwargs = util.split_kwargs(torch_em.default_segmentation_dataset, **kwargs) 209 dataset = get_topbrain_dataset(path, patch_shape, modality, resize_inputs, download, version, **ds_kwargs) 210 return torch_em.get_data_loader(dataset, batch_size, **loader_kwargs)
57def get_topbrain_data( 58 path: Union[os.PathLike, str], download: bool = False, version: Literal["v1", "v2"] = "v1" 59) -> str: 60 """Download the TopBrain dataset. 61 62 Args: 63 path: Filepath to a folder where the data is downloaded for further processing. 64 download: Whether to download the data if it is not present. 65 version: The version of the dataset. Either 'v1' (2025 batch-1 release) or 'v2' (2026 release with 66 the 'TA36' ground-truth for the combined batch-1 and batch-2 angiographies). 67 68 Returns: 69 Filepath where the data is downloaded. 70 """ 71 if version not in VERSIONS: 72 raise ValueError(f"'{version}' is not a valid version. Please choose one of {VERSIONS}.") 73 74 data_dir = os.path.join(path, DATA_DIR_NAME[version]) 75 if os.path.exists(data_dir): 76 return data_dir 77 78 os.makedirs(path, exist_ok=True) 79 80 zip_path = os.path.join(path, f"{DATA_DIR_NAME[version]}.zip") 81 util.download_source(path=zip_path, url=URL[version], download=download, checksum=CHECKSUM[version]) 82 util.unzip(zip_path=zip_path, dst=path) 83 84 return data_dir
Download the TopBrain dataset.
Arguments:
- path: Filepath to a folder where the data is downloaded for further processing.
- download: Whether to download the data if it is not present.
- version: The version of the dataset. Either 'v1' (2025 batch-1 release) or 'v2' (2026 release with the 'TA36' ground-truth for the combined batch-1 and batch-2 angiographies).
Returns:
Filepath where the data is downloaded.
87def get_topbrain_paths( 88 path: Union[os.PathLike, str], 89 modality: Optional[Literal["ct", "mr"]] = None, 90 download: bool = False, 91 version: Literal["v1", "v2"] = "v1", 92) -> Tuple[List[str], List[str]]: 93 """Get paths to the TopBrain data. 94 95 Args: 96 path: Filepath to a folder where the data is downloaded for further processing. 97 modality: The angiography modality. Either 'ct' or 'mr'. If None, both modalities are returned. 98 download: Whether to download the data if it is not present. 99 version: The version of the dataset. Either 'v1' (2025 batch-1 release) or 'v2' (2026 release with 100 the 'TA36' ground-truth for the combined batch-1 and batch-2 angiographies). 101 102 Returns: 103 List of filepaths for the image data. 104 List of filepaths for the label data. 105 """ 106 data_dir = get_topbrain_data(path, download, version) 107 108 if modality is not None and modality not in MODALITIES: 109 raise ValueError(f"'{modality}' is not a valid modality. Please choose one of {MODALITIES}.") 110 111 if version == "v1": 112 modalities = MODALITIES if modality is None else [modality] 113 114 raw_paths, label_paths = [], [] 115 for mod in modalities: 116 this_label_paths = natsorted(glob(os.path.join(data_dir, f"labelsTr_topbrain_{mod}", "*.nii.gz"))) 117 # The images carry the channel suffix '_0000' of the nnU-Net format, the labels do not. 118 this_raw_paths = [ 119 os.path.join( 120 data_dir, f"imagesTr_topbrain_{mod}", os.path.basename(p).replace(".nii.gz", "_0000.nii.gz") 121 ) 122 for p in this_label_paths 123 ] 124 assert len(this_raw_paths) > 0 and all(os.path.exists(p) for p in this_raw_paths) 125 raw_paths.extend(this_raw_paths) 126 label_paths.extend(this_label_paths) 127 else: 128 pattern = "topcow_*.nii.gz" if modality is None else f"topcow_{modality}_*.nii.gz" 129 label_paths = natsorted(glob(os.path.join(data_dir, "labelsTr_topbrain_v2_topaneu36class", pattern))) 130 # The images carry the channel suffix '_0000' of the nnU-Net format, the labels do not. 131 raw_paths = [ 132 os.path.join(data_dir, "imagesTr_topbrain", os.path.basename(p).replace(".nii.gz", "_0000.nii.gz")) 133 for p in label_paths 134 ] 135 assert len(raw_paths) > 0 and all(os.path.exists(p) for p in raw_paths) 136 137 return raw_paths, label_paths
Get paths to the TopBrain data.
Arguments:
- path: Filepath to a folder where the data is downloaded for further processing.
- modality: The angiography modality. Either 'ct' or 'mr'. If None, both modalities are returned.
- download: Whether to download the data if it is not present.
- version: The version of the dataset. Either 'v1' (2025 batch-1 release) or 'v2' (2026 release with the 'TA36' ground-truth for the combined batch-1 and batch-2 angiographies).
Returns:
List of filepaths for the image data. List of filepaths for the label data.
140def get_topbrain_dataset( 141 path: Union[os.PathLike, str], 142 patch_shape: Tuple[int, ...], 143 modality: Optional[Literal["ct", "mr"]] = None, 144 resize_inputs: bool = False, 145 download: bool = False, 146 version: Literal["v1", "v2"] = "v1", 147 **kwargs 148) -> Dataset: 149 """Get the TopBrain dataset for whole brain vessel anatomy segmentation. 150 151 Args: 152 path: Filepath to a folder where the data is downloaded for further processing. 153 patch_shape: The patch shape to use for training. 154 modality: The angiography modality. Either 'ct' or 'mr'. If None, both modalities are returned. 155 resize_inputs: Whether to resize inputs to the desired patch shape. 156 download: Whether to download the data if it is not present. 157 version: The version of the dataset. Either 'v1' (2025 batch-1 release) or 'v2' (2026 release with 158 the 'TA36' ground-truth for the combined batch-1 and batch-2 angiographies). 159 kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset`. 160 161 Returns: 162 The segmentation dataset. 163 """ 164 raw_paths, label_paths = get_topbrain_paths(path, modality, download, version) 165 166 if resize_inputs: 167 resize_kwargs = {"patch_shape": patch_shape, "is_rgb": False} 168 kwargs, patch_shape = util.update_kwargs_for_resize_trafo( 169 kwargs=kwargs, patch_shape=patch_shape, resize_inputs=resize_inputs, resize_kwargs=resize_kwargs 170 ) 171 172 return torch_em.default_segmentation_dataset( 173 raw_paths=raw_paths, 174 raw_key="data", 175 label_paths=label_paths, 176 label_key="data", 177 patch_shape=patch_shape, 178 is_seg_dataset=True, 179 **kwargs 180 )
Get the TopBrain dataset for whole brain vessel anatomy segmentation.
Arguments:
- path: Filepath to a folder where the data is downloaded for further processing.
- patch_shape: The patch shape to use for training.
- modality: The angiography modality. Either 'ct' or 'mr'. If None, both modalities are returned.
- resize_inputs: Whether to resize inputs to the desired patch shape.
- download: Whether to download the data if it is not present.
- version: The version of the dataset. Either 'v1' (2025 batch-1 release) or 'v2' (2026 release with the 'TA36' ground-truth for the combined batch-1 and batch-2 angiographies).
- kwargs: Additional keyword arguments for
torch_em.default_segmentation_dataset.
Returns:
The segmentation dataset.
183def get_topbrain_loader( 184 path: Union[os.PathLike, str], 185 batch_size: int, 186 patch_shape: Tuple[int, ...], 187 modality: Optional[Literal["ct", "mr"]] = None, 188 resize_inputs: bool = False, 189 download: bool = False, 190 version: Literal["v1", "v2"] = "v1", 191 **kwargs 192) -> DataLoader: 193 """Get the TopBrain dataloader for whole brain vessel anatomy segmentation. 194 195 Args: 196 path: Filepath to a folder where the data is downloaded for further processing. 197 batch_size: The batch size for training. 198 patch_shape: The patch shape to use for training. 199 modality: The angiography modality. Either 'ct' or 'mr'. If None, both modalities are returned. 200 resize_inputs: Whether to resize inputs to the desired patch shape. 201 download: Whether to download the data if it is not present. 202 version: The version of the dataset. Either 'v1' (2025 batch-1 release) or 'v2' (2026 release with 203 the 'TA36' ground-truth for the combined batch-1 and batch-2 angiographies). 204 kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset` or for the PyTorch DataLoader. 205 206 Returns: 207 The DataLoader. 208 """ 209 ds_kwargs, loader_kwargs = util.split_kwargs(torch_em.default_segmentation_dataset, **kwargs) 210 dataset = get_topbrain_dataset(path, patch_shape, modality, resize_inputs, download, version, **ds_kwargs) 211 return torch_em.get_data_loader(dataset, batch_size, **loader_kwargs)
Get the TopBrain dataloader for whole brain vessel anatomy segmentation.
Arguments:
- path: Filepath to a folder where the data is downloaded for further processing.
- batch_size: The batch size for training.
- patch_shape: The patch shape to use for training.
- modality: The angiography modality. Either 'ct' or 'mr'. If None, both modalities are returned.
- resize_inputs: Whether to resize inputs to the desired patch shape.
- download: Whether to download the data if it is not present.
- version: The version of the dataset. Either 'v1' (2025 batch-1 release) or 'v2' (2026 release with the 'TA36' ground-truth for the combined batch-1 and batch-2 angiographies).
- kwargs: Additional keyword arguments for
torch_em.default_segmentation_datasetor for the PyTorch DataLoader.
Returns:
The DataLoader.