torch_em.data.datasets.medical.mindboggle101
The Mindboggle-101 dataset contains annotations for cortical parcellation in T1-weighted brain MRI.
The dataset comprises 101 healthy subjects pooled from five cohorts (OASIS-TRT-20, NKI-TRT-20, NKI-RS-22, MMRR-21 and Extra-18, the last of which bundles the smaller HLN-12, Twins-2, MMRR-3T7T-2, Colin27 and Afterthought cohorts) that were manually labelled according to the Desikan-Killiany-Tourville (DKT) cortical labeling protocol.
The data is hosted at https://zenodo.org/records/22070005 (DOI 10.5281/zenodo.22070005) and distributed under the CC BY 4.0 license, no signup required. This module downloads and uses only the 'Mindboggle101_release3.zip' archive, which stores the labelled volumes as one tar.gz file per cohort (the archive also ships unlabelled templates, atlases and cortical surfaces, which are not used here).
Two label sets are available per subject and can be selected with the 'label_choice' argument:
- 'cortical':
labels.DKT31.manual.nii.gz, the manually corrected DKT31 cortical parcellation (background plus 31 cortical regions per hemisphere). - 'full' (default):
labels.DKT31.manual+aseg.nii.gz, the same cortical labels combined with FreeSurfer's automated subcortical segmentation ('aseg'), giving a whole-brain parcellation into over 100 structures. NOTE: Only the 'cortical' labels are the result of manual editing; the subcortical structures added in the 'full' label set come from FreeSurfer's automated 'aseg' pipeline and are not manually verified.
The T1 volume and the chosen label volume of each subject are bundled into one hdf5 file by this module, with the keys 'raw' and 'labels'.
This dataset is from the publication https://doi.org/10.3389/fnins.2012.00171. Please cite it if you use this dataset in your research.
1"""The Mindboggle-101 dataset contains annotations for cortical parcellation in T1-weighted brain MRI. 2 3The dataset comprises 101 healthy subjects pooled from five cohorts (OASIS-TRT-20, NKI-TRT-20, NKI-RS-22, 4MMRR-21 and Extra-18, the last of which bundles the smaller HLN-12, Twins-2, MMRR-3T7T-2, Colin27 and 5Afterthought cohorts) that were manually labelled according to the Desikan-Killiany-Tourville (DKT) 6cortical labeling protocol. 7 8The data is hosted at https://zenodo.org/records/22070005 (DOI 10.5281/zenodo.22070005) and distributed 9under the CC BY 4.0 license, no signup required. This module downloads and uses only the 10'Mindboggle101_release3.zip' archive, which stores the labelled volumes as one tar.gz file per cohort 11(the archive also ships unlabelled templates, atlases and cortical surfaces, which are not used here). 12 13Two label sets are available per subject and can be selected with the 'label_choice' argument: 14- 'cortical': `labels.DKT31.manual.nii.gz`, the manually corrected DKT31 cortical parcellation 15 (background plus 31 cortical regions per hemisphere). 16- 'full' (default): `labels.DKT31.manual+aseg.nii.gz`, the same cortical labels combined with FreeSurfer's 17 automated subcortical segmentation ('aseg'), giving a whole-brain parcellation into over 100 structures. 18NOTE: Only the 'cortical' labels are the result of manual editing; the subcortical structures added in the 19'full' label set come from FreeSurfer's automated 'aseg' pipeline and are not manually verified. 20 21The T1 volume and the chosen label volume of each subject are bundled into one hdf5 file by this module, 22with the keys 'raw' and 'labels'. 23 24This dataset is from the publication https://doi.org/10.3389/fnins.2012.00171. 25Please cite it if you use this dataset in your research. 26""" 27 28import os 29from glob import glob 30from tqdm import tqdm 31from natsort import natsorted 32from typing import Union, Tuple, Literal, List 33 34import numpy as np 35 36from torch.utils.data import Dataset, DataLoader 37 38import torch_em 39 40from .. import util 41 42 43URL = "https://zenodo.org/records/22070005/files/Mindboggle101_release3.zip?download=1" 44 45CHECKSUM = "56dfca8fed2f80740c763c0bf05e48cc736eb793ee1511f621ef5f82eb08f26e" 46 47COHORTS = ["OASIS-TRT-20", "NKI-TRT-20", "NKI-RS-22", "MMRR-21", "Extra-18"] 48 49LABEL_FILES = {"cortical": "labels.DKT31.manual.nii.gz", "full": "labels.DKT31.manual+aseg.nii.gz"} 50 51 52def _get_volumes_dir(path: Union[os.PathLike, str], download: bool) -> str: 53 import zipfile 54 55 volumes_dir = os.path.join(path, "Mindboggle101_release3", "Mindboggle101_volumes") 56 if all(os.path.exists(os.path.join(volumes_dir, f"{cohort}_volumes")) for cohort in COHORTS): 57 return volumes_dir 58 59 os.makedirs(path, exist_ok=True) 60 61 tar_paths = {cohort: os.path.join(volumes_dir, f"{cohort}_volumes.tar.gz") for cohort in COHORTS} 62 if not all(os.path.exists(p) for p in tar_paths.values()): 63 zip_path = os.path.join(path, "Mindboggle101_release3.zip") 64 util.download_source(path=zip_path, url=URL, download=download, checksum=CHECKSUM) 65 66 tar_names = [f"Mindboggle101_release3/Mindboggle101_volumes/{cohort}_volumes.tar.gz" for cohort in COHORTS] 67 with zipfile.ZipFile(zip_path) as f: 68 f.extractall(path, members=tar_names) 69 70 os.remove(zip_path) 71 72 for cohort in COHORTS: 73 target_dir = os.path.join(volumes_dir, f"{cohort}_volumes") 74 if os.path.exists(target_dir): 75 continue 76 util.unzip_tarfile(tar_paths[cohort], dst=volumes_dir, remove=True) 77 78 return volumes_dir 79 80 81def _preprocess_inputs(volumes_dir, label_choice, preprocessed_dir): 82 import h5py 83 import nibabel as nib 84 85 os.makedirs(preprocessed_dir, exist_ok=True) 86 label_file = LABEL_FILES[label_choice] 87 88 subject_dirs = natsorted(glob(os.path.join(volumes_dir, "*_volumes", "*"))) 89 for subject_dir in tqdm(subject_dirs, desc="Preprocessing the Mindboggle-101 subjects"): 90 subject_id = os.path.basename(subject_dir) 91 volume_path = os.path.join(preprocessed_dir, f"{subject_id}.h5") 92 if os.path.exists(volume_path): 93 continue 94 95 raw_path = os.path.join(subject_dir, "t1weighted.nii.gz") 96 label_path = os.path.join(subject_dir, label_file) 97 if not (os.path.exists(raw_path) and os.path.exists(label_path)): 98 continue 99 100 raw = np.asarray(nib.load(raw_path).dataobj) 101 labels = np.asarray(nib.load(label_path).dataobj).astype("uint16") 102 assert raw.shape == labels.shape, f"Shape mismatch for {subject_id}: {raw.shape} vs. {labels.shape}." 103 104 # The file is written to a temporary path first, so that an interrupted run leaves no corrupt file. 105 with h5py.File(f"{volume_path}.tmp", "w") as f: 106 f.create_dataset("raw", data=raw, compression="gzip") 107 f.create_dataset("labels", data=labels, compression="gzip") 108 109 os.rename(f"{volume_path}.tmp", volume_path) 110 111 # Marks that preprocessing has finished for all subjects, so that a partially preprocessed 112 # directory (e.g. from an interrupted previous run) is not mistaken for a complete one. 113 open(os.path.join(preprocessed_dir, ".preprocessing_done"), "w").close() 114 115 116def get_mindboggle101_data( 117 path: Union[os.PathLike, str], label_choice: Literal["cortical", "full"] = "full", download: bool = False 118) -> str: 119 """Download the Mindboggle-101 dataset. 120 121 Args: 122 path: Filepath to a folder where the data is downloaded for further processing. 123 label_choice: The choice of label set. Either 'cortical' or 'full'. See the module docstring. 124 download: Whether to download the data if it is not present. 125 126 Returns: 127 Filepath to the folder where the preprocessed data is stored. 128 """ 129 if label_choice not in LABEL_FILES: 130 raise ValueError(f"'{label_choice}' is not a valid label choice. Please choose one of {list(LABEL_FILES)}.") 131 132 preprocessed_dir = os.path.join(path, "preprocessed", label_choice) 133 if os.path.exists(os.path.join(preprocessed_dir, ".preprocessing_done")): 134 return preprocessed_dir 135 136 volumes_dir = _get_volumes_dir(path, download) 137 _preprocess_inputs(volumes_dir, label_choice, preprocessed_dir) 138 return preprocessed_dir 139 140 141def get_mindboggle101_paths( 142 path: Union[os.PathLike, str], 143 label_choice: Literal["cortical", "full"] = "full", 144 download: bool = False, 145) -> List[str]: 146 """Get paths to the Mindboggle-101 data. 147 148 Args: 149 path: Filepath to a folder where the data is downloaded for further processing. 150 label_choice: The choice of label set. Either 'cortical' or 'full'. See the module docstring. 151 download: Whether to download the data if it is not present. 152 153 Returns: 154 List of filepaths for the hdf5 files, which contain the image data ('raw') and the label data ('labels'). 155 """ 156 preprocessed_dir = get_mindboggle101_data(path, label_choice, download) 157 volume_paths = natsorted(glob(os.path.join(preprocessed_dir, "*.h5"))) 158 assert len(volume_paths) > 0, f"Could not find any preprocessed Mindboggle-101 volumes at '{path}'." 159 return volume_paths 160 161 162def get_mindboggle101_dataset( 163 path: Union[os.PathLike, str], 164 patch_shape: Tuple[int, ...], 165 label_choice: Literal["cortical", "full"] = "full", 166 resize_inputs: bool = False, 167 download: bool = False, 168 **kwargs 169) -> Dataset: 170 """Get the Mindboggle-101 dataset for cortical parcellation. 171 172 Args: 173 path: Filepath to a folder where the data is downloaded for further processing. 174 patch_shape: The patch shape to use for training. 175 label_choice: The choice of label set. Either 'cortical' or 'full'. See the module docstring. 176 resize_inputs: Whether to resize inputs to the desired patch shape. 177 download: Whether to download the data if it is not present. 178 kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset`. 179 180 Returns: 181 The segmentation dataset. 182 """ 183 volume_paths = get_mindboggle101_paths(path, label_choice, download) 184 185 if resize_inputs: 186 resize_kwargs = {"patch_shape": patch_shape, "is_rgb": False} 187 kwargs, patch_shape = util.update_kwargs_for_resize_trafo( 188 kwargs=kwargs, patch_shape=patch_shape, resize_inputs=resize_inputs, resize_kwargs=resize_kwargs 189 ) 190 191 return torch_em.default_segmentation_dataset( 192 raw_paths=volume_paths, 193 raw_key="raw", 194 label_paths=volume_paths, 195 label_key="labels", 196 patch_shape=patch_shape, 197 is_seg_dataset=True, 198 **kwargs 199 ) 200 201 202def get_mindboggle101_loader( 203 path: Union[os.PathLike, str], 204 batch_size: int, 205 patch_shape: Tuple[int, ...], 206 label_choice: Literal["cortical", "full"] = "full", 207 resize_inputs: bool = False, 208 download: bool = False, 209 **kwargs 210) -> DataLoader: 211 """Get the Mindboggle-101 dataloader for cortical parcellation. 212 213 Args: 214 path: Filepath to a folder where the data is downloaded for further processing. 215 batch_size: The batch size for training. 216 patch_shape: The patch shape to use for training. 217 label_choice: The choice of label set. Either 'cortical' or 'full'. See the module docstring. 218 resize_inputs: Whether to resize inputs to the desired patch shape. 219 download: Whether to download the data if it is not present. 220 kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset` or for the PyTorch DataLoader. 221 222 Returns: 223 The DataLoader. 224 """ 225 ds_kwargs, loader_kwargs = util.split_kwargs(torch_em.default_segmentation_dataset, **kwargs) 226 dataset = get_mindboggle101_dataset(path, patch_shape, label_choice, resize_inputs, download, **ds_kwargs) 227 return torch_em.get_data_loader(dataset, batch_size, **loader_kwargs)
117def get_mindboggle101_data( 118 path: Union[os.PathLike, str], label_choice: Literal["cortical", "full"] = "full", download: bool = False 119) -> str: 120 """Download the Mindboggle-101 dataset. 121 122 Args: 123 path: Filepath to a folder where the data is downloaded for further processing. 124 label_choice: The choice of label set. Either 'cortical' or 'full'. See the module docstring. 125 download: Whether to download the data if it is not present. 126 127 Returns: 128 Filepath to the folder where the preprocessed data is stored. 129 """ 130 if label_choice not in LABEL_FILES: 131 raise ValueError(f"'{label_choice}' is not a valid label choice. Please choose one of {list(LABEL_FILES)}.") 132 133 preprocessed_dir = os.path.join(path, "preprocessed", label_choice) 134 if os.path.exists(os.path.join(preprocessed_dir, ".preprocessing_done")): 135 return preprocessed_dir 136 137 volumes_dir = _get_volumes_dir(path, download) 138 _preprocess_inputs(volumes_dir, label_choice, preprocessed_dir) 139 return preprocessed_dir
Download the Mindboggle-101 dataset.
Arguments:
- path: Filepath to a folder where the data is downloaded for further processing.
- label_choice: The choice of label set. Either 'cortical' or 'full'. See the module docstring.
- download: Whether to download the data if it is not present.
Returns:
Filepath to the folder where the preprocessed data is stored.
142def get_mindboggle101_paths( 143 path: Union[os.PathLike, str], 144 label_choice: Literal["cortical", "full"] = "full", 145 download: bool = False, 146) -> List[str]: 147 """Get paths to the Mindboggle-101 data. 148 149 Args: 150 path: Filepath to a folder where the data is downloaded for further processing. 151 label_choice: The choice of label set. Either 'cortical' or 'full'. See the module docstring. 152 download: Whether to download the data if it is not present. 153 154 Returns: 155 List of filepaths for the hdf5 files, which contain the image data ('raw') and the label data ('labels'). 156 """ 157 preprocessed_dir = get_mindboggle101_data(path, label_choice, download) 158 volume_paths = natsorted(glob(os.path.join(preprocessed_dir, "*.h5"))) 159 assert len(volume_paths) > 0, f"Could not find any preprocessed Mindboggle-101 volumes at '{path}'." 160 return volume_paths
Get paths to the Mindboggle-101 data.
Arguments:
- path: Filepath to a folder where the data is downloaded for further processing.
- label_choice: The choice of label set. Either 'cortical' or 'full'. See the module docstring.
- download: Whether to download the data if it is not present.
Returns:
List of filepaths for the hdf5 files, which contain the image data ('raw') and the label data ('labels').
163def get_mindboggle101_dataset( 164 path: Union[os.PathLike, str], 165 patch_shape: Tuple[int, ...], 166 label_choice: Literal["cortical", "full"] = "full", 167 resize_inputs: bool = False, 168 download: bool = False, 169 **kwargs 170) -> Dataset: 171 """Get the Mindboggle-101 dataset for cortical parcellation. 172 173 Args: 174 path: Filepath to a folder where the data is downloaded for further processing. 175 patch_shape: The patch shape to use for training. 176 label_choice: The choice of label set. Either 'cortical' or 'full'. See the module docstring. 177 resize_inputs: Whether to resize inputs to the desired patch shape. 178 download: Whether to download the data if it is not present. 179 kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset`. 180 181 Returns: 182 The segmentation dataset. 183 """ 184 volume_paths = get_mindboggle101_paths(path, label_choice, download) 185 186 if resize_inputs: 187 resize_kwargs = {"patch_shape": patch_shape, "is_rgb": False} 188 kwargs, patch_shape = util.update_kwargs_for_resize_trafo( 189 kwargs=kwargs, patch_shape=patch_shape, resize_inputs=resize_inputs, resize_kwargs=resize_kwargs 190 ) 191 192 return torch_em.default_segmentation_dataset( 193 raw_paths=volume_paths, 194 raw_key="raw", 195 label_paths=volume_paths, 196 label_key="labels", 197 patch_shape=patch_shape, 198 is_seg_dataset=True, 199 **kwargs 200 )
Get the Mindboggle-101 dataset for cortical parcellation.
Arguments:
- path: Filepath to a folder where the data is downloaded for further processing.
- patch_shape: The patch shape to use for training.
- label_choice: The choice of label set. Either 'cortical' or 'full'. See the module docstring.
- resize_inputs: Whether to resize inputs to the desired patch shape.
- download: Whether to download the data if it is not present.
- kwargs: Additional keyword arguments for
torch_em.default_segmentation_dataset.
Returns:
The segmentation dataset.
203def get_mindboggle101_loader( 204 path: Union[os.PathLike, str], 205 batch_size: int, 206 patch_shape: Tuple[int, ...], 207 label_choice: Literal["cortical", "full"] = "full", 208 resize_inputs: bool = False, 209 download: bool = False, 210 **kwargs 211) -> DataLoader: 212 """Get the Mindboggle-101 dataloader for cortical parcellation. 213 214 Args: 215 path: Filepath to a folder where the data is downloaded for further processing. 216 batch_size: The batch size for training. 217 patch_shape: The patch shape to use for training. 218 label_choice: The choice of label set. Either 'cortical' or 'full'. See the module docstring. 219 resize_inputs: Whether to resize inputs to the desired patch shape. 220 download: Whether to download the data if it is not present. 221 kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset` or for the PyTorch DataLoader. 222 223 Returns: 224 The DataLoader. 225 """ 226 ds_kwargs, loader_kwargs = util.split_kwargs(torch_em.default_segmentation_dataset, **kwargs) 227 dataset = get_mindboggle101_dataset(path, patch_shape, label_choice, resize_inputs, download, **ds_kwargs) 228 return torch_em.get_data_loader(dataset, batch_size, **loader_kwargs)
Get the Mindboggle-101 dataloader for cortical parcellation.
Arguments:
- path: Filepath to a folder where the data is downloaded for further processing.
- batch_size: The batch size for training.
- patch_shape: The patch shape to use for training.
- label_choice: The choice of label set. Either 'cortical' or 'full'. See the module docstring.
- resize_inputs: Whether to resize inputs to the desired patch shape.
- download: Whether to download the data if it is not present.
- kwargs: Additional keyword arguments for
torch_em.default_segmentation_datasetor for the PyTorch DataLoader.
Returns:
The DataLoader.