torch_em.data.datasets.medical.mindboggle101

The Mindboggle-101 dataset contains annotations for cortical parcellation in T1-weighted brain MRI.

The dataset comprises 101 healthy subjects pooled from five cohorts (OASIS-TRT-20, NKI-TRT-20, NKI-RS-22, MMRR-21 and Extra-18, the last of which bundles the smaller HLN-12, Twins-2, MMRR-3T7T-2, Colin27 and Afterthought cohorts) that were manually labelled according to the Desikan-Killiany-Tourville (DKT) cortical labeling protocol.

The data is hosted at https://zenodo.org/records/22070005 (DOI 10.5281/zenodo.22070005) and distributed under the CC BY 4.0 license, no signup required. This module downloads and uses only the 'Mindboggle101_release3.zip' archive, which stores the labelled volumes as one tar.gz file per cohort (the archive also ships unlabelled templates, atlases and cortical surfaces, which are not used here).

Two label sets are available per subject and can be selected with the 'label_choice' argument:

  • 'cortical': labels.DKT31.manual.nii.gz, the manually corrected DKT31 cortical parcellation (background plus 31 cortical regions per hemisphere).
  • 'full' (default): labels.DKT31.manual+aseg.nii.gz, the same cortical labels combined with FreeSurfer's automated subcortical segmentation ('aseg'), giving a whole-brain parcellation into over 100 structures. NOTE: Only the 'cortical' labels are the result of manual editing; the subcortical structures added in the 'full' label set come from FreeSurfer's automated 'aseg' pipeline and are not manually verified.

The T1 volume and the chosen label volume of each subject are bundled into one hdf5 file by this module, with the keys 'raw' and 'labels'.

This dataset is from the publication https://doi.org/10.3389/fnins.2012.00171. Please cite it if you use this dataset in your research.

  1"""The Mindboggle-101 dataset contains annotations for cortical parcellation in T1-weighted brain MRI.
  2
  3The dataset comprises 101 healthy subjects pooled from five cohorts (OASIS-TRT-20, NKI-TRT-20, NKI-RS-22,
  4MMRR-21 and Extra-18, the last of which bundles the smaller HLN-12, Twins-2, MMRR-3T7T-2, Colin27 and
  5Afterthought cohorts) that were manually labelled according to the Desikan-Killiany-Tourville (DKT)
  6cortical labeling protocol.
  7
  8The data is hosted at https://zenodo.org/records/22070005 (DOI 10.5281/zenodo.22070005) and distributed
  9under the CC BY 4.0 license, no signup required. This module downloads and uses only the
 10'Mindboggle101_release3.zip' archive, which stores the labelled volumes as one tar.gz file per cohort
 11(the archive also ships unlabelled templates, atlases and cortical surfaces, which are not used here).
 12
 13Two label sets are available per subject and can be selected with the 'label_choice' argument:
 14- 'cortical': `labels.DKT31.manual.nii.gz`, the manually corrected DKT31 cortical parcellation
 15  (background plus 31 cortical regions per hemisphere).
 16- 'full' (default): `labels.DKT31.manual+aseg.nii.gz`, the same cortical labels combined with FreeSurfer's
 17  automated subcortical segmentation ('aseg'), giving a whole-brain parcellation into over 100 structures.
 18NOTE: Only the 'cortical' labels are the result of manual editing; the subcortical structures added in the
 19'full' label set come from FreeSurfer's automated 'aseg' pipeline and are not manually verified.
 20
 21The T1 volume and the chosen label volume of each subject are bundled into one hdf5 file by this module,
 22with the keys 'raw' and 'labels'.
 23
 24This dataset is from the publication https://doi.org/10.3389/fnins.2012.00171.
 25Please cite it if you use this dataset in your research.
 26"""
 27
 28import os
 29from glob import glob
 30from tqdm import tqdm
 31from natsort import natsorted
 32from typing import Union, Tuple, Literal, List
 33
 34import numpy as np
 35
 36from torch.utils.data import Dataset, DataLoader
 37
 38import torch_em
 39
 40from .. import util
 41
 42
 43URL = "https://zenodo.org/records/22070005/files/Mindboggle101_release3.zip?download=1"
 44
 45CHECKSUM = "56dfca8fed2f80740c763c0bf05e48cc736eb793ee1511f621ef5f82eb08f26e"
 46
 47COHORTS = ["OASIS-TRT-20", "NKI-TRT-20", "NKI-RS-22", "MMRR-21", "Extra-18"]
 48
 49LABEL_FILES = {"cortical": "labels.DKT31.manual.nii.gz", "full": "labels.DKT31.manual+aseg.nii.gz"}
 50
 51
 52def _get_volumes_dir(path: Union[os.PathLike, str], download: bool) -> str:
 53    import zipfile
 54
 55    volumes_dir = os.path.join(path, "Mindboggle101_release3", "Mindboggle101_volumes")
 56    if all(os.path.exists(os.path.join(volumes_dir, f"{cohort}_volumes")) for cohort in COHORTS):
 57        return volumes_dir
 58
 59    os.makedirs(path, exist_ok=True)
 60
 61    tar_paths = {cohort: os.path.join(volumes_dir, f"{cohort}_volumes.tar.gz") for cohort in COHORTS}
 62    if not all(os.path.exists(p) for p in tar_paths.values()):
 63        zip_path = os.path.join(path, "Mindboggle101_release3.zip")
 64        util.download_source(path=zip_path, url=URL, download=download, checksum=CHECKSUM)
 65
 66        tar_names = [f"Mindboggle101_release3/Mindboggle101_volumes/{cohort}_volumes.tar.gz" for cohort in COHORTS]
 67        with zipfile.ZipFile(zip_path) as f:
 68            f.extractall(path, members=tar_names)
 69
 70        os.remove(zip_path)
 71
 72    for cohort in COHORTS:
 73        target_dir = os.path.join(volumes_dir, f"{cohort}_volumes")
 74        if os.path.exists(target_dir):
 75            continue
 76        util.unzip_tarfile(tar_paths[cohort], dst=volumes_dir, remove=True)
 77
 78    return volumes_dir
 79
 80
 81def _preprocess_inputs(volumes_dir, label_choice, preprocessed_dir):
 82    import h5py
 83    import nibabel as nib
 84
 85    os.makedirs(preprocessed_dir, exist_ok=True)
 86    label_file = LABEL_FILES[label_choice]
 87
 88    subject_dirs = natsorted(glob(os.path.join(volumes_dir, "*_volumes", "*")))
 89    for subject_dir in tqdm(subject_dirs, desc="Preprocessing the Mindboggle-101 subjects"):
 90        subject_id = os.path.basename(subject_dir)
 91        volume_path = os.path.join(preprocessed_dir, f"{subject_id}.h5")
 92        if os.path.exists(volume_path):
 93            continue
 94
 95        raw_path = os.path.join(subject_dir, "t1weighted.nii.gz")
 96        label_path = os.path.join(subject_dir, label_file)
 97        if not (os.path.exists(raw_path) and os.path.exists(label_path)):
 98            continue
 99
100        raw = np.asarray(nib.load(raw_path).dataobj)
101        labels = np.asarray(nib.load(label_path).dataobj).astype("uint16")
102        assert raw.shape == labels.shape, f"Shape mismatch for {subject_id}: {raw.shape} vs. {labels.shape}."
103
104        # The file is written to a temporary path first, so that an interrupted run leaves no corrupt file.
105        with h5py.File(f"{volume_path}.tmp", "w") as f:
106            f.create_dataset("raw", data=raw, compression="gzip")
107            f.create_dataset("labels", data=labels, compression="gzip")
108
109        os.rename(f"{volume_path}.tmp", volume_path)
110
111    # Marks that preprocessing has finished for all subjects, so that a partially preprocessed
112    # directory (e.g. from an interrupted previous run) is not mistaken for a complete one.
113    open(os.path.join(preprocessed_dir, ".preprocessing_done"), "w").close()
114
115
116def get_mindboggle101_data(
117    path: Union[os.PathLike, str], label_choice: Literal["cortical", "full"] = "full", download: bool = False
118) -> str:
119    """Download the Mindboggle-101 dataset.
120
121    Args:
122        path: Filepath to a folder where the data is downloaded for further processing.
123        label_choice: The choice of label set. Either 'cortical' or 'full'. See the module docstring.
124        download: Whether to download the data if it is not present.
125
126    Returns:
127        Filepath to the folder where the preprocessed data is stored.
128    """
129    if label_choice not in LABEL_FILES:
130        raise ValueError(f"'{label_choice}' is not a valid label choice. Please choose one of {list(LABEL_FILES)}.")
131
132    preprocessed_dir = os.path.join(path, "preprocessed", label_choice)
133    if os.path.exists(os.path.join(preprocessed_dir, ".preprocessing_done")):
134        return preprocessed_dir
135
136    volumes_dir = _get_volumes_dir(path, download)
137    _preprocess_inputs(volumes_dir, label_choice, preprocessed_dir)
138    return preprocessed_dir
139
140
141def get_mindboggle101_paths(
142    path: Union[os.PathLike, str],
143    label_choice: Literal["cortical", "full"] = "full",
144    download: bool = False,
145) -> List[str]:
146    """Get paths to the Mindboggle-101 data.
147
148    Args:
149        path: Filepath to a folder where the data is downloaded for further processing.
150        label_choice: The choice of label set. Either 'cortical' or 'full'. See the module docstring.
151        download: Whether to download the data if it is not present.
152
153    Returns:
154        List of filepaths for the hdf5 files, which contain the image data ('raw') and the label data ('labels').
155    """
156    preprocessed_dir = get_mindboggle101_data(path, label_choice, download)
157    volume_paths = natsorted(glob(os.path.join(preprocessed_dir, "*.h5")))
158    assert len(volume_paths) > 0, f"Could not find any preprocessed Mindboggle-101 volumes at '{path}'."
159    return volume_paths
160
161
162def get_mindboggle101_dataset(
163    path: Union[os.PathLike, str],
164    patch_shape: Tuple[int, ...],
165    label_choice: Literal["cortical", "full"] = "full",
166    resize_inputs: bool = False,
167    download: bool = False,
168    **kwargs
169) -> Dataset:
170    """Get the Mindboggle-101 dataset for cortical parcellation.
171
172    Args:
173        path: Filepath to a folder where the data is downloaded for further processing.
174        patch_shape: The patch shape to use for training.
175        label_choice: The choice of label set. Either 'cortical' or 'full'. See the module docstring.
176        resize_inputs: Whether to resize inputs to the desired patch shape.
177        download: Whether to download the data if it is not present.
178        kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset`.
179
180    Returns:
181        The segmentation dataset.
182    """
183    volume_paths = get_mindboggle101_paths(path, label_choice, download)
184
185    if resize_inputs:
186        resize_kwargs = {"patch_shape": patch_shape, "is_rgb": False}
187        kwargs, patch_shape = util.update_kwargs_for_resize_trafo(
188            kwargs=kwargs, patch_shape=patch_shape, resize_inputs=resize_inputs, resize_kwargs=resize_kwargs
189        )
190
191    return torch_em.default_segmentation_dataset(
192        raw_paths=volume_paths,
193        raw_key="raw",
194        label_paths=volume_paths,
195        label_key="labels",
196        patch_shape=patch_shape,
197        is_seg_dataset=True,
198        **kwargs
199    )
200
201
202def get_mindboggle101_loader(
203    path: Union[os.PathLike, str],
204    batch_size: int,
205    patch_shape: Tuple[int, ...],
206    label_choice: Literal["cortical", "full"] = "full",
207    resize_inputs: bool = False,
208    download: bool = False,
209    **kwargs
210) -> DataLoader:
211    """Get the Mindboggle-101 dataloader for cortical parcellation.
212
213    Args:
214        path: Filepath to a folder where the data is downloaded for further processing.
215        batch_size: The batch size for training.
216        patch_shape: The patch shape to use for training.
217        label_choice: The choice of label set. Either 'cortical' or 'full'. See the module docstring.
218        resize_inputs: Whether to resize inputs to the desired patch shape.
219        download: Whether to download the data if it is not present.
220        kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset` or for the PyTorch DataLoader.
221
222    Returns:
223        The DataLoader.
224    """
225    ds_kwargs, loader_kwargs = util.split_kwargs(torch_em.default_segmentation_dataset, **kwargs)
226    dataset = get_mindboggle101_dataset(path, patch_shape, label_choice, resize_inputs, download, **ds_kwargs)
227    return torch_em.get_data_loader(dataset, batch_size, **loader_kwargs)
URL = 'https://zenodo.org/records/22070005/files/Mindboggle101_release3.zip?download=1'
CHECKSUM = '56dfca8fed2f80740c763c0bf05e48cc736eb793ee1511f621ef5f82eb08f26e'
COHORTS = ['OASIS-TRT-20', 'NKI-TRT-20', 'NKI-RS-22', 'MMRR-21', 'Extra-18']
LABEL_FILES = {'cortical': 'labels.DKT31.manual.nii.gz', 'full': 'labels.DKT31.manual+aseg.nii.gz'}
def get_mindboggle101_data( path: Union[os.PathLike, str], label_choice: Literal['cortical', 'full'] = 'full', download: bool = False) -> str:
117def get_mindboggle101_data(
118    path: Union[os.PathLike, str], label_choice: Literal["cortical", "full"] = "full", download: bool = False
119) -> str:
120    """Download the Mindboggle-101 dataset.
121
122    Args:
123        path: Filepath to a folder where the data is downloaded for further processing.
124        label_choice: The choice of label set. Either 'cortical' or 'full'. See the module docstring.
125        download: Whether to download the data if it is not present.
126
127    Returns:
128        Filepath to the folder where the preprocessed data is stored.
129    """
130    if label_choice not in LABEL_FILES:
131        raise ValueError(f"'{label_choice}' is not a valid label choice. Please choose one of {list(LABEL_FILES)}.")
132
133    preprocessed_dir = os.path.join(path, "preprocessed", label_choice)
134    if os.path.exists(os.path.join(preprocessed_dir, ".preprocessing_done")):
135        return preprocessed_dir
136
137    volumes_dir = _get_volumes_dir(path, download)
138    _preprocess_inputs(volumes_dir, label_choice, preprocessed_dir)
139    return preprocessed_dir

Download the Mindboggle-101 dataset.

Arguments:
  • path: Filepath to a folder where the data is downloaded for further processing.
  • label_choice: The choice of label set. Either 'cortical' or 'full'. See the module docstring.
  • download: Whether to download the data if it is not present.
Returns:

Filepath to the folder where the preprocessed data is stored.

def get_mindboggle101_paths( path: Union[os.PathLike, str], label_choice: Literal['cortical', 'full'] = 'full', download: bool = False) -> List[str]:
142def get_mindboggle101_paths(
143    path: Union[os.PathLike, str],
144    label_choice: Literal["cortical", "full"] = "full",
145    download: bool = False,
146) -> List[str]:
147    """Get paths to the Mindboggle-101 data.
148
149    Args:
150        path: Filepath to a folder where the data is downloaded for further processing.
151        label_choice: The choice of label set. Either 'cortical' or 'full'. See the module docstring.
152        download: Whether to download the data if it is not present.
153
154    Returns:
155        List of filepaths for the hdf5 files, which contain the image data ('raw') and the label data ('labels').
156    """
157    preprocessed_dir = get_mindboggle101_data(path, label_choice, download)
158    volume_paths = natsorted(glob(os.path.join(preprocessed_dir, "*.h5")))
159    assert len(volume_paths) > 0, f"Could not find any preprocessed Mindboggle-101 volumes at '{path}'."
160    return volume_paths

Get paths to the Mindboggle-101 data.

Arguments:
  • path: Filepath to a folder where the data is downloaded for further processing.
  • label_choice: The choice of label set. Either 'cortical' or 'full'. See the module docstring.
  • download: Whether to download the data if it is not present.
Returns:

List of filepaths for the hdf5 files, which contain the image data ('raw') and the label data ('labels').

def get_mindboggle101_dataset( path: Union[os.PathLike, str], patch_shape: Tuple[int, ...], label_choice: Literal['cortical', 'full'] = 'full', resize_inputs: bool = False, download: bool = False, **kwargs) -> torch.utils.data.dataset.Dataset:
163def get_mindboggle101_dataset(
164    path: Union[os.PathLike, str],
165    patch_shape: Tuple[int, ...],
166    label_choice: Literal["cortical", "full"] = "full",
167    resize_inputs: bool = False,
168    download: bool = False,
169    **kwargs
170) -> Dataset:
171    """Get the Mindboggle-101 dataset for cortical parcellation.
172
173    Args:
174        path: Filepath to a folder where the data is downloaded for further processing.
175        patch_shape: The patch shape to use for training.
176        label_choice: The choice of label set. Either 'cortical' or 'full'. See the module docstring.
177        resize_inputs: Whether to resize inputs to the desired patch shape.
178        download: Whether to download the data if it is not present.
179        kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset`.
180
181    Returns:
182        The segmentation dataset.
183    """
184    volume_paths = get_mindboggle101_paths(path, label_choice, download)
185
186    if resize_inputs:
187        resize_kwargs = {"patch_shape": patch_shape, "is_rgb": False}
188        kwargs, patch_shape = util.update_kwargs_for_resize_trafo(
189            kwargs=kwargs, patch_shape=patch_shape, resize_inputs=resize_inputs, resize_kwargs=resize_kwargs
190        )
191
192    return torch_em.default_segmentation_dataset(
193        raw_paths=volume_paths,
194        raw_key="raw",
195        label_paths=volume_paths,
196        label_key="labels",
197        patch_shape=patch_shape,
198        is_seg_dataset=True,
199        **kwargs
200    )

Get the Mindboggle-101 dataset for cortical parcellation.

Arguments:
  • path: Filepath to a folder where the data is downloaded for further processing.
  • patch_shape: The patch shape to use for training.
  • label_choice: The choice of label set. Either 'cortical' or 'full'. See the module docstring.
  • resize_inputs: Whether to resize inputs to the desired patch shape.
  • download: Whether to download the data if it is not present.
  • kwargs: Additional keyword arguments for torch_em.default_segmentation_dataset.
Returns:

The segmentation dataset.

def get_mindboggle101_loader( path: Union[os.PathLike, str], batch_size: int, patch_shape: Tuple[int, ...], label_choice: Literal['cortical', 'full'] = 'full', resize_inputs: bool = False, download: bool = False, **kwargs) -> torch.utils.data.dataloader.DataLoader:
203def get_mindboggle101_loader(
204    path: Union[os.PathLike, str],
205    batch_size: int,
206    patch_shape: Tuple[int, ...],
207    label_choice: Literal["cortical", "full"] = "full",
208    resize_inputs: bool = False,
209    download: bool = False,
210    **kwargs
211) -> DataLoader:
212    """Get the Mindboggle-101 dataloader for cortical parcellation.
213
214    Args:
215        path: Filepath to a folder where the data is downloaded for further processing.
216        batch_size: The batch size for training.
217        patch_shape: The patch shape to use for training.
218        label_choice: The choice of label set. Either 'cortical' or 'full'. See the module docstring.
219        resize_inputs: Whether to resize inputs to the desired patch shape.
220        download: Whether to download the data if it is not present.
221        kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset` or for the PyTorch DataLoader.
222
223    Returns:
224        The DataLoader.
225    """
226    ds_kwargs, loader_kwargs = util.split_kwargs(torch_em.default_segmentation_dataset, **kwargs)
227    dataset = get_mindboggle101_dataset(path, patch_shape, label_choice, resize_inputs, download, **ds_kwargs)
228    return torch_em.get_data_loader(dataset, batch_size, **loader_kwargs)

Get the Mindboggle-101 dataloader for cortical parcellation.

Arguments:
  • path: Filepath to a folder where the data is downloaded for further processing.
  • batch_size: The batch size for training.
  • patch_shape: The patch shape to use for training.
  • label_choice: The choice of label set. Either 'cortical' or 'full'. See the module docstring.
  • resize_inputs: Whether to resize inputs to the desired patch shape.
  • download: Whether to download the data if it is not present.
  • kwargs: Additional keyword arguments for torch_em.default_segmentation_dataset or for the PyTorch DataLoader.
Returns:

The DataLoader.