torch_em.data.datasets.medical.cc359

The Calgary-Campinas-359 (CC-359) dataset contains annotations for brain extraction (skull-stripping) in T1-weighted brain MRI.

The dataset comprises 359 volumetric T1-weighted scans of healthy older adults (29-80 years), acquired on scanners from three vendors (Siemens, Philips, General Electric) at both 1.5T and 3T field strengths, with approximately 60 subjects per vendor / field strength combination. The scans are described at https://www.ccdataset.com; this module downloads the data from its mirror on the Canada Open Neuroscience Platform (CONP, https://github.com/conpdatasets/calgary-campinas), which serves the files directly over https with no signup or click-through form (unlike the ccdataset.com website, which asks for a filled-in usage form before handing out the same files).

NOTE ON LABEL QUALITY: The brain masks are "silver-standard" labels, not manual ground truth. They were created by STAPLE consensus fusion of several automated skull-stripping methods (Souza et al., NeuroImage 2018, https://doi.org/10.1016/j.neuroimage.2017.08.021). The raw STAPLE output is a per-voxel consensus probability in [0, 1] rather than a hard label; this module binarizes it at a probability of 0.5 to obtain the 0/1 mask used for training. Do not treat these labels as expert manual annotations; they are best used for pretraining, benchmarking segmentation agreement, or as weak supervision.

NOTE: The dataset also ships silver-standard hippocampus masks (also STAPLE-derived), but those are stored on a cropped bounding box with a different image origin/shape than the corresponding raw scan (i.e. they are not defined on the same voxel grid), so they cannot be paired directly as same-shape raw/label volumes and are not exposed by this module.

The data is distributed under the CC BY-ND 4.0 license.

This dataset is from the publication https://doi.org/10.1016/j.neuroimage.2017.08.021. Please cite it if you use this dataset in your research.

  1"""The Calgary-Campinas-359 (CC-359) dataset contains annotations for brain extraction (skull-stripping)
  2in T1-weighted brain MRI.
  3
  4The dataset comprises 359 volumetric T1-weighted scans of healthy older adults (29-80 years), acquired on
  5scanners from three vendors (Siemens, Philips, General Electric) at both 1.5T and 3T field strengths, with
  6approximately 60 subjects per vendor / field strength combination. The scans are described at
  7https://www.ccdataset.com; this module downloads the data from its mirror on the Canada Open Neuroscience
  8Platform (CONP, https://github.com/conpdatasets/calgary-campinas), which serves the files directly over
  9https with no signup or click-through form (unlike the ccdataset.com website, which asks for a filled-in
 10usage form before handing out the same files).
 11
 12NOTE ON LABEL QUALITY: The brain masks are "silver-standard" labels, not manual ground truth. They were
 13created by STAPLE consensus fusion of several automated skull-stripping methods (Souza et al., NeuroImage
 142018, https://doi.org/10.1016/j.neuroimage.2017.08.021). The raw STAPLE output is a per-voxel consensus
 15probability in [0, 1] rather than a hard label; this module binarizes it at a probability of 0.5 to obtain
 16the 0/1 mask used for training. Do not treat these labels as expert manual annotations; they are best used
 17for pretraining, benchmarking segmentation agreement, or as weak supervision.
 18
 19NOTE: The dataset also ships silver-standard hippocampus masks (also STAPLE-derived), but those are stored
 20on a cropped bounding box with a different image origin/shape than the corresponding raw scan (i.e. they
 21are not defined on the same voxel grid), so they cannot be paired directly as same-shape raw/label volumes
 22and are not exposed by this module.
 23
 24The data is distributed under the CC BY-ND 4.0 license.
 25
 26This dataset is from the publication https://doi.org/10.1016/j.neuroimage.2017.08.021.
 27Please cite it if you use this dataset in your research.
 28"""
 29
 30import os
 31from glob import glob
 32from tqdm import tqdm
 33from natsort import natsorted
 34from typing import Union, Tuple, List
 35
 36import numpy as np
 37
 38from torch.utils.data import Dataset, DataLoader
 39
 40import torch_em
 41
 42from .. import util
 43
 44
 45BASE_URL = "https://sftp.conp.ca/users/calgary-campinas/CC359/Reconstructed"
 46
 47URLS = {"raw": f"{BASE_URL}/Original.zip", "brain": f"{BASE_URL}/Silver-standard-STAPLE.zip"}
 48
 49CHECKSUMS = {
 50    "raw": "c3108b56e4d2c950537b04b7b3f597c5511111dbfe02cc629610ae6538c91211",
 51    "brain": "8779fb6cd4b403d2cafe9bb8bcf72d14ed59f6b73dc8c1d39be9675fffc5a03e",
 52}
 53
 54LABEL_IDS = {"background": 0, "brain": 1}
 55
 56STAPLE_THRESHOLD = 0.5
 57
 58
 59def _preprocess_inputs(path, raw_dir, label_dir):
 60    import h5py
 61    import nibabel as nib
 62
 63    preprocessed_dir = os.path.join(path, "preprocessed")
 64    os.makedirs(preprocessed_dir, exist_ok=True)
 65
 66    raw_paths = natsorted(glob(os.path.join(raw_dir, "*.nii.gz")))
 67    for raw_path in tqdm(raw_paths, desc="Preprocessing the CC-359 volumes"):
 68        subject_id = os.path.basename(raw_path)[:-len(".nii.gz")]
 69        volume_path = os.path.join(preprocessed_dir, f"{subject_id}.h5")
 70        if os.path.exists(volume_path):
 71            continue
 72
 73        label_path = os.path.join(label_dir, f"{subject_id}_staple.nii.gz")
 74        if not os.path.exists(label_path):
 75            continue
 76
 77        raw = np.asarray(nib.load(raw_path).dataobj)
 78        label = np.asarray(nib.load(label_path).dataobj)
 79        assert raw.shape == label.shape, f"Shape mismatch for {subject_id}: {raw.shape} vs. {label.shape}."
 80        label = (label > STAPLE_THRESHOLD).astype("uint8")
 81
 82        # The file is written to a temporary path first, so that an interrupted run leaves no corrupt file.
 83        with h5py.File(f"{volume_path}.tmp", "w") as f:
 84            f.create_dataset("raw", data=raw, compression="gzip")
 85            f.create_dataset("labels", data=label, compression="gzip")
 86
 87        os.rename(f"{volume_path}.tmp", volume_path)
 88
 89    # Marks that preprocessing has finished for all subjects, so that a partially preprocessed
 90    # directory (e.g. from an interrupted previous run) is not mistaken for a complete one.
 91    open(os.path.join(preprocessed_dir, ".preprocessing_done"), "w").close()
 92
 93    return preprocessed_dir
 94
 95
 96def get_cc359_data(path: Union[os.PathLike, str], download: bool = False) -> str:
 97    """Download the CC-359 dataset.
 98
 99    Args:
100        path: Filepath to a folder where the data is downloaded for further processing.
101        download: Whether to download the data if it is not present.
102
103    Returns:
104        Filepath to the folder where the preprocessed data is stored.
105    """
106    preprocessed_dir = os.path.join(path, "preprocessed")
107    if os.path.exists(os.path.join(preprocessed_dir, ".preprocessing_done")):
108        return preprocessed_dir
109
110    os.makedirs(path, exist_ok=True)
111
112    raw_dir = os.path.join(path, "Original")
113    if not os.path.exists(raw_dir):
114        zip_path = os.path.join(path, "raw.zip")
115        util.download_source(path=zip_path, url=URLS["raw"], download=download, checksum=CHECKSUMS["raw"])
116        util.unzip(zip_path=zip_path, dst=path)
117
118    label_dir = os.path.join(path, "STAPLE")
119    if not os.path.exists(label_dir):
120        zip_path = os.path.join(path, "brain.zip")
121        util.download_source(path=zip_path, url=URLS["brain"], download=download, checksum=CHECKSUMS["brain"])
122        util.unzip(zip_path=zip_path, dst=path)
123
124    return _preprocess_inputs(path, raw_dir, label_dir)
125
126
127def get_cc359_paths(path: Union[os.PathLike, str], download: bool = False) -> Tuple[List[str], List[str]]:
128    """Get paths to the CC-359 data.
129
130    Args:
131        path: Filepath to a folder where the data is downloaded for further processing.
132        download: Whether to download the data if it is not present.
133
134    Returns:
135        List of filepaths for the hdf5 files, which contain the image data ('raw') and the label data ('labels').
136    """
137    preprocessed_dir = get_cc359_data(path, download)
138    volume_paths = natsorted(glob(os.path.join(preprocessed_dir, "*.h5")))
139    assert len(volume_paths) > 0, f"Could not find any preprocessed CC-359 volumes at '{path}'."
140    return volume_paths, volume_paths
141
142
143def get_cc359_dataset(
144    path: Union[os.PathLike, str],
145    patch_shape: Tuple[int, ...],
146    resize_inputs: bool = False,
147    download: bool = False,
148    **kwargs
149) -> Dataset:
150    """Get the CC-359 dataset for brain extraction.
151
152    Args:
153        path: Filepath to a folder where the data is downloaded for further processing.
154        patch_shape: The patch shape to use for training.
155        resize_inputs: Whether to resize inputs to the desired patch shape.
156        download: Whether to download the data if it is not present.
157        kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset`.
158
159    Returns:
160        The segmentation dataset.
161    """
162    raw_paths, label_paths = get_cc359_paths(path, download)
163
164    if resize_inputs:
165        resize_kwargs = {"patch_shape": patch_shape, "is_rgb": False}
166        kwargs, patch_shape = util.update_kwargs_for_resize_trafo(
167            kwargs=kwargs, patch_shape=patch_shape, resize_inputs=resize_inputs, resize_kwargs=resize_kwargs
168        )
169
170    return torch_em.default_segmentation_dataset(
171        raw_paths=raw_paths,
172        raw_key="raw",
173        label_paths=label_paths,
174        label_key="labels",
175        patch_shape=patch_shape,
176        is_seg_dataset=True,
177        **kwargs
178    )
179
180
181def get_cc359_loader(
182    path: Union[os.PathLike, str],
183    batch_size: int,
184    patch_shape: Tuple[int, ...],
185    resize_inputs: bool = False,
186    download: bool = False,
187    **kwargs
188) -> DataLoader:
189    """Get the CC-359 dataloader for brain extraction.
190
191    Args:
192        path: Filepath to a folder where the data is downloaded for further processing.
193        batch_size: The batch size for training.
194        patch_shape: The patch shape to use for training.
195        resize_inputs: Whether to resize inputs to the desired patch shape.
196        download: Whether to download the data if it is not present.
197        kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset` or for the PyTorch DataLoader.
198
199    Returns:
200        The DataLoader.
201    """
202    ds_kwargs, loader_kwargs = util.split_kwargs(torch_em.default_segmentation_dataset, **kwargs)
203    dataset = get_cc359_dataset(path, patch_shape, resize_inputs, download, **ds_kwargs)
204    return torch_em.get_data_loader(dataset, batch_size, **loader_kwargs)
BASE_URL = 'https://sftp.conp.ca/users/calgary-campinas/CC359/Reconstructed'
URLS = {'raw': 'https://sftp.conp.ca/users/calgary-campinas/CC359/Reconstructed/Original.zip', 'brain': 'https://sftp.conp.ca/users/calgary-campinas/CC359/Reconstructed/Silver-standard-STAPLE.zip'}
CHECKSUMS = {'raw': 'c3108b56e4d2c950537b04b7b3f597c5511111dbfe02cc629610ae6538c91211', 'brain': '8779fb6cd4b403d2cafe9bb8bcf72d14ed59f6b73dc8c1d39be9675fffc5a03e'}
LABEL_IDS = {'background': 0, 'brain': 1}
STAPLE_THRESHOLD = 0.5
def get_cc359_data(path: Union[os.PathLike, str], download: bool = False) -> str:
 97def get_cc359_data(path: Union[os.PathLike, str], download: bool = False) -> str:
 98    """Download the CC-359 dataset.
 99
100    Args:
101        path: Filepath to a folder where the data is downloaded for further processing.
102        download: Whether to download the data if it is not present.
103
104    Returns:
105        Filepath to the folder where the preprocessed data is stored.
106    """
107    preprocessed_dir = os.path.join(path, "preprocessed")
108    if os.path.exists(os.path.join(preprocessed_dir, ".preprocessing_done")):
109        return preprocessed_dir
110
111    os.makedirs(path, exist_ok=True)
112
113    raw_dir = os.path.join(path, "Original")
114    if not os.path.exists(raw_dir):
115        zip_path = os.path.join(path, "raw.zip")
116        util.download_source(path=zip_path, url=URLS["raw"], download=download, checksum=CHECKSUMS["raw"])
117        util.unzip(zip_path=zip_path, dst=path)
118
119    label_dir = os.path.join(path, "STAPLE")
120    if not os.path.exists(label_dir):
121        zip_path = os.path.join(path, "brain.zip")
122        util.download_source(path=zip_path, url=URLS["brain"], download=download, checksum=CHECKSUMS["brain"])
123        util.unzip(zip_path=zip_path, dst=path)
124
125    return _preprocess_inputs(path, raw_dir, label_dir)

Download the CC-359 dataset.

Arguments:
  • path: Filepath to a folder where the data is downloaded for further processing.
  • download: Whether to download the data if it is not present.
Returns:

Filepath to the folder where the preprocessed data is stored.

def get_cc359_paths( path: Union[os.PathLike, str], download: bool = False) -> Tuple[List[str], List[str]]:
128def get_cc359_paths(path: Union[os.PathLike, str], download: bool = False) -> Tuple[List[str], List[str]]:
129    """Get paths to the CC-359 data.
130
131    Args:
132        path: Filepath to a folder where the data is downloaded for further processing.
133        download: Whether to download the data if it is not present.
134
135    Returns:
136        List of filepaths for the hdf5 files, which contain the image data ('raw') and the label data ('labels').
137    """
138    preprocessed_dir = get_cc359_data(path, download)
139    volume_paths = natsorted(glob(os.path.join(preprocessed_dir, "*.h5")))
140    assert len(volume_paths) > 0, f"Could not find any preprocessed CC-359 volumes at '{path}'."
141    return volume_paths, volume_paths

Get paths to the CC-359 data.

Arguments:
  • path: Filepath to a folder where the data is downloaded for further processing.
  • download: Whether to download the data if it is not present.
Returns:

List of filepaths for the hdf5 files, which contain the image data ('raw') and the label data ('labels').

def get_cc359_dataset( path: Union[os.PathLike, str], patch_shape: Tuple[int, ...], resize_inputs: bool = False, download: bool = False, **kwargs) -> torch.utils.data.dataset.Dataset:
144def get_cc359_dataset(
145    path: Union[os.PathLike, str],
146    patch_shape: Tuple[int, ...],
147    resize_inputs: bool = False,
148    download: bool = False,
149    **kwargs
150) -> Dataset:
151    """Get the CC-359 dataset for brain extraction.
152
153    Args:
154        path: Filepath to a folder where the data is downloaded for further processing.
155        patch_shape: The patch shape to use for training.
156        resize_inputs: Whether to resize inputs to the desired patch shape.
157        download: Whether to download the data if it is not present.
158        kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset`.
159
160    Returns:
161        The segmentation dataset.
162    """
163    raw_paths, label_paths = get_cc359_paths(path, download)
164
165    if resize_inputs:
166        resize_kwargs = {"patch_shape": patch_shape, "is_rgb": False}
167        kwargs, patch_shape = util.update_kwargs_for_resize_trafo(
168            kwargs=kwargs, patch_shape=patch_shape, resize_inputs=resize_inputs, resize_kwargs=resize_kwargs
169        )
170
171    return torch_em.default_segmentation_dataset(
172        raw_paths=raw_paths,
173        raw_key="raw",
174        label_paths=label_paths,
175        label_key="labels",
176        patch_shape=patch_shape,
177        is_seg_dataset=True,
178        **kwargs
179    )

Get the CC-359 dataset for brain extraction.

Arguments:
  • path: Filepath to a folder where the data is downloaded for further processing.
  • patch_shape: The patch shape to use for training.
  • resize_inputs: Whether to resize inputs to the desired patch shape.
  • download: Whether to download the data if it is not present.
  • kwargs: Additional keyword arguments for torch_em.default_segmentation_dataset.
Returns:

The segmentation dataset.

def get_cc359_loader( path: Union[os.PathLike, str], batch_size: int, patch_shape: Tuple[int, ...], resize_inputs: bool = False, download: bool = False, **kwargs) -> torch.utils.data.dataloader.DataLoader:
182def get_cc359_loader(
183    path: Union[os.PathLike, str],
184    batch_size: int,
185    patch_shape: Tuple[int, ...],
186    resize_inputs: bool = False,
187    download: bool = False,
188    **kwargs
189) -> DataLoader:
190    """Get the CC-359 dataloader for brain extraction.
191
192    Args:
193        path: Filepath to a folder where the data is downloaded for further processing.
194        batch_size: The batch size for training.
195        patch_shape: The patch shape to use for training.
196        resize_inputs: Whether to resize inputs to the desired patch shape.
197        download: Whether to download the data if it is not present.
198        kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset` or for the PyTorch DataLoader.
199
200    Returns:
201        The DataLoader.
202    """
203    ds_kwargs, loader_kwargs = util.split_kwargs(torch_em.default_segmentation_dataset, **kwargs)
204    dataset = get_cc359_dataset(path, patch_shape, resize_inputs, download, **ds_kwargs)
205    return torch_em.get_data_loader(dataset, batch_size, **loader_kwargs)

Get the CC-359 dataloader for brain extraction.

Arguments:
  • path: Filepath to a folder where the data is downloaded for further processing.
  • batch_size: The batch size for training.
  • patch_shape: The patch shape to use for training.
  • resize_inputs: Whether to resize inputs to the desired patch shape.
  • download: Whether to download the data if it is not present.
  • kwargs: Additional keyword arguments for torch_em.default_segmentation_dataset or for the PyTorch DataLoader.
Returns:

The DataLoader.