torch_em.data.datasets.medical.cc359
The Calgary-Campinas-359 (CC-359) dataset contains annotations for brain extraction (skull-stripping) in T1-weighted brain MRI.
The dataset comprises 359 volumetric T1-weighted scans of healthy older adults (29-80 years), acquired on scanners from three vendors (Siemens, Philips, General Electric) at both 1.5T and 3T field strengths, with approximately 60 subjects per vendor / field strength combination. The scans are described at https://www.ccdataset.com; this module downloads the data from its mirror on the Canada Open Neuroscience Platform (CONP, https://github.com/conpdatasets/calgary-campinas), which serves the files directly over https with no signup or click-through form (unlike the ccdataset.com website, which asks for a filled-in usage form before handing out the same files).
NOTE ON LABEL QUALITY: The brain masks are "silver-standard" labels, not manual ground truth. They were created by STAPLE consensus fusion of several automated skull-stripping methods (Souza et al., NeuroImage 2018, https://doi.org/10.1016/j.neuroimage.2017.08.021). The raw STAPLE output is a per-voxel consensus probability in [0, 1] rather than a hard label; this module binarizes it at a probability of 0.5 to obtain the 0/1 mask used for training. Do not treat these labels as expert manual annotations; they are best used for pretraining, benchmarking segmentation agreement, or as weak supervision.
NOTE: The dataset also ships silver-standard hippocampus masks (also STAPLE-derived), but those are stored on a cropped bounding box with a different image origin/shape than the corresponding raw scan (i.e. they are not defined on the same voxel grid), so they cannot be paired directly as same-shape raw/label volumes and are not exposed by this module.
The data is distributed under the CC BY-ND 4.0 license.
This dataset is from the publication https://doi.org/10.1016/j.neuroimage.2017.08.021. Please cite it if you use this dataset in your research.
1"""The Calgary-Campinas-359 (CC-359) dataset contains annotations for brain extraction (skull-stripping) 2in T1-weighted brain MRI. 3 4The dataset comprises 359 volumetric T1-weighted scans of healthy older adults (29-80 years), acquired on 5scanners from three vendors (Siemens, Philips, General Electric) at both 1.5T and 3T field strengths, with 6approximately 60 subjects per vendor / field strength combination. The scans are described at 7https://www.ccdataset.com; this module downloads the data from its mirror on the Canada Open Neuroscience 8Platform (CONP, https://github.com/conpdatasets/calgary-campinas), which serves the files directly over 9https with no signup or click-through form (unlike the ccdataset.com website, which asks for a filled-in 10usage form before handing out the same files). 11 12NOTE ON LABEL QUALITY: The brain masks are "silver-standard" labels, not manual ground truth. They were 13created by STAPLE consensus fusion of several automated skull-stripping methods (Souza et al., NeuroImage 142018, https://doi.org/10.1016/j.neuroimage.2017.08.021). The raw STAPLE output is a per-voxel consensus 15probability in [0, 1] rather than a hard label; this module binarizes it at a probability of 0.5 to obtain 16the 0/1 mask used for training. Do not treat these labels as expert manual annotations; they are best used 17for pretraining, benchmarking segmentation agreement, or as weak supervision. 18 19NOTE: The dataset also ships silver-standard hippocampus masks (also STAPLE-derived), but those are stored 20on a cropped bounding box with a different image origin/shape than the corresponding raw scan (i.e. they 21are not defined on the same voxel grid), so they cannot be paired directly as same-shape raw/label volumes 22and are not exposed by this module. 23 24The data is distributed under the CC BY-ND 4.0 license. 25 26This dataset is from the publication https://doi.org/10.1016/j.neuroimage.2017.08.021. 27Please cite it if you use this dataset in your research. 28""" 29 30import os 31from glob import glob 32from tqdm import tqdm 33from natsort import natsorted 34from typing import Union, Tuple, List 35 36import numpy as np 37 38from torch.utils.data import Dataset, DataLoader 39 40import torch_em 41 42from .. import util 43 44 45BASE_URL = "https://sftp.conp.ca/users/calgary-campinas/CC359/Reconstructed" 46 47URLS = {"raw": f"{BASE_URL}/Original.zip", "brain": f"{BASE_URL}/Silver-standard-STAPLE.zip"} 48 49CHECKSUMS = { 50 "raw": "c3108b56e4d2c950537b04b7b3f597c5511111dbfe02cc629610ae6538c91211", 51 "brain": "8779fb6cd4b403d2cafe9bb8bcf72d14ed59f6b73dc8c1d39be9675fffc5a03e", 52} 53 54LABEL_IDS = {"background": 0, "brain": 1} 55 56STAPLE_THRESHOLD = 0.5 57 58 59def _preprocess_inputs(path, raw_dir, label_dir): 60 import h5py 61 import nibabel as nib 62 63 preprocessed_dir = os.path.join(path, "preprocessed") 64 os.makedirs(preprocessed_dir, exist_ok=True) 65 66 raw_paths = natsorted(glob(os.path.join(raw_dir, "*.nii.gz"))) 67 for raw_path in tqdm(raw_paths, desc="Preprocessing the CC-359 volumes"): 68 subject_id = os.path.basename(raw_path)[:-len(".nii.gz")] 69 volume_path = os.path.join(preprocessed_dir, f"{subject_id}.h5") 70 if os.path.exists(volume_path): 71 continue 72 73 label_path = os.path.join(label_dir, f"{subject_id}_staple.nii.gz") 74 if not os.path.exists(label_path): 75 continue 76 77 raw = np.asarray(nib.load(raw_path).dataobj) 78 label = np.asarray(nib.load(label_path).dataobj) 79 assert raw.shape == label.shape, f"Shape mismatch for {subject_id}: {raw.shape} vs. {label.shape}." 80 label = (label > STAPLE_THRESHOLD).astype("uint8") 81 82 # The file is written to a temporary path first, so that an interrupted run leaves no corrupt file. 83 with h5py.File(f"{volume_path}.tmp", "w") as f: 84 f.create_dataset("raw", data=raw, compression="gzip") 85 f.create_dataset("labels", data=label, compression="gzip") 86 87 os.rename(f"{volume_path}.tmp", volume_path) 88 89 # Marks that preprocessing has finished for all subjects, so that a partially preprocessed 90 # directory (e.g. from an interrupted previous run) is not mistaken for a complete one. 91 open(os.path.join(preprocessed_dir, ".preprocessing_done"), "w").close() 92 93 return preprocessed_dir 94 95 96def get_cc359_data(path: Union[os.PathLike, str], download: bool = False) -> str: 97 """Download the CC-359 dataset. 98 99 Args: 100 path: Filepath to a folder where the data is downloaded for further processing. 101 download: Whether to download the data if it is not present. 102 103 Returns: 104 Filepath to the folder where the preprocessed data is stored. 105 """ 106 preprocessed_dir = os.path.join(path, "preprocessed") 107 if os.path.exists(os.path.join(preprocessed_dir, ".preprocessing_done")): 108 return preprocessed_dir 109 110 os.makedirs(path, exist_ok=True) 111 112 raw_dir = os.path.join(path, "Original") 113 if not os.path.exists(raw_dir): 114 zip_path = os.path.join(path, "raw.zip") 115 util.download_source(path=zip_path, url=URLS["raw"], download=download, checksum=CHECKSUMS["raw"]) 116 util.unzip(zip_path=zip_path, dst=path) 117 118 label_dir = os.path.join(path, "STAPLE") 119 if not os.path.exists(label_dir): 120 zip_path = os.path.join(path, "brain.zip") 121 util.download_source(path=zip_path, url=URLS["brain"], download=download, checksum=CHECKSUMS["brain"]) 122 util.unzip(zip_path=zip_path, dst=path) 123 124 return _preprocess_inputs(path, raw_dir, label_dir) 125 126 127def get_cc359_paths(path: Union[os.PathLike, str], download: bool = False) -> Tuple[List[str], List[str]]: 128 """Get paths to the CC-359 data. 129 130 Args: 131 path: Filepath to a folder where the data is downloaded for further processing. 132 download: Whether to download the data if it is not present. 133 134 Returns: 135 List of filepaths for the hdf5 files, which contain the image data ('raw') and the label data ('labels'). 136 """ 137 preprocessed_dir = get_cc359_data(path, download) 138 volume_paths = natsorted(glob(os.path.join(preprocessed_dir, "*.h5"))) 139 assert len(volume_paths) > 0, f"Could not find any preprocessed CC-359 volumes at '{path}'." 140 return volume_paths, volume_paths 141 142 143def get_cc359_dataset( 144 path: Union[os.PathLike, str], 145 patch_shape: Tuple[int, ...], 146 resize_inputs: bool = False, 147 download: bool = False, 148 **kwargs 149) -> Dataset: 150 """Get the CC-359 dataset for brain extraction. 151 152 Args: 153 path: Filepath to a folder where the data is downloaded for further processing. 154 patch_shape: The patch shape to use for training. 155 resize_inputs: Whether to resize inputs to the desired patch shape. 156 download: Whether to download the data if it is not present. 157 kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset`. 158 159 Returns: 160 The segmentation dataset. 161 """ 162 raw_paths, label_paths = get_cc359_paths(path, download) 163 164 if resize_inputs: 165 resize_kwargs = {"patch_shape": patch_shape, "is_rgb": False} 166 kwargs, patch_shape = util.update_kwargs_for_resize_trafo( 167 kwargs=kwargs, patch_shape=patch_shape, resize_inputs=resize_inputs, resize_kwargs=resize_kwargs 168 ) 169 170 return torch_em.default_segmentation_dataset( 171 raw_paths=raw_paths, 172 raw_key="raw", 173 label_paths=label_paths, 174 label_key="labels", 175 patch_shape=patch_shape, 176 is_seg_dataset=True, 177 **kwargs 178 ) 179 180 181def get_cc359_loader( 182 path: Union[os.PathLike, str], 183 batch_size: int, 184 patch_shape: Tuple[int, ...], 185 resize_inputs: bool = False, 186 download: bool = False, 187 **kwargs 188) -> DataLoader: 189 """Get the CC-359 dataloader for brain extraction. 190 191 Args: 192 path: Filepath to a folder where the data is downloaded for further processing. 193 batch_size: The batch size for training. 194 patch_shape: The patch shape to use for training. 195 resize_inputs: Whether to resize inputs to the desired patch shape. 196 download: Whether to download the data if it is not present. 197 kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset` or for the PyTorch DataLoader. 198 199 Returns: 200 The DataLoader. 201 """ 202 ds_kwargs, loader_kwargs = util.split_kwargs(torch_em.default_segmentation_dataset, **kwargs) 203 dataset = get_cc359_dataset(path, patch_shape, resize_inputs, download, **ds_kwargs) 204 return torch_em.get_data_loader(dataset, batch_size, **loader_kwargs)
97def get_cc359_data(path: Union[os.PathLike, str], download: bool = False) -> str: 98 """Download the CC-359 dataset. 99 100 Args: 101 path: Filepath to a folder where the data is downloaded for further processing. 102 download: Whether to download the data if it is not present. 103 104 Returns: 105 Filepath to the folder where the preprocessed data is stored. 106 """ 107 preprocessed_dir = os.path.join(path, "preprocessed") 108 if os.path.exists(os.path.join(preprocessed_dir, ".preprocessing_done")): 109 return preprocessed_dir 110 111 os.makedirs(path, exist_ok=True) 112 113 raw_dir = os.path.join(path, "Original") 114 if not os.path.exists(raw_dir): 115 zip_path = os.path.join(path, "raw.zip") 116 util.download_source(path=zip_path, url=URLS["raw"], download=download, checksum=CHECKSUMS["raw"]) 117 util.unzip(zip_path=zip_path, dst=path) 118 119 label_dir = os.path.join(path, "STAPLE") 120 if not os.path.exists(label_dir): 121 zip_path = os.path.join(path, "brain.zip") 122 util.download_source(path=zip_path, url=URLS["brain"], download=download, checksum=CHECKSUMS["brain"]) 123 util.unzip(zip_path=zip_path, dst=path) 124 125 return _preprocess_inputs(path, raw_dir, label_dir)
Download the CC-359 dataset.
Arguments:
- path: Filepath to a folder where the data is downloaded for further processing.
- download: Whether to download the data if it is not present.
Returns:
Filepath to the folder where the preprocessed data is stored.
128def get_cc359_paths(path: Union[os.PathLike, str], download: bool = False) -> Tuple[List[str], List[str]]: 129 """Get paths to the CC-359 data. 130 131 Args: 132 path: Filepath to a folder where the data is downloaded for further processing. 133 download: Whether to download the data if it is not present. 134 135 Returns: 136 List of filepaths for the hdf5 files, which contain the image data ('raw') and the label data ('labels'). 137 """ 138 preprocessed_dir = get_cc359_data(path, download) 139 volume_paths = natsorted(glob(os.path.join(preprocessed_dir, "*.h5"))) 140 assert len(volume_paths) > 0, f"Could not find any preprocessed CC-359 volumes at '{path}'." 141 return volume_paths, volume_paths
Get paths to the CC-359 data.
Arguments:
- path: Filepath to a folder where the data is downloaded for further processing.
- download: Whether to download the data if it is not present.
Returns:
List of filepaths for the hdf5 files, which contain the image data ('raw') and the label data ('labels').
144def get_cc359_dataset( 145 path: Union[os.PathLike, str], 146 patch_shape: Tuple[int, ...], 147 resize_inputs: bool = False, 148 download: bool = False, 149 **kwargs 150) -> Dataset: 151 """Get the CC-359 dataset for brain extraction. 152 153 Args: 154 path: Filepath to a folder where the data is downloaded for further processing. 155 patch_shape: The patch shape to use for training. 156 resize_inputs: Whether to resize inputs to the desired patch shape. 157 download: Whether to download the data if it is not present. 158 kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset`. 159 160 Returns: 161 The segmentation dataset. 162 """ 163 raw_paths, label_paths = get_cc359_paths(path, download) 164 165 if resize_inputs: 166 resize_kwargs = {"patch_shape": patch_shape, "is_rgb": False} 167 kwargs, patch_shape = util.update_kwargs_for_resize_trafo( 168 kwargs=kwargs, patch_shape=patch_shape, resize_inputs=resize_inputs, resize_kwargs=resize_kwargs 169 ) 170 171 return torch_em.default_segmentation_dataset( 172 raw_paths=raw_paths, 173 raw_key="raw", 174 label_paths=label_paths, 175 label_key="labels", 176 patch_shape=patch_shape, 177 is_seg_dataset=True, 178 **kwargs 179 )
Get the CC-359 dataset for brain extraction.
Arguments:
- path: Filepath to a folder where the data is downloaded for further processing.
- patch_shape: The patch shape to use for training.
- resize_inputs: Whether to resize inputs to the desired patch shape.
- download: Whether to download the data if it is not present.
- kwargs: Additional keyword arguments for
torch_em.default_segmentation_dataset.
Returns:
The segmentation dataset.
182def get_cc359_loader( 183 path: Union[os.PathLike, str], 184 batch_size: int, 185 patch_shape: Tuple[int, ...], 186 resize_inputs: bool = False, 187 download: bool = False, 188 **kwargs 189) -> DataLoader: 190 """Get the CC-359 dataloader for brain extraction. 191 192 Args: 193 path: Filepath to a folder where the data is downloaded for further processing. 194 batch_size: The batch size for training. 195 patch_shape: The patch shape to use for training. 196 resize_inputs: Whether to resize inputs to the desired patch shape. 197 download: Whether to download the data if it is not present. 198 kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset` or for the PyTorch DataLoader. 199 200 Returns: 201 The DataLoader. 202 """ 203 ds_kwargs, loader_kwargs = util.split_kwargs(torch_em.default_segmentation_dataset, **kwargs) 204 dataset = get_cc359_dataset(path, patch_shape, resize_inputs, download, **ds_kwargs) 205 return torch_em.get_data_loader(dataset, batch_size, **loader_kwargs)
Get the CC-359 dataloader for brain extraction.
Arguments:
- path: Filepath to a folder where the data is downloaded for further processing.
- batch_size: The batch size for training.
- patch_shape: The patch shape to use for training.
- resize_inputs: Whether to resize inputs to the desired patch shape.
- download: Whether to download the data if it is not present.
- kwargs: Additional keyword arguments for
torch_em.default_segmentation_datasetor for the PyTorch DataLoader.
Returns:
The DataLoader.