torch_em.data.datasets.medical.stanford_coca
The COCA dataset contains annotations for coronary artery calcium segmentation in ECG-gated cardiac CT.
The dataset is released by the Stanford Center for Artificial Intelligence in Medicine and Imaging (AIMI) as "COCA - Coronary Calcium and chest CT's". It comprises two parts: gated coronary CT DICOM series with per-slice coronary artery calcium ROI annotations (stored as a binary property list '.xml' file per patient), and non-gated chest CT DICOM series with only a coronary artery calcium score (no pixel-level annotation). This module only supports the gated part, since only that part has pixel-level segmentation masks; the non-gated part is score-only and does not fit torch_em's segmentation dataset pattern.
The ROI file lists, per slice ('ImageIndex') and per coronary artery, one or more closed polygons ('Point_px', in pixel coordinates) outlining the calcified regions. The label ids used here are:
- 0: background, 1: right coronary artery, 2: left anterior descending artery,
- 3: left coronary artery, 4: left circumflex artery
(see
LABEL_IDS; the artery names and this correspondence were taken from a community reimplementation of the dataset loading, https://github.com/msingh9/cs230-Coronary-Calcium-Scoring-/blob/master/code/my_lib.py, since the ROI file format itself is undocumented by Stanford AIMI).
NOTE: The slice that a set of ROIs belongs to is identified in the ROI file only by an integer 'ImageIndex', with no reference to a DICOM SOP instance UID or file name. This module assumes that 'ImageIndex' is the 0-based position of the slice when the series is sorted by 'InstanceNumber'. This assumption could not be verified against the real data, since the dataset requires registration and a signed data use agreement, so it was not accessible while writing this module. Please verify the resulting masks against the DICOM series (e.g. by overlaying them) before relying on this dataset.
NOTE: The dataset requires registration and cannot be downloaded automatically. Please follow these steps:
- Visit https://aimi.stanford.edu/datasets/coca-coronary-calcium-chest-ct and follow the link to the dataset on Stanford's Redivis platform (https://stanford.redivis.com/datasets/1vm9-30b7p5srg).
- Log in with your institutional account, fill in your contact details and accept the Stanford University Dataset Research Use Agreement.
- Download the 'Gated_release_final' folder (e.g. with the Redivis download tool) and place it such that
'
/Gated_release_final/patient/ /.../*.dcm' (the DICOM series) and ' /Gated_release_final/calcium_xml/ .xml' (the ROI annotations) exist.
The dataset is located at https://aimi.stanford.edu/datasets/coca-coronary-calcium-chest-ct (DOI: https://doi.org/10.71718/ge5g-ds80) and is released under the Stanford University Dataset Research Use Agreement.
This dataset is from the publications https://doi.org/10.1007/s10554-019-01710-8 and https://doi.org/10.1148/ryai.2020190004. Please cite them if you use this dataset in your research.
NOTE: The DICOM and ROI parsing requires 'pydicom'. Install it with 'pip install pydicom'.
1"""The COCA dataset contains annotations for coronary artery calcium segmentation in ECG-gated 2cardiac CT. 3 4The dataset is released by the Stanford Center for Artificial Intelligence in Medicine and Imaging 5(AIMI) as "COCA - Coronary Calcium and chest CT's". It comprises two parts: gated coronary CT DICOM 6series with per-slice coronary artery calcium ROI annotations (stored as a binary property list 7'.xml' file per patient), and non-gated chest CT DICOM series with only a coronary artery calcium 8score (no pixel-level annotation). This module only supports the gated part, since only that part 9has pixel-level segmentation masks; the non-gated part is score-only and does not fit torch_em's 10segmentation dataset pattern. 11 12The ROI file lists, per slice ('ImageIndex') and per coronary artery, one or more closed polygons 13('Point_px', in pixel coordinates) outlining the calcified regions. The label ids used here are: 14- 0: background, 1: right coronary artery, 2: left anterior descending artery, 15- 3: left coronary artery, 4: left circumflex artery 16(see `LABEL_IDS`; the artery names and this correspondence were taken from a community reimplementation 17of the dataset loading, https://github.com/msingh9/cs230-Coronary-Calcium-Scoring-/blob/master/code/my_lib.py, 18since the ROI file format itself is undocumented by Stanford AIMI). 19 20NOTE: The slice that a set of ROIs belongs to is identified in the ROI file only by an integer 21'ImageIndex', with no reference to a DICOM SOP instance UID or file name. This module assumes that 22'ImageIndex' is the 0-based position of the slice when the series is sorted by 'InstanceNumber'. 23This assumption could not be verified against the real data, since the dataset requires registration 24and a signed data use agreement, so it was not accessible while writing this module. Please verify 25the resulting masks against the DICOM series (e.g. by overlaying them) before relying on this dataset. 26 27NOTE: The dataset requires registration and cannot be downloaded automatically. Please follow these steps: 28- Visit https://aimi.stanford.edu/datasets/coca-coronary-calcium-chest-ct and follow the link to the 29 dataset on Stanford's Redivis platform (https://stanford.redivis.com/datasets/1vm9-30b7p5srg). 30- Log in with your institutional account, fill in your contact details and accept the Stanford University 31 Dataset Research Use Agreement. 32- Download the 'Gated_release_final' folder (e.g. with the Redivis download tool) and place it such that 33 '<path>/Gated_release_final/patient/<patient_id>/.../*.dcm' (the DICOM series) and 34 '<path>/Gated_release_final/calcium_xml/<patient_id>.xml' (the ROI annotations) exist. 35 36The dataset is located at https://aimi.stanford.edu/datasets/coca-coronary-calcium-chest-ct 37(DOI: https://doi.org/10.71718/ge5g-ds80) and is released under the Stanford University Dataset 38Research Use Agreement. 39 40This dataset is from the publications https://doi.org/10.1007/s10554-019-01710-8 and 41https://doi.org/10.1148/ryai.2020190004. Please cite them if you use this dataset in your research. 42 43NOTE: The DICOM and ROI parsing requires 'pydicom'. Install it with 'pip install pydicom'. 44""" 45 46import os 47import plistlib 48from glob import glob 49from tqdm import tqdm 50from natsort import natsorted 51from typing import Union, Tuple, List 52 53import numpy as np 54 55from torch.utils.data import Dataset, DataLoader 56 57import torch_em 58 59from .. import util 60 61 62LABEL_IDS = { 63 "background": 0, 64 "right_coronary_artery": 1, 65 "left_anterior_descending_artery": 2, 66 "left_coronary_artery": 3, 67 "left_circumflex_artery": 4, 68} 69 70ROI_NAME_TO_LABEL_ID = { 71 "Right Coronary Artery": LABEL_IDS["right_coronary_artery"], 72 "Left Anterior Descending Artery": LABEL_IDS["left_anterior_descending_artery"], 73 "Left Coronary Artery": LABEL_IDS["left_coronary_artery"], 74 "Left Circumflex Artery": LABEL_IDS["left_circumflex_artery"], 75} 76 77 78def _load_gated_series(patient_dir): 79 import pydicom 80 81 dcm_paths = natsorted(glob(os.path.join(patient_dir, "**", "*.dcm"), recursive=True)) 82 assert len(dcm_paths) > 0, f"Could not find any DICOM files in '{patient_dir}'." 83 84 slices = [pydicom.dcmread(p) for p in dcm_paths] 85 slices.sort(key=lambda dcm: int(dcm.InstanceNumber)) 86 87 volume = np.stack([dcm.pixel_array for dcm in slices]).astype("float32") 88 slope = float(getattr(slices[0], "RescaleSlope", 1.0)) 89 intercept = float(getattr(slices[0], "RescaleIntercept", 0.0)) 90 volume = np.round(volume * slope + intercept).astype("int16") 91 92 return volume 93 94 95def _parse_calcium_xml(xml_path): 96 with open(xml_path, "rb") as f: 97 annotations = plistlib.load(f) 98 99 rois_per_slice = {} 100 for image in annotations["Images"]: 101 image_index = image["ImageIndex"] 102 for roi in image["ROIs"]: 103 points = roi.get("Point_px", []) 104 if len(points) == 0: 105 continue 106 107 label_id = ROI_NAME_TO_LABEL_ID.get(roi["Name"]) 108 if label_id is None: 109 continue 110 111 pixels = np.array([[float(v) for v in point.strip("()").split(",")] for point in points]) 112 rois_per_slice.setdefault(image_index, []).append((label_id, pixels)) 113 114 return rois_per_slice 115 116 117def _rasterize_labels(shape, rois_per_slice): 118 from skimage.draw import polygon 119 120 labels = np.zeros(shape, dtype="uint8") 121 for slice_id, rois in rois_per_slice.items(): 122 if slice_id >= shape[0]: 123 continue 124 for label_id, pixels in rois: 125 rr, cc = polygon(pixels[:, 1], pixels[:, 0], shape=shape[1:]) 126 labels[slice_id, rr, cc] = label_id 127 128 return labels 129 130 131def _preprocess_inputs(path, gated_dir): 132 import h5py 133 134 preprocessed_dir = os.path.join(path, "preprocessed") 135 os.makedirs(preprocessed_dir, exist_ok=True) 136 137 xml_paths = natsorted(glob(os.path.join(gated_dir, "calcium_xml", "*.xml"))) 138 for xml_path in tqdm(xml_paths, desc="Preprocessing the COCA gated CT scans"): 139 patient_id = os.path.basename(xml_path)[:-len(".xml")] 140 volume_path = os.path.join(preprocessed_dir, f"{patient_id}.h5") 141 if os.path.exists(volume_path): 142 continue 143 144 patient_dir = os.path.join(gated_dir, "patient", patient_id) 145 if not os.path.isdir(patient_dir): 146 continue 147 148 raw = _load_gated_series(patient_dir) 149 rois_per_slice = _parse_calcium_xml(xml_path) 150 labels = _rasterize_labels(raw.shape, rois_per_slice) 151 152 # The file is written to a temporary path first, so that an interrupted run does not leave a corrupt file. 153 with h5py.File(f"{volume_path}.tmp", "w") as f: 154 f.create_dataset("raw", data=raw, compression="gzip") 155 f.create_dataset("labels", data=labels, compression="gzip") 156 157 os.rename(f"{volume_path}.tmp", volume_path) 158 159 return preprocessed_dir 160 161 162def get_stanford_coca_data(path: Union[os.PathLike, str], download: bool = False) -> str: 163 """Obtain the COCA dataset. 164 165 Args: 166 path: Filepath to a folder where the data is downloaded for further processing. 167 download: Whether to download the data if it is not present. 168 169 Returns: 170 Filepath where the data is preprocessed. 171 """ 172 preprocessed_dir = os.path.join(path, "preprocessed") 173 if os.path.exists(preprocessed_dir) and len(glob(os.path.join(preprocessed_dir, "*.h5"))) > 0: 174 return preprocessed_dir 175 176 if download: 177 raise RuntimeError( 178 "Download is set to True, but 'torch_em' cannot download the COCA dataset automatically: " 179 "it requires registration and a signed Stanford University Dataset Research Use Agreement. " 180 "See 'torch_em.data.datasets.medical.stanford_coca' for the manual download instructions." 181 ) 182 183 xml_dirs = [p for p in glob(os.path.join(path, "**", "calcium_xml"), recursive=True) if os.path.isdir(p)] 184 if len(xml_dirs) == 0: 185 raise RuntimeError( 186 f"Could not find the COCA 'Gated_release_final' folder at '{path}'. This dataset requires " 187 "registration and a signed Stanford University Dataset Research Use Agreement, so it cannot be " 188 "downloaded automatically. Please follow these steps: " 189 "1) Visit https://aimi.stanford.edu/datasets/coca-coronary-calcium-chest-ct and follow the link " 190 "to the dataset on Stanford's Redivis platform (https://stanford.redivis.com/datasets/1vm9-30b7p5srg). " 191 "2) Log in with your institutional account, fill in your contact details and accept the Stanford " 192 "University Dataset Research Use Agreement. " 193 f"3) Download the 'Gated_release_final' folder and place it such that " 194 f"'{path}/Gated_release_final/patient/<patient_id>/.../*.dcm' and " 195 f"'{path}/Gated_release_final/calcium_xml/<patient_id>.xml' exist." 196 ) 197 198 gated_dir = os.path.dirname(xml_dirs[0]) 199 return _preprocess_inputs(path, gated_dir) 200 201 202def get_stanford_coca_paths(path: Union[os.PathLike, str], download: bool = False) -> Tuple[List[str], List[str]]: 203 """Get paths to the COCA data. 204 205 Args: 206 path: Filepath to a folder where the data is downloaded for further processing. 207 download: Whether to download the data if it is not present. 208 209 Returns: 210 List of filepaths for the image data. 211 List of filepaths for the label data. 212 """ 213 data_dir = get_stanford_coca_data(path, download) 214 volume_paths = natsorted(glob(os.path.join(data_dir, "*.h5"))) 215 assert len(volume_paths) > 0, f"Could not find any preprocessed volumes in '{data_dir}'." 216 return volume_paths, volume_paths 217 218 219def get_stanford_coca_dataset( 220 path: Union[os.PathLike, str], 221 patch_shape: Tuple[int, ...], 222 resize_inputs: bool = False, 223 download: bool = False, 224 **kwargs 225) -> Dataset: 226 """Get the COCA dataset for coronary artery calcium segmentation. 227 228 Args: 229 path: Filepath to a folder where the data is downloaded for further processing. 230 patch_shape: The patch shape to use for training. 231 resize_inputs: Whether to resize inputs to the desired patch shape. 232 download: Whether to download the data if it is not present. 233 kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset`. 234 235 Returns: 236 The segmentation dataset. 237 """ 238 raw_paths, label_paths = get_stanford_coca_paths(path, download) 239 240 if resize_inputs: 241 resize_kwargs = {"patch_shape": patch_shape, "is_rgb": False} 242 kwargs, patch_shape = util.update_kwargs_for_resize_trafo( 243 kwargs=kwargs, patch_shape=patch_shape, resize_inputs=resize_inputs, resize_kwargs=resize_kwargs 244 ) 245 246 return torch_em.default_segmentation_dataset( 247 raw_paths=raw_paths, 248 raw_key="raw", 249 label_paths=label_paths, 250 label_key="labels", 251 patch_shape=patch_shape, 252 is_seg_dataset=True, 253 **kwargs 254 ) 255 256 257def get_stanford_coca_loader( 258 path: Union[os.PathLike, str], 259 batch_size: int, 260 patch_shape: Tuple[int, ...], 261 resize_inputs: bool = False, 262 download: bool = False, 263 **kwargs 264) -> DataLoader: 265 """Get the COCA dataloader for coronary artery calcium segmentation. 266 267 Args: 268 path: Filepath to a folder where the data is downloaded for further processing. 269 batch_size: The batch size for training. 270 patch_shape: The patch shape to use for training. 271 resize_inputs: Whether to resize inputs to the desired patch shape. 272 download: Whether to download the data if it is not present. 273 kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset` or for the PyTorch DataLoader. 274 275 Returns: 276 The DataLoader. 277 """ 278 ds_kwargs, loader_kwargs = util.split_kwargs(torch_em.default_segmentation_dataset, **kwargs) 279 dataset = get_stanford_coca_dataset(path, patch_shape, resize_inputs, download, **ds_kwargs) 280 return torch_em.get_data_loader(dataset, batch_size, **loader_kwargs)
163def get_stanford_coca_data(path: Union[os.PathLike, str], download: bool = False) -> str: 164 """Obtain the COCA dataset. 165 166 Args: 167 path: Filepath to a folder where the data is downloaded for further processing. 168 download: Whether to download the data if it is not present. 169 170 Returns: 171 Filepath where the data is preprocessed. 172 """ 173 preprocessed_dir = os.path.join(path, "preprocessed") 174 if os.path.exists(preprocessed_dir) and len(glob(os.path.join(preprocessed_dir, "*.h5"))) > 0: 175 return preprocessed_dir 176 177 if download: 178 raise RuntimeError( 179 "Download is set to True, but 'torch_em' cannot download the COCA dataset automatically: " 180 "it requires registration and a signed Stanford University Dataset Research Use Agreement. " 181 "See 'torch_em.data.datasets.medical.stanford_coca' for the manual download instructions." 182 ) 183 184 xml_dirs = [p for p in glob(os.path.join(path, "**", "calcium_xml"), recursive=True) if os.path.isdir(p)] 185 if len(xml_dirs) == 0: 186 raise RuntimeError( 187 f"Could not find the COCA 'Gated_release_final' folder at '{path}'. This dataset requires " 188 "registration and a signed Stanford University Dataset Research Use Agreement, so it cannot be " 189 "downloaded automatically. Please follow these steps: " 190 "1) Visit https://aimi.stanford.edu/datasets/coca-coronary-calcium-chest-ct and follow the link " 191 "to the dataset on Stanford's Redivis platform (https://stanford.redivis.com/datasets/1vm9-30b7p5srg). " 192 "2) Log in with your institutional account, fill in your contact details and accept the Stanford " 193 "University Dataset Research Use Agreement. " 194 f"3) Download the 'Gated_release_final' folder and place it such that " 195 f"'{path}/Gated_release_final/patient/<patient_id>/.../*.dcm' and " 196 f"'{path}/Gated_release_final/calcium_xml/<patient_id>.xml' exist." 197 ) 198 199 gated_dir = os.path.dirname(xml_dirs[0]) 200 return _preprocess_inputs(path, gated_dir)
Obtain the COCA dataset.
Arguments:
- path: Filepath to a folder where the data is downloaded for further processing.
- download: Whether to download the data if it is not present.
Returns:
Filepath where the data is preprocessed.
203def get_stanford_coca_paths(path: Union[os.PathLike, str], download: bool = False) -> Tuple[List[str], List[str]]: 204 """Get paths to the COCA data. 205 206 Args: 207 path: Filepath to a folder where the data is downloaded for further processing. 208 download: Whether to download the data if it is not present. 209 210 Returns: 211 List of filepaths for the image data. 212 List of filepaths for the label data. 213 """ 214 data_dir = get_stanford_coca_data(path, download) 215 volume_paths = natsorted(glob(os.path.join(data_dir, "*.h5"))) 216 assert len(volume_paths) > 0, f"Could not find any preprocessed volumes in '{data_dir}'." 217 return volume_paths, volume_paths
Get paths to the COCA data.
Arguments:
- path: Filepath to a folder where the data is downloaded for further processing.
- download: Whether to download the data if it is not present.
Returns:
List of filepaths for the image data. List of filepaths for the label data.
220def get_stanford_coca_dataset( 221 path: Union[os.PathLike, str], 222 patch_shape: Tuple[int, ...], 223 resize_inputs: bool = False, 224 download: bool = False, 225 **kwargs 226) -> Dataset: 227 """Get the COCA dataset for coronary artery calcium segmentation. 228 229 Args: 230 path: Filepath to a folder where the data is downloaded for further processing. 231 patch_shape: The patch shape to use for training. 232 resize_inputs: Whether to resize inputs to the desired patch shape. 233 download: Whether to download the data if it is not present. 234 kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset`. 235 236 Returns: 237 The segmentation dataset. 238 """ 239 raw_paths, label_paths = get_stanford_coca_paths(path, download) 240 241 if resize_inputs: 242 resize_kwargs = {"patch_shape": patch_shape, "is_rgb": False} 243 kwargs, patch_shape = util.update_kwargs_for_resize_trafo( 244 kwargs=kwargs, patch_shape=patch_shape, resize_inputs=resize_inputs, resize_kwargs=resize_kwargs 245 ) 246 247 return torch_em.default_segmentation_dataset( 248 raw_paths=raw_paths, 249 raw_key="raw", 250 label_paths=label_paths, 251 label_key="labels", 252 patch_shape=patch_shape, 253 is_seg_dataset=True, 254 **kwargs 255 )
Get the COCA dataset for coronary artery calcium segmentation.
Arguments:
- path: Filepath to a folder where the data is downloaded for further processing.
- patch_shape: The patch shape to use for training.
- resize_inputs: Whether to resize inputs to the desired patch shape.
- download: Whether to download the data if it is not present.
- kwargs: Additional keyword arguments for
torch_em.default_segmentation_dataset.
Returns:
The segmentation dataset.
258def get_stanford_coca_loader( 259 path: Union[os.PathLike, str], 260 batch_size: int, 261 patch_shape: Tuple[int, ...], 262 resize_inputs: bool = False, 263 download: bool = False, 264 **kwargs 265) -> DataLoader: 266 """Get the COCA dataloader for coronary artery calcium segmentation. 267 268 Args: 269 path: Filepath to a folder where the data is downloaded for further processing. 270 batch_size: The batch size for training. 271 patch_shape: The patch shape to use for training. 272 resize_inputs: Whether to resize inputs to the desired patch shape. 273 download: Whether to download the data if it is not present. 274 kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset` or for the PyTorch DataLoader. 275 276 Returns: 277 The DataLoader. 278 """ 279 ds_kwargs, loader_kwargs = util.split_kwargs(torch_em.default_segmentation_dataset, **kwargs) 280 dataset = get_stanford_coca_dataset(path, patch_shape, resize_inputs, download, **ds_kwargs) 281 return torch_em.get_data_loader(dataset, batch_size, **loader_kwargs)
Get the COCA dataloader for coronary artery calcium segmentation.
Arguments:
- path: Filepath to a folder where the data is downloaded for further processing.
- batch_size: The batch size for training.
- patch_shape: The patch shape to use for training.
- resize_inputs: Whether to resize inputs to the desired patch shape.
- download: Whether to download the data if it is not present.
- kwargs: Additional keyword arguments for
torch_em.default_segmentation_datasetor for the PyTorch DataLoader.
Returns:
The DataLoader.