torch_em.data.datasets.medical.stanford_coca

The COCA dataset contains annotations for coronary artery calcium segmentation in ECG-gated cardiac CT.

The dataset is released by the Stanford Center for Artificial Intelligence in Medicine and Imaging (AIMI) as "COCA - Coronary Calcium and chest CT's". It comprises two parts: gated coronary CT DICOM series with per-slice coronary artery calcium ROI annotations (stored as a binary property list '.xml' file per patient), and non-gated chest CT DICOM series with only a coronary artery calcium score (no pixel-level annotation). This module only supports the gated part, since only that part has pixel-level segmentation masks; the non-gated part is score-only and does not fit torch_em's segmentation dataset pattern.

The ROI file lists, per slice ('ImageIndex') and per coronary artery, one or more closed polygons ('Point_px', in pixel coordinates) outlining the calcified regions. The label ids used here are:

NOTE: The slice that a set of ROIs belongs to is identified in the ROI file only by an integer 'ImageIndex', with no reference to a DICOM SOP instance UID or file name. This module assumes that 'ImageIndex' is the 0-based position of the slice when the series is sorted by 'InstanceNumber'. This assumption could not be verified against the real data, since the dataset requires registration and a signed data use agreement, so it was not accessible while writing this module. Please verify the resulting masks against the DICOM series (e.g. by overlaying them) before relying on this dataset.

NOTE: The dataset requires registration and cannot be downloaded automatically. Please follow these steps:

  • Visit https://aimi.stanford.edu/datasets/coca-coronary-calcium-chest-ct and follow the link to the dataset on Stanford's Redivis platform (https://stanford.redivis.com/datasets/1vm9-30b7p5srg).
  • Log in with your institutional account, fill in your contact details and accept the Stanford University Dataset Research Use Agreement.
  • Download the 'Gated_release_final' folder (e.g. with the Redivis download tool) and place it such that '/Gated_release_final/patient//.../*.dcm' (the DICOM series) and '/Gated_release_final/calcium_xml/.xml' (the ROI annotations) exist.

The dataset is located at https://aimi.stanford.edu/datasets/coca-coronary-calcium-chest-ct (DOI: https://doi.org/10.71718/ge5g-ds80) and is released under the Stanford University Dataset Research Use Agreement.

This dataset is from the publications https://doi.org/10.1007/s10554-019-01710-8 and https://doi.org/10.1148/ryai.2020190004. Please cite them if you use this dataset in your research.

NOTE: The DICOM and ROI parsing requires 'pydicom'. Install it with 'pip install pydicom'.

  1"""The COCA dataset contains annotations for coronary artery calcium segmentation in ECG-gated
  2cardiac CT.
  3
  4The dataset is released by the Stanford Center for Artificial Intelligence in Medicine and Imaging
  5(AIMI) as "COCA - Coronary Calcium and chest CT's". It comprises two parts: gated coronary CT DICOM
  6series with per-slice coronary artery calcium ROI annotations (stored as a binary property list
  7'.xml' file per patient), and non-gated chest CT DICOM series with only a coronary artery calcium
  8score (no pixel-level annotation). This module only supports the gated part, since only that part
  9has pixel-level segmentation masks; the non-gated part is score-only and does not fit torch_em's
 10segmentation dataset pattern.
 11
 12The ROI file lists, per slice ('ImageIndex') and per coronary artery, one or more closed polygons
 13('Point_px', in pixel coordinates) outlining the calcified regions. The label ids used here are:
 14- 0: background, 1: right coronary artery, 2: left anterior descending artery,
 15- 3: left coronary artery, 4: left circumflex artery
 16(see `LABEL_IDS`; the artery names and this correspondence were taken from a community reimplementation
 17of the dataset loading, https://github.com/msingh9/cs230-Coronary-Calcium-Scoring-/blob/master/code/my_lib.py,
 18since the ROI file format itself is undocumented by Stanford AIMI).
 19
 20NOTE: The slice that a set of ROIs belongs to is identified in the ROI file only by an integer
 21'ImageIndex', with no reference to a DICOM SOP instance UID or file name. This module assumes that
 22'ImageIndex' is the 0-based position of the slice when the series is sorted by 'InstanceNumber'.
 23This assumption could not be verified against the real data, since the dataset requires registration
 24and a signed data use agreement, so it was not accessible while writing this module. Please verify
 25the resulting masks against the DICOM series (e.g. by overlaying them) before relying on this dataset.
 26
 27NOTE: The dataset requires registration and cannot be downloaded automatically. Please follow these steps:
 28- Visit https://aimi.stanford.edu/datasets/coca-coronary-calcium-chest-ct and follow the link to the
 29  dataset on Stanford's Redivis platform (https://stanford.redivis.com/datasets/1vm9-30b7p5srg).
 30- Log in with your institutional account, fill in your contact details and accept the Stanford University
 31  Dataset Research Use Agreement.
 32- Download the 'Gated_release_final' folder (e.g. with the Redivis download tool) and place it such that
 33  '<path>/Gated_release_final/patient/<patient_id>/.../*.dcm' (the DICOM series) and
 34  '<path>/Gated_release_final/calcium_xml/<patient_id>.xml' (the ROI annotations) exist.
 35
 36The dataset is located at https://aimi.stanford.edu/datasets/coca-coronary-calcium-chest-ct
 37(DOI: https://doi.org/10.71718/ge5g-ds80) and is released under the Stanford University Dataset
 38Research Use Agreement.
 39
 40This dataset is from the publications https://doi.org/10.1007/s10554-019-01710-8 and
 41https://doi.org/10.1148/ryai.2020190004. Please cite them if you use this dataset in your research.
 42
 43NOTE: The DICOM and ROI parsing requires 'pydicom'. Install it with 'pip install pydicom'.
 44"""
 45
 46import os
 47import plistlib
 48from glob import glob
 49from tqdm import tqdm
 50from natsort import natsorted
 51from typing import Union, Tuple, List
 52
 53import numpy as np
 54
 55from torch.utils.data import Dataset, DataLoader
 56
 57import torch_em
 58
 59from .. import util
 60
 61
 62LABEL_IDS = {
 63    "background": 0,
 64    "right_coronary_artery": 1,
 65    "left_anterior_descending_artery": 2,
 66    "left_coronary_artery": 3,
 67    "left_circumflex_artery": 4,
 68}
 69
 70ROI_NAME_TO_LABEL_ID = {
 71    "Right Coronary Artery": LABEL_IDS["right_coronary_artery"],
 72    "Left Anterior Descending Artery": LABEL_IDS["left_anterior_descending_artery"],
 73    "Left Coronary Artery": LABEL_IDS["left_coronary_artery"],
 74    "Left Circumflex Artery": LABEL_IDS["left_circumflex_artery"],
 75}
 76
 77
 78def _load_gated_series(patient_dir):
 79    import pydicom
 80
 81    dcm_paths = natsorted(glob(os.path.join(patient_dir, "**", "*.dcm"), recursive=True))
 82    assert len(dcm_paths) > 0, f"Could not find any DICOM files in '{patient_dir}'."
 83
 84    slices = [pydicom.dcmread(p) for p in dcm_paths]
 85    slices.sort(key=lambda dcm: int(dcm.InstanceNumber))
 86
 87    volume = np.stack([dcm.pixel_array for dcm in slices]).astype("float32")
 88    slope = float(getattr(slices[0], "RescaleSlope", 1.0))
 89    intercept = float(getattr(slices[0], "RescaleIntercept", 0.0))
 90    volume = np.round(volume * slope + intercept).astype("int16")
 91
 92    return volume
 93
 94
 95def _parse_calcium_xml(xml_path):
 96    with open(xml_path, "rb") as f:
 97        annotations = plistlib.load(f)
 98
 99    rois_per_slice = {}
100    for image in annotations["Images"]:
101        image_index = image["ImageIndex"]
102        for roi in image["ROIs"]:
103            points = roi.get("Point_px", [])
104            if len(points) == 0:
105                continue
106
107            label_id = ROI_NAME_TO_LABEL_ID.get(roi["Name"])
108            if label_id is None:
109                continue
110
111            pixels = np.array([[float(v) for v in point.strip("()").split(",")] for point in points])
112            rois_per_slice.setdefault(image_index, []).append((label_id, pixels))
113
114    return rois_per_slice
115
116
117def _rasterize_labels(shape, rois_per_slice):
118    from skimage.draw import polygon
119
120    labels = np.zeros(shape, dtype="uint8")
121    for slice_id, rois in rois_per_slice.items():
122        if slice_id >= shape[0]:
123            continue
124        for label_id, pixels in rois:
125            rr, cc = polygon(pixels[:, 1], pixels[:, 0], shape=shape[1:])
126            labels[slice_id, rr, cc] = label_id
127
128    return labels
129
130
131def _preprocess_inputs(path, gated_dir):
132    import h5py
133
134    preprocessed_dir = os.path.join(path, "preprocessed")
135    os.makedirs(preprocessed_dir, exist_ok=True)
136
137    xml_paths = natsorted(glob(os.path.join(gated_dir, "calcium_xml", "*.xml")))
138    for xml_path in tqdm(xml_paths, desc="Preprocessing the COCA gated CT scans"):
139        patient_id = os.path.basename(xml_path)[:-len(".xml")]
140        volume_path = os.path.join(preprocessed_dir, f"{patient_id}.h5")
141        if os.path.exists(volume_path):
142            continue
143
144        patient_dir = os.path.join(gated_dir, "patient", patient_id)
145        if not os.path.isdir(patient_dir):
146            continue
147
148        raw = _load_gated_series(patient_dir)
149        rois_per_slice = _parse_calcium_xml(xml_path)
150        labels = _rasterize_labels(raw.shape, rois_per_slice)
151
152        # The file is written to a temporary path first, so that an interrupted run does not leave a corrupt file.
153        with h5py.File(f"{volume_path}.tmp", "w") as f:
154            f.create_dataset("raw", data=raw, compression="gzip")
155            f.create_dataset("labels", data=labels, compression="gzip")
156
157        os.rename(f"{volume_path}.tmp", volume_path)
158
159    return preprocessed_dir
160
161
162def get_stanford_coca_data(path: Union[os.PathLike, str], download: bool = False) -> str:
163    """Obtain the COCA dataset.
164
165    Args:
166        path: Filepath to a folder where the data is downloaded for further processing.
167        download: Whether to download the data if it is not present.
168
169    Returns:
170        Filepath where the data is preprocessed.
171    """
172    preprocessed_dir = os.path.join(path, "preprocessed")
173    if os.path.exists(preprocessed_dir) and len(glob(os.path.join(preprocessed_dir, "*.h5"))) > 0:
174        return preprocessed_dir
175
176    if download:
177        raise RuntimeError(
178            "Download is set to True, but 'torch_em' cannot download the COCA dataset automatically: "
179            "it requires registration and a signed Stanford University Dataset Research Use Agreement. "
180            "See 'torch_em.data.datasets.medical.stanford_coca' for the manual download instructions."
181        )
182
183    xml_dirs = [p for p in glob(os.path.join(path, "**", "calcium_xml"), recursive=True) if os.path.isdir(p)]
184    if len(xml_dirs) == 0:
185        raise RuntimeError(
186            f"Could not find the COCA 'Gated_release_final' folder at '{path}'. This dataset requires "
187            "registration and a signed Stanford University Dataset Research Use Agreement, so it cannot be "
188            "downloaded automatically. Please follow these steps: "
189            "1) Visit https://aimi.stanford.edu/datasets/coca-coronary-calcium-chest-ct and follow the link "
190            "to the dataset on Stanford's Redivis platform (https://stanford.redivis.com/datasets/1vm9-30b7p5srg). "
191            "2) Log in with your institutional account, fill in your contact details and accept the Stanford "
192            "University Dataset Research Use Agreement. "
193            f"3) Download the 'Gated_release_final' folder and place it such that "
194            f"'{path}/Gated_release_final/patient/<patient_id>/.../*.dcm' and "
195            f"'{path}/Gated_release_final/calcium_xml/<patient_id>.xml' exist."
196        )
197
198    gated_dir = os.path.dirname(xml_dirs[0])
199    return _preprocess_inputs(path, gated_dir)
200
201
202def get_stanford_coca_paths(path: Union[os.PathLike, str], download: bool = False) -> Tuple[List[str], List[str]]:
203    """Get paths to the COCA data.
204
205    Args:
206        path: Filepath to a folder where the data is downloaded for further processing.
207        download: Whether to download the data if it is not present.
208
209    Returns:
210        List of filepaths for the image data.
211        List of filepaths for the label data.
212    """
213    data_dir = get_stanford_coca_data(path, download)
214    volume_paths = natsorted(glob(os.path.join(data_dir, "*.h5")))
215    assert len(volume_paths) > 0, f"Could not find any preprocessed volumes in '{data_dir}'."
216    return volume_paths, volume_paths
217
218
219def get_stanford_coca_dataset(
220    path: Union[os.PathLike, str],
221    patch_shape: Tuple[int, ...],
222    resize_inputs: bool = False,
223    download: bool = False,
224    **kwargs
225) -> Dataset:
226    """Get the COCA dataset for coronary artery calcium segmentation.
227
228    Args:
229        path: Filepath to a folder where the data is downloaded for further processing.
230        patch_shape: The patch shape to use for training.
231        resize_inputs: Whether to resize inputs to the desired patch shape.
232        download: Whether to download the data if it is not present.
233        kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset`.
234
235    Returns:
236        The segmentation dataset.
237    """
238    raw_paths, label_paths = get_stanford_coca_paths(path, download)
239
240    if resize_inputs:
241        resize_kwargs = {"patch_shape": patch_shape, "is_rgb": False}
242        kwargs, patch_shape = util.update_kwargs_for_resize_trafo(
243            kwargs=kwargs, patch_shape=patch_shape, resize_inputs=resize_inputs, resize_kwargs=resize_kwargs
244        )
245
246    return torch_em.default_segmentation_dataset(
247        raw_paths=raw_paths,
248        raw_key="raw",
249        label_paths=label_paths,
250        label_key="labels",
251        patch_shape=patch_shape,
252        is_seg_dataset=True,
253        **kwargs
254    )
255
256
257def get_stanford_coca_loader(
258    path: Union[os.PathLike, str],
259    batch_size: int,
260    patch_shape: Tuple[int, ...],
261    resize_inputs: bool = False,
262    download: bool = False,
263    **kwargs
264) -> DataLoader:
265    """Get the COCA dataloader for coronary artery calcium segmentation.
266
267    Args:
268        path: Filepath to a folder where the data is downloaded for further processing.
269        batch_size: The batch size for training.
270        patch_shape: The patch shape to use for training.
271        resize_inputs: Whether to resize inputs to the desired patch shape.
272        download: Whether to download the data if it is not present.
273        kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset` or for the PyTorch DataLoader.
274
275    Returns:
276        The DataLoader.
277    """
278    ds_kwargs, loader_kwargs = util.split_kwargs(torch_em.default_segmentation_dataset, **kwargs)
279    dataset = get_stanford_coca_dataset(path, patch_shape, resize_inputs, download, **ds_kwargs)
280    return torch_em.get_data_loader(dataset, batch_size, **loader_kwargs)
LABEL_IDS = {'background': 0, 'right_coronary_artery': 1, 'left_anterior_descending_artery': 2, 'left_coronary_artery': 3, 'left_circumflex_artery': 4}
ROI_NAME_TO_LABEL_ID = {'Right Coronary Artery': 1, 'Left Anterior Descending Artery': 2, 'Left Coronary Artery': 3, 'Left Circumflex Artery': 4}
def get_stanford_coca_data(path: Union[os.PathLike, str], download: bool = False) -> str:
163def get_stanford_coca_data(path: Union[os.PathLike, str], download: bool = False) -> str:
164    """Obtain the COCA dataset.
165
166    Args:
167        path: Filepath to a folder where the data is downloaded for further processing.
168        download: Whether to download the data if it is not present.
169
170    Returns:
171        Filepath where the data is preprocessed.
172    """
173    preprocessed_dir = os.path.join(path, "preprocessed")
174    if os.path.exists(preprocessed_dir) and len(glob(os.path.join(preprocessed_dir, "*.h5"))) > 0:
175        return preprocessed_dir
176
177    if download:
178        raise RuntimeError(
179            "Download is set to True, but 'torch_em' cannot download the COCA dataset automatically: "
180            "it requires registration and a signed Stanford University Dataset Research Use Agreement. "
181            "See 'torch_em.data.datasets.medical.stanford_coca' for the manual download instructions."
182        )
183
184    xml_dirs = [p for p in glob(os.path.join(path, "**", "calcium_xml"), recursive=True) if os.path.isdir(p)]
185    if len(xml_dirs) == 0:
186        raise RuntimeError(
187            f"Could not find the COCA 'Gated_release_final' folder at '{path}'. This dataset requires "
188            "registration and a signed Stanford University Dataset Research Use Agreement, so it cannot be "
189            "downloaded automatically. Please follow these steps: "
190            "1) Visit https://aimi.stanford.edu/datasets/coca-coronary-calcium-chest-ct and follow the link "
191            "to the dataset on Stanford's Redivis platform (https://stanford.redivis.com/datasets/1vm9-30b7p5srg). "
192            "2) Log in with your institutional account, fill in your contact details and accept the Stanford "
193            "University Dataset Research Use Agreement. "
194            f"3) Download the 'Gated_release_final' folder and place it such that "
195            f"'{path}/Gated_release_final/patient/<patient_id>/.../*.dcm' and "
196            f"'{path}/Gated_release_final/calcium_xml/<patient_id>.xml' exist."
197        )
198
199    gated_dir = os.path.dirname(xml_dirs[0])
200    return _preprocess_inputs(path, gated_dir)

Obtain the COCA dataset.

Arguments:
  • path: Filepath to a folder where the data is downloaded for further processing.
  • download: Whether to download the data if it is not present.
Returns:

Filepath where the data is preprocessed.

def get_stanford_coca_paths( path: Union[os.PathLike, str], download: bool = False) -> Tuple[List[str], List[str]]:
203def get_stanford_coca_paths(path: Union[os.PathLike, str], download: bool = False) -> Tuple[List[str], List[str]]:
204    """Get paths to the COCA data.
205
206    Args:
207        path: Filepath to a folder where the data is downloaded for further processing.
208        download: Whether to download the data if it is not present.
209
210    Returns:
211        List of filepaths for the image data.
212        List of filepaths for the label data.
213    """
214    data_dir = get_stanford_coca_data(path, download)
215    volume_paths = natsorted(glob(os.path.join(data_dir, "*.h5")))
216    assert len(volume_paths) > 0, f"Could not find any preprocessed volumes in '{data_dir}'."
217    return volume_paths, volume_paths

Get paths to the COCA data.

Arguments:
  • path: Filepath to a folder where the data is downloaded for further processing.
  • download: Whether to download the data if it is not present.
Returns:

List of filepaths for the image data. List of filepaths for the label data.

def get_stanford_coca_dataset( path: Union[os.PathLike, str], patch_shape: Tuple[int, ...], resize_inputs: bool = False, download: bool = False, **kwargs) -> torch.utils.data.dataset.Dataset:
220def get_stanford_coca_dataset(
221    path: Union[os.PathLike, str],
222    patch_shape: Tuple[int, ...],
223    resize_inputs: bool = False,
224    download: bool = False,
225    **kwargs
226) -> Dataset:
227    """Get the COCA dataset for coronary artery calcium segmentation.
228
229    Args:
230        path: Filepath to a folder where the data is downloaded for further processing.
231        patch_shape: The patch shape to use for training.
232        resize_inputs: Whether to resize inputs to the desired patch shape.
233        download: Whether to download the data if it is not present.
234        kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset`.
235
236    Returns:
237        The segmentation dataset.
238    """
239    raw_paths, label_paths = get_stanford_coca_paths(path, download)
240
241    if resize_inputs:
242        resize_kwargs = {"patch_shape": patch_shape, "is_rgb": False}
243        kwargs, patch_shape = util.update_kwargs_for_resize_trafo(
244            kwargs=kwargs, patch_shape=patch_shape, resize_inputs=resize_inputs, resize_kwargs=resize_kwargs
245        )
246
247    return torch_em.default_segmentation_dataset(
248        raw_paths=raw_paths,
249        raw_key="raw",
250        label_paths=label_paths,
251        label_key="labels",
252        patch_shape=patch_shape,
253        is_seg_dataset=True,
254        **kwargs
255    )

Get the COCA dataset for coronary artery calcium segmentation.

Arguments:
  • path: Filepath to a folder where the data is downloaded for further processing.
  • patch_shape: The patch shape to use for training.
  • resize_inputs: Whether to resize inputs to the desired patch shape.
  • download: Whether to download the data if it is not present.
  • kwargs: Additional keyword arguments for torch_em.default_segmentation_dataset.
Returns:

The segmentation dataset.

def get_stanford_coca_loader( path: Union[os.PathLike, str], batch_size: int, patch_shape: Tuple[int, ...], resize_inputs: bool = False, download: bool = False, **kwargs) -> torch.utils.data.dataloader.DataLoader:
258def get_stanford_coca_loader(
259    path: Union[os.PathLike, str],
260    batch_size: int,
261    patch_shape: Tuple[int, ...],
262    resize_inputs: bool = False,
263    download: bool = False,
264    **kwargs
265) -> DataLoader:
266    """Get the COCA dataloader for coronary artery calcium segmentation.
267
268    Args:
269        path: Filepath to a folder where the data is downloaded for further processing.
270        batch_size: The batch size for training.
271        patch_shape: The patch shape to use for training.
272        resize_inputs: Whether to resize inputs to the desired patch shape.
273        download: Whether to download the data if it is not present.
274        kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset` or for the PyTorch DataLoader.
275
276    Returns:
277        The DataLoader.
278    """
279    ds_kwargs, loader_kwargs = util.split_kwargs(torch_em.default_segmentation_dataset, **kwargs)
280    dataset = get_stanford_coca_dataset(path, patch_shape, resize_inputs, download, **ds_kwargs)
281    return torch_em.get_data_loader(dataset, batch_size, **loader_kwargs)

Get the COCA dataloader for coronary artery calcium segmentation.

Arguments:
  • path: Filepath to a folder where the data is downloaded for further processing.
  • batch_size: The batch size for training.
  • patch_shape: The patch shape to use for training.
  • resize_inputs: Whether to resize inputs to the desired patch shape.
  • download: Whether to download the data if it is not present.
  • kwargs: Additional keyword arguments for torch_em.default_segmentation_dataset or for the PyTorch DataLoader.
Returns:

The DataLoader.