torch_em.data.datasets.medical.lascarqs
The LAScarQS dataset contains annotations for left atrium and atrial scar segmentation in late gadolinium enhanced (LGE) cardiac MRI of patients with atrial fibrillation.
The data was curated for the LAScarQS 2022 challenge (https://zmiclab.github.io/projects/lascarqs22/), which was held together with MICCAI 2022. It comprises 194 LGE-MRI scans acquired at three clinical centers (University of Utah, Beth Israel Deaconess Medical Center and King's College London) and is split into two tasks, selected with the 'task' argument:
- 'task1': 60 labeled post-ablation scans with annotations of the left atrium cavity and the atrial scar.
- 'task2': 130 labeled pre- and post-ablation scans with annotations of the left atrium cavity only. Each task also ships additional unlabeled test scans, which are not exposed by this module.
The label ids are described in LABEL_IDS: 0 = background, 1 = left atrium cavity, 2 = atrial scar
(only present for 'task1'). This is a different annotation target than the other cardiac MRI datasets
in this library (e.g. torch_em.data.datasets.medical.atriaseg), which do not provide scar annotations.
The LGE-MRI are stored as nifti volumes, which are converted to hdf5 volumes with the keys 'raw' and 'labels' by this module.
NOTE: The data is not available for automatic download. It requires registering with the organizers,
see get_lascarqs_data for the manual download steps. The data is distributed under the
CC BY-NC-ND license, see https://zmiclab.github.io/projects/lascarqs22/data.html for details.
This dataset is from the publication https://doi.org/10.1007/978-3-031-31778-1. Please cite it if you use this dataset in your research.
1"""The LAScarQS dataset contains annotations for left atrium and atrial scar segmentation 2in late gadolinium enhanced (LGE) cardiac MRI of patients with atrial fibrillation. 3 4The data was curated for the LAScarQS 2022 challenge (https://zmiclab.github.io/projects/lascarqs22/), 5which was held together with MICCAI 2022. It comprises 194 LGE-MRI scans acquired at three clinical 6centers (University of Utah, Beth Israel Deaconess Medical Center and King's College London) and is split 7into two tasks, selected with the 'task' argument: 8- 'task1': 60 labeled post-ablation scans with annotations of the left atrium cavity and the atrial scar. 9- 'task2': 130 labeled pre- and post-ablation scans with annotations of the left atrium cavity only. 10Each task also ships additional unlabeled test scans, which are not exposed by this module. 11 12The label ids are described in `LABEL_IDS`: 0 = background, 1 = left atrium cavity, 2 = atrial scar 13(only present for 'task1'). This is a different annotation target than the other cardiac MRI datasets 14in this library (e.g. `torch_em.data.datasets.medical.atriaseg`), which do not provide scar annotations. 15 16The LGE-MRI are stored as nifti volumes, which are converted to hdf5 volumes with the keys 'raw' and 17'labels' by this module. 18 19NOTE: The data is not available for automatic download. It requires registering with the organizers, 20see `get_lascarqs_data` for the manual download steps. The data is distributed under the 21CC BY-NC-ND license, see https://zmiclab.github.io/projects/lascarqs22/data.html for details. 22 23This dataset is from the publication https://doi.org/10.1007/978-3-031-31778-1. 24Please cite it if you use this dataset in your research. 25""" 26 27import os 28from glob import glob 29from tqdm import tqdm 30from natsort import natsorted 31from typing import Union, Tuple, List, Literal 32 33import numpy as np 34 35from torch.utils.data import Dataset, DataLoader 36 37import torch_em 38 39from .. import util 40 41 42LABEL_IDS = {"background": 0, "la_cavity": 1, "la_scar": 2} 43 44TASKS = {"task1": 60, "task2": 130} 45 46 47def get_lascarqs_data(path: Union[os.PathLike, str], download: bool = False) -> str: 48 """Obtain the LAScarQS 2022 dataset. 49 50 Args: 51 path: Filepath to a folder where the data is downloaded for further processing. 52 download: Whether to download the data if it is not present. 53 54 Returns: 55 Filepath to the folder with the downloaded and extracted dataset. 56 """ 57 data_dir = os.path.join(path, "LAScarQS2022") 58 if os.path.exists(data_dir): 59 return data_dir 60 61 if download: 62 raise NotImplementedError( 63 "The LAScarQS 2022 dataset cannot be downloaded automatically. See 'get_lascarqs_data' for details." 64 ) 65 66 zip_path = os.path.join(path, "LAScarQS2022.zip") 67 if not os.path.exists(zip_path): 68 raise RuntimeError( 69 f"Could not find the LAScarQS 2022 dataset at '{path}'. This dataset is not available for automatic " 70 "download. To obtain it, please follow these steps:\n" 71 "- Visit https://zmiclab.github.io/projects/lascarqs22/data.html and read the data usage agreement.\n" 72 "- Register by sending the requested information to LAScarQS2022@outlook.com or LAScarQS2022@163.com.\n" 73 "- Once you receive the download link from the organizers, download and extract the archive.\n" 74 f"- Place the extracted 'LAScarQS2022' folder (with the 'task1' and 'task2' subfolders) at '{path}', " 75 f"or place the downloaded zip archive at '{zip_path}'." 76 ) 77 78 util.unzip(zip_path=zip_path, dst=path, remove=False) 79 return data_dir 80 81 82def _preprocess_inputs(data_dir, task, preprocessed_dir): 83 import h5py 84 import nibabel as nib 85 86 case_dirs = natsorted(glob(os.path.join(data_dir, task, "train_data", "train_*"))) 87 os.makedirs(preprocessed_dir, exist_ok=True) 88 89 for case_dir in tqdm(case_dirs, desc=f"Preprocessing the LAScarQS '{task}' cases"): 90 case_id = os.path.basename(case_dir) 91 volume_path = os.path.join(preprocessed_dir, f"{case_id}.h5") 92 if os.path.exists(volume_path): 93 continue 94 95 # The transpose maps the nifti axis order (X, Y, Z) to the (Z, Y, X) order used for the volumes. 96 raw = np.asarray(nib.load(os.path.join(case_dir, "enhanced.nii.gz")).dataobj).T 97 cavity = np.asarray(nib.load(os.path.join(case_dir, "atriumSegImgMO.nii.gz")).dataobj).T 98 99 labels = np.zeros(raw.shape, dtype="uint8") 100 labels[cavity > 0] = LABEL_IDS["la_cavity"] 101 102 if task == "task1": 103 scar = np.asarray(nib.load(os.path.join(case_dir, "scarSegImgM.nii.gz")).dataobj).T 104 labels[scar > 0] = LABEL_IDS["la_scar"] 105 106 # The file is written to a temporary path first, so that an interrupted run leaves no corrupt file. 107 with h5py.File(f"{volume_path}.tmp", "w") as f: 108 f.create_dataset("raw", data=raw, compression="gzip") 109 f.create_dataset("labels", data=labels, compression="gzip") 110 111 os.rename(f"{volume_path}.tmp", volume_path) 112 113 114def get_lascarqs_paths( 115 path: Union[os.PathLike, str], task: Literal["task1", "task2"], download: bool = False 116) -> List[str]: 117 """Get paths to the LAScarQS data. 118 119 Args: 120 path: Filepath to a folder where the data is downloaded for further processing. 121 task: The choice of task. Either 'task1' (LA cavity and scar) or 'task2' (LA cavity only). 122 download: Whether to download the data if it is not present. 123 124 Returns: 125 List of filepaths for the hdf5 files, which contain the image data ('raw') and the label data ('labels'). 126 """ 127 if task not in TASKS: 128 raise ValueError(f"'{task}' is not a valid task. Please choose one of {list(TASKS.keys())}.") 129 130 data_dir = get_lascarqs_data(path, download) 131 132 preprocessed_dir = os.path.join(path, "preprocessed", task) 133 if len(glob(os.path.join(preprocessed_dir, "*.h5"))) != TASKS[task]: 134 _preprocess_inputs(data_dir, task, preprocessed_dir) 135 136 volume_paths = natsorted(glob(os.path.join(preprocessed_dir, "*.h5"))) 137 assert len(volume_paths) > 0, f"Could not find any preprocessed volumes in '{preprocessed_dir}'." 138 139 return volume_paths 140 141 142def get_lascarqs_dataset( 143 path: Union[os.PathLike, str], 144 patch_shape: Tuple[int, ...], 145 task: Literal["task1", "task2"], 146 resize_inputs: bool = False, 147 download: bool = False, 148 **kwargs 149) -> Dataset: 150 """Get the LAScarQS dataset for left atrium and atrial scar segmentation. 151 152 Args: 153 path: Filepath to a folder where the data is downloaded for further processing. 154 patch_shape: The patch shape to use for training. 155 task: The choice of task. Either 'task1' (LA cavity and scar) or 'task2' (LA cavity only). 156 resize_inputs: Whether to resize inputs to the desired patch shape. 157 download: Whether to download the data if it is not present. 158 kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset`. 159 160 Returns: 161 The segmentation dataset. 162 """ 163 volume_paths = get_lascarqs_paths(path, task, download) 164 165 if resize_inputs: 166 resize_kwargs = {"patch_shape": patch_shape, "is_rgb": False} 167 kwargs, patch_shape = util.update_kwargs_for_resize_trafo( 168 kwargs=kwargs, patch_shape=patch_shape, resize_inputs=resize_inputs, resize_kwargs=resize_kwargs 169 ) 170 171 return torch_em.default_segmentation_dataset( 172 raw_paths=volume_paths, 173 raw_key="raw", 174 label_paths=volume_paths, 175 label_key="labels", 176 patch_shape=patch_shape, 177 is_seg_dataset=True, 178 **kwargs 179 ) 180 181 182def get_lascarqs_loader( 183 path: Union[os.PathLike, str], 184 batch_size: int, 185 patch_shape: Tuple[int, ...], 186 task: Literal["task1", "task2"], 187 resize_inputs: bool = False, 188 download: bool = False, 189 **kwargs 190) -> DataLoader: 191 """Get the LAScarQS dataloader for left atrium and atrial scar segmentation. 192 193 Args: 194 path: Filepath to a folder where the data is downloaded for further processing. 195 batch_size: The batch size for training. 196 patch_shape: The patch shape to use for training. 197 task: The choice of task. Either 'task1' (LA cavity and scar) or 'task2' (LA cavity only). 198 resize_inputs: Whether to resize inputs to the desired patch shape. 199 download: Whether to download the data if it is not present. 200 kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset` or for the PyTorch DataLoader. 201 202 Returns: 203 The DataLoader. 204 """ 205 ds_kwargs, loader_kwargs = util.split_kwargs(torch_em.default_segmentation_dataset, **kwargs) 206 dataset = get_lascarqs_dataset(path, patch_shape, task, resize_inputs, download, **ds_kwargs) 207 return torch_em.get_data_loader(dataset, batch_size, **loader_kwargs)
48def get_lascarqs_data(path: Union[os.PathLike, str], download: bool = False) -> str: 49 """Obtain the LAScarQS 2022 dataset. 50 51 Args: 52 path: Filepath to a folder where the data is downloaded for further processing. 53 download: Whether to download the data if it is not present. 54 55 Returns: 56 Filepath to the folder with the downloaded and extracted dataset. 57 """ 58 data_dir = os.path.join(path, "LAScarQS2022") 59 if os.path.exists(data_dir): 60 return data_dir 61 62 if download: 63 raise NotImplementedError( 64 "The LAScarQS 2022 dataset cannot be downloaded automatically. See 'get_lascarqs_data' for details." 65 ) 66 67 zip_path = os.path.join(path, "LAScarQS2022.zip") 68 if not os.path.exists(zip_path): 69 raise RuntimeError( 70 f"Could not find the LAScarQS 2022 dataset at '{path}'. This dataset is not available for automatic " 71 "download. To obtain it, please follow these steps:\n" 72 "- Visit https://zmiclab.github.io/projects/lascarqs22/data.html and read the data usage agreement.\n" 73 "- Register by sending the requested information to LAScarQS2022@outlook.com or LAScarQS2022@163.com.\n" 74 "- Once you receive the download link from the organizers, download and extract the archive.\n" 75 f"- Place the extracted 'LAScarQS2022' folder (with the 'task1' and 'task2' subfolders) at '{path}', " 76 f"or place the downloaded zip archive at '{zip_path}'." 77 ) 78 79 util.unzip(zip_path=zip_path, dst=path, remove=False) 80 return data_dir
Obtain the LAScarQS 2022 dataset.
Arguments:
- path: Filepath to a folder where the data is downloaded for further processing.
- download: Whether to download the data if it is not present.
Returns:
Filepath to the folder with the downloaded and extracted dataset.
115def get_lascarqs_paths( 116 path: Union[os.PathLike, str], task: Literal["task1", "task2"], download: bool = False 117) -> List[str]: 118 """Get paths to the LAScarQS data. 119 120 Args: 121 path: Filepath to a folder where the data is downloaded for further processing. 122 task: The choice of task. Either 'task1' (LA cavity and scar) or 'task2' (LA cavity only). 123 download: Whether to download the data if it is not present. 124 125 Returns: 126 List of filepaths for the hdf5 files, which contain the image data ('raw') and the label data ('labels'). 127 """ 128 if task not in TASKS: 129 raise ValueError(f"'{task}' is not a valid task. Please choose one of {list(TASKS.keys())}.") 130 131 data_dir = get_lascarqs_data(path, download) 132 133 preprocessed_dir = os.path.join(path, "preprocessed", task) 134 if len(glob(os.path.join(preprocessed_dir, "*.h5"))) != TASKS[task]: 135 _preprocess_inputs(data_dir, task, preprocessed_dir) 136 137 volume_paths = natsorted(glob(os.path.join(preprocessed_dir, "*.h5"))) 138 assert len(volume_paths) > 0, f"Could not find any preprocessed volumes in '{preprocessed_dir}'." 139 140 return volume_paths
Get paths to the LAScarQS data.
Arguments:
- path: Filepath to a folder where the data is downloaded for further processing.
- task: The choice of task. Either 'task1' (LA cavity and scar) or 'task2' (LA cavity only).
- download: Whether to download the data if it is not present.
Returns:
List of filepaths for the hdf5 files, which contain the image data ('raw') and the label data ('labels').
143def get_lascarqs_dataset( 144 path: Union[os.PathLike, str], 145 patch_shape: Tuple[int, ...], 146 task: Literal["task1", "task2"], 147 resize_inputs: bool = False, 148 download: bool = False, 149 **kwargs 150) -> Dataset: 151 """Get the LAScarQS dataset for left atrium and atrial scar segmentation. 152 153 Args: 154 path: Filepath to a folder where the data is downloaded for further processing. 155 patch_shape: The patch shape to use for training. 156 task: The choice of task. Either 'task1' (LA cavity and scar) or 'task2' (LA cavity only). 157 resize_inputs: Whether to resize inputs to the desired patch shape. 158 download: Whether to download the data if it is not present. 159 kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset`. 160 161 Returns: 162 The segmentation dataset. 163 """ 164 volume_paths = get_lascarqs_paths(path, task, download) 165 166 if resize_inputs: 167 resize_kwargs = {"patch_shape": patch_shape, "is_rgb": False} 168 kwargs, patch_shape = util.update_kwargs_for_resize_trafo( 169 kwargs=kwargs, patch_shape=patch_shape, resize_inputs=resize_inputs, resize_kwargs=resize_kwargs 170 ) 171 172 return torch_em.default_segmentation_dataset( 173 raw_paths=volume_paths, 174 raw_key="raw", 175 label_paths=volume_paths, 176 label_key="labels", 177 patch_shape=patch_shape, 178 is_seg_dataset=True, 179 **kwargs 180 )
Get the LAScarQS dataset for left atrium and atrial scar segmentation.
Arguments:
- path: Filepath to a folder where the data is downloaded for further processing.
- patch_shape: The patch shape to use for training.
- task: The choice of task. Either 'task1' (LA cavity and scar) or 'task2' (LA cavity only).
- resize_inputs: Whether to resize inputs to the desired patch shape.
- download: Whether to download the data if it is not present.
- kwargs: Additional keyword arguments for
torch_em.default_segmentation_dataset.
Returns:
The segmentation dataset.
183def get_lascarqs_loader( 184 path: Union[os.PathLike, str], 185 batch_size: int, 186 patch_shape: Tuple[int, ...], 187 task: Literal["task1", "task2"], 188 resize_inputs: bool = False, 189 download: bool = False, 190 **kwargs 191) -> DataLoader: 192 """Get the LAScarQS dataloader for left atrium and atrial scar segmentation. 193 194 Args: 195 path: Filepath to a folder where the data is downloaded for further processing. 196 batch_size: The batch size for training. 197 patch_shape: The patch shape to use for training. 198 task: The choice of task. Either 'task1' (LA cavity and scar) or 'task2' (LA cavity only). 199 resize_inputs: Whether to resize inputs to the desired patch shape. 200 download: Whether to download the data if it is not present. 201 kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset` or for the PyTorch DataLoader. 202 203 Returns: 204 The DataLoader. 205 """ 206 ds_kwargs, loader_kwargs = util.split_kwargs(torch_em.default_segmentation_dataset, **kwargs) 207 dataset = get_lascarqs_dataset(path, patch_shape, task, resize_inputs, download, **ds_kwargs) 208 return torch_em.get_data_loader(dataset, batch_size, **loader_kwargs)
Get the LAScarQS dataloader for left atrium and atrial scar segmentation.
Arguments:
- path: Filepath to a folder where the data is downloaded for further processing.
- batch_size: The batch size for training.
- patch_shape: The patch shape to use for training.
- task: The choice of task. Either 'task1' (LA cavity and scar) or 'task2' (LA cavity only).
- resize_inputs: Whether to resize inputs to the desired patch shape.
- download: Whether to download the data if it is not present.
- kwargs: Additional keyword arguments for
torch_em.default_segmentation_datasetor for the PyTorch DataLoader.
Returns:
The DataLoader.