torch_em.data.datasets.medical.lascarqs

The LAScarQS dataset contains annotations for left atrium and atrial scar segmentation in late gadolinium enhanced (LGE) cardiac MRI of patients with atrial fibrillation.

The data was curated for the LAScarQS 2022 challenge (https://zmiclab.github.io/projects/lascarqs22/), which was held together with MICCAI 2022. It comprises 194 LGE-MRI scans acquired at three clinical centers (University of Utah, Beth Israel Deaconess Medical Center and King's College London) and is split into two tasks, selected with the 'task' argument:

  • 'task1': 60 labeled post-ablation scans with annotations of the left atrium cavity and the atrial scar.
  • 'task2': 130 labeled pre- and post-ablation scans with annotations of the left atrium cavity only. Each task also ships additional unlabeled test scans, which are not exposed by this module.

The label ids are described in LABEL_IDS: 0 = background, 1 = left atrium cavity, 2 = atrial scar (only present for 'task1'). This is a different annotation target than the other cardiac MRI datasets in this library (e.g. torch_em.data.datasets.medical.atriaseg), which do not provide scar annotations.

The LGE-MRI are stored as nifti volumes, which are converted to hdf5 volumes with the keys 'raw' and 'labels' by this module.

NOTE: The data is not available for automatic download. It requires registering with the organizers, see get_lascarqs_data for the manual download steps. The data is distributed under the CC BY-NC-ND license, see https://zmiclab.github.io/projects/lascarqs22/data.html for details.

This dataset is from the publication https://doi.org/10.1007/978-3-031-31778-1. Please cite it if you use this dataset in your research.

  1"""The LAScarQS dataset contains annotations for left atrium and atrial scar segmentation
  2in late gadolinium enhanced (LGE) cardiac MRI of patients with atrial fibrillation.
  3
  4The data was curated for the LAScarQS 2022 challenge (https://zmiclab.github.io/projects/lascarqs22/),
  5which was held together with MICCAI 2022. It comprises 194 LGE-MRI scans acquired at three clinical
  6centers (University of Utah, Beth Israel Deaconess Medical Center and King's College London) and is split
  7into two tasks, selected with the 'task' argument:
  8- 'task1': 60 labeled post-ablation scans with annotations of the left atrium cavity and the atrial scar.
  9- 'task2': 130 labeled pre- and post-ablation scans with annotations of the left atrium cavity only.
 10Each task also ships additional unlabeled test scans, which are not exposed by this module.
 11
 12The label ids are described in `LABEL_IDS`: 0 = background, 1 = left atrium cavity, 2 = atrial scar
 13(only present for 'task1'). This is a different annotation target than the other cardiac MRI datasets
 14in this library (e.g. `torch_em.data.datasets.medical.atriaseg`), which do not provide scar annotations.
 15
 16The LGE-MRI are stored as nifti volumes, which are converted to hdf5 volumes with the keys 'raw' and
 17'labels' by this module.
 18
 19NOTE: The data is not available for automatic download. It requires registering with the organizers,
 20see `get_lascarqs_data` for the manual download steps. The data is distributed under the
 21CC BY-NC-ND license, see https://zmiclab.github.io/projects/lascarqs22/data.html for details.
 22
 23This dataset is from the publication https://doi.org/10.1007/978-3-031-31778-1.
 24Please cite it if you use this dataset in your research.
 25"""
 26
 27import os
 28from glob import glob
 29from tqdm import tqdm
 30from natsort import natsorted
 31from typing import Union, Tuple, List, Literal
 32
 33import numpy as np
 34
 35from torch.utils.data import Dataset, DataLoader
 36
 37import torch_em
 38
 39from .. import util
 40
 41
 42LABEL_IDS = {"background": 0, "la_cavity": 1, "la_scar": 2}
 43
 44TASKS = {"task1": 60, "task2": 130}
 45
 46
 47def get_lascarqs_data(path: Union[os.PathLike, str], download: bool = False) -> str:
 48    """Obtain the LAScarQS 2022 dataset.
 49
 50    Args:
 51        path: Filepath to a folder where the data is downloaded for further processing.
 52        download: Whether to download the data if it is not present.
 53
 54    Returns:
 55        Filepath to the folder with the downloaded and extracted dataset.
 56    """
 57    data_dir = os.path.join(path, "LAScarQS2022")
 58    if os.path.exists(data_dir):
 59        return data_dir
 60
 61    if download:
 62        raise NotImplementedError(
 63            "The LAScarQS 2022 dataset cannot be downloaded automatically. See 'get_lascarqs_data' for details."
 64        )
 65
 66    zip_path = os.path.join(path, "LAScarQS2022.zip")
 67    if not os.path.exists(zip_path):
 68        raise RuntimeError(
 69            f"Could not find the LAScarQS 2022 dataset at '{path}'. This dataset is not available for automatic "
 70            "download. To obtain it, please follow these steps:\n"
 71            "- Visit https://zmiclab.github.io/projects/lascarqs22/data.html and read the data usage agreement.\n"
 72            "- Register by sending the requested information to LAScarQS2022@outlook.com or LAScarQS2022@163.com.\n"
 73            "- Once you receive the download link from the organizers, download and extract the archive.\n"
 74            f"- Place the extracted 'LAScarQS2022' folder (with the 'task1' and 'task2' subfolders) at '{path}', "
 75            f"or place the downloaded zip archive at '{zip_path}'."
 76        )
 77
 78    util.unzip(zip_path=zip_path, dst=path, remove=False)
 79    return data_dir
 80
 81
 82def _preprocess_inputs(data_dir, task, preprocessed_dir):
 83    import h5py
 84    import nibabel as nib
 85
 86    case_dirs = natsorted(glob(os.path.join(data_dir, task, "train_data", "train_*")))
 87    os.makedirs(preprocessed_dir, exist_ok=True)
 88
 89    for case_dir in tqdm(case_dirs, desc=f"Preprocessing the LAScarQS '{task}' cases"):
 90        case_id = os.path.basename(case_dir)
 91        volume_path = os.path.join(preprocessed_dir, f"{case_id}.h5")
 92        if os.path.exists(volume_path):
 93            continue
 94
 95        # The transpose maps the nifti axis order (X, Y, Z) to the (Z, Y, X) order used for the volumes.
 96        raw = np.asarray(nib.load(os.path.join(case_dir, "enhanced.nii.gz")).dataobj).T
 97        cavity = np.asarray(nib.load(os.path.join(case_dir, "atriumSegImgMO.nii.gz")).dataobj).T
 98
 99        labels = np.zeros(raw.shape, dtype="uint8")
100        labels[cavity > 0] = LABEL_IDS["la_cavity"]
101
102        if task == "task1":
103            scar = np.asarray(nib.load(os.path.join(case_dir, "scarSegImgM.nii.gz")).dataobj).T
104            labels[scar > 0] = LABEL_IDS["la_scar"]
105
106        # The file is written to a temporary path first, so that an interrupted run leaves no corrupt file.
107        with h5py.File(f"{volume_path}.tmp", "w") as f:
108            f.create_dataset("raw", data=raw, compression="gzip")
109            f.create_dataset("labels", data=labels, compression="gzip")
110
111        os.rename(f"{volume_path}.tmp", volume_path)
112
113
114def get_lascarqs_paths(
115    path: Union[os.PathLike, str], task: Literal["task1", "task2"], download: bool = False
116) -> List[str]:
117    """Get paths to the LAScarQS data.
118
119    Args:
120        path: Filepath to a folder where the data is downloaded for further processing.
121        task: The choice of task. Either 'task1' (LA cavity and scar) or 'task2' (LA cavity only).
122        download: Whether to download the data if it is not present.
123
124    Returns:
125        List of filepaths for the hdf5 files, which contain the image data ('raw') and the label data ('labels').
126    """
127    if task not in TASKS:
128        raise ValueError(f"'{task}' is not a valid task. Please choose one of {list(TASKS.keys())}.")
129
130    data_dir = get_lascarqs_data(path, download)
131
132    preprocessed_dir = os.path.join(path, "preprocessed", task)
133    if len(glob(os.path.join(preprocessed_dir, "*.h5"))) != TASKS[task]:
134        _preprocess_inputs(data_dir, task, preprocessed_dir)
135
136    volume_paths = natsorted(glob(os.path.join(preprocessed_dir, "*.h5")))
137    assert len(volume_paths) > 0, f"Could not find any preprocessed volumes in '{preprocessed_dir}'."
138
139    return volume_paths
140
141
142def get_lascarqs_dataset(
143    path: Union[os.PathLike, str],
144    patch_shape: Tuple[int, ...],
145    task: Literal["task1", "task2"],
146    resize_inputs: bool = False,
147    download: bool = False,
148    **kwargs
149) -> Dataset:
150    """Get the LAScarQS dataset for left atrium and atrial scar segmentation.
151
152    Args:
153        path: Filepath to a folder where the data is downloaded for further processing.
154        patch_shape: The patch shape to use for training.
155        task: The choice of task. Either 'task1' (LA cavity and scar) or 'task2' (LA cavity only).
156        resize_inputs: Whether to resize inputs to the desired patch shape.
157        download: Whether to download the data if it is not present.
158        kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset`.
159
160    Returns:
161        The segmentation dataset.
162    """
163    volume_paths = get_lascarqs_paths(path, task, download)
164
165    if resize_inputs:
166        resize_kwargs = {"patch_shape": patch_shape, "is_rgb": False}
167        kwargs, patch_shape = util.update_kwargs_for_resize_trafo(
168            kwargs=kwargs, patch_shape=patch_shape, resize_inputs=resize_inputs, resize_kwargs=resize_kwargs
169        )
170
171    return torch_em.default_segmentation_dataset(
172        raw_paths=volume_paths,
173        raw_key="raw",
174        label_paths=volume_paths,
175        label_key="labels",
176        patch_shape=patch_shape,
177        is_seg_dataset=True,
178        **kwargs
179    )
180
181
182def get_lascarqs_loader(
183    path: Union[os.PathLike, str],
184    batch_size: int,
185    patch_shape: Tuple[int, ...],
186    task: Literal["task1", "task2"],
187    resize_inputs: bool = False,
188    download: bool = False,
189    **kwargs
190) -> DataLoader:
191    """Get the LAScarQS dataloader for left atrium and atrial scar segmentation.
192
193    Args:
194        path: Filepath to a folder where the data is downloaded for further processing.
195        batch_size: The batch size for training.
196        patch_shape: The patch shape to use for training.
197        task: The choice of task. Either 'task1' (LA cavity and scar) or 'task2' (LA cavity only).
198        resize_inputs: Whether to resize inputs to the desired patch shape.
199        download: Whether to download the data if it is not present.
200        kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset` or for the PyTorch DataLoader.
201
202    Returns:
203        The DataLoader.
204    """
205    ds_kwargs, loader_kwargs = util.split_kwargs(torch_em.default_segmentation_dataset, **kwargs)
206    dataset = get_lascarqs_dataset(path, patch_shape, task, resize_inputs, download, **ds_kwargs)
207    return torch_em.get_data_loader(dataset, batch_size, **loader_kwargs)
LABEL_IDS = {'background': 0, 'la_cavity': 1, 'la_scar': 2}
TASKS = {'task1': 60, 'task2': 130}
def get_lascarqs_data(path: Union[os.PathLike, str], download: bool = False) -> str:
48def get_lascarqs_data(path: Union[os.PathLike, str], download: bool = False) -> str:
49    """Obtain the LAScarQS 2022 dataset.
50
51    Args:
52        path: Filepath to a folder where the data is downloaded for further processing.
53        download: Whether to download the data if it is not present.
54
55    Returns:
56        Filepath to the folder with the downloaded and extracted dataset.
57    """
58    data_dir = os.path.join(path, "LAScarQS2022")
59    if os.path.exists(data_dir):
60        return data_dir
61
62    if download:
63        raise NotImplementedError(
64            "The LAScarQS 2022 dataset cannot be downloaded automatically. See 'get_lascarqs_data' for details."
65        )
66
67    zip_path = os.path.join(path, "LAScarQS2022.zip")
68    if not os.path.exists(zip_path):
69        raise RuntimeError(
70            f"Could not find the LAScarQS 2022 dataset at '{path}'. This dataset is not available for automatic "
71            "download. To obtain it, please follow these steps:\n"
72            "- Visit https://zmiclab.github.io/projects/lascarqs22/data.html and read the data usage agreement.\n"
73            "- Register by sending the requested information to LAScarQS2022@outlook.com or LAScarQS2022@163.com.\n"
74            "- Once you receive the download link from the organizers, download and extract the archive.\n"
75            f"- Place the extracted 'LAScarQS2022' folder (with the 'task1' and 'task2' subfolders) at '{path}', "
76            f"or place the downloaded zip archive at '{zip_path}'."
77        )
78
79    util.unzip(zip_path=zip_path, dst=path, remove=False)
80    return data_dir

Obtain the LAScarQS 2022 dataset.

Arguments:
  • path: Filepath to a folder where the data is downloaded for further processing.
  • download: Whether to download the data if it is not present.
Returns:

Filepath to the folder with the downloaded and extracted dataset.

def get_lascarqs_paths( path: Union[os.PathLike, str], task: Literal['task1', 'task2'], download: bool = False) -> List[str]:
115def get_lascarqs_paths(
116    path: Union[os.PathLike, str], task: Literal["task1", "task2"], download: bool = False
117) -> List[str]:
118    """Get paths to the LAScarQS data.
119
120    Args:
121        path: Filepath to a folder where the data is downloaded for further processing.
122        task: The choice of task. Either 'task1' (LA cavity and scar) or 'task2' (LA cavity only).
123        download: Whether to download the data if it is not present.
124
125    Returns:
126        List of filepaths for the hdf5 files, which contain the image data ('raw') and the label data ('labels').
127    """
128    if task not in TASKS:
129        raise ValueError(f"'{task}' is not a valid task. Please choose one of {list(TASKS.keys())}.")
130
131    data_dir = get_lascarqs_data(path, download)
132
133    preprocessed_dir = os.path.join(path, "preprocessed", task)
134    if len(glob(os.path.join(preprocessed_dir, "*.h5"))) != TASKS[task]:
135        _preprocess_inputs(data_dir, task, preprocessed_dir)
136
137    volume_paths = natsorted(glob(os.path.join(preprocessed_dir, "*.h5")))
138    assert len(volume_paths) > 0, f"Could not find any preprocessed volumes in '{preprocessed_dir}'."
139
140    return volume_paths

Get paths to the LAScarQS data.

Arguments:
  • path: Filepath to a folder where the data is downloaded for further processing.
  • task: The choice of task. Either 'task1' (LA cavity and scar) or 'task2' (LA cavity only).
  • download: Whether to download the data if it is not present.
Returns:

List of filepaths for the hdf5 files, which contain the image data ('raw') and the label data ('labels').

def get_lascarqs_dataset( path: Union[os.PathLike, str], patch_shape: Tuple[int, ...], task: Literal['task1', 'task2'], resize_inputs: bool = False, download: bool = False, **kwargs) -> torch.utils.data.dataset.Dataset:
143def get_lascarqs_dataset(
144    path: Union[os.PathLike, str],
145    patch_shape: Tuple[int, ...],
146    task: Literal["task1", "task2"],
147    resize_inputs: bool = False,
148    download: bool = False,
149    **kwargs
150) -> Dataset:
151    """Get the LAScarQS dataset for left atrium and atrial scar segmentation.
152
153    Args:
154        path: Filepath to a folder where the data is downloaded for further processing.
155        patch_shape: The patch shape to use for training.
156        task: The choice of task. Either 'task1' (LA cavity and scar) or 'task2' (LA cavity only).
157        resize_inputs: Whether to resize inputs to the desired patch shape.
158        download: Whether to download the data if it is not present.
159        kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset`.
160
161    Returns:
162        The segmentation dataset.
163    """
164    volume_paths = get_lascarqs_paths(path, task, download)
165
166    if resize_inputs:
167        resize_kwargs = {"patch_shape": patch_shape, "is_rgb": False}
168        kwargs, patch_shape = util.update_kwargs_for_resize_trafo(
169            kwargs=kwargs, patch_shape=patch_shape, resize_inputs=resize_inputs, resize_kwargs=resize_kwargs
170        )
171
172    return torch_em.default_segmentation_dataset(
173        raw_paths=volume_paths,
174        raw_key="raw",
175        label_paths=volume_paths,
176        label_key="labels",
177        patch_shape=patch_shape,
178        is_seg_dataset=True,
179        **kwargs
180    )

Get the LAScarQS dataset for left atrium and atrial scar segmentation.

Arguments:
  • path: Filepath to a folder where the data is downloaded for further processing.
  • patch_shape: The patch shape to use for training.
  • task: The choice of task. Either 'task1' (LA cavity and scar) or 'task2' (LA cavity only).
  • resize_inputs: Whether to resize inputs to the desired patch shape.
  • download: Whether to download the data if it is not present.
  • kwargs: Additional keyword arguments for torch_em.default_segmentation_dataset.
Returns:

The segmentation dataset.

def get_lascarqs_loader( path: Union[os.PathLike, str], batch_size: int, patch_shape: Tuple[int, ...], task: Literal['task1', 'task2'], resize_inputs: bool = False, download: bool = False, **kwargs) -> torch.utils.data.dataloader.DataLoader:
183def get_lascarqs_loader(
184    path: Union[os.PathLike, str],
185    batch_size: int,
186    patch_shape: Tuple[int, ...],
187    task: Literal["task1", "task2"],
188    resize_inputs: bool = False,
189    download: bool = False,
190    **kwargs
191) -> DataLoader:
192    """Get the LAScarQS dataloader for left atrium and atrial scar segmentation.
193
194    Args:
195        path: Filepath to a folder where the data is downloaded for further processing.
196        batch_size: The batch size for training.
197        patch_shape: The patch shape to use for training.
198        task: The choice of task. Either 'task1' (LA cavity and scar) or 'task2' (LA cavity only).
199        resize_inputs: Whether to resize inputs to the desired patch shape.
200        download: Whether to download the data if it is not present.
201        kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset` or for the PyTorch DataLoader.
202
203    Returns:
204        The DataLoader.
205    """
206    ds_kwargs, loader_kwargs = util.split_kwargs(torch_em.default_segmentation_dataset, **kwargs)
207    dataset = get_lascarqs_dataset(path, patch_shape, task, resize_inputs, download, **ds_kwargs)
208    return torch_em.get_data_loader(dataset, batch_size, **loader_kwargs)

Get the LAScarQS dataloader for left atrium and atrial scar segmentation.

Arguments:
  • path: Filepath to a folder where the data is downloaded for further processing.
  • batch_size: The batch size for training.
  • patch_shape: The patch shape to use for training.
  • task: The choice of task. Either 'task1' (LA cavity and scar) or 'task2' (LA cavity only).
  • resize_inputs: Whether to resize inputs to the desired patch shape.
  • download: Whether to download the data if it is not present.
  • kwargs: Additional keyword arguments for torch_em.default_segmentation_dataset or for the PyTorch DataLoader.
Returns:

The DataLoader.