torch_em.data.datasets.medical.liver_hcc_seg

The LiverHccSeg dataset contains annotations for liver and hepatocellular carcinoma (HCC) tumor segmentation in multiphasic contrast-enhanced MRI.

The dataset is built from the multi-parametric MRI arm of TCGA-LIHC. It ships 4 native acquisition phases ('pre', 'art', 'pv', 'del' - pre-contrast, arterial, portal-venous and delayed) plus 3 phases registered onto the arterial phase ('art_pre', 'art_pv', 'art_del') for each of 18 exams. Whole-liver masks are available for 17 exams and HCC tumor masks (up to 3 lesions per exam) for 14 exams, each independently annotated by two board-certified abdominal radiologists ('rater1' and 'rater2'). This module merges the per-lesion tumor masks of an exam into a single instance label volume.

NOTE: One exam ships a raw volume and liver mask with a mismatched slice count; this loader skips that pair, so 16 (not 17) liver exams are usable.

NOTE: This is MRI data and is not the same as the already-integrated torch_em.data.datasets.medical.waw_tace or torch_em.data.datasets.medical.hcc_tace, which are contrast-enhanced CT datasets, nor torch_em.data.datasets.medical.openswisshcc, which is a different, larger multiphasic liver/HCC MRI cohort.

The data is located at https://doi.org/10.5281/zenodo.7957516, released under a CC-BY-4.0 license.

This dataset is from the publication https://doi.org/10.1016/j.dib.2023.109607. Please cite it if you use this dataset for your research.

  1"""The LiverHccSeg dataset contains annotations for liver and hepatocellular carcinoma (HCC)
  2tumor segmentation in multiphasic contrast-enhanced MRI.
  3
  4The dataset is built from the multi-parametric MRI arm of TCGA-LIHC. It ships 4 native
  5acquisition phases ('pre', 'art', 'pv', 'del' - pre-contrast, arterial, portal-venous and
  6delayed) plus 3 phases registered onto the arterial phase ('art_pre', 'art_pv', 'art_del')
  7for each of 18 exams. Whole-liver masks are available for 17 exams and HCC tumor masks
  8(up to 3 lesions per exam) for 14 exams, each independently annotated by two board-certified
  9abdominal radiologists ('rater1' and 'rater2'). This module merges the per-lesion tumor masks
 10of an exam into a single instance label volume.
 11
 12NOTE: One exam ships a raw volume and liver mask with a mismatched slice count; this loader skips
 13that pair, so 16 (not 17) liver exams are usable.
 14
 15NOTE: This is MRI data and is not the same as the already-integrated
 16`torch_em.data.datasets.medical.waw_tace` or `torch_em.data.datasets.medical.hcc_tace`, which
 17are contrast-enhanced CT datasets, nor `torch_em.data.datasets.medical.openswisshcc`, which is a
 18different, larger multiphasic liver/HCC MRI cohort.
 19
 20The data is located at https://doi.org/10.5281/zenodo.7957516, released under a CC-BY-4.0 license.
 21
 22This dataset is from the publication https://doi.org/10.1016/j.dib.2023.109607.
 23Please cite it if you use this dataset for your research.
 24"""
 25
 26import os
 27from glob import glob
 28from natsort import natsorted
 29from typing import Union, Tuple, Literal, List
 30
 31import numpy as np
 32
 33from torch.utils.data import Dataset, DataLoader
 34
 35import torch_em
 36
 37from .. import util
 38
 39
 40URL = "https://zenodo.org/records/7957516/files/nifti_and_segms.zip"
 41CHECKSUM = "dba5f89a95b9c0cc4fec466fa8105663b754e7039ebee996b708ecf3ee114d7d"
 42
 43PHASES = ("pre", "art", "pv", "del", "art_pre", "art_pv", "art_del")
 44RATERS = (1, 2)
 45
 46
 47def get_liver_hcc_seg_data(path: Union[os.PathLike, str], download: bool = False) -> str:
 48    """Download the LiverHccSeg dataset.
 49
 50    Args:
 51        path: Filepath to a folder where the data is downloaded for further processing.
 52        download: Whether to download the data if it is not present.
 53
 54    Returns:
 55        Filepath where the data is downloaded.
 56    """
 57    data_dir = os.path.join(path, "nifti_and_segms")
 58    if os.path.exists(data_dir):
 59        return data_dir
 60
 61    os.makedirs(path, exist_ok=True)
 62
 63    zip_path = os.path.join(path, "nifti_and_segms.zip")
 64    util.download_source(path=zip_path, url=URL, download=download, checksum=CHECKSUM)
 65    util.unzip(zip_path=zip_path, dst=path)
 66
 67    return data_dir
 68
 69
 70def _exam_dirs(data_dir):
 71    return natsorted(glob(os.path.join(data_dir, "TCGA-*", "*")))
 72
 73
 74def _merged_tumor_mask_path(exam_dir, rater):
 75    return os.path.join(exam_dir, f"rater{rater}_tumor_merged.nii.gz")
 76
 77
 78def _merge_tumor_masks(exam_dir, rater):
 79    """Merge the per-lesion tumor masks of one rater into a single instance label volume."""
 80    out_path = _merged_tumor_mask_path(exam_dir, rater)
 81    if os.path.exists(out_path):
 82        return out_path
 83
 84    lesion_paths = natsorted(glob(os.path.join(exam_dir, f"rater{rater}_tumor*.nii.gz")))
 85    if not lesion_paths:
 86        return None
 87
 88    import nibabel as nib
 89
 90    merged = None
 91    reference = None
 92    for lesion_id, lesion_path in enumerate(lesion_paths, start=1):
 93        lesion_img = nib.load(lesion_path)
 94        if merged is None:
 95            reference = lesion_img
 96            merged = np.zeros(lesion_img.shape, dtype="uint8")
 97        merged[lesion_img.get_fdata() > 0] = lesion_id
 98
 99    nib.save(nib.Nifti1Image(merged, reference.affine, reference.header), out_path)
100    return out_path
101
102
103def get_liver_hcc_seg_paths(
104    path: Union[os.PathLike, str],
105    phase: str = "pre",
106    target: Literal["liver", "tumor"] = "liver",
107    rater: Literal[1, 2] = 1,
108    download: bool = False,
109) -> Tuple[List[str], List[str]]:
110    """Get paths to the LiverHccSeg data.
111
112    Args:
113        path: Filepath to a folder where the data is downloaded for further processing.
114        phase: The choice of MRI phase. One of 'pre', 'art', 'pv', 'del', 'art_pre', 'art_pv', 'art_del'.
115        target: The choice of segmentation target. Either 'liver' (whole-liver mask) or
116            'tumor' (merged instance mask of the annotated HCC lesions).
117        rater: The choice of annotator. Either 1 or 2.
118        download: Whether to download the data if it is not present.
119
120    Returns:
121        List of filepaths for the image data.
122        List of filepaths for the label data.
123    """
124    if phase not in PHASES:
125        raise ValueError(f"'{phase}' is not a valid phase. Choose one of {PHASES}.")
126    if target not in ("liver", "tumor"):
127        raise ValueError(f"'{target}' is not a valid target. Choose 'liver' or 'tumor'.")
128    if rater not in RATERS:
129        raise ValueError(f"'{rater}' is not a valid rater. Choose one of {RATERS}.")
130
131    import nibabel as nib
132
133    data_dir = get_liver_hcc_seg_data(path, download)
134
135    raw_paths, label_paths = [], []
136    for exam_dir in _exam_dirs(data_dir):
137        raw_path = os.path.join(exam_dir, f"{phase}.nii.gz")
138        if not os.path.exists(raw_path):
139            continue
140
141        if target == "liver":
142            label_path = os.path.join(exam_dir, f"rater{rater}_liver.nii.gz")
143            if not os.path.exists(label_path):
144                continue
145        else:
146            label_path = _merge_tumor_masks(exam_dir, rater)
147            if label_path is None:
148                continue
149
150        # One exam (TCGA-DD-A4NH) ships a raw volume and liver mask with a mismatched slice count.
151        # Skip such pairs rather than failing the whole dataset.
152        if nib.load(raw_path).shape != nib.load(label_path).shape:
153            continue
154
155        raw_paths.append(raw_path)
156        label_paths.append(label_path)
157
158    assert len(raw_paths) == len(label_paths) and len(raw_paths) > 0
159    return raw_paths, label_paths
160
161
162def get_liver_hcc_seg_dataset(
163    path: Union[os.PathLike, str],
164    patch_shape: Tuple[int, ...],
165    phase: str = "pre",
166    target: Literal["liver", "tumor"] = "liver",
167    rater: Literal[1, 2] = 1,
168    resize_inputs: bool = False,
169    download: bool = False,
170    **kwargs
171) -> Dataset:
172    """Get the LiverHccSeg dataset for liver and HCC tumor segmentation in multiphasic MRI.
173
174    Args:
175        path: Filepath to a folder where the data is downloaded for further processing.
176        patch_shape: The patch shape to use for training.
177        phase: The choice of MRI phase. One of 'pre', 'art', 'pv', 'del', 'art_pre', 'art_pv', 'art_del'.
178        target: The choice of segmentation target. Either 'liver' or 'tumor'.
179        rater: The choice of annotator. Either 1 or 2.
180        resize_inputs: Whether to resize the inputs to the patch shape.
181        download: Whether to download the data if it is not present.
182        kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset`.
183
184    Returns:
185        The segmentation dataset.
186    """
187    raw_paths, label_paths = get_liver_hcc_seg_paths(path, phase, target, rater, download)
188
189    if resize_inputs:
190        resize_kwargs = {"patch_shape": patch_shape, "is_rgb": False}
191        kwargs, patch_shape = util.update_kwargs_for_resize_trafo(
192            kwargs=kwargs, patch_shape=patch_shape, resize_inputs=resize_inputs, resize_kwargs=resize_kwargs
193        )
194
195    return torch_em.default_segmentation_dataset(
196        raw_paths=raw_paths,
197        raw_key="data",
198        label_paths=label_paths,
199        label_key="data",
200        patch_shape=patch_shape,
201        is_seg_dataset=True,
202        **kwargs
203    )
204
205
206def get_liver_hcc_seg_loader(
207    path: Union[os.PathLike, str],
208    batch_size: int,
209    patch_shape: Tuple[int, ...],
210    phase: str = "pre",
211    target: Literal["liver", "tumor"] = "liver",
212    rater: Literal[1, 2] = 1,
213    resize_inputs: bool = False,
214    download: bool = False,
215    **kwargs
216) -> DataLoader:
217    """Get the LiverHccSeg dataloader for liver and HCC tumor segmentation in multiphasic MRI.
218
219    Args:
220        path: Filepath to a folder where the data is downloaded for further processing.
221        batch_size: The batch size for training.
222        patch_shape: The patch shape to use for training.
223        phase: The choice of MRI phase. One of 'pre', 'art', 'pv', 'del', 'art_pre', 'art_pv', 'art_del'.
224        target: The choice of segmentation target. Either 'liver' or 'tumor'.
225        rater: The choice of annotator. Either 1 or 2.
226        resize_inputs: Whether to resize the inputs to the patch shape.
227        download: Whether to download the data if it is not present.
228        kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset` or for the PyTorch DataLoader.
229
230    Returns:
231        The DataLoader.
232    """
233    ds_kwargs, loader_kwargs = util.split_kwargs(torch_em.default_segmentation_dataset, **kwargs)
234    dataset = get_liver_hcc_seg_dataset(path, patch_shape, phase, target, rater, resize_inputs, download, **ds_kwargs)
235    return torch_em.get_data_loader(dataset, batch_size, **loader_kwargs)
URL = 'https://zenodo.org/records/7957516/files/nifti_and_segms.zip'
CHECKSUM = 'dba5f89a95b9c0cc4fec466fa8105663b754e7039ebee996b708ecf3ee114d7d'
PHASES = ('pre', 'art', 'pv', 'del', 'art_pre', 'art_pv', 'art_del')
RATERS = (1, 2)
def get_liver_hcc_seg_data(path: Union[os.PathLike, str], download: bool = False) -> str:
48def get_liver_hcc_seg_data(path: Union[os.PathLike, str], download: bool = False) -> str:
49    """Download the LiverHccSeg dataset.
50
51    Args:
52        path: Filepath to a folder where the data is downloaded for further processing.
53        download: Whether to download the data if it is not present.
54
55    Returns:
56        Filepath where the data is downloaded.
57    """
58    data_dir = os.path.join(path, "nifti_and_segms")
59    if os.path.exists(data_dir):
60        return data_dir
61
62    os.makedirs(path, exist_ok=True)
63
64    zip_path = os.path.join(path, "nifti_and_segms.zip")
65    util.download_source(path=zip_path, url=URL, download=download, checksum=CHECKSUM)
66    util.unzip(zip_path=zip_path, dst=path)
67
68    return data_dir

Download the LiverHccSeg dataset.

Arguments:
  • path: Filepath to a folder where the data is downloaded for further processing.
  • download: Whether to download the data if it is not present.
Returns:

Filepath where the data is downloaded.

def get_liver_hcc_seg_paths( path: Union[os.PathLike, str], phase: str = 'pre', target: Literal['liver', 'tumor'] = 'liver', rater: Literal[1, 2] = 1, download: bool = False) -> Tuple[List[str], List[str]]:
104def get_liver_hcc_seg_paths(
105    path: Union[os.PathLike, str],
106    phase: str = "pre",
107    target: Literal["liver", "tumor"] = "liver",
108    rater: Literal[1, 2] = 1,
109    download: bool = False,
110) -> Tuple[List[str], List[str]]:
111    """Get paths to the LiverHccSeg data.
112
113    Args:
114        path: Filepath to a folder where the data is downloaded for further processing.
115        phase: The choice of MRI phase. One of 'pre', 'art', 'pv', 'del', 'art_pre', 'art_pv', 'art_del'.
116        target: The choice of segmentation target. Either 'liver' (whole-liver mask) or
117            'tumor' (merged instance mask of the annotated HCC lesions).
118        rater: The choice of annotator. Either 1 or 2.
119        download: Whether to download the data if it is not present.
120
121    Returns:
122        List of filepaths for the image data.
123        List of filepaths for the label data.
124    """
125    if phase not in PHASES:
126        raise ValueError(f"'{phase}' is not a valid phase. Choose one of {PHASES}.")
127    if target not in ("liver", "tumor"):
128        raise ValueError(f"'{target}' is not a valid target. Choose 'liver' or 'tumor'.")
129    if rater not in RATERS:
130        raise ValueError(f"'{rater}' is not a valid rater. Choose one of {RATERS}.")
131
132    import nibabel as nib
133
134    data_dir = get_liver_hcc_seg_data(path, download)
135
136    raw_paths, label_paths = [], []
137    for exam_dir in _exam_dirs(data_dir):
138        raw_path = os.path.join(exam_dir, f"{phase}.nii.gz")
139        if not os.path.exists(raw_path):
140            continue
141
142        if target == "liver":
143            label_path = os.path.join(exam_dir, f"rater{rater}_liver.nii.gz")
144            if not os.path.exists(label_path):
145                continue
146        else:
147            label_path = _merge_tumor_masks(exam_dir, rater)
148            if label_path is None:
149                continue
150
151        # One exam (TCGA-DD-A4NH) ships a raw volume and liver mask with a mismatched slice count.
152        # Skip such pairs rather than failing the whole dataset.
153        if nib.load(raw_path).shape != nib.load(label_path).shape:
154            continue
155
156        raw_paths.append(raw_path)
157        label_paths.append(label_path)
158
159    assert len(raw_paths) == len(label_paths) and len(raw_paths) > 0
160    return raw_paths, label_paths

Get paths to the LiverHccSeg data.

Arguments:
  • path: Filepath to a folder where the data is downloaded for further processing.
  • phase: The choice of MRI phase. One of 'pre', 'art', 'pv', 'del', 'art_pre', 'art_pv', 'art_del'.
  • target: The choice of segmentation target. Either 'liver' (whole-liver mask) or 'tumor' (merged instance mask of the annotated HCC lesions).
  • rater: The choice of annotator. Either 1 or 2.
  • download: Whether to download the data if it is not present.
Returns:

List of filepaths for the image data. List of filepaths for the label data.

def get_liver_hcc_seg_dataset( path: Union[os.PathLike, str], patch_shape: Tuple[int, ...], phase: str = 'pre', target: Literal['liver', 'tumor'] = 'liver', rater: Literal[1, 2] = 1, resize_inputs: bool = False, download: bool = False, **kwargs) -> torch.utils.data.dataset.Dataset:
163def get_liver_hcc_seg_dataset(
164    path: Union[os.PathLike, str],
165    patch_shape: Tuple[int, ...],
166    phase: str = "pre",
167    target: Literal["liver", "tumor"] = "liver",
168    rater: Literal[1, 2] = 1,
169    resize_inputs: bool = False,
170    download: bool = False,
171    **kwargs
172) -> Dataset:
173    """Get the LiverHccSeg dataset for liver and HCC tumor segmentation in multiphasic MRI.
174
175    Args:
176        path: Filepath to a folder where the data is downloaded for further processing.
177        patch_shape: The patch shape to use for training.
178        phase: The choice of MRI phase. One of 'pre', 'art', 'pv', 'del', 'art_pre', 'art_pv', 'art_del'.
179        target: The choice of segmentation target. Either 'liver' or 'tumor'.
180        rater: The choice of annotator. Either 1 or 2.
181        resize_inputs: Whether to resize the inputs to the patch shape.
182        download: Whether to download the data if it is not present.
183        kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset`.
184
185    Returns:
186        The segmentation dataset.
187    """
188    raw_paths, label_paths = get_liver_hcc_seg_paths(path, phase, target, rater, download)
189
190    if resize_inputs:
191        resize_kwargs = {"patch_shape": patch_shape, "is_rgb": False}
192        kwargs, patch_shape = util.update_kwargs_for_resize_trafo(
193            kwargs=kwargs, patch_shape=patch_shape, resize_inputs=resize_inputs, resize_kwargs=resize_kwargs
194        )
195
196    return torch_em.default_segmentation_dataset(
197        raw_paths=raw_paths,
198        raw_key="data",
199        label_paths=label_paths,
200        label_key="data",
201        patch_shape=patch_shape,
202        is_seg_dataset=True,
203        **kwargs
204    )

Get the LiverHccSeg dataset for liver and HCC tumor segmentation in multiphasic MRI.

Arguments:
  • path: Filepath to a folder where the data is downloaded for further processing.
  • patch_shape: The patch shape to use for training.
  • phase: The choice of MRI phase. One of 'pre', 'art', 'pv', 'del', 'art_pre', 'art_pv', 'art_del'.
  • target: The choice of segmentation target. Either 'liver' or 'tumor'.
  • rater: The choice of annotator. Either 1 or 2.
  • resize_inputs: Whether to resize the inputs to the patch shape.
  • download: Whether to download the data if it is not present.
  • kwargs: Additional keyword arguments for torch_em.default_segmentation_dataset.
Returns:

The segmentation dataset.

def get_liver_hcc_seg_loader( path: Union[os.PathLike, str], batch_size: int, patch_shape: Tuple[int, ...], phase: str = 'pre', target: Literal['liver', 'tumor'] = 'liver', rater: Literal[1, 2] = 1, resize_inputs: bool = False, download: bool = False, **kwargs) -> torch.utils.data.dataloader.DataLoader:
207def get_liver_hcc_seg_loader(
208    path: Union[os.PathLike, str],
209    batch_size: int,
210    patch_shape: Tuple[int, ...],
211    phase: str = "pre",
212    target: Literal["liver", "tumor"] = "liver",
213    rater: Literal[1, 2] = 1,
214    resize_inputs: bool = False,
215    download: bool = False,
216    **kwargs
217) -> DataLoader:
218    """Get the LiverHccSeg dataloader for liver and HCC tumor segmentation in multiphasic MRI.
219
220    Args:
221        path: Filepath to a folder where the data is downloaded for further processing.
222        batch_size: The batch size for training.
223        patch_shape: The patch shape to use for training.
224        phase: The choice of MRI phase. One of 'pre', 'art', 'pv', 'del', 'art_pre', 'art_pv', 'art_del'.
225        target: The choice of segmentation target. Either 'liver' or 'tumor'.
226        rater: The choice of annotator. Either 1 or 2.
227        resize_inputs: Whether to resize the inputs to the patch shape.
228        download: Whether to download the data if it is not present.
229        kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset` or for the PyTorch DataLoader.
230
231    Returns:
232        The DataLoader.
233    """
234    ds_kwargs, loader_kwargs = util.split_kwargs(torch_em.default_segmentation_dataset, **kwargs)
235    dataset = get_liver_hcc_seg_dataset(path, patch_shape, phase, target, rater, resize_inputs, download, **ds_kwargs)
236    return torch_em.get_data_loader(dataset, batch_size, **loader_kwargs)

Get the LiverHccSeg dataloader for liver and HCC tumor segmentation in multiphasic MRI.

Arguments:
  • path: Filepath to a folder where the data is downloaded for further processing.
  • batch_size: The batch size for training.
  • patch_shape: The patch shape to use for training.
  • phase: The choice of MRI phase. One of 'pre', 'art', 'pv', 'del', 'art_pre', 'art_pv', 'art_del'.
  • target: The choice of segmentation target. Either 'liver' or 'tumor'.
  • rater: The choice of annotator. Either 1 or 2.
  • resize_inputs: Whether to resize the inputs to the patch shape.
  • download: Whether to download the data if it is not present.
  • kwargs: Additional keyword arguments for torch_em.default_segmentation_dataset or for the PyTorch DataLoader.
Returns:

The DataLoader.