torch_em.data.datasets.medical.hvm

HVM (Hepatic Vessel Map) is a dual-center dataset for the segmentation of hepatic veins, portal veins (to third-order branches) and liver tumors in contrast-enhanced abdominal CT.

The dataset comprises 282 patients: 170 scans from Center 1 (The University of Hong Kong-Shenzhen Hospital, of which 106 have hepatic and portal vein annotations) and 176 scans from Center 2 (Peking University Shenzhen Hospital, of which 174 have hepatic and portal vein annotations). Liver tumor annotations (96 cases) are only provided for a subset of Center 1 (the archive also ships an identical copy of these files under Center 2, which does not correspond to any Center 2 scan and is therefore not used by this loader).

NOTE: The raw scans and the annotations in each folder are not matched by their filename id (e.g. 'Image/1.nii.gz' does not necessarily correspond to 'Annotation_Hepatic veins/001.nii.gz'). They are instead matched here by comparing the NIfTI header geometry (shape, voxel spacing and origin), which uniquely identifies the corresponding scan for (almost) every annotation.

This dataset is NOT the same as the 'HVA-CT' dataset (see hva_ct.py), which re-annotates the 61 scans of the Medical Segmentation Decathlon hepatic vessel task.

The dataset is located at https://doi.org/10.5281/zenodo.19885789 and is distributed under the CC BY 4.0 license. The dataset is from the publication https://doi.org/10.1038/s41597-026-07550-3. Please cite it if you use this dataset for your research.

  1"""HVM (Hepatic Vessel Map) is a dual-center dataset for the segmentation of hepatic veins, portal
  2veins (to third-order branches) and liver tumors in contrast-enhanced abdominal CT.
  3
  4The dataset comprises 282 patients: 170 scans from Center 1 (The University of Hong Kong-Shenzhen
  5Hospital, of which 106 have hepatic and portal vein annotations) and 176 scans from Center 2
  6(Peking University Shenzhen Hospital, of which 174 have hepatic and portal vein annotations). Liver
  7tumor annotations (96 cases) are only provided for a subset of Center 1 (the archive also ships an
  8identical copy of these files under Center 2, which does not correspond to any Center 2 scan and is
  9therefore not used by this loader).
 10
 11NOTE: The raw scans and the annotations in each folder are not matched by their filename id (e.g.
 12'Image/1.nii.gz' does not necessarily correspond to 'Annotation_Hepatic veins/001.nii.gz'). They are
 13instead matched here by comparing the NIfTI header geometry (shape, voxel spacing and origin), which
 14uniquely identifies the corresponding scan for (almost) every annotation.
 15
 16This dataset is NOT the same as the 'HVA-CT' dataset (see `hva_ct.py`), which re-annotates the 61
 17scans of the Medical Segmentation Decathlon hepatic vessel task.
 18
 19The dataset is located at https://doi.org/10.5281/zenodo.19885789 and is distributed under the
 20CC BY 4.0 license. The dataset is from the publication https://doi.org/10.1038/s41597-026-07550-3.
 21Please cite it if you use this dataset for your research.
 22"""
 23
 24import os
 25import warnings
 26from glob import glob
 27from natsort import natsorted
 28from typing import Union, Tuple, Literal, List
 29
 30from torch.utils.data import Dataset, DataLoader
 31
 32import torch_em
 33
 34from .. import util
 35
 36
 37URL = "https://zenodo.org/api/records/19885789/files/HVM%20Dataset.zip/content"
 38CHECKSUM = "7fe2a00bc8a40658d45bedd27cd70df22e1a58a92af20037c22b05065407a6b8"
 39
 40CENTERS = ["1", "2"]
 41
 42ANNOTATIONS = {
 43    "hepatic_veins": "Annotation_Hepatic veins",
 44    "portal_veins": "Annotation_Portal veins",
 45    "liver_tumor": "Annotation_Liver tumors",
 46}
 47
 48
 49def get_hvm_data(path: Union[os.PathLike, str], download: bool = False) -> str:
 50    """Download the HVM dataset.
 51
 52    Args:
 53        path: Filepath to a folder where the data is downloaded for further processing.
 54        download: Whether to download the data if it is not present.
 55
 56    Returns:
 57        Filepath where the data is downloaded.
 58    """
 59    data_dir = os.path.join(path, "HVM Dataset")
 60    if os.path.exists(data_dir):
 61        return data_dir
 62
 63    os.makedirs(path, exist_ok=True)
 64
 65    zip_path = os.path.join(path, "HVM_Dataset.zip")
 66    util.download_source(path=zip_path, url=URL, download=download, checksum=CHECKSUM)
 67    util.unzip(zip_path=zip_path, dst=path)
 68
 69    return data_dir
 70
 71
 72def _header_signature(path):
 73    import nibabel as nib
 74
 75    header = nib.load(path).header
 76    shape = header.get_data_shape()
 77    zooms = tuple(header.get_zooms())
 78    origin = tuple(header.get_best_affine()[:3, 3])
 79    return shape, zooms, origin
 80
 81
 82def _match_images_to_annotations(image_paths, annotation_paths):
 83    image_sigs = {p: _header_signature(p) for p in image_paths}
 84
 85    matched_images, matched_annotations = [], []
 86    for annotation_path in annotation_paths:
 87        annotation_sig = _header_signature(annotation_path)
 88        candidates = [p for p, sig in image_sigs.items() if sig == annotation_sig]
 89        if not candidates:
 90            warnings.warn(f"Could not find a matching scan for the annotation at '{annotation_path}'. Skipping it.")
 91            continue
 92        matched_images.append(natsorted(candidates)[0])
 93        matched_annotations.append(annotation_path)
 94
 95    return matched_images, matched_annotations
 96
 97
 98def get_hvm_paths(
 99    path: Union[os.PathLike, str],
100    annotation: Literal["hepatic_veins", "portal_veins", "liver_tumor"] = "hepatic_veins",
101    center: Literal["1", "2", "both"] = "both",
102    download: bool = False,
103) -> Tuple[List[str], List[str]]:
104    """Get paths to the HVM data.
105
106    Args:
107        path: Filepath to a folder where the data is downloaded for further processing.
108        annotation: The choice of annotations. One of 'hepatic_veins', 'portal_veins' or 'liver_tumor'.
109        center: The choice of source center. One of '1', '2' or 'both'. Ignored (forced to '1') when
110            `annotation` is 'liver_tumor', as the tumor annotations shipped under 'Center 2' are an
111            identical copy of the 'Center 1' ones and do not correspond to any 'Center 2' scan.
112        download: Whether to download the data if it is not present.
113
114    Returns:
115        List of filepaths for the image data.
116        List of filepaths for the label data.
117    """
118    if annotation not in ANNOTATIONS:
119        raise ValueError(f"'{annotation}' is not a valid annotation. Choose from {list(ANNOTATIONS.keys())}.")
120    if center not in CENTERS + ["both"]:
121        raise ValueError(f"'{center}' is not a valid center. Choose from {CENTERS + ['both']}.")
122
123    data_dir = get_hvm_data(path=path, download=download)
124
125    centers = ["1"] if annotation == "liver_tumor" else (CENTERS if center == "both" else [center])
126
127    image_paths, gt_paths = [], []
128    for c in centers:
129        center_dir = os.path.join(data_dir, f"Center {c}")
130        image_dir = os.path.join(center_dir, "Image")
131        annotation_dir = os.path.join(center_dir, ANNOTATIONS[annotation])
132
133        this_images = natsorted(glob(os.path.join(image_dir, "*.nii.gz")))
134        this_annotations = natsorted(glob(os.path.join(annotation_dir, "*.nii.gz")))
135
136        matched_images, matched_annotations = _match_images_to_annotations(this_images, this_annotations)
137        image_paths.extend(matched_images)
138        gt_paths.extend(matched_annotations)
139
140    assert len(image_paths) == len(gt_paths) and len(image_paths) > 0
141
142    return image_paths, gt_paths
143
144
145def get_hvm_dataset(
146    path: Union[os.PathLike, str],
147    patch_shape: Tuple[int, ...],
148    annotation: Literal["hepatic_veins", "portal_veins", "liver_tumor"] = "hepatic_veins",
149    center: Literal["1", "2", "both"] = "both",
150    resize_inputs: bool = False,
151    download: bool = False,
152    **kwargs
153) -> Dataset:
154    """Get the HVM dataset for hepatic vein, portal vein and liver tumor segmentation in abdominal CT.
155
156    Args:
157        path: Filepath to a folder where the data is downloaded for further processing.
158        patch_shape: The patch shape to use for training.
159        annotation: The choice of annotations. One of 'hepatic_veins', 'portal_veins' or 'liver_tumor'.
160        center: The choice of source center. One of '1', '2' or 'both'. Ignored (forced to '1') when
161            `annotation` is 'liver_tumor'.
162        resize_inputs: Whether to resize inputs to the desired patch shape.
163        download: Whether to download the data if it is not present.
164        kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset`.
165
166    Returns:
167        The segmentation dataset.
168    """
169    raw_paths, label_paths = get_hvm_paths(path, annotation, center, download)
170
171    if resize_inputs:
172        resize_kwargs = {"patch_shape": patch_shape, "is_rgb": False}
173        kwargs, patch_shape = util.update_kwargs_for_resize_trafo(
174            kwargs=kwargs, patch_shape=patch_shape, resize_inputs=resize_inputs, resize_kwargs=resize_kwargs
175        )
176
177    return torch_em.default_segmentation_dataset(
178        raw_paths=raw_paths,
179        raw_key="data",
180        label_paths=label_paths,
181        label_key="data",
182        patch_shape=patch_shape,
183        is_seg_dataset=True,
184        **kwargs
185    )
186
187
188def get_hvm_loader(
189    path: Union[os.PathLike, str],
190    batch_size: int,
191    patch_shape: Tuple[int, ...],
192    annotation: Literal["hepatic_veins", "portal_veins", "liver_tumor"] = "hepatic_veins",
193    center: Literal["1", "2", "both"] = "both",
194    resize_inputs: bool = False,
195    download: bool = False,
196    **kwargs
197) -> DataLoader:
198    """Get the HVM dataloader for hepatic vein, portal vein and liver tumor segmentation in abdominal CT.
199
200    Args:
201        path: Filepath to a folder where the data is downloaded for further processing.
202        batch_size: The batch size for training.
203        patch_shape: The patch shape to use for training.
204        annotation: The choice of annotations. One of 'hepatic_veins', 'portal_veins' or 'liver_tumor'.
205        center: The choice of source center. One of '1', '2' or 'both'. Ignored (forced to '1') when
206            `annotation` is 'liver_tumor'.
207        resize_inputs: Whether to resize inputs to the desired patch shape.
208        download: Whether to download the data if it is not present.
209        kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset` or for the PyTorch DataLoader.
210
211    Returns:
212        The DataLoader.
213    """
214    ds_kwargs, loader_kwargs = util.split_kwargs(torch_em.default_segmentation_dataset, **kwargs)
215    dataset = get_hvm_dataset(path, patch_shape, annotation, center, resize_inputs, download, **ds_kwargs)
216    return torch_em.get_data_loader(dataset, batch_size, **loader_kwargs)
URL = 'https://zenodo.org/api/records/19885789/files/HVM%20Dataset.zip/content'
CHECKSUM = '7fe2a00bc8a40658d45bedd27cd70df22e1a58a92af20037c22b05065407a6b8'
CENTERS = ['1', '2']
ANNOTATIONS = {'hepatic_veins': 'Annotation_Hepatic veins', 'portal_veins': 'Annotation_Portal veins', 'liver_tumor': 'Annotation_Liver tumors'}
def get_hvm_data(path: Union[os.PathLike, str], download: bool = False) -> str:
50def get_hvm_data(path: Union[os.PathLike, str], download: bool = False) -> str:
51    """Download the HVM dataset.
52
53    Args:
54        path: Filepath to a folder where the data is downloaded for further processing.
55        download: Whether to download the data if it is not present.
56
57    Returns:
58        Filepath where the data is downloaded.
59    """
60    data_dir = os.path.join(path, "HVM Dataset")
61    if os.path.exists(data_dir):
62        return data_dir
63
64    os.makedirs(path, exist_ok=True)
65
66    zip_path = os.path.join(path, "HVM_Dataset.zip")
67    util.download_source(path=zip_path, url=URL, download=download, checksum=CHECKSUM)
68    util.unzip(zip_path=zip_path, dst=path)
69
70    return data_dir

Download the HVM dataset.

Arguments:
  • path: Filepath to a folder where the data is downloaded for further processing.
  • download: Whether to download the data if it is not present.
Returns:

Filepath where the data is downloaded.

def get_hvm_paths( path: Union[os.PathLike, str], annotation: Literal['hepatic_veins', 'portal_veins', 'liver_tumor'] = 'hepatic_veins', center: Literal['1', '2', 'both'] = 'both', download: bool = False) -> Tuple[List[str], List[str]]:
 99def get_hvm_paths(
100    path: Union[os.PathLike, str],
101    annotation: Literal["hepatic_veins", "portal_veins", "liver_tumor"] = "hepatic_veins",
102    center: Literal["1", "2", "both"] = "both",
103    download: bool = False,
104) -> Tuple[List[str], List[str]]:
105    """Get paths to the HVM data.
106
107    Args:
108        path: Filepath to a folder where the data is downloaded for further processing.
109        annotation: The choice of annotations. One of 'hepatic_veins', 'portal_veins' or 'liver_tumor'.
110        center: The choice of source center. One of '1', '2' or 'both'. Ignored (forced to '1') when
111            `annotation` is 'liver_tumor', as the tumor annotations shipped under 'Center 2' are an
112            identical copy of the 'Center 1' ones and do not correspond to any 'Center 2' scan.
113        download: Whether to download the data if it is not present.
114
115    Returns:
116        List of filepaths for the image data.
117        List of filepaths for the label data.
118    """
119    if annotation not in ANNOTATIONS:
120        raise ValueError(f"'{annotation}' is not a valid annotation. Choose from {list(ANNOTATIONS.keys())}.")
121    if center not in CENTERS + ["both"]:
122        raise ValueError(f"'{center}' is not a valid center. Choose from {CENTERS + ['both']}.")
123
124    data_dir = get_hvm_data(path=path, download=download)
125
126    centers = ["1"] if annotation == "liver_tumor" else (CENTERS if center == "both" else [center])
127
128    image_paths, gt_paths = [], []
129    for c in centers:
130        center_dir = os.path.join(data_dir, f"Center {c}")
131        image_dir = os.path.join(center_dir, "Image")
132        annotation_dir = os.path.join(center_dir, ANNOTATIONS[annotation])
133
134        this_images = natsorted(glob(os.path.join(image_dir, "*.nii.gz")))
135        this_annotations = natsorted(glob(os.path.join(annotation_dir, "*.nii.gz")))
136
137        matched_images, matched_annotations = _match_images_to_annotations(this_images, this_annotations)
138        image_paths.extend(matched_images)
139        gt_paths.extend(matched_annotations)
140
141    assert len(image_paths) == len(gt_paths) and len(image_paths) > 0
142
143    return image_paths, gt_paths

Get paths to the HVM data.

Arguments:
  • path: Filepath to a folder where the data is downloaded for further processing.
  • annotation: The choice of annotations. One of 'hepatic_veins', 'portal_veins' or 'liver_tumor'.
  • center: The choice of source center. One of '1', '2' or 'both'. Ignored (forced to '1') when annotation is 'liver_tumor', as the tumor annotations shipped under 'Center 2' are an identical copy of the 'Center 1' ones and do not correspond to any 'Center 2' scan.
  • download: Whether to download the data if it is not present.
Returns:

List of filepaths for the image data. List of filepaths for the label data.

def get_hvm_dataset( path: Union[os.PathLike, str], patch_shape: Tuple[int, ...], annotation: Literal['hepatic_veins', 'portal_veins', 'liver_tumor'] = 'hepatic_veins', center: Literal['1', '2', 'both'] = 'both', resize_inputs: bool = False, download: bool = False, **kwargs) -> torch.utils.data.dataset.Dataset:
146def get_hvm_dataset(
147    path: Union[os.PathLike, str],
148    patch_shape: Tuple[int, ...],
149    annotation: Literal["hepatic_veins", "portal_veins", "liver_tumor"] = "hepatic_veins",
150    center: Literal["1", "2", "both"] = "both",
151    resize_inputs: bool = False,
152    download: bool = False,
153    **kwargs
154) -> Dataset:
155    """Get the HVM dataset for hepatic vein, portal vein and liver tumor segmentation in abdominal CT.
156
157    Args:
158        path: Filepath to a folder where the data is downloaded for further processing.
159        patch_shape: The patch shape to use for training.
160        annotation: The choice of annotations. One of 'hepatic_veins', 'portal_veins' or 'liver_tumor'.
161        center: The choice of source center. One of '1', '2' or 'both'. Ignored (forced to '1') when
162            `annotation` is 'liver_tumor'.
163        resize_inputs: Whether to resize inputs to the desired patch shape.
164        download: Whether to download the data if it is not present.
165        kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset`.
166
167    Returns:
168        The segmentation dataset.
169    """
170    raw_paths, label_paths = get_hvm_paths(path, annotation, center, download)
171
172    if resize_inputs:
173        resize_kwargs = {"patch_shape": patch_shape, "is_rgb": False}
174        kwargs, patch_shape = util.update_kwargs_for_resize_trafo(
175            kwargs=kwargs, patch_shape=patch_shape, resize_inputs=resize_inputs, resize_kwargs=resize_kwargs
176        )
177
178    return torch_em.default_segmentation_dataset(
179        raw_paths=raw_paths,
180        raw_key="data",
181        label_paths=label_paths,
182        label_key="data",
183        patch_shape=patch_shape,
184        is_seg_dataset=True,
185        **kwargs
186    )

Get the HVM dataset for hepatic vein, portal vein and liver tumor segmentation in abdominal CT.

Arguments:
  • path: Filepath to a folder where the data is downloaded for further processing.
  • patch_shape: The patch shape to use for training.
  • annotation: The choice of annotations. One of 'hepatic_veins', 'portal_veins' or 'liver_tumor'.
  • center: The choice of source center. One of '1', '2' or 'both'. Ignored (forced to '1') when annotation is 'liver_tumor'.
  • resize_inputs: Whether to resize inputs to the desired patch shape.
  • download: Whether to download the data if it is not present.
  • kwargs: Additional keyword arguments for torch_em.default_segmentation_dataset.
Returns:

The segmentation dataset.

def get_hvm_loader( path: Union[os.PathLike, str], batch_size: int, patch_shape: Tuple[int, ...], annotation: Literal['hepatic_veins', 'portal_veins', 'liver_tumor'] = 'hepatic_veins', center: Literal['1', '2', 'both'] = 'both', resize_inputs: bool = False, download: bool = False, **kwargs) -> torch.utils.data.dataloader.DataLoader:
189def get_hvm_loader(
190    path: Union[os.PathLike, str],
191    batch_size: int,
192    patch_shape: Tuple[int, ...],
193    annotation: Literal["hepatic_veins", "portal_veins", "liver_tumor"] = "hepatic_veins",
194    center: Literal["1", "2", "both"] = "both",
195    resize_inputs: bool = False,
196    download: bool = False,
197    **kwargs
198) -> DataLoader:
199    """Get the HVM dataloader for hepatic vein, portal vein and liver tumor segmentation in abdominal CT.
200
201    Args:
202        path: Filepath to a folder where the data is downloaded for further processing.
203        batch_size: The batch size for training.
204        patch_shape: The patch shape to use for training.
205        annotation: The choice of annotations. One of 'hepatic_veins', 'portal_veins' or 'liver_tumor'.
206        center: The choice of source center. One of '1', '2' or 'both'. Ignored (forced to '1') when
207            `annotation` is 'liver_tumor'.
208        resize_inputs: Whether to resize inputs to the desired patch shape.
209        download: Whether to download the data if it is not present.
210        kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset` or for the PyTorch DataLoader.
211
212    Returns:
213        The DataLoader.
214    """
215    ds_kwargs, loader_kwargs = util.split_kwargs(torch_em.default_segmentation_dataset, **kwargs)
216    dataset = get_hvm_dataset(path, patch_shape, annotation, center, resize_inputs, download, **ds_kwargs)
217    return torch_em.get_data_loader(dataset, batch_size, **loader_kwargs)

Get the HVM dataloader for hepatic vein, portal vein and liver tumor segmentation in abdominal CT.

Arguments:
  • path: Filepath to a folder where the data is downloaded for further processing.
  • batch_size: The batch size for training.
  • patch_shape: The patch shape to use for training.
  • annotation: The choice of annotations. One of 'hepatic_veins', 'portal_veins' or 'liver_tumor'.
  • center: The choice of source center. One of '1', '2' or 'both'. Ignored (forced to '1') when annotation is 'liver_tumor'.
  • resize_inputs: Whether to resize inputs to the desired patch shape.
  • download: Whether to download the data if it is not present.
  • kwargs: Additional keyword arguments for torch_em.default_segmentation_dataset or for the PyTorch DataLoader.
Returns:

The DataLoader.