torch_em.data.datasets.medical.hvm
HVM (Hepatic Vessel Map) is a dual-center dataset for the segmentation of hepatic veins, portal veins (to third-order branches) and liver tumors in contrast-enhanced abdominal CT.
The dataset comprises 282 patients: 170 scans from Center 1 (The University of Hong Kong-Shenzhen Hospital, of which 106 have hepatic and portal vein annotations) and 176 scans from Center 2 (Peking University Shenzhen Hospital, of which 174 have hepatic and portal vein annotations). Liver tumor annotations (96 cases) are only provided for a subset of Center 1 (the archive also ships an identical copy of these files under Center 2, which does not correspond to any Center 2 scan and is therefore not used by this loader).
NOTE: The raw scans and the annotations in each folder are not matched by their filename id (e.g. 'Image/1.nii.gz' does not necessarily correspond to 'Annotation_Hepatic veins/001.nii.gz'). They are instead matched here by comparing the NIfTI header geometry (shape, voxel spacing and origin), which uniquely identifies the corresponding scan for (almost) every annotation.
This dataset is NOT the same as the 'HVA-CT' dataset (see hva_ct.py), which re-annotates the 61
scans of the Medical Segmentation Decathlon hepatic vessel task.
The dataset is located at https://doi.org/10.5281/zenodo.19885789 and is distributed under the CC BY 4.0 license. The dataset is from the publication https://doi.org/10.1038/s41597-026-07550-3. Please cite it if you use this dataset for your research.
1"""HVM (Hepatic Vessel Map) is a dual-center dataset for the segmentation of hepatic veins, portal 2veins (to third-order branches) and liver tumors in contrast-enhanced abdominal CT. 3 4The dataset comprises 282 patients: 170 scans from Center 1 (The University of Hong Kong-Shenzhen 5Hospital, of which 106 have hepatic and portal vein annotations) and 176 scans from Center 2 6(Peking University Shenzhen Hospital, of which 174 have hepatic and portal vein annotations). Liver 7tumor annotations (96 cases) are only provided for a subset of Center 1 (the archive also ships an 8identical copy of these files under Center 2, which does not correspond to any Center 2 scan and is 9therefore not used by this loader). 10 11NOTE: The raw scans and the annotations in each folder are not matched by their filename id (e.g. 12'Image/1.nii.gz' does not necessarily correspond to 'Annotation_Hepatic veins/001.nii.gz'). They are 13instead matched here by comparing the NIfTI header geometry (shape, voxel spacing and origin), which 14uniquely identifies the corresponding scan for (almost) every annotation. 15 16This dataset is NOT the same as the 'HVA-CT' dataset (see `hva_ct.py`), which re-annotates the 61 17scans of the Medical Segmentation Decathlon hepatic vessel task. 18 19The dataset is located at https://doi.org/10.5281/zenodo.19885789 and is distributed under the 20CC BY 4.0 license. The dataset is from the publication https://doi.org/10.1038/s41597-026-07550-3. 21Please cite it if you use this dataset for your research. 22""" 23 24import os 25import warnings 26from glob import glob 27from natsort import natsorted 28from typing import Union, Tuple, Literal, List 29 30from torch.utils.data import Dataset, DataLoader 31 32import torch_em 33 34from .. import util 35 36 37URL = "https://zenodo.org/api/records/19885789/files/HVM%20Dataset.zip/content" 38CHECKSUM = "7fe2a00bc8a40658d45bedd27cd70df22e1a58a92af20037c22b05065407a6b8" 39 40CENTERS = ["1", "2"] 41 42ANNOTATIONS = { 43 "hepatic_veins": "Annotation_Hepatic veins", 44 "portal_veins": "Annotation_Portal veins", 45 "liver_tumor": "Annotation_Liver tumors", 46} 47 48 49def get_hvm_data(path: Union[os.PathLike, str], download: bool = False) -> str: 50 """Download the HVM dataset. 51 52 Args: 53 path: Filepath to a folder where the data is downloaded for further processing. 54 download: Whether to download the data if it is not present. 55 56 Returns: 57 Filepath where the data is downloaded. 58 """ 59 data_dir = os.path.join(path, "HVM Dataset") 60 if os.path.exists(data_dir): 61 return data_dir 62 63 os.makedirs(path, exist_ok=True) 64 65 zip_path = os.path.join(path, "HVM_Dataset.zip") 66 util.download_source(path=zip_path, url=URL, download=download, checksum=CHECKSUM) 67 util.unzip(zip_path=zip_path, dst=path) 68 69 return data_dir 70 71 72def _header_signature(path): 73 import nibabel as nib 74 75 header = nib.load(path).header 76 shape = header.get_data_shape() 77 zooms = tuple(header.get_zooms()) 78 origin = tuple(header.get_best_affine()[:3, 3]) 79 return shape, zooms, origin 80 81 82def _match_images_to_annotations(image_paths, annotation_paths): 83 image_sigs = {p: _header_signature(p) for p in image_paths} 84 85 matched_images, matched_annotations = [], [] 86 for annotation_path in annotation_paths: 87 annotation_sig = _header_signature(annotation_path) 88 candidates = [p for p, sig in image_sigs.items() if sig == annotation_sig] 89 if not candidates: 90 warnings.warn(f"Could not find a matching scan for the annotation at '{annotation_path}'. Skipping it.") 91 continue 92 matched_images.append(natsorted(candidates)[0]) 93 matched_annotations.append(annotation_path) 94 95 return matched_images, matched_annotations 96 97 98def get_hvm_paths( 99 path: Union[os.PathLike, str], 100 annotation: Literal["hepatic_veins", "portal_veins", "liver_tumor"] = "hepatic_veins", 101 center: Literal["1", "2", "both"] = "both", 102 download: bool = False, 103) -> Tuple[List[str], List[str]]: 104 """Get paths to the HVM data. 105 106 Args: 107 path: Filepath to a folder where the data is downloaded for further processing. 108 annotation: The choice of annotations. One of 'hepatic_veins', 'portal_veins' or 'liver_tumor'. 109 center: The choice of source center. One of '1', '2' or 'both'. Ignored (forced to '1') when 110 `annotation` is 'liver_tumor', as the tumor annotations shipped under 'Center 2' are an 111 identical copy of the 'Center 1' ones and do not correspond to any 'Center 2' scan. 112 download: Whether to download the data if it is not present. 113 114 Returns: 115 List of filepaths for the image data. 116 List of filepaths for the label data. 117 """ 118 if annotation not in ANNOTATIONS: 119 raise ValueError(f"'{annotation}' is not a valid annotation. Choose from {list(ANNOTATIONS.keys())}.") 120 if center not in CENTERS + ["both"]: 121 raise ValueError(f"'{center}' is not a valid center. Choose from {CENTERS + ['both']}.") 122 123 data_dir = get_hvm_data(path=path, download=download) 124 125 centers = ["1"] if annotation == "liver_tumor" else (CENTERS if center == "both" else [center]) 126 127 image_paths, gt_paths = [], [] 128 for c in centers: 129 center_dir = os.path.join(data_dir, f"Center {c}") 130 image_dir = os.path.join(center_dir, "Image") 131 annotation_dir = os.path.join(center_dir, ANNOTATIONS[annotation]) 132 133 this_images = natsorted(glob(os.path.join(image_dir, "*.nii.gz"))) 134 this_annotations = natsorted(glob(os.path.join(annotation_dir, "*.nii.gz"))) 135 136 matched_images, matched_annotations = _match_images_to_annotations(this_images, this_annotations) 137 image_paths.extend(matched_images) 138 gt_paths.extend(matched_annotations) 139 140 assert len(image_paths) == len(gt_paths) and len(image_paths) > 0 141 142 return image_paths, gt_paths 143 144 145def get_hvm_dataset( 146 path: Union[os.PathLike, str], 147 patch_shape: Tuple[int, ...], 148 annotation: Literal["hepatic_veins", "portal_veins", "liver_tumor"] = "hepatic_veins", 149 center: Literal["1", "2", "both"] = "both", 150 resize_inputs: bool = False, 151 download: bool = False, 152 **kwargs 153) -> Dataset: 154 """Get the HVM dataset for hepatic vein, portal vein and liver tumor segmentation in abdominal CT. 155 156 Args: 157 path: Filepath to a folder where the data is downloaded for further processing. 158 patch_shape: The patch shape to use for training. 159 annotation: The choice of annotations. One of 'hepatic_veins', 'portal_veins' or 'liver_tumor'. 160 center: The choice of source center. One of '1', '2' or 'both'. Ignored (forced to '1') when 161 `annotation` is 'liver_tumor'. 162 resize_inputs: Whether to resize inputs to the desired patch shape. 163 download: Whether to download the data if it is not present. 164 kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset`. 165 166 Returns: 167 The segmentation dataset. 168 """ 169 raw_paths, label_paths = get_hvm_paths(path, annotation, center, download) 170 171 if resize_inputs: 172 resize_kwargs = {"patch_shape": patch_shape, "is_rgb": False} 173 kwargs, patch_shape = util.update_kwargs_for_resize_trafo( 174 kwargs=kwargs, patch_shape=patch_shape, resize_inputs=resize_inputs, resize_kwargs=resize_kwargs 175 ) 176 177 return torch_em.default_segmentation_dataset( 178 raw_paths=raw_paths, 179 raw_key="data", 180 label_paths=label_paths, 181 label_key="data", 182 patch_shape=patch_shape, 183 is_seg_dataset=True, 184 **kwargs 185 ) 186 187 188def get_hvm_loader( 189 path: Union[os.PathLike, str], 190 batch_size: int, 191 patch_shape: Tuple[int, ...], 192 annotation: Literal["hepatic_veins", "portal_veins", "liver_tumor"] = "hepatic_veins", 193 center: Literal["1", "2", "both"] = "both", 194 resize_inputs: bool = False, 195 download: bool = False, 196 **kwargs 197) -> DataLoader: 198 """Get the HVM dataloader for hepatic vein, portal vein and liver tumor segmentation in abdominal CT. 199 200 Args: 201 path: Filepath to a folder where the data is downloaded for further processing. 202 batch_size: The batch size for training. 203 patch_shape: The patch shape to use for training. 204 annotation: The choice of annotations. One of 'hepatic_veins', 'portal_veins' or 'liver_tumor'. 205 center: The choice of source center. One of '1', '2' or 'both'. Ignored (forced to '1') when 206 `annotation` is 'liver_tumor'. 207 resize_inputs: Whether to resize inputs to the desired patch shape. 208 download: Whether to download the data if it is not present. 209 kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset` or for the PyTorch DataLoader. 210 211 Returns: 212 The DataLoader. 213 """ 214 ds_kwargs, loader_kwargs = util.split_kwargs(torch_em.default_segmentation_dataset, **kwargs) 215 dataset = get_hvm_dataset(path, patch_shape, annotation, center, resize_inputs, download, **ds_kwargs) 216 return torch_em.get_data_loader(dataset, batch_size, **loader_kwargs)
50def get_hvm_data(path: Union[os.PathLike, str], download: bool = False) -> str: 51 """Download the HVM dataset. 52 53 Args: 54 path: Filepath to a folder where the data is downloaded for further processing. 55 download: Whether to download the data if it is not present. 56 57 Returns: 58 Filepath where the data is downloaded. 59 """ 60 data_dir = os.path.join(path, "HVM Dataset") 61 if os.path.exists(data_dir): 62 return data_dir 63 64 os.makedirs(path, exist_ok=True) 65 66 zip_path = os.path.join(path, "HVM_Dataset.zip") 67 util.download_source(path=zip_path, url=URL, download=download, checksum=CHECKSUM) 68 util.unzip(zip_path=zip_path, dst=path) 69 70 return data_dir
Download the HVM dataset.
Arguments:
- path: Filepath to a folder where the data is downloaded for further processing.
- download: Whether to download the data if it is not present.
Returns:
Filepath where the data is downloaded.
99def get_hvm_paths( 100 path: Union[os.PathLike, str], 101 annotation: Literal["hepatic_veins", "portal_veins", "liver_tumor"] = "hepatic_veins", 102 center: Literal["1", "2", "both"] = "both", 103 download: bool = False, 104) -> Tuple[List[str], List[str]]: 105 """Get paths to the HVM data. 106 107 Args: 108 path: Filepath to a folder where the data is downloaded for further processing. 109 annotation: The choice of annotations. One of 'hepatic_veins', 'portal_veins' or 'liver_tumor'. 110 center: The choice of source center. One of '1', '2' or 'both'. Ignored (forced to '1') when 111 `annotation` is 'liver_tumor', as the tumor annotations shipped under 'Center 2' are an 112 identical copy of the 'Center 1' ones and do not correspond to any 'Center 2' scan. 113 download: Whether to download the data if it is not present. 114 115 Returns: 116 List of filepaths for the image data. 117 List of filepaths for the label data. 118 """ 119 if annotation not in ANNOTATIONS: 120 raise ValueError(f"'{annotation}' is not a valid annotation. Choose from {list(ANNOTATIONS.keys())}.") 121 if center not in CENTERS + ["both"]: 122 raise ValueError(f"'{center}' is not a valid center. Choose from {CENTERS + ['both']}.") 123 124 data_dir = get_hvm_data(path=path, download=download) 125 126 centers = ["1"] if annotation == "liver_tumor" else (CENTERS if center == "both" else [center]) 127 128 image_paths, gt_paths = [], [] 129 for c in centers: 130 center_dir = os.path.join(data_dir, f"Center {c}") 131 image_dir = os.path.join(center_dir, "Image") 132 annotation_dir = os.path.join(center_dir, ANNOTATIONS[annotation]) 133 134 this_images = natsorted(glob(os.path.join(image_dir, "*.nii.gz"))) 135 this_annotations = natsorted(glob(os.path.join(annotation_dir, "*.nii.gz"))) 136 137 matched_images, matched_annotations = _match_images_to_annotations(this_images, this_annotations) 138 image_paths.extend(matched_images) 139 gt_paths.extend(matched_annotations) 140 141 assert len(image_paths) == len(gt_paths) and len(image_paths) > 0 142 143 return image_paths, gt_paths
Get paths to the HVM data.
Arguments:
- path: Filepath to a folder where the data is downloaded for further processing.
- annotation: The choice of annotations. One of 'hepatic_veins', 'portal_veins' or 'liver_tumor'.
- center: The choice of source center. One of '1', '2' or 'both'. Ignored (forced to '1') when
annotationis 'liver_tumor', as the tumor annotations shipped under 'Center 2' are an identical copy of the 'Center 1' ones and do not correspond to any 'Center 2' scan. - download: Whether to download the data if it is not present.
Returns:
List of filepaths for the image data. List of filepaths for the label data.
146def get_hvm_dataset( 147 path: Union[os.PathLike, str], 148 patch_shape: Tuple[int, ...], 149 annotation: Literal["hepatic_veins", "portal_veins", "liver_tumor"] = "hepatic_veins", 150 center: Literal["1", "2", "both"] = "both", 151 resize_inputs: bool = False, 152 download: bool = False, 153 **kwargs 154) -> Dataset: 155 """Get the HVM dataset for hepatic vein, portal vein and liver tumor segmentation in abdominal CT. 156 157 Args: 158 path: Filepath to a folder where the data is downloaded for further processing. 159 patch_shape: The patch shape to use for training. 160 annotation: The choice of annotations. One of 'hepatic_veins', 'portal_veins' or 'liver_tumor'. 161 center: The choice of source center. One of '1', '2' or 'both'. Ignored (forced to '1') when 162 `annotation` is 'liver_tumor'. 163 resize_inputs: Whether to resize inputs to the desired patch shape. 164 download: Whether to download the data if it is not present. 165 kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset`. 166 167 Returns: 168 The segmentation dataset. 169 """ 170 raw_paths, label_paths = get_hvm_paths(path, annotation, center, download) 171 172 if resize_inputs: 173 resize_kwargs = {"patch_shape": patch_shape, "is_rgb": False} 174 kwargs, patch_shape = util.update_kwargs_for_resize_trafo( 175 kwargs=kwargs, patch_shape=patch_shape, resize_inputs=resize_inputs, resize_kwargs=resize_kwargs 176 ) 177 178 return torch_em.default_segmentation_dataset( 179 raw_paths=raw_paths, 180 raw_key="data", 181 label_paths=label_paths, 182 label_key="data", 183 patch_shape=patch_shape, 184 is_seg_dataset=True, 185 **kwargs 186 )
Get the HVM dataset for hepatic vein, portal vein and liver tumor segmentation in abdominal CT.
Arguments:
- path: Filepath to a folder where the data is downloaded for further processing.
- patch_shape: The patch shape to use for training.
- annotation: The choice of annotations. One of 'hepatic_veins', 'portal_veins' or 'liver_tumor'.
- center: The choice of source center. One of '1', '2' or 'both'. Ignored (forced to '1') when
annotationis 'liver_tumor'. - resize_inputs: Whether to resize inputs to the desired patch shape.
- download: Whether to download the data if it is not present.
- kwargs: Additional keyword arguments for
torch_em.default_segmentation_dataset.
Returns:
The segmentation dataset.
189def get_hvm_loader( 190 path: Union[os.PathLike, str], 191 batch_size: int, 192 patch_shape: Tuple[int, ...], 193 annotation: Literal["hepatic_veins", "portal_veins", "liver_tumor"] = "hepatic_veins", 194 center: Literal["1", "2", "both"] = "both", 195 resize_inputs: bool = False, 196 download: bool = False, 197 **kwargs 198) -> DataLoader: 199 """Get the HVM dataloader for hepatic vein, portal vein and liver tumor segmentation in abdominal CT. 200 201 Args: 202 path: Filepath to a folder where the data is downloaded for further processing. 203 batch_size: The batch size for training. 204 patch_shape: The patch shape to use for training. 205 annotation: The choice of annotations. One of 'hepatic_veins', 'portal_veins' or 'liver_tumor'. 206 center: The choice of source center. One of '1', '2' or 'both'. Ignored (forced to '1') when 207 `annotation` is 'liver_tumor'. 208 resize_inputs: Whether to resize inputs to the desired patch shape. 209 download: Whether to download the data if it is not present. 210 kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset` or for the PyTorch DataLoader. 211 212 Returns: 213 The DataLoader. 214 """ 215 ds_kwargs, loader_kwargs = util.split_kwargs(torch_em.default_segmentation_dataset, **kwargs) 216 dataset = get_hvm_dataset(path, patch_shape, annotation, center, resize_inputs, download, **ds_kwargs) 217 return torch_em.get_data_loader(dataset, batch_size, **loader_kwargs)
Get the HVM dataloader for hepatic vein, portal vein and liver tumor segmentation in abdominal CT.
Arguments:
- path: Filepath to a folder where the data is downloaded for further processing.
- batch_size: The batch size for training.
- patch_shape: The patch shape to use for training.
- annotation: The choice of annotations. One of 'hepatic_veins', 'portal_veins' or 'liver_tumor'.
- center: The choice of source center. One of '1', '2' or 'both'. Ignored (forced to '1') when
annotationis 'liver_tumor'. - resize_inputs: Whether to resize inputs to the desired patch shape.
- download: Whether to download the data if it is not present.
- kwargs: Additional keyword arguments for
torch_em.default_segmentation_datasetor for the PyTorch DataLoader.
Returns:
The DataLoader.