torch_em.data.datasets.medical.uls23
ULS23 (Universal Lesion Segmentation Challenge 2023) provides a unified 3D lesion segmentation task that combines novel lesion annotations with volumes-of-interest (VOIs) processed from several existing CT lesion datasets. This module covers the 'processed_data' part of the challenge's training data, which re-uses the source images of three existing lesion datasets and adds new 3D VOI-level lesion masks for them:
- 'kits21': 332 kidney lesion VOIs, images taken from the KiTS21 kidney tumor segmentation challenge.
- 'lits': 888 liver lesion VOIs, images taken from the LiTS liver tumor segmentation challenge.
- 'lidc-idri': 2246 lung lesion VOIs, images taken from the LIDC-IDRI lung nodule dataset.
Each VOI is a 3D CT sub-volume cropped around one lesion, with a binary lesion segmentation mask.
The novel lesion annotations of ULS23 (eg. the DeepLesion3D subset) are covered by the separate
torch_em.data.datasets.medical.deeplesion module.
The images (~44 GB, split into a multi-part zip archive) are located at https://doi.org/10.5281/zenodo.10050960 (part 2 of the ULS23 training data). The labels are located at https://github.com/DIAGNijmegen/ULS23. NOTE: The images are distributed as (nested) multi-part zip archives, so the '7z' CLI is required to extract them (install it via 'conda install -c conda-forge p7zip'). Both the images and labels are stored with a trailing singleton dimension (eg. shape (256, 256, 128, 1)), which this module squeezes in-place after extraction so that the volumes are plain 3D arrays. The data is licensed under CC BY-NC-SA 4.0.
This dataset is from the publication https://doi.org/10.1016/j.media.2025.103525. Please cite it (and the publication of the source dataset you use: KiTS21 https://doi.org/10.48550/arXiv.2307.01984, LiTS https://doi.org/10.1016/j.media.2022.102680, or LIDC-IDRI https://doi.org/10.1118/1.3528204) if you use this dataset in your research.
1"""ULS23 (Universal Lesion Segmentation Challenge 2023) provides a unified 3D lesion segmentation task 2that combines novel lesion annotations with volumes-of-interest (VOIs) processed from several existing 3CT lesion datasets. This module covers the 'processed_data' part of the challenge's training data, which 4re-uses the source images of three existing lesion datasets and adds new 3D VOI-level lesion masks for them: 5 6- 'kits21': 332 kidney lesion VOIs, images taken from the KiTS21 kidney tumor segmentation challenge. 7- 'lits': 888 liver lesion VOIs, images taken from the LiTS liver tumor segmentation challenge. 8- 'lidc-idri': 2246 lung lesion VOIs, images taken from the LIDC-IDRI lung nodule dataset. 9 10Each VOI is a 3D CT sub-volume cropped around one lesion, with a binary lesion segmentation mask. 11The novel lesion annotations of ULS23 (eg. the DeepLesion3D subset) are covered by the separate 12`torch_em.data.datasets.medical.deeplesion` module. 13 14The images (~44 GB, split into a multi-part zip archive) are located at 15https://doi.org/10.5281/zenodo.10050960 (part 2 of the ULS23 training data). 16The labels are located at https://github.com/DIAGNijmegen/ULS23. 17NOTE: The images are distributed as (nested) multi-part zip archives, so the '7z' CLI is required to extract 18them (install it via 'conda install -c conda-forge p7zip'). Both the images and labels are stored with a 19trailing singleton dimension (eg. shape (256, 256, 128, 1)), which this module squeezes in-place after 20extraction so that the volumes are plain 3D arrays. 21The data is licensed under CC BY-NC-SA 4.0. 22 23This dataset is from the publication https://doi.org/10.1016/j.media.2025.103525. 24Please cite it (and the publication of the source dataset you use: KiTS21 https://doi.org/10.48550/arXiv.2307.01984, 25LiTS https://doi.org/10.1016/j.media.2022.102680, or LIDC-IDRI https://doi.org/10.1118/1.3528204) if you use 26this dataset in your research. 27""" 28 29import os 30from glob import glob 31from tqdm import tqdm 32from shutil import which 33from subprocess import run 34from natsort import natsorted 35from typing import Union, Tuple, Literal, List 36 37import numpy as np 38 39from torch.utils.data import Dataset, DataLoader 40 41import torch_em 42 43from .. import util 44 45 46URLS = { 47 "ULS23_Part2.zip": "https://zenodo.org/records/10050960/files/ULS23_Part2.zip?download=1", 48 "ULS23_Part2.z01": "https://zenodo.org/records/10050960/files/ULS23_Part2.z01?download=1", 49 "ULS23_Part2.z02": "https://zenodo.org/records/10050960/files/ULS23_Part2.z02?download=1", 50 "ULS23_Part2.z03": "https://zenodo.org/records/10050960/files/ULS23_Part2.z03?download=1", 51 "ULS23_Part2.z04": "https://zenodo.org/records/10050960/files/ULS23_Part2.z04?download=1", 52 "ULS23_Part2.z05": "https://zenodo.org/records/10050960/files/ULS23_Part2.z05?download=1", 53 "ULS23_Part2.z06": "https://zenodo.org/records/10050960/files/ULS23_Part2.z06?download=1", 54 "ULS23_Part2.z07": "https://zenodo.org/records/10050960/files/ULS23_Part2.z07?download=1", 55 "ULS23_Part2.z08": "https://zenodo.org/records/10050960/files/ULS23_Part2.z08?download=1", 56 "ULS23_Part2.z09": "https://zenodo.org/records/10050960/files/ULS23_Part2.z09?download=1", 57 "ULS23_Part2.z10": "https://zenodo.org/records/10050960/files/ULS23_Part2.z10?download=1", 58 "annotations": "https://github.com/DIAGNijmegen/ULS23/archive/06a2bffc433418f72d04f7ecbb23b28694c81e6b.zip", 59} 60 61CHECKSUMS = { 62 "ULS23_Part2.zip": "67e9343c58b25ba871ae35aeac5db9cc89acde01d2ce64b3fa0fdf9d17b19bd5", 63 "ULS23_Part2.z01": "803b18b4ddd24175589e52dcf847000daa80c1e471bd955bb1f27624f51ace05", 64 "ULS23_Part2.z02": "e743e4fc733ae7608fb24f4057399e14aea89aef761ccc275492167b53861529", 65 "ULS23_Part2.z03": "44d3092e8fe8f1bf7c8e23fdd93b035967d0771ba6905af0e61cb294d56d6c08", 66 "ULS23_Part2.z04": "56d0a40862d82d5f4fce39baeca32714563b2c1ba0768e43d7eb9af9331c2fc6", 67 "ULS23_Part2.z05": "dc9696532cad2bd3ed6d3c31fcf4afa044a2c5b0bda804d8fd2bb647cd56d102", 68 "ULS23_Part2.z06": "f13ff8402e7e08cc0989b47f115255339061d09d628d06fc5d09fee4983f146d", 69 "ULS23_Part2.z07": "d491fd122425add97a951dd7e8edb2cc79abab94e9be9c94b6f1e70a1eaf6e68", 70 "ULS23_Part2.z08": "669811801c36514faa906feee953d5b0e970554ad7f0ca382a80996a65b45c18", 71 "ULS23_Part2.z09": "0d880b1e22686becbccbbf6debe12076a6ed5b87d4fd64651967e4b1463ef456", 72 "ULS23_Part2.z10": "efd22fcff15049cdaac5c002c9e90d15afaae62a9fe5a60fadc79f65d369a8cd", 73 "annotations": "19ae6b84aae1a94aa8329a0ce4c6586d1e7cf28532b27367d923773adc3ebe28", 74} 75 76SOURCES = { 77 "kits21": os.path.join("ULS23", "processed_data", "fully_annotated", "kits21"), 78 "lits": os.path.join("ULS23", "processed_data", "fully_annotated", "LiTS"), 79 "lidc-idri": os.path.join("ULS23", "processed_data", "fully_annotated", "LIDC-IDRI"), 80} 81"""Mapping from the source dataset name to its folder in the ULS23 archives.""" 82 83 84def _run_7z_x(archive, dst, members=None): 85 if which("7z") is None: 86 raise RuntimeError( 87 "The ULS23 images are distributed as (nested) multi-part zip archives, which require the '7z' CLI. " 88 "You can install it via 'conda install -c conda-forge p7zip'." 89 ) 90 cmd = ["7z", "x", f"-o{dst}", "-y", archive] 91 if members is not None: 92 cmd += members 93 run(cmd, check=True) 94 95 96def _squeeze_trailing_singleton_dims(root): 97 """The ULS23 VOIs are stored with a trailing singleton dimension (eg. shape (256, 256, 128, 1)), which 98 is squeezed here in-place so that the volumes are plain 3D arrays.""" 99 import nibabel as nib 100 101 for dirpath, _, filenames in os.walk(root): 102 for fname in filenames: 103 if not fname.endswith(".nii.gz"): 104 continue 105 fpath = os.path.join(dirpath, fname) 106 nifti = nib.load(fpath) 107 if nifti.shape[-1] != 1 or len(nifti.shape) <= 3: 108 continue 109 data = np.squeeze(np.asarray(nifti.dataobj), axis=-1) 110 nib.save(nib.Nifti1Image(data, nifti.affine), fpath) 111 112 113def _index_by_basename(root): 114 """Build a mapping from filename to filepath for all '*.nii.gz' files under 'root' (searched recursively, 115 since the inner per-source archives do not consistently nest their images under the same sub-folder).""" 116 index = {} 117 for dirpath, _, filenames in os.walk(root): 118 for fname in filenames: 119 if fname.endswith(".nii.gz"): 120 index[fname] = os.path.join(dirpath, fname) 121 return index 122 123 124def _extract_uls23_images(path, source, download): 125 image_dir = os.path.join(path, SOURCES[source], "images") 126 if os.path.exists(image_dir) and _index_by_basename(image_dir): 127 return image_dir 128 129 outer_zip = os.path.join(path, "ULS23_Part2.zip") 130 for name, url in URLS.items(): 131 if not name.startswith("ULS23_Part2"): 132 continue 133 util.download_source(path=os.path.join(path, name), url=url, download=download, checksum=CHECKSUMS[name]) 134 135 inner_zip = os.path.join(SOURCES[source], "images.zip") 136 _run_7z_x(outer_zip, path, members=[inner_zip + "*"]) 137 138 inner_zip_path = os.path.join(path, inner_zip) 139 _run_7z_x(inner_zip_path, image_dir) 140 141 # Remove the (potentially multi-part) inner zip archive, but not the extracted 'images.zip' folder itself. 142 inner_zip_stem = inner_zip_path[:-len(".zip")] 143 for part in glob(inner_zip_stem + ".z[0-9][0-9]"): 144 os.remove(part) 145 os.remove(inner_zip_path) 146 147 _squeeze_trailing_singleton_dims(image_dir) 148 149 return image_dir 150 151 152def _extract_uls23_labels(path, source, download): 153 label_dir = os.path.join(path, SOURCES[source], "labels") 154 if os.path.exists(label_dir) and glob(os.path.join(label_dir, "*.nii.gz")): 155 return label_dir 156 157 annotation_zip = os.path.join(path, "ULS23_annotations.zip") 158 util.download_source( 159 path=annotation_zip, url=URLS["annotations"], download=download, checksum=CHECKSUMS["annotations"] 160 ) 161 extracted_dir = os.path.join(path, "ULS23_annotations") 162 if not os.path.exists(extracted_dir): 163 util.unzip(zip_path=annotation_zip, dst=extracted_dir, remove=False) 164 165 src_label_dir = glob(os.path.join(extracted_dir, "*", "annotations", SOURCES[source], "labels"))[0] 166 for label_zip in tqdm(natsorted(glob(os.path.join(src_label_dir, "*.zip"))), desc=f"Preparing {source} labels"): 167 util.unzip(zip_path=label_zip, dst=label_dir, remove=False) 168 169 _squeeze_trailing_singleton_dims(label_dir) 170 171 return label_dir 172 173 174def get_uls23_data( 175 path: Union[os.PathLike, str], source: Literal["kits21", "lits", "lidc-idri"], download: bool = False 176) -> Tuple[str, str]: 177 """Download the ULS23 processed-data images and labels for one source dataset. 178 179 Args: 180 path: Filepath to a folder where the data is downloaded for further processing. 181 source: The source dataset to fetch the VOIs for. One of 'kits21', 'lits', 'lidc-idri'. 182 download: Whether to download the data if it is not present. 183 184 Returns: 185 Filepath to the folder with the image data. 186 Filepath to the folder with the label data. 187 """ 188 if source not in SOURCES: 189 raise ValueError(f"'{source}' is not a supported source. Please choose one of {list(SOURCES.keys())}.") 190 191 os.makedirs(path, exist_ok=True) 192 image_dir = _extract_uls23_images(path, source, download) 193 label_dir = _extract_uls23_labels(path, source, download) 194 195 return image_dir, label_dir 196 197 198def get_uls23_paths( 199 path: Union[os.PathLike, str], source: Literal["kits21", "lits", "lidc-idri"], download: bool = False 200) -> Tuple[List[str], List[str]]: 201 """Get paths to the ULS23 processed-data VOIs for one source dataset. 202 203 Args: 204 path: Filepath to a folder where the data is downloaded for further processing. 205 source: The source dataset to fetch the VOIs for. One of 'kits21', 'lits', 'lidc-idri'. 206 download: Whether to download the data if it is not present. 207 208 Returns: 209 List of filepaths for the image data. 210 List of filepaths for the label data. 211 """ 212 image_dir, label_dir = get_uls23_data(path, source, download) 213 image_index = _index_by_basename(image_dir) 214 215 raw_paths, label_paths = [], [] 216 for label_path in natsorted(glob(os.path.join(label_dir, "*.nii.gz"))): 217 fname = os.path.basename(label_path) 218 raw_path = image_index.get(fname) 219 if raw_path is None: 220 continue 221 raw_paths.append(raw_path) 222 label_paths.append(label_path) 223 224 if len(raw_paths) == 0 or len(raw_paths) != len(label_paths): 225 raise RuntimeError("Something went wrong with fetching the image and label paths.") 226 227 return raw_paths, label_paths 228 229 230def get_uls23_dataset( 231 path: Union[os.PathLike, str], 232 patch_shape: Tuple[int, ...], 233 source: Literal["kits21", "lits", "lidc-idri"], 234 resize_inputs: bool = False, 235 download: bool = False, 236 **kwargs 237) -> Dataset: 238 """Get the ULS23 dataset for universal lesion segmentation. 239 240 Args: 241 path: Filepath to a folder where the data is downloaded for further processing. 242 patch_shape: The patch shape to use for training. 243 source: The source dataset to fetch the VOIs for. One of 'kits21', 'lits', 'lidc-idri'. 244 resize_inputs: Whether to resize inputs to the desired patch shape. 245 download: Whether to download the data if it is not present. 246 kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset`. 247 248 Returns: 249 The segmentation dataset. 250 """ 251 raw_paths, label_paths = get_uls23_paths(path, source, download) 252 253 if resize_inputs: 254 resize_kwargs = {"patch_shape": patch_shape, "is_rgb": False} 255 kwargs, patch_shape = util.update_kwargs_for_resize_trafo( 256 kwargs=kwargs, patch_shape=patch_shape, resize_inputs=resize_inputs, resize_kwargs=resize_kwargs 257 ) 258 259 return torch_em.default_segmentation_dataset( 260 raw_paths=raw_paths, 261 raw_key="data", 262 label_paths=label_paths, 263 label_key="data", 264 patch_shape=patch_shape, 265 is_seg_dataset=True, 266 **kwargs 267 ) 268 269 270def get_uls23_loader( 271 path: Union[os.PathLike, str], 272 batch_size: int, 273 patch_shape: Tuple[int, ...], 274 source: Literal["kits21", "lits", "lidc-idri"], 275 resize_inputs: bool = False, 276 download: bool = False, 277 **kwargs 278) -> DataLoader: 279 """Get the ULS23 dataloader for universal lesion segmentation. 280 281 Args: 282 path: Filepath to a folder where the data is downloaded for further processing. 283 batch_size: The batch size for training. 284 patch_shape: The patch shape to use for training. 285 source: The source dataset to fetch the VOIs for. One of 'kits21', 'lits', 'lidc-idri'. 286 resize_inputs: Whether to resize inputs to the desired patch shape. 287 download: Whether to download the data if it is not present. 288 kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset` or for the PyTorch DataLoader. 289 290 Returns: 291 The DataLoader. 292 """ 293 ds_kwargs, loader_kwargs = util.split_kwargs(torch_em.default_segmentation_dataset, **kwargs) 294 dataset = get_uls23_dataset(path, patch_shape, source, resize_inputs, download, **ds_kwargs) 295 return torch_em.get_data_loader(dataset, batch_size, **loader_kwargs)
Mapping from the source dataset name to its folder in the ULS23 archives.
175def get_uls23_data( 176 path: Union[os.PathLike, str], source: Literal["kits21", "lits", "lidc-idri"], download: bool = False 177) -> Tuple[str, str]: 178 """Download the ULS23 processed-data images and labels for one source dataset. 179 180 Args: 181 path: Filepath to a folder where the data is downloaded for further processing. 182 source: The source dataset to fetch the VOIs for. One of 'kits21', 'lits', 'lidc-idri'. 183 download: Whether to download the data if it is not present. 184 185 Returns: 186 Filepath to the folder with the image data. 187 Filepath to the folder with the label data. 188 """ 189 if source not in SOURCES: 190 raise ValueError(f"'{source}' is not a supported source. Please choose one of {list(SOURCES.keys())}.") 191 192 os.makedirs(path, exist_ok=True) 193 image_dir = _extract_uls23_images(path, source, download) 194 label_dir = _extract_uls23_labels(path, source, download) 195 196 return image_dir, label_dir
Download the ULS23 processed-data images and labels for one source dataset.
Arguments:
- path: Filepath to a folder where the data is downloaded for further processing.
- source: The source dataset to fetch the VOIs for. One of 'kits21', 'lits', 'lidc-idri'.
- download: Whether to download the data if it is not present.
Returns:
Filepath to the folder with the image data. Filepath to the folder with the label data.
199def get_uls23_paths( 200 path: Union[os.PathLike, str], source: Literal["kits21", "lits", "lidc-idri"], download: bool = False 201) -> Tuple[List[str], List[str]]: 202 """Get paths to the ULS23 processed-data VOIs for one source dataset. 203 204 Args: 205 path: Filepath to a folder where the data is downloaded for further processing. 206 source: The source dataset to fetch the VOIs for. One of 'kits21', 'lits', 'lidc-idri'. 207 download: Whether to download the data if it is not present. 208 209 Returns: 210 List of filepaths for the image data. 211 List of filepaths for the label data. 212 """ 213 image_dir, label_dir = get_uls23_data(path, source, download) 214 image_index = _index_by_basename(image_dir) 215 216 raw_paths, label_paths = [], [] 217 for label_path in natsorted(glob(os.path.join(label_dir, "*.nii.gz"))): 218 fname = os.path.basename(label_path) 219 raw_path = image_index.get(fname) 220 if raw_path is None: 221 continue 222 raw_paths.append(raw_path) 223 label_paths.append(label_path) 224 225 if len(raw_paths) == 0 or len(raw_paths) != len(label_paths): 226 raise RuntimeError("Something went wrong with fetching the image and label paths.") 227 228 return raw_paths, label_paths
Get paths to the ULS23 processed-data VOIs for one source dataset.
Arguments:
- path: Filepath to a folder where the data is downloaded for further processing.
- source: The source dataset to fetch the VOIs for. One of 'kits21', 'lits', 'lidc-idri'.
- download: Whether to download the data if it is not present.
Returns:
List of filepaths for the image data. List of filepaths for the label data.
231def get_uls23_dataset( 232 path: Union[os.PathLike, str], 233 patch_shape: Tuple[int, ...], 234 source: Literal["kits21", "lits", "lidc-idri"], 235 resize_inputs: bool = False, 236 download: bool = False, 237 **kwargs 238) -> Dataset: 239 """Get the ULS23 dataset for universal lesion segmentation. 240 241 Args: 242 path: Filepath to a folder where the data is downloaded for further processing. 243 patch_shape: The patch shape to use for training. 244 source: The source dataset to fetch the VOIs for. One of 'kits21', 'lits', 'lidc-idri'. 245 resize_inputs: Whether to resize inputs to the desired patch shape. 246 download: Whether to download the data if it is not present. 247 kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset`. 248 249 Returns: 250 The segmentation dataset. 251 """ 252 raw_paths, label_paths = get_uls23_paths(path, source, download) 253 254 if resize_inputs: 255 resize_kwargs = {"patch_shape": patch_shape, "is_rgb": False} 256 kwargs, patch_shape = util.update_kwargs_for_resize_trafo( 257 kwargs=kwargs, patch_shape=patch_shape, resize_inputs=resize_inputs, resize_kwargs=resize_kwargs 258 ) 259 260 return torch_em.default_segmentation_dataset( 261 raw_paths=raw_paths, 262 raw_key="data", 263 label_paths=label_paths, 264 label_key="data", 265 patch_shape=patch_shape, 266 is_seg_dataset=True, 267 **kwargs 268 )
Get the ULS23 dataset for universal lesion segmentation.
Arguments:
- path: Filepath to a folder where the data is downloaded for further processing.
- patch_shape: The patch shape to use for training.
- source: The source dataset to fetch the VOIs for. One of 'kits21', 'lits', 'lidc-idri'.
- resize_inputs: Whether to resize inputs to the desired patch shape.
- download: Whether to download the data if it is not present.
- kwargs: Additional keyword arguments for
torch_em.default_segmentation_dataset.
Returns:
The segmentation dataset.
271def get_uls23_loader( 272 path: Union[os.PathLike, str], 273 batch_size: int, 274 patch_shape: Tuple[int, ...], 275 source: Literal["kits21", "lits", "lidc-idri"], 276 resize_inputs: bool = False, 277 download: bool = False, 278 **kwargs 279) -> DataLoader: 280 """Get the ULS23 dataloader for universal lesion segmentation. 281 282 Args: 283 path: Filepath to a folder where the data is downloaded for further processing. 284 batch_size: The batch size for training. 285 patch_shape: The patch shape to use for training. 286 source: The source dataset to fetch the VOIs for. One of 'kits21', 'lits', 'lidc-idri'. 287 resize_inputs: Whether to resize inputs to the desired patch shape. 288 download: Whether to download the data if it is not present. 289 kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset` or for the PyTorch DataLoader. 290 291 Returns: 292 The DataLoader. 293 """ 294 ds_kwargs, loader_kwargs = util.split_kwargs(torch_em.default_segmentation_dataset, **kwargs) 295 dataset = get_uls23_dataset(path, patch_shape, source, resize_inputs, download, **ds_kwargs) 296 return torch_em.get_data_loader(dataset, batch_size, **loader_kwargs)
Get the ULS23 dataloader for universal lesion segmentation.
Arguments:
- path: Filepath to a folder where the data is downloaded for further processing.
- batch_size: The batch size for training.
- patch_shape: The patch shape to use for training.
- source: The source dataset to fetch the VOIs for. One of 'kits21', 'lits', 'lidc-idri'.
- resize_inputs: Whether to resize inputs to the desired patch shape.
- download: Whether to download the data if it is not present.
- kwargs: Additional keyword arguments for
torch_em.default_segmentation_datasetor for the PyTorch DataLoader.
Returns:
The DataLoader.