torch_em.data.datasets.medical.airrc
AirRC contains annotations for pulmonary artery, vein and airway (lumen and wall) segmentation in chest CT.
The dataset consists of 3D masks for 254 CT scans, manually corrected and multi-expert verified. The masks are distributed on figshare, while the raw CT volumes are the original scans of the LUNA16 dataset (https://luna16.grand-challenge.org/, itself a re-release of a subset of LIDC-IDRI), which are downloaded from Zenodo and matched to the AirRC masks by their DICOM Series Instance UID.
NOTE: The AirRC masks are semantic label volumes with 4 foreground classes, following
LABEL_IDS: 1 = pulmonary artery, 2 = pulmonary vein, 3 = airway lumen, 4 = airway wall. The
masks are on their own grid (a resampled, isotropic 1mm crop around the lungs), which does not
match the grid of the original LUNA16 CT. This module therefore resamples the matching LUNA16
CT volume onto the mask grid (trilinear interpolation) and stores the pair in a single hdf5
file per case.
The AirRC masks are located at https://doi.org/10.6084/m9.figshare.26878867 (figshare, CC BY 4.0). The LUNA16 CT volumes are located at https://doi.org/10.5281/zenodo.3723295 (subsets 0-6) and https://doi.org/10.5281/zenodo.4121926 (subsets 7-9), both public domain / CC BY 3.0.
This dataset is from the publication https://doi.org/10.1038/s41597-025-06074-6. The LUNA16 CT volumes are from https://doi.org/10.1016/j.media.2017.06.015 (via LIDC-IDRI, https://doi.org/10.1118/1.3528204). Please cite them if you use this dataset in your research.
1"""AirRC contains annotations for pulmonary artery, vein and airway (lumen and wall) 2segmentation in chest CT. 3 4The dataset consists of 3D masks for 254 CT scans, manually corrected and multi-expert 5verified. The masks are distributed on figshare, while the raw CT volumes are the original 6scans of the LUNA16 dataset (https://luna16.grand-challenge.org/, itself a re-release of a 7subset of LIDC-IDRI), which are downloaded from Zenodo and matched to the AirRC masks by their 8DICOM Series Instance UID. 9 10NOTE: The AirRC masks are semantic label volumes with 4 foreground classes, following 11`LABEL_IDS`: 1 = pulmonary artery, 2 = pulmonary vein, 3 = airway lumen, 4 = airway wall. The 12masks are on their own grid (a resampled, isotropic 1mm crop around the lungs), which does not 13match the grid of the original LUNA16 CT. This module therefore resamples the matching LUNA16 14CT volume onto the mask grid (trilinear interpolation) and stores the pair in a single hdf5 15file per case. 16 17The AirRC masks are located at https://doi.org/10.6084/m9.figshare.26878867 (figshare, CC BY 184.0). The LUNA16 CT volumes are located at https://doi.org/10.5281/zenodo.3723295 (subsets 0-6) 19and https://doi.org/10.5281/zenodo.4121926 (subsets 7-9), both public domain / CC BY 3.0. 20 21This dataset is from the publication https://doi.org/10.1038/s41597-025-06074-6. The LUNA16 CT 22volumes are from https://doi.org/10.1016/j.media.2017.06.015 (via LIDC-IDRI, 23https://doi.org/10.1118/1.3528204). Please cite them if you use this dataset in your research. 24""" 25 26import os 27from glob import glob 28from tqdm import tqdm 29from natsort import natsorted 30from typing import Union, Tuple, List, Optional 31 32from torch.utils.data import Dataset, DataLoader 33 34import torch_em 35 36from .. import util 37 38 39LABELS_URLS = { 40 "metadata": "https://ndownloader.figshare.com/files/57305609", 41 "labelsTr.zip": "https://ndownloader.figshare.com/files/57334925", 42 "labelsTr.z01": "https://ndownloader.figshare.com/files/57334910", 43 "labelsTr.z02": "https://ndownloader.figshare.com/files/57334913", 44 "labelsTr.z03": "https://ndownloader.figshare.com/files/57334916", 45 "labelsTr.z04": "https://ndownloader.figshare.com/files/57334919", 46 "labelsTr.z05": "https://ndownloader.figshare.com/files/57334922", 47 "labelsTr.z06": "https://ndownloader.figshare.com/files/57334928", 48 "labelsTr.z07": "https://ndownloader.figshare.com/files/57334931", 49} 50 51LABELS_CHECKSUMS = { 52 "metadata": "77a0e10d152517d8840ff1298af18f25ce7830ff13d665f60449ed17ce0836ed", 53 "labelsTr.zip": "6a3c5c39c14da117e9aea2836309cee528657cd130974e47acc5cffc6649b098", 54 "labelsTr.z01": "cd36b2e58a9b37d38ef3a55ea7b333cb21516a8d55df2c80900f7f1072f945ba", 55 "labelsTr.z02": "9d88bb8a6dbd1371d5bd17e8d2c48d2ddaac9c5cfdb4863d3f0a9fa162077f2e", 56 "labelsTr.z03": "828a919e6f835da98615bee9f0b06ea93f310830c9988b2c350f87d10e35ea5b", 57 "labelsTr.z04": "174418141b2f069c6fa4e70b79ccbdb44a3d786c60d1a29120b90f6748198f9b", 58 "labelsTr.z05": "50f0441d66d6db6c92d6b43986aac8438e2f566003a5b893b590e664193fbb59", 59 "labelsTr.z06": "f7fb0b38cc2746f993c508a3ba24e0f4934379efdc0fa3a9cc943fc6c5325339", 60 "labelsTr.z07": "fce0ee59d6426699ebda7bfc1acb84a786c09989564f4d2479577fdd70f771df", 61} 62 63# The LUNA16 CT subsets are hosted across two Zenodo records, split because the full dataset 64# exceeds the size Zenodo allows in a single record. Only sha256 checksums are recorded for 65# the subsets that were used to validate this module; the Zenodo records only provide md5 66# checksums for the rest, which `download_source` (sha256-only) cannot verify against. 67LUNA16_URLS = { 68 "subset0": "https://zenodo.org/records/3723295/files/subset0.zip", 69 "subset1": "https://zenodo.org/records/3723295/files/subset1.zip", 70 "subset2": "https://zenodo.org/records/3723295/files/subset2.zip", 71 "subset3": "https://zenodo.org/records/3723295/files/subset3.zip", 72 "subset4": "https://zenodo.org/records/3723295/files/subset4.zip", 73 "subset5": "https://zenodo.org/records/3723295/files/subset5.zip", 74 "subset6": "https://zenodo.org/records/3723295/files/subset6.zip", 75 "subset7": "https://zenodo.org/records/4121926/files/subset7.zip", 76 "subset8": "https://zenodo.org/records/4121926/files/subset8.zip", 77 "subset9": "https://zenodo.org/records/4121926/files/subset9.zip", 78} 79 80LUNA16_CHECKSUMS = {f"subset{i}": None for i in range(10)} 81 82LABEL_IDS = {"background": 0, "artery": 1, "vein": 2, "airway_lumen": 3, "airway_wall": 4} 83"""The semantic label ids of the AirRC classes.""" 84 85 86def _get_airrc_labels(path, download): 87 label_dir = os.path.join(path, "labelsTr") 88 if os.path.exists(label_dir): 89 return label_dir 90 91 os.makedirs(path, exist_ok=True) 92 93 metadata_path = os.path.join(path, "metadata.xlsx") 94 util.download_source( 95 path=metadata_path, url=LABELS_URLS["metadata"], download=download, checksum=LABELS_CHECKSUMS["metadata"] 96 ) 97 98 # The labels are distributed as a split zip archive (labelsTr.zip + labelsTr.z01 - z07), 99 # which have to be joined into a single archive before they can be extracted. 100 for name in ["labelsTr.zip"] + [f"labelsTr.z{i:02d}" for i in range(1, 8)]: 101 part_path = os.path.join(path, name) 102 util.download_source(path=part_path, url=LABELS_URLS[name], download=download, checksum=LABELS_CHECKSUMS[name]) 103 104 combined_path = os.path.join(path, "labelsTr_combined.zip") 105 if not os.path.exists(combined_path): 106 import subprocess 107 subprocess.run( 108 ["zip", "-FF", os.path.join(path, "labelsTr.zip"), "--out", combined_path], check=True, cwd=path 109 ) 110 111 util.unzip(zip_path=combined_path, dst=path, remove=False) 112 return label_dir 113 114 115def _get_luna16_raw(path, download, subset_ids): 116 raw_dir = os.path.join(path, "luna16") 117 os.makedirs(raw_dir, exist_ok=True) 118 119 for i in subset_ids: 120 name = f"subset{i}" 121 marker_path = os.path.join(path, f"{name}.extracted") 122 if os.path.exists(marker_path): 123 continue 124 125 zip_path = os.path.join(path, f"{name}.zip") 126 util.download_source(path=zip_path, url=LUNA16_URLS[name], download=download, checksum=LUNA16_CHECKSUMS[name]) 127 util.unzip(zip_path=zip_path, dst=raw_dir, remove=False) 128 open(marker_path, "w").close() 129 130 return raw_dir 131 132 133def _preprocess_airrc(label_dir, raw_dir, preprocessed_dir): 134 import h5py 135 import SimpleITK as sitk 136 137 os.makedirs(preprocessed_dir, exist_ok=True) 138 139 # Each LUNA16 subset zip unpacks into its own 'subsetN' sub-folder, so the raw volumes have 140 # to be looked up recursively rather than directly under `raw_dir`. 141 raw_paths_by_uid = { 142 os.path.basename(p)[:-len(".mhd")]: p for p in glob(os.path.join(raw_dir, "**", "*.mhd"), recursive=True) 143 } 144 145 label_paths = natsorted(glob(os.path.join(label_dir, "*.nii.gz"))) 146 for label_path in tqdm(label_paths, desc="Preprocessing the AirRC scans"): 147 series_uid = os.path.basename(label_path)[:-len(".nii.gz")] 148 149 out_path = os.path.join(preprocessed_dir, f"{series_uid}.h5") 150 if os.path.exists(out_path): 151 continue 152 153 raw_path = raw_paths_by_uid.get(series_uid) 154 if raw_path is None: # The matching LUNA16 subset was not downloaded. 155 continue 156 157 label_img = sitk.ReadImage(label_path) 158 raw_img = sitk.ReadImage(raw_path) 159 160 resampler = sitk.ResampleImageFilter() 161 resampler.SetReferenceImage(label_img) 162 resampler.SetInterpolator(sitk.sitkLinear) 163 resampler.SetDefaultPixelValue(-1024) # air / outside-scan HU value 164 resampled_raw = resampler.Execute(raw_img) 165 166 raw = sitk.GetArrayFromImage(resampled_raw).astype("float32") 167 labels = sitk.GetArrayFromImage(label_img).astype("uint8") 168 assert raw.shape == labels.shape, f"Shape mismatch for {series_uid}: {raw.shape} vs. {labels.shape}." 169 170 with h5py.File(f"{out_path}.tmp", "w") as f: 171 f.create_dataset("raw", data=raw, compression="gzip") 172 f.create_dataset("labels", data=labels, compression="gzip") 173 os.rename(f"{out_path}.tmp", out_path) 174 175 176def get_airrc_data( 177 path: Union[os.PathLike, str], subset_ids: Optional[List[int]] = None, download: bool = False 178) -> str: 179 """Download the AirRC dataset and the matching LUNA16 CT volumes. 180 181 Args: 182 path: Filepath to a folder where the data is downloaded for further processing. 183 subset_ids: The LUNA16 subset ids (0-9) to download and match against the AirRC masks. 184 By default, all ten subsets are downloaded (around 65 GB). Restricting this list 185 only yields the AirRC cases whose CT volume happens to be in the requested subsets. 186 download: Whether to download the data if it is not present. 187 188 Returns: 189 Filepath where the preprocessed data is stored. 190 """ 191 if subset_ids is None: 192 subset_ids = list(range(10)) 193 194 preprocessed_dir = os.path.join(path, "preprocessed") 195 196 os.makedirs(path, exist_ok=True) 197 198 label_dir = _get_airrc_labels(path, download) 199 raw_dir = _get_luna16_raw(path, download, subset_ids) 200 _preprocess_airrc(label_dir, raw_dir, preprocessed_dir) 201 202 return preprocessed_dir 203 204 205def get_airrc_paths( 206 path: Union[os.PathLike, str], subset_ids: Optional[List[int]] = None, download: bool = False 207) -> List[str]: 208 """Get paths to the AirRC data. 209 210 Args: 211 path: Filepath to a folder where the data is downloaded for further processing. 212 subset_ids: The LUNA16 subset ids (0-9) to download and match against the AirRC masks. 213 download: Whether to download the data if it is not present. 214 215 Returns: 216 List of filepaths for the hdf5 files, which contain the image data ('raw') and the 217 label data ('labels'). 218 """ 219 preprocessed_dir = get_airrc_data(path, subset_ids, download) 220 volume_paths = natsorted(glob(os.path.join(preprocessed_dir, "*.h5"))) 221 assert len(volume_paths) > 0, f"Could not find any preprocessed volumes in '{preprocessed_dir}'." 222 return volume_paths 223 224 225def get_airrc_dataset( 226 path: Union[os.PathLike, str], 227 patch_shape: Tuple[int, ...], 228 subset_ids: Optional[List[int]] = None, 229 resize_inputs: bool = False, 230 download: bool = False, 231 **kwargs 232) -> Dataset: 233 """Get the AirRC dataset for pulmonary artery, vein and airway segmentation. 234 235 Args: 236 path: Filepath to a folder where the data is downloaded for further processing. 237 patch_shape: The patch shape to use for training. 238 subset_ids: The LUNA16 subset ids (0-9) to download and match against the AirRC masks. 239 resize_inputs: Whether to resize inputs to the desired patch shape. 240 download: Whether to download the data if it is not present. 241 kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset`. 242 243 Returns: 244 The segmentation dataset. 245 """ 246 volume_paths = get_airrc_paths(path, subset_ids, download) 247 248 if resize_inputs: 249 resize_kwargs = {"patch_shape": patch_shape, "is_rgb": False} 250 kwargs, patch_shape = util.update_kwargs_for_resize_trafo( 251 kwargs=kwargs, patch_shape=patch_shape, resize_inputs=resize_inputs, resize_kwargs=resize_kwargs 252 ) 253 254 return torch_em.default_segmentation_dataset( 255 raw_paths=volume_paths, 256 raw_key="raw", 257 label_paths=volume_paths, 258 label_key="labels", 259 patch_shape=patch_shape, 260 is_seg_dataset=True, 261 **kwargs 262 ) 263 264 265def get_airrc_loader( 266 path: Union[os.PathLike, str], 267 batch_size: int, 268 patch_shape: Tuple[int, ...], 269 subset_ids: Optional[List[int]] = None, 270 resize_inputs: bool = False, 271 download: bool = False, 272 **kwargs 273) -> DataLoader: 274 """Get the AirRC dataloader for pulmonary artery, vein and airway segmentation. 275 276 Args: 277 path: Filepath to a folder where the data is downloaded for further processing. 278 batch_size: The batch size for training. 279 patch_shape: The patch shape to use for training. 280 subset_ids: The LUNA16 subset ids (0-9) to download and match against the AirRC masks. 281 resize_inputs: Whether to resize inputs to the desired patch shape. 282 download: Whether to download the data if it is not present. 283 kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset` or for the PyTorch DataLoader. 284 285 Returns: 286 The DataLoader. 287 """ 288 ds_kwargs, loader_kwargs = util.split_kwargs(torch_em.default_segmentation_dataset, **kwargs) 289 dataset = get_airrc_dataset(path, patch_shape, subset_ids, resize_inputs, download, **ds_kwargs) 290 return torch_em.get_data_loader(dataset, batch_size, **loader_kwargs)
The semantic label ids of the AirRC classes.
177def get_airrc_data( 178 path: Union[os.PathLike, str], subset_ids: Optional[List[int]] = None, download: bool = False 179) -> str: 180 """Download the AirRC dataset and the matching LUNA16 CT volumes. 181 182 Args: 183 path: Filepath to a folder where the data is downloaded for further processing. 184 subset_ids: The LUNA16 subset ids (0-9) to download and match against the AirRC masks. 185 By default, all ten subsets are downloaded (around 65 GB). Restricting this list 186 only yields the AirRC cases whose CT volume happens to be in the requested subsets. 187 download: Whether to download the data if it is not present. 188 189 Returns: 190 Filepath where the preprocessed data is stored. 191 """ 192 if subset_ids is None: 193 subset_ids = list(range(10)) 194 195 preprocessed_dir = os.path.join(path, "preprocessed") 196 197 os.makedirs(path, exist_ok=True) 198 199 label_dir = _get_airrc_labels(path, download) 200 raw_dir = _get_luna16_raw(path, download, subset_ids) 201 _preprocess_airrc(label_dir, raw_dir, preprocessed_dir) 202 203 return preprocessed_dir
Download the AirRC dataset and the matching LUNA16 CT volumes.
Arguments:
- path: Filepath to a folder where the data is downloaded for further processing.
- subset_ids: The LUNA16 subset ids (0-9) to download and match against the AirRC masks. By default, all ten subsets are downloaded (around 65 GB). Restricting this list only yields the AirRC cases whose CT volume happens to be in the requested subsets.
- download: Whether to download the data if it is not present.
Returns:
Filepath where the preprocessed data is stored.
206def get_airrc_paths( 207 path: Union[os.PathLike, str], subset_ids: Optional[List[int]] = None, download: bool = False 208) -> List[str]: 209 """Get paths to the AirRC data. 210 211 Args: 212 path: Filepath to a folder where the data is downloaded for further processing. 213 subset_ids: The LUNA16 subset ids (0-9) to download and match against the AirRC masks. 214 download: Whether to download the data if it is not present. 215 216 Returns: 217 List of filepaths for the hdf5 files, which contain the image data ('raw') and the 218 label data ('labels'). 219 """ 220 preprocessed_dir = get_airrc_data(path, subset_ids, download) 221 volume_paths = natsorted(glob(os.path.join(preprocessed_dir, "*.h5"))) 222 assert len(volume_paths) > 0, f"Could not find any preprocessed volumes in '{preprocessed_dir}'." 223 return volume_paths
Get paths to the AirRC data.
Arguments:
- path: Filepath to a folder where the data is downloaded for further processing.
- subset_ids: The LUNA16 subset ids (0-9) to download and match against the AirRC masks.
- download: Whether to download the data if it is not present.
Returns:
List of filepaths for the hdf5 files, which contain the image data ('raw') and the label data ('labels').
226def get_airrc_dataset( 227 path: Union[os.PathLike, str], 228 patch_shape: Tuple[int, ...], 229 subset_ids: Optional[List[int]] = None, 230 resize_inputs: bool = False, 231 download: bool = False, 232 **kwargs 233) -> Dataset: 234 """Get the AirRC dataset for pulmonary artery, vein and airway segmentation. 235 236 Args: 237 path: Filepath to a folder where the data is downloaded for further processing. 238 patch_shape: The patch shape to use for training. 239 subset_ids: The LUNA16 subset ids (0-9) to download and match against the AirRC masks. 240 resize_inputs: Whether to resize inputs to the desired patch shape. 241 download: Whether to download the data if it is not present. 242 kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset`. 243 244 Returns: 245 The segmentation dataset. 246 """ 247 volume_paths = get_airrc_paths(path, subset_ids, download) 248 249 if resize_inputs: 250 resize_kwargs = {"patch_shape": patch_shape, "is_rgb": False} 251 kwargs, patch_shape = util.update_kwargs_for_resize_trafo( 252 kwargs=kwargs, patch_shape=patch_shape, resize_inputs=resize_inputs, resize_kwargs=resize_kwargs 253 ) 254 255 return torch_em.default_segmentation_dataset( 256 raw_paths=volume_paths, 257 raw_key="raw", 258 label_paths=volume_paths, 259 label_key="labels", 260 patch_shape=patch_shape, 261 is_seg_dataset=True, 262 **kwargs 263 )
Get the AirRC dataset for pulmonary artery, vein and airway segmentation.
Arguments:
- path: Filepath to a folder where the data is downloaded for further processing.
- patch_shape: The patch shape to use for training.
- subset_ids: The LUNA16 subset ids (0-9) to download and match against the AirRC masks.
- resize_inputs: Whether to resize inputs to the desired patch shape.
- download: Whether to download the data if it is not present.
- kwargs: Additional keyword arguments for
torch_em.default_segmentation_dataset.
Returns:
The segmentation dataset.
266def get_airrc_loader( 267 path: Union[os.PathLike, str], 268 batch_size: int, 269 patch_shape: Tuple[int, ...], 270 subset_ids: Optional[List[int]] = None, 271 resize_inputs: bool = False, 272 download: bool = False, 273 **kwargs 274) -> DataLoader: 275 """Get the AirRC dataloader for pulmonary artery, vein and airway segmentation. 276 277 Args: 278 path: Filepath to a folder where the data is downloaded for further processing. 279 batch_size: The batch size for training. 280 patch_shape: The patch shape to use for training. 281 subset_ids: The LUNA16 subset ids (0-9) to download and match against the AirRC masks. 282 resize_inputs: Whether to resize inputs to the desired patch shape. 283 download: Whether to download the data if it is not present. 284 kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset` or for the PyTorch DataLoader. 285 286 Returns: 287 The DataLoader. 288 """ 289 ds_kwargs, loader_kwargs = util.split_kwargs(torch_em.default_segmentation_dataset, **kwargs) 290 dataset = get_airrc_dataset(path, patch_shape, subset_ids, resize_inputs, download, **ds_kwargs) 291 return torch_em.get_data_loader(dataset, batch_size, **loader_kwargs)
Get the AirRC dataloader for pulmonary artery, vein and airway segmentation.
Arguments:
- path: Filepath to a folder where the data is downloaded for further processing.
- batch_size: The batch size for training.
- patch_shape: The patch shape to use for training.
- subset_ids: The LUNA16 subset ids (0-9) to download and match against the AirRC masks.
- resize_inputs: Whether to resize inputs to the desired patch shape.
- download: Whether to download the data if it is not present.
- kwargs: Additional keyword arguments for
torch_em.default_segmentation_datasetor for the PyTorch DataLoader.
Returns:
The DataLoader.