torch_em.data.datasets.medical.hubmap_kidney
The HuBMAP - Hacking the Kidney dataset contains annotations for glomeruli functional tissue unit (FTU) segmentation in human kidney whole-slide histopathology images (PAS-stained).
This is the dataset for the "HuBMAP - Hacking the Kidney" Kaggle competition, located at https://www.kaggle.com/competitions/hubmap-kidney-segmentation. It provides 8 PAS-stained kidney whole-slide images (fresh-frozen and FFPE tissue preparations, contributed by the Human BioMolecular Atlas Program / HuBMAP through the BIOMIC team at Vanderbilt University), each paired with expert-annotated glomerulus segmentations, given as run-length encoded (RLE) masks in 'train.csv'. The competition data is released under the CC BY 4.0 license.
NOTE: This is a DIFFERENT Kaggle competition from "HuBMAP + HPA - Hacking the Human Body" (see 'hubmap_hpa.py' for that dataset), even though both are hosted by HuBMAP and target FTU segmentation.
NOTE: Downloading this dataset requires a Kaggle account that has accepted the competition rules at https://www.kaggle.com/competitions/hubmap-kidney-segmentation/rules. Without that, the Kaggle API download fails with an HTTP 403 error, even with valid API credentials.
NOTE: The whole-slide images are very large single-resolution BigTIFFs (tens of thousands of pixels per side, several hundred MB to multiple GB per file). As in 'histopathology/camelyon.py', each image is read lazily (via a 'zarr' view over the tiled TIFF where available, otherwise a memory-mapped read) and converted once into a chunked HDF5 file that stores the raw image together with the binary glomerulus mask decoded from the RLE encoding. Downstream patch extraction for training then happens lazily via 'torch_em.default_segmentation_dataset' from these HDF5 files, so no manual tiling logic is required here.
This dataset is described in the publication https://doi.org/10.1038/s41467-023-40291-0 ("Segmenting functional tissue units across human organs using community-driven development of generalizable machine learning algorithms", Nature Communications, 2023). Please cite it if you use this dataset in your research.
1"""The HuBMAP - Hacking the Kidney dataset contains annotations for glomeruli functional tissue unit (FTU) 2segmentation in human kidney whole-slide histopathology images (PAS-stained). 3 4This is the dataset for the "HuBMAP - Hacking the Kidney" Kaggle competition, located at 5https://www.kaggle.com/competitions/hubmap-kidney-segmentation. It provides 8 PAS-stained kidney whole-slide 6images (fresh-frozen and FFPE tissue preparations, contributed by the Human BioMolecular Atlas Program / 7HuBMAP through the BIOMIC team at Vanderbilt University), each paired with expert-annotated glomerulus 8segmentations, given as run-length encoded (RLE) masks in 'train.csv'. The competition data is released 9under the CC BY 4.0 license. 10 11NOTE: This is a DIFFERENT Kaggle competition from "HuBMAP + HPA - Hacking the Human Body" (see 12'hubmap_hpa.py' for that dataset), even though both are hosted by HuBMAP and target FTU segmentation. 13 14NOTE: Downloading this dataset requires a Kaggle account that has accepted the competition rules at 15https://www.kaggle.com/competitions/hubmap-kidney-segmentation/rules. Without that, the Kaggle API 16download fails with an HTTP 403 error, even with valid API credentials. 17 18NOTE: The whole-slide images are very large single-resolution BigTIFFs (tens of thousands of pixels per 19side, several hundred MB to multiple GB per file). As in 'histopathology/camelyon.py', each image is read 20lazily (via a 'zarr' view over the tiled TIFF where available, otherwise a memory-mapped read) and converted 21once into a chunked HDF5 file that stores the raw image together with the binary glomerulus mask decoded 22from the RLE encoding. Downstream patch extraction for training then happens lazily via 23'torch_em.default_segmentation_dataset' from these HDF5 files, so no manual tiling logic is required here. 24 25This dataset is described in the publication https://doi.org/10.1038/s41467-023-40291-0 ("Segmenting 26functional tissue units across human organs using community-driven development of generalizable machine 27learning algorithms", Nature Communications, 2023). Please cite it if you use this dataset in your research. 28""" 29 30import os 31import csv 32from pathlib import Path 33from typing import List, Tuple, Union 34 35import numpy as np 36from tqdm import tqdm 37 38from torch.utils.data import Dataset, DataLoader 39 40import torch_em 41 42from .. import util 43 44# The RLE-encoded masks in 'train.csv' exceed the default field size limit for these large WSIs. 45csv.field_size_limit(10**8) 46 47 48def _rle_decode(rle: str, shape: Tuple[int, int]) -> np.ndarray: 49 """Decode a Kaggle-style run-length-encoded mask into a binary array. 50 51 The encoding lists 1-indexed (start, length) pairs over a column-major (Fortran order) flattening of 52 the image, i.e. pixels are numbered from top to bottom, then left to right. 53 """ 54 height, width = shape 55 mask = np.zeros(height * width, dtype=np.uint8) 56 57 values = [int(v) for v in rle.split()] 58 starts, lengths = values[0::2], values[1::2] 59 for start, length in zip(starts, lengths): 60 start -= 1 61 mask[start:start + length] = 1 62 63 return mask.reshape((width, height)).T 64 65 66def _open_raw_array(tiff_path): 67 import tifffile 68 69 tiff = tifffile.TiffFile(tiff_path) 70 # Some slides are stored as several series (e.g. a full-resolution image plus a thumbnail); 71 # pick the one with the most pixels rather than assuming series[0] is the full-resolution one. 72 series = max(tiff.series, key=lambda s: np.prod(s.shape)) 73 74 try: 75 import zarr 76 array = zarr.open(series.aszarr(), mode="r") 77 if not hasattr(array, "shape"): # The tiff is pyramidal: pick the largest resolution level. 78 array = max(array.values(), key=lambda level: np.prod(level.shape)) 79 except Exception: 80 array = series.asarray() 81 82 # Some slides carry extra singleton axes (e.g. an OME-style (T, Z, C, H, W) layout), which have 83 # to be dropped before the channel axis can be identified. Indexed rather than via '.squeeze()', 84 # so a lazy zarr array is not fully materialized just to drop size-1 axes. 85 squeeze_axes = tuple(i for i, size in enumerate(array.shape) if size == 1) 86 if squeeze_axes: 87 index = tuple(0 if i in squeeze_axes else slice(None) for i in range(array.ndim)) 88 array = array[index] 89 90 # Some HuBMAP kidney tiffs store the channel axis first, i.e. (3, height, width) instead of the usual 91 # channel-last (height, width, 3) layout. Normalize to channel-last so downstream code is uniform. 92 if array.ndim == 3 and array.shape[0] == 3 and array.shape[-1] != 3: 93 array = np.transpose(array, (1, 2, 0)) 94 95 assert array.ndim == 3 and array.shape[-1] == 3, f"Unexpected array shape {array.shape} for {tiff_path}." 96 return array 97 98 99def _convert_slide(tiff_path: str, rle: str, output_path: str, tile: int = 4096) -> None: 100 import h5py 101 102 array = _open_raw_array(tiff_path) 103 height, width = array.shape[0], array.shape[1] 104 mask = _rle_decode(rle, (height, width)) 105 106 tmp_path = output_path + ".tmp" 107 with h5py.File(tmp_path, "w") as f: 108 raw = f.create_dataset( 109 "raw", shape=(3, height, width), dtype="uint8", compression="gzip", chunks=(1, 512, 512) 110 ) 111 labels = f.create_dataset( 112 "labels", shape=(height, width), dtype="uint8", compression="gzip", chunks=(512, 512) 113 ) 114 for y in tqdm(range(0, height, tile), desc=f"Converting {Path(tiff_path).stem}"): 115 for x in range(0, width, tile): 116 th, tw = min(tile, height - y), min(tile, width - x) 117 tile_data = np.asarray(array[y:y + th, x:x + tw]) 118 raw[:, y:y + th, x:x + tw] = tile_data.transpose(2, 0, 1) 119 labels[y:y + th, x:x + tw] = mask[y:y + th, x:x + tw] 120 121 os.replace(tmp_path, output_path) 122 123 124def _load_train_annotations(data_dir): 125 csv_path = os.path.join(data_dir, "train.csv") 126 if not os.path.exists(csv_path): 127 raise RuntimeError(f"Could not find 'train.csv' at {csv_path}.") 128 with open(csv_path, "r") as f: 129 return list(csv.DictReader(f)) 130 131 132def get_hubmap_kidney_data(path: Union[os.PathLike, str], download: bool = False) -> str: 133 """Download the HuBMAP - Hacking the Kidney dataset. 134 135 Args: 136 path: Filepath to a folder where the data will be saved. 137 download: Whether to download the data if it is not present. 138 139 Returns: 140 Filepath to the folder where the raw data is stored. 141 """ 142 data_dir = os.path.join(path, "train") 143 if os.path.exists(data_dir): 144 return path 145 146 os.makedirs(path, exist_ok=True) 147 zip_path = os.path.join(path, "hubmap-kidney-segmentation.zip") 148 util.download_source_kaggle( 149 path=path, dataset_name="hubmap-kidney-segmentation", download=download, competition=True 150 ) 151 util.unzip(zip_path=zip_path, dst=path) 152 153 if not os.path.exists(data_dir): 154 raise RuntimeError(f"The expected 'train' folder is missing after extraction at {path}.") 155 return path 156 157 158def get_hubmap_kidney_paths(path: Union[os.PathLike, str], download: bool = False) -> List[str]: 159 """Get paths to the preprocessed HuBMAP - Hacking the Kidney data. 160 161 Each returned HDF5 file stores the raw image under 'raw' (channels, height, width) and the 162 binary glomerulus segmentation mask under 'labels' (height, width). 163 164 Args: 165 path: Filepath to a folder where the data will be saved. 166 download: Whether to download the data if it is not present. 167 168 Returns: 169 List of filepaths to the preprocessed HDF5 files. 170 """ 171 data_dir = get_hubmap_kidney_data(path, download) 172 rows = _load_train_annotations(data_dir) 173 174 preprocessed_dir = os.path.join(path, "preprocessed") 175 os.makedirs(preprocessed_dir, exist_ok=True) 176 177 volume_paths = [] 178 for row in rows: 179 image_id = row["id"] 180 tiff_path = os.path.join(data_dir, "train", f"{image_id}.tiff") 181 if not os.path.exists(tiff_path): 182 continue 183 184 output_path = os.path.join(preprocessed_dir, f"{image_id}.h5") 185 if not os.path.exists(output_path): 186 _convert_slide(tiff_path, row["encoding"], output_path) 187 volume_paths.append(output_path) 188 189 if not volume_paths: 190 raise RuntimeError(f"No annotated images were found at {data_dir}.") 191 192 return volume_paths 193 194 195def get_hubmap_kidney_dataset( 196 path: Union[os.PathLike, str], 197 patch_shape: Tuple[int, int], 198 download: bool = False, 199 resize_inputs: bool = False, 200 **kwargs, 201) -> Dataset: 202 """Get the HuBMAP - Hacking the Kidney dataset for glomerulus segmentation in kidney whole-slide images. 203 204 Args: 205 path: Filepath to a folder where the data will be saved. 206 patch_shape: The patch shape to use for training. 207 download: Whether to download the data if it is not present. 208 resize_inputs: Whether to resize the inputs. 209 kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset`. 210 211 Returns: 212 The segmentation dataset. 213 """ 214 volume_paths = get_hubmap_kidney_paths(path, download) 215 216 if resize_inputs: 217 resize_kwargs = {"patch_shape": patch_shape, "is_rgb": True} 218 kwargs, patch_shape = util.update_kwargs_for_resize_trafo( 219 kwargs=kwargs, patch_shape=patch_shape, resize_inputs=resize_inputs, resize_kwargs=resize_kwargs 220 ) 221 222 return torch_em.default_segmentation_dataset( 223 raw_paths=volume_paths, 224 raw_key="raw", 225 label_paths=volume_paths, 226 label_key="labels", 227 patch_shape=patch_shape, 228 is_seg_dataset=True, 229 with_channels=True, 230 ndim=2, 231 **kwargs, 232 ) 233 234 235def get_hubmap_kidney_loader( 236 path: Union[os.PathLike, str], 237 batch_size: int, 238 patch_shape: Tuple[int, int], 239 download: bool = False, 240 resize_inputs: bool = False, 241 **kwargs, 242) -> DataLoader: 243 """Get the HuBMAP - Hacking the Kidney dataloader for glomerulus segmentation in kidney whole-slide images. 244 245 Args: 246 path: Filepath to a folder where the data will be saved. 247 batch_size: The batch size for training. 248 patch_shape: The patch shape to use for training. 249 download: Whether to download the data if it is not present. 250 resize_inputs: Whether to resize the inputs. 251 kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset` or the PyTorch DataLoader. 252 253 Returns: 254 The DataLoader. 255 """ 256 ds_kwargs, loader_kwargs = util.split_kwargs(torch_em.default_segmentation_dataset, **kwargs) 257 dataset = get_hubmap_kidney_dataset(path, patch_shape, download, resize_inputs, **ds_kwargs) 258 return torch_em.get_data_loader(dataset, batch_size, **loader_kwargs)
133def get_hubmap_kidney_data(path: Union[os.PathLike, str], download: bool = False) -> str: 134 """Download the HuBMAP - Hacking the Kidney dataset. 135 136 Args: 137 path: Filepath to a folder where the data will be saved. 138 download: Whether to download the data if it is not present. 139 140 Returns: 141 Filepath to the folder where the raw data is stored. 142 """ 143 data_dir = os.path.join(path, "train") 144 if os.path.exists(data_dir): 145 return path 146 147 os.makedirs(path, exist_ok=True) 148 zip_path = os.path.join(path, "hubmap-kidney-segmentation.zip") 149 util.download_source_kaggle( 150 path=path, dataset_name="hubmap-kidney-segmentation", download=download, competition=True 151 ) 152 util.unzip(zip_path=zip_path, dst=path) 153 154 if not os.path.exists(data_dir): 155 raise RuntimeError(f"The expected 'train' folder is missing after extraction at {path}.") 156 return path
Download the HuBMAP - Hacking the Kidney dataset.
Arguments:
- path: Filepath to a folder where the data will be saved.
- download: Whether to download the data if it is not present.
Returns:
Filepath to the folder where the raw data is stored.
159def get_hubmap_kidney_paths(path: Union[os.PathLike, str], download: bool = False) -> List[str]: 160 """Get paths to the preprocessed HuBMAP - Hacking the Kidney data. 161 162 Each returned HDF5 file stores the raw image under 'raw' (channels, height, width) and the 163 binary glomerulus segmentation mask under 'labels' (height, width). 164 165 Args: 166 path: Filepath to a folder where the data will be saved. 167 download: Whether to download the data if it is not present. 168 169 Returns: 170 List of filepaths to the preprocessed HDF5 files. 171 """ 172 data_dir = get_hubmap_kidney_data(path, download) 173 rows = _load_train_annotations(data_dir) 174 175 preprocessed_dir = os.path.join(path, "preprocessed") 176 os.makedirs(preprocessed_dir, exist_ok=True) 177 178 volume_paths = [] 179 for row in rows: 180 image_id = row["id"] 181 tiff_path = os.path.join(data_dir, "train", f"{image_id}.tiff") 182 if not os.path.exists(tiff_path): 183 continue 184 185 output_path = os.path.join(preprocessed_dir, f"{image_id}.h5") 186 if not os.path.exists(output_path): 187 _convert_slide(tiff_path, row["encoding"], output_path) 188 volume_paths.append(output_path) 189 190 if not volume_paths: 191 raise RuntimeError(f"No annotated images were found at {data_dir}.") 192 193 return volume_paths
Get paths to the preprocessed HuBMAP - Hacking the Kidney data.
Each returned HDF5 file stores the raw image under 'raw' (channels, height, width) and the binary glomerulus segmentation mask under 'labels' (height, width).
Arguments:
- path: Filepath to a folder where the data will be saved.
- download: Whether to download the data if it is not present.
Returns:
List of filepaths to the preprocessed HDF5 files.
196def get_hubmap_kidney_dataset( 197 path: Union[os.PathLike, str], 198 patch_shape: Tuple[int, int], 199 download: bool = False, 200 resize_inputs: bool = False, 201 **kwargs, 202) -> Dataset: 203 """Get the HuBMAP - Hacking the Kidney dataset for glomerulus segmentation in kidney whole-slide images. 204 205 Args: 206 path: Filepath to a folder where the data will be saved. 207 patch_shape: The patch shape to use for training. 208 download: Whether to download the data if it is not present. 209 resize_inputs: Whether to resize the inputs. 210 kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset`. 211 212 Returns: 213 The segmentation dataset. 214 """ 215 volume_paths = get_hubmap_kidney_paths(path, download) 216 217 if resize_inputs: 218 resize_kwargs = {"patch_shape": patch_shape, "is_rgb": True} 219 kwargs, patch_shape = util.update_kwargs_for_resize_trafo( 220 kwargs=kwargs, patch_shape=patch_shape, resize_inputs=resize_inputs, resize_kwargs=resize_kwargs 221 ) 222 223 return torch_em.default_segmentation_dataset( 224 raw_paths=volume_paths, 225 raw_key="raw", 226 label_paths=volume_paths, 227 label_key="labels", 228 patch_shape=patch_shape, 229 is_seg_dataset=True, 230 with_channels=True, 231 ndim=2, 232 **kwargs, 233 )
Get the HuBMAP - Hacking the Kidney dataset for glomerulus segmentation in kidney whole-slide images.
Arguments:
- path: Filepath to a folder where the data will be saved.
- patch_shape: The patch shape to use for training.
- download: Whether to download the data if it is not present.
- resize_inputs: Whether to resize the inputs.
- kwargs: Additional keyword arguments for
torch_em.default_segmentation_dataset.
Returns:
The segmentation dataset.
236def get_hubmap_kidney_loader( 237 path: Union[os.PathLike, str], 238 batch_size: int, 239 patch_shape: Tuple[int, int], 240 download: bool = False, 241 resize_inputs: bool = False, 242 **kwargs, 243) -> DataLoader: 244 """Get the HuBMAP - Hacking the Kidney dataloader for glomerulus segmentation in kidney whole-slide images. 245 246 Args: 247 path: Filepath to a folder where the data will be saved. 248 batch_size: The batch size for training. 249 patch_shape: The patch shape to use for training. 250 download: Whether to download the data if it is not present. 251 resize_inputs: Whether to resize the inputs. 252 kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset` or the PyTorch DataLoader. 253 254 Returns: 255 The DataLoader. 256 """ 257 ds_kwargs, loader_kwargs = util.split_kwargs(torch_em.default_segmentation_dataset, **kwargs) 258 dataset = get_hubmap_kidney_dataset(path, patch_shape, download, resize_inputs, **ds_kwargs) 259 return torch_em.get_data_loader(dataset, batch_size, **loader_kwargs)
Get the HuBMAP - Hacking the Kidney dataloader for glomerulus segmentation in kidney whole-slide images.
Arguments:
- path: Filepath to a folder where the data will be saved.
- batch_size: The batch size for training.
- patch_shape: The patch shape to use for training.
- download: Whether to download the data if it is not present.
- resize_inputs: Whether to resize the inputs.
- kwargs: Additional keyword arguments for
torch_em.default_segmentation_datasetor the PyTorch DataLoader.
Returns:
The DataLoader.