torch_em.data.datasets.histopathology.prostate_cnb
The Prostate CNB dataset contains annotations for the semantic segmentation of prostate tissue types and histopathological structures in H&E whole-slide images of image-guided prostate core needle biopsies (CNBs).
The dataset consists of 37 whole-slide images (WSIs) with 60 core needle biopsies from 32 patients of a single institution, scanned at 0.243 micrometer per pixel. There are 15,480 manual annotations of two uropathologists (with a consensus review), which are provided as multiclass semantic masks (and as GeoJSON vectors, which are not used by this module). Regions that were not annotated are labeled as stroma automatically.
The label ids of the masks are:
- 0: background, 1: tumor, 2: benign gland, 3: blood vessels, 4: fibromuscular bundles, 5: abnormal secretions,
6: contamination with another tissue, 7: prominent nucleolus, 8: immune cells, 9: nerve, 10: artifact,
11: seminal vesicle, 12: adipose tissue, 13: normal secretions, 14: stromal retraction spaces, 15: muscle,
16: foreign body contamination, 17: high-grade prostatic intraepithelial neoplasia (HGPIN), 18: calcifications,
19: intestinal glands and mucus, 20: perineural invasion, 21: hemorrhage, 22: intraductal carcinoma,
23: necrosis, 24: mitosis, 25: nerve ganglion, 26: atypical intraductal proliferation, 27: red blood cells,
28: stroma
(see
CLASS_NAMES). Not every class occurs in every slide.
NOTE: This is a different dataset than the already integrated torch_em.data.datasets.histopathology.precise
(PRECISE), which consists of paired H&E and IHC slides of 25 other patients with 7 classes.
The data is located at https://doi.org/10.5281/zenodo.18930299 (the latest version is
https://zenodo.org/records/21780258, the first version https://zenodo.org/records/18930300) as a single 34.5 GB zip
archive with pyramidal OME-TIFF images and masks, released under a CC BY 4.0 license. To avoid downloading the whole
archive, this module reads its directory from the server and fetches only the requested slides and masks. Each slide
is converted once into a chunked HDF5 file at a chosen pyramid level (pixel size 0.243 * 2 ** level micrometer, the
default level 1 is 0.486 micrometer), the downloaded OME-TIFFs are removed afterwards. Use n_cases to only prepare
a subset.
This dataset is from the publication https://doi.org/10.1038/s41597-026-08349-y. Please cite it if you use this dataset in your research.
1"""The Prostate CNB dataset contains annotations for the semantic segmentation of prostate tissue types and 2histopathological structures in H&E whole-slide images of image-guided prostate core needle biopsies (CNBs). 3 4The dataset consists of 37 whole-slide images (WSIs) with 60 core needle biopsies from 32 patients of a single 5institution, scanned at 0.243 micrometer per pixel. There are 15,480 manual annotations of two uropathologists 6(with a consensus review), which are provided as multiclass semantic masks (and as GeoJSON vectors, which are not used 7by this module). Regions that were not annotated are labeled as stroma automatically. 8 9The label ids of the masks are: 10- 0: background, 1: tumor, 2: benign gland, 3: blood vessels, 4: fibromuscular bundles, 5: abnormal secretions, 11 6: contamination with another tissue, 7: prominent nucleolus, 8: immune cells, 9: nerve, 10: artifact, 12 11: seminal vesicle, 12: adipose tissue, 13: normal secretions, 14: stromal retraction spaces, 15: muscle, 13 16: foreign body contamination, 17: high-grade prostatic intraepithelial neoplasia (HGPIN), 18: calcifications, 14 19: intestinal glands and mucus, 20: perineural invasion, 21: hemorrhage, 22: intraductal carcinoma, 15 23: necrosis, 24: mitosis, 25: nerve ganglion, 26: atypical intraductal proliferation, 27: red blood cells, 16 28: stroma 17(see `CLASS_NAMES`). Not every class occurs in every slide. 18 19NOTE: This is a different dataset than the already integrated `torch_em.data.datasets.histopathology.precise` 20(PRECISE), which consists of paired H&E and IHC slides of 25 other patients with 7 classes. 21 22The data is located at https://doi.org/10.5281/zenodo.18930299 (the latest version is 23https://zenodo.org/records/21780258, the first version https://zenodo.org/records/18930300) as a single 34.5 GB zip 24archive with pyramidal OME-TIFF images and masks, released under a CC BY 4.0 license. To avoid downloading the whole 25archive, this module reads its directory from the server and fetches only the requested slides and masks. Each slide 26is converted once into a chunked HDF5 file at a chosen pyramid level (pixel size 0.243 * 2 ** level micrometer, the 27default level 1 is 0.486 micrometer), the downloaded OME-TIFFs are removed afterwards. Use `n_cases` to only prepare 28a subset. 29 30This dataset is from the publication https://doi.org/10.1038/s41597-026-08349-y. 31Please cite it if you use this dataset in your research. 32""" 33 34import os 35import re 36import json 37import uuid 38import zlib 39import struct 40from typing import List, Optional, Tuple, Union 41 42from tqdm import tqdm 43 44import torch 45 46from torch.utils.data import Dataset, DataLoader 47 48import torch_em 49 50from .. import util 51 52 53ZIP_URL = "https://zenodo.org/api/records/21780258/files/dataset.zip/content" 54 55CLASS_NAMES = { 56 0: "background", 57 1: "tumor", 58 2: "benign gland", 59 3: "blood vessels", 60 4: "fibromuscular bundles", 61 5: "abnormal secretions", 62 6: "contamination with another tissue", 63 7: "prominent nucleolus", 64 8: "immune cells", 65 9: "nerve", 66 10: "artifact", 67 11: "seminal vesicle", 68 12: "adipose tissue", 69 13: "normal secretions", 70 14: "stromal retraction spaces", 71 15: "muscle", 72 16: "foreign body contamination", 73 17: "high-grade prostatic intraepithelial neoplasia", 74 18: "calcifications", 75 19: "intestinal glands and mucus", 76 20: "perineural invasion", 77 21: "hemorrhage", 78 22: "intraductal carcinoma", 79 23: "necrosis", 80 24: "mitosis", 81 25: "nerve ganglion", 82 26: "atypical intraductal proliferation", 83 27: "red blood cells", 84 28: "stroma", 85} 86 87N_LEVELS = 6 88 89 90def _range_get(url, start, end, **kwargs): 91 import requests 92 93 response = requests.get(url, headers={"Range": f"bytes={start}-{end}"}, **kwargs) 94 response.raise_for_status() 95 return response 96 97 98def _read_zip_entries(path): 99 cache_path = os.path.join(path, "zip_entries.json") 100 if os.path.exists(cache_path): 101 with open(cache_path) as f: 102 return json.load(f) 103 104 import requests 105 106 size = int(requests.head(ZIP_URL, allow_redirects=True).headers["Content-Length"]) 107 tail = _range_get(ZIP_URL, size - 65557, size - 1).content 108 eocd = tail[tail.rfind(b"PK\x05\x06"):] 109 n_entries, cd_size, cd_offset = struct.unpack("<HII", eocd[10:20]) 110 if cd_offset == 0xFFFFFFFF or n_entries == 0xFFFF: 111 locator = tail[tail.rfind(b"PK\x06\x07"):][:20] 112 zip64_offset = struct.unpack("<Q", locator[8:16])[0] 113 record = _range_get(ZIP_URL, zip64_offset, zip64_offset + 55).content 114 n_entries, cd_size, cd_offset = struct.unpack("<QQQ", record[32:56]) 115 116 directory = _range_get(ZIP_URL, cd_offset, cd_offset + cd_size - 1).content 117 entries, pos = {}, 0 118 while pos < len(directory): 119 (signature, _, _, _, _, _, _, crc, comp_size, size, name_len, extra_len, comment_len, _, _, _, 120 header_offset) = struct.unpack("<IHHHHHHIIIHHHHHII", directory[pos:pos + 46]) 121 assert signature == 0x02014B50, "The zip directory could not be parsed." 122 name = directory[pos + 46:pos + 46 + name_len].decode() 123 extra = directory[pos + 46 + name_len:pos + 46 + name_len + extra_len] 124 extra_pos = 0 125 while extra_pos < len(extra): 126 field_id, field_size = struct.unpack("<HH", extra[extra_pos:extra_pos + 4]) 127 if field_id == 1: 128 field, k = extra[extra_pos + 4:extra_pos + 4 + field_size], 0 129 if size == 0xFFFFFFFF: 130 size, k = struct.unpack("<Q", field[k:k + 8])[0], k + 8 131 if comp_size == 0xFFFFFFFF: 132 comp_size, k = struct.unpack("<Q", field[k:k + 8])[0], k + 8 133 if header_offset == 0xFFFFFFFF: 134 header_offset = struct.unpack("<Q", field[k:k + 8])[0] 135 extra_pos += 4 + field_size 136 if size > 0: 137 entries[name] = {"crc": crc, "comp_size": comp_size, "size": size, "offset": header_offset} 138 pos += 46 + name_len + extra_len + comment_len 139 140 tmp_path = f"{cache_path}.{uuid.uuid4().hex}.tmp" 141 with open(tmp_path, "w") as f: 142 json.dump(entries, f) 143 os.replace(tmp_path, cache_path) 144 return entries 145 146 147def _fetch_member(entry, out_path): 148 header = _range_get(ZIP_URL, entry["offset"], entry["offset"] + 29).content 149 name_len, extra_len = struct.unpack("<HH", header[26:30]) 150 start = entry["offset"] + 30 + name_len + extra_len 151 152 decompressor, crc, n_bytes = zlib.decompressobj(-15), 0, 0 153 tmp_path = f"{out_path}.{uuid.uuid4().hex}.tmp" 154 with _range_get(ZIP_URL, start, start + entry["comp_size"] - 1, stream=True) as response, \ 155 open(tmp_path, "wb") as f: 156 for chunk in response.iter_content(1 << 20): 157 data = decompressor.decompress(chunk) 158 crc, n_bytes = zlib.crc32(data, crc), n_bytes + len(data) 159 f.write(data) 160 data = decompressor.flush() 161 crc, n_bytes = zlib.crc32(data, crc), n_bytes + len(data) 162 f.write(data) 163 164 if crc != entry["crc"] or n_bytes != entry["size"]: 165 os.remove(tmp_path) 166 raise RuntimeError(f"The download of {os.path.basename(out_path)} is corrupted, please try again.") 167 os.replace(tmp_path, out_path) 168 169 170def _get_cases(entries): 171 pattern = re.compile(r"^images/(.+)\.ome\.tif$") 172 cases = {} 173 for name in entries: 174 match = pattern.match(name) 175 if match: 176 case = match.group(1) 177 mask_name = f"semantic_masks/{case}__mask_multiclass.ome.tif" 178 assert mask_name in entries, f"Cannot find the mask for {name}." 179 cases[case] = (name, mask_name) 180 return dict(sorted(cases.items())) 181 182 183def _open_level(series, level_index): 184 import zarr 185 186 # The pyramidal TIFFs are tiled, so the zarr view reads only the requested tiles. 187 array = zarr.open(series.aszarr(), mode="r") 188 return array if hasattr(array, "shape") else array[str(level_index)] 189 190 191def _convert_case(image_path, mask_path, out_path, level, tile=4096): 192 import h5py 193 import tifffile 194 import numpy as np 195 196 image_series = tifffile.TiffFile(image_path).series[0] 197 mask_series = tifffile.TiffFile(mask_path).series[0] 198 image, mask = _open_level(image_series, level), _open_level(mask_series, level) 199 200 # The pyramid levels of the image and mask are rounded differently, so the mask level can be a few pixels off. 201 # It is mapped to the image grid by nearest neighbor sampling, unless the shapes really differ. 202 height, width = image.shape[:2] 203 if any(abs(image.shape[i] - mask.shape[i]) > 1e-3 * image.shape[i] + 1 for i in range(2)): 204 raise RuntimeError(f"The image {image.shape} and mask {mask.shape} shapes of {image_path} do not match.") 205 rows = np.minimum(((np.arange(height) + 0.5) * mask.shape[0] / height).astype(int), mask.shape[0] - 1) 206 cols = np.minimum(((np.arange(width) + 0.5) * mask.shape[1] / width).astype(int), mask.shape[1] - 1) 207 208 tmp_path = f"{out_path}.{uuid.uuid4().hex}.tmp" 209 with h5py.File(tmp_path, "w") as f: 210 raw = f.create_dataset( 211 "raw", shape=(3, height, width), dtype="uint8", compression="gzip", chunks=(3, 512, 512) 212 ) 213 labels = f.create_dataset( 214 "labels", shape=(height, width), dtype="uint8", compression="gzip", chunks=(512, 512) 215 ) 216 for y in range(0, height, tile): 217 for x in range(0, width, tile): 218 bb = (slice(y, min(y + tile, height)), slice(x, min(x + tile, width))) 219 raw[(slice(None),) + bb] = image[bb].transpose(2, 0, 1) 220 r, c = rows[bb[0]], cols[bb[1]] 221 mask_block = mask[r[0]:r[-1] + 1, c[0]:c[-1] + 1] 222 labels[bb] = mask_block[np.ix_(r - r[0], c - c[0])] 223 os.replace(tmp_path, out_path) 224 225 226def get_prostate_cnb_data( 227 path: Union[os.PathLike, str], n_cases: Optional[int] = None, level: int = 1, download: bool = False, 228) -> str: 229 """Download and preprocess the Prostate CNB dataset. 230 231 NOTE: The full archive is 34.5 GB and is never downloaded as a whole, only the requested slides are fetched. 232 Use `n_cases` to only prepare a subset, e.g. a slide takes about 0.2 to 5.5 GB to download. 233 234 Args: 235 path: Filepath to a folder where the data is downloaded for further processing. 236 n_cases: The number of slides to use, sorted by slide id ('B<batch>-<patient>-<block>-<section>-<scan>'). 237 By default all 37 slides are used. 238 level: The pyramid level of the preprocessed data (0 to 5), with a pixel size of 0.243 * 2 ** level 239 micrometer. The preprocessed data is stored separately for each level. 240 download: Whether to download the data if it is not present. 241 242 Returns: 243 Filepath to the folder where the preprocessed data is stored. 244 """ 245 if level not in range(N_LEVELS): 246 raise ValueError(f"'{level}' is not a valid pyramid level. Choose a level from 0 to {N_LEVELS - 1}.") 247 248 os.makedirs(path, exist_ok=True) 249 if not download and not os.path.exists(os.path.join(path, "zip_entries.json")): 250 raise RuntimeError(f"Cannot find the data at {path}, but download was set to False.") 251 entries = _read_zip_entries(path) 252 cases = _get_cases(entries) 253 case_ids = list(cases)[:n_cases] 254 255 preprocessed_dir = os.path.join(path, "preprocessed", f"level{level}") 256 raw_dir = os.path.join(path, "raw") 257 os.makedirs(preprocessed_dir, exist_ok=True) 258 os.makedirs(raw_dir, exist_ok=True) 259 260 missing = [case for case in case_ids if not os.path.exists(os.path.join(preprocessed_dir, f"{case}.h5"))] 261 if missing and not download: 262 raise RuntimeError(f"Cannot find the data at {path}, but download was set to False.") 263 264 for case in tqdm(missing, desc="Prepare Prostate CNB slides"): 265 image_name, mask_name = cases[case] 266 image_path, mask_path = (os.path.join(raw_dir, os.path.basename(name)) for name in (image_name, mask_name)) 267 for name, out_path in ((image_name, image_path), (mask_name, mask_path)): 268 if not os.path.exists(out_path): 269 _fetch_member(entries[name], out_path) 270 271 _convert_case(image_path, mask_path, os.path.join(preprocessed_dir, f"{case}.h5"), level) 272 os.remove(image_path) 273 os.remove(mask_path) 274 275 return preprocessed_dir 276 277 278def get_prostate_cnb_paths( 279 path: Union[os.PathLike, str], n_cases: Optional[int] = None, level: int = 1, download: bool = False, 280) -> List[str]: 281 """Get paths to the Prostate CNB data. 282 283 Args: 284 path: Filepath to a folder where the data is downloaded for further processing. 285 n_cases: The number of slides to use, sorted by slide id ('B<batch>-<patient>-<block>-<section>-<scan>'). 286 By default all 37 slides are used. 287 level: The pyramid level of the preprocessed data (0 to 5), with a pixel size of 0.243 * 2 ** level 288 micrometer. 289 download: Whether to download the data if it is not present. 290 291 Returns: 292 List of filepaths to the preprocessed HDF5 files, which contain the image data ('raw') and the 293 label data ('labels'). 294 """ 295 preprocessed_dir = get_prostate_cnb_data(path, n_cases, level, download) 296 case_ids = list(_get_cases(_read_zip_entries(path)))[:n_cases] 297 return [os.path.join(preprocessed_dir, f"{case}.h5") for case in case_ids] 298 299 300def get_prostate_cnb_dataset( 301 path: Union[os.PathLike, str], 302 patch_shape: Tuple[int, int], 303 n_cases: Optional[int] = None, 304 level: int = 1, 305 download: bool = False, 306 label_dtype: torch.dtype = torch.int64, 307 resize_inputs: bool = False, 308 **kwargs 309) -> Dataset: 310 """Get the Prostate CNB dataset for semantic segmentation of prostate tissue in whole-slide images. 311 312 Args: 313 path: Filepath to a folder where the data is downloaded for further processing. 314 patch_shape: The patch shape to use for training. 315 n_cases: The number of slides to use, sorted by slide id ('B<batch>-<patient>-<block>-<section>-<scan>'). 316 By default all 37 slides are used. 317 level: The pyramid level of the preprocessed data (0 to 5), with a pixel size of 0.243 * 2 ** level 318 micrometer. 319 download: Whether to download the data if it is not present. 320 label_dtype: The datatype of the labels. 321 resize_inputs: Whether to resize the input images. 322 kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset`. 323 324 Returns: 325 The segmentation dataset. 326 """ 327 volume_paths = get_prostate_cnb_paths(path, n_cases, level, download) 328 329 if resize_inputs: 330 resize_kwargs = {"patch_shape": patch_shape, "is_rgb": True} 331 kwargs, patch_shape = util.update_kwargs_for_resize_trafo( 332 kwargs=kwargs, patch_shape=patch_shape, resize_inputs=resize_inputs, resize_kwargs=resize_kwargs 333 ) 334 335 return torch_em.default_segmentation_dataset( 336 raw_paths=volume_paths, 337 raw_key="raw", 338 label_paths=volume_paths, 339 label_key="labels", 340 patch_shape=patch_shape, 341 label_dtype=label_dtype, 342 is_seg_dataset=True, 343 with_channels=True, 344 ndim=2, 345 **kwargs 346 ) 347 348 349def get_prostate_cnb_loader( 350 path: Union[os.PathLike, str], 351 batch_size: int, 352 patch_shape: Tuple[int, int], 353 n_cases: Optional[int] = None, 354 level: int = 1, 355 download: bool = False, 356 label_dtype: torch.dtype = torch.int64, 357 resize_inputs: bool = False, 358 **kwargs 359) -> DataLoader: 360 """Get the Prostate CNB dataloader for semantic segmentation of prostate tissue in whole-slide images. 361 362 Args: 363 path: Filepath to a folder where the data is downloaded for further processing. 364 batch_size: The batch size for training. 365 patch_shape: The patch shape to use for training. 366 n_cases: The number of slides to use, sorted by slide id ('B<batch>-<patient>-<block>-<section>-<scan>'). 367 By default all 37 slides are used. 368 level: The pyramid level of the preprocessed data (0 to 5), with a pixel size of 0.243 * 2 ** level 369 micrometer. 370 download: Whether to download the data if it is not present. 371 label_dtype: The datatype of the labels. 372 resize_inputs: Whether to resize the input images. 373 kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset` or for the PyTorch DataLoader. 374 375 Returns: 376 The DataLoader. 377 """ 378 ds_kwargs, loader_kwargs = util.split_kwargs(torch_em.default_segmentation_dataset, **kwargs) 379 dataset = get_prostate_cnb_dataset( 380 path=path, patch_shape=patch_shape, n_cases=n_cases, level=level, download=download, 381 label_dtype=label_dtype, resize_inputs=resize_inputs, **ds_kwargs 382 ) 383 return torch_em.get_data_loader(dataset, batch_size, **loader_kwargs)
227def get_prostate_cnb_data( 228 path: Union[os.PathLike, str], n_cases: Optional[int] = None, level: int = 1, download: bool = False, 229) -> str: 230 """Download and preprocess the Prostate CNB dataset. 231 232 NOTE: The full archive is 34.5 GB and is never downloaded as a whole, only the requested slides are fetched. 233 Use `n_cases` to only prepare a subset, e.g. a slide takes about 0.2 to 5.5 GB to download. 234 235 Args: 236 path: Filepath to a folder where the data is downloaded for further processing. 237 n_cases: The number of slides to use, sorted by slide id ('B<batch>-<patient>-<block>-<section>-<scan>'). 238 By default all 37 slides are used. 239 level: The pyramid level of the preprocessed data (0 to 5), with a pixel size of 0.243 * 2 ** level 240 micrometer. The preprocessed data is stored separately for each level. 241 download: Whether to download the data if it is not present. 242 243 Returns: 244 Filepath to the folder where the preprocessed data is stored. 245 """ 246 if level not in range(N_LEVELS): 247 raise ValueError(f"'{level}' is not a valid pyramid level. Choose a level from 0 to {N_LEVELS - 1}.") 248 249 os.makedirs(path, exist_ok=True) 250 if not download and not os.path.exists(os.path.join(path, "zip_entries.json")): 251 raise RuntimeError(f"Cannot find the data at {path}, but download was set to False.") 252 entries = _read_zip_entries(path) 253 cases = _get_cases(entries) 254 case_ids = list(cases)[:n_cases] 255 256 preprocessed_dir = os.path.join(path, "preprocessed", f"level{level}") 257 raw_dir = os.path.join(path, "raw") 258 os.makedirs(preprocessed_dir, exist_ok=True) 259 os.makedirs(raw_dir, exist_ok=True) 260 261 missing = [case for case in case_ids if not os.path.exists(os.path.join(preprocessed_dir, f"{case}.h5"))] 262 if missing and not download: 263 raise RuntimeError(f"Cannot find the data at {path}, but download was set to False.") 264 265 for case in tqdm(missing, desc="Prepare Prostate CNB slides"): 266 image_name, mask_name = cases[case] 267 image_path, mask_path = (os.path.join(raw_dir, os.path.basename(name)) for name in (image_name, mask_name)) 268 for name, out_path in ((image_name, image_path), (mask_name, mask_path)): 269 if not os.path.exists(out_path): 270 _fetch_member(entries[name], out_path) 271 272 _convert_case(image_path, mask_path, os.path.join(preprocessed_dir, f"{case}.h5"), level) 273 os.remove(image_path) 274 os.remove(mask_path) 275 276 return preprocessed_dir
Download and preprocess the Prostate CNB dataset.
NOTE: The full archive is 34.5 GB and is never downloaded as a whole, only the requested slides are fetched.
Use n_cases to only prepare a subset, e.g. a slide takes about 0.2 to 5.5 GB to download.
Arguments:
- path: Filepath to a folder where the data is downloaded for further processing.
- n_cases: The number of slides to use, sorted by slide id ('B
- - - - '). By default all 37 slides are used. - level: The pyramid level of the preprocessed data (0 to 5), with a pixel size of 0.243 * 2 ** level micrometer. The preprocessed data is stored separately for each level.
- download: Whether to download the data if it is not present.
Returns:
Filepath to the folder where the preprocessed data is stored.
279def get_prostate_cnb_paths( 280 path: Union[os.PathLike, str], n_cases: Optional[int] = None, level: int = 1, download: bool = False, 281) -> List[str]: 282 """Get paths to the Prostate CNB data. 283 284 Args: 285 path: Filepath to a folder where the data is downloaded for further processing. 286 n_cases: The number of slides to use, sorted by slide id ('B<batch>-<patient>-<block>-<section>-<scan>'). 287 By default all 37 slides are used. 288 level: The pyramid level of the preprocessed data (0 to 5), with a pixel size of 0.243 * 2 ** level 289 micrometer. 290 download: Whether to download the data if it is not present. 291 292 Returns: 293 List of filepaths to the preprocessed HDF5 files, which contain the image data ('raw') and the 294 label data ('labels'). 295 """ 296 preprocessed_dir = get_prostate_cnb_data(path, n_cases, level, download) 297 case_ids = list(_get_cases(_read_zip_entries(path)))[:n_cases] 298 return [os.path.join(preprocessed_dir, f"{case}.h5") for case in case_ids]
Get paths to the Prostate CNB data.
Arguments:
- path: Filepath to a folder where the data is downloaded for further processing.
- n_cases: The number of slides to use, sorted by slide id ('B
- - - - '). By default all 37 slides are used. - level: The pyramid level of the preprocessed data (0 to 5), with a pixel size of 0.243 * 2 ** level micrometer.
- download: Whether to download the data if it is not present.
Returns:
List of filepaths to the preprocessed HDF5 files, which contain the image data ('raw') and the label data ('labels').
301def get_prostate_cnb_dataset( 302 path: Union[os.PathLike, str], 303 patch_shape: Tuple[int, int], 304 n_cases: Optional[int] = None, 305 level: int = 1, 306 download: bool = False, 307 label_dtype: torch.dtype = torch.int64, 308 resize_inputs: bool = False, 309 **kwargs 310) -> Dataset: 311 """Get the Prostate CNB dataset for semantic segmentation of prostate tissue in whole-slide images. 312 313 Args: 314 path: Filepath to a folder where the data is downloaded for further processing. 315 patch_shape: The patch shape to use for training. 316 n_cases: The number of slides to use, sorted by slide id ('B<batch>-<patient>-<block>-<section>-<scan>'). 317 By default all 37 slides are used. 318 level: The pyramid level of the preprocessed data (0 to 5), with a pixel size of 0.243 * 2 ** level 319 micrometer. 320 download: Whether to download the data if it is not present. 321 label_dtype: The datatype of the labels. 322 resize_inputs: Whether to resize the input images. 323 kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset`. 324 325 Returns: 326 The segmentation dataset. 327 """ 328 volume_paths = get_prostate_cnb_paths(path, n_cases, level, download) 329 330 if resize_inputs: 331 resize_kwargs = {"patch_shape": patch_shape, "is_rgb": True} 332 kwargs, patch_shape = util.update_kwargs_for_resize_trafo( 333 kwargs=kwargs, patch_shape=patch_shape, resize_inputs=resize_inputs, resize_kwargs=resize_kwargs 334 ) 335 336 return torch_em.default_segmentation_dataset( 337 raw_paths=volume_paths, 338 raw_key="raw", 339 label_paths=volume_paths, 340 label_key="labels", 341 patch_shape=patch_shape, 342 label_dtype=label_dtype, 343 is_seg_dataset=True, 344 with_channels=True, 345 ndim=2, 346 **kwargs 347 )
Get the Prostate CNB dataset for semantic segmentation of prostate tissue in whole-slide images.
Arguments:
- path: Filepath to a folder where the data is downloaded for further processing.
- patch_shape: The patch shape to use for training.
- n_cases: The number of slides to use, sorted by slide id ('B
- - - - '). By default all 37 slides are used. - level: The pyramid level of the preprocessed data (0 to 5), with a pixel size of 0.243 * 2 ** level micrometer.
- download: Whether to download the data if it is not present.
- label_dtype: The datatype of the labels.
- resize_inputs: Whether to resize the input images.
- kwargs: Additional keyword arguments for
torch_em.default_segmentation_dataset.
Returns:
The segmentation dataset.
350def get_prostate_cnb_loader( 351 path: Union[os.PathLike, str], 352 batch_size: int, 353 patch_shape: Tuple[int, int], 354 n_cases: Optional[int] = None, 355 level: int = 1, 356 download: bool = False, 357 label_dtype: torch.dtype = torch.int64, 358 resize_inputs: bool = False, 359 **kwargs 360) -> DataLoader: 361 """Get the Prostate CNB dataloader for semantic segmentation of prostate tissue in whole-slide images. 362 363 Args: 364 path: Filepath to a folder where the data is downloaded for further processing. 365 batch_size: The batch size for training. 366 patch_shape: The patch shape to use for training. 367 n_cases: The number of slides to use, sorted by slide id ('B<batch>-<patient>-<block>-<section>-<scan>'). 368 By default all 37 slides are used. 369 level: The pyramid level of the preprocessed data (0 to 5), with a pixel size of 0.243 * 2 ** level 370 micrometer. 371 download: Whether to download the data if it is not present. 372 label_dtype: The datatype of the labels. 373 resize_inputs: Whether to resize the input images. 374 kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset` or for the PyTorch DataLoader. 375 376 Returns: 377 The DataLoader. 378 """ 379 ds_kwargs, loader_kwargs = util.split_kwargs(torch_em.default_segmentation_dataset, **kwargs) 380 dataset = get_prostate_cnb_dataset( 381 path=path, patch_shape=patch_shape, n_cases=n_cases, level=level, download=download, 382 label_dtype=label_dtype, resize_inputs=resize_inputs, **ds_kwargs 383 ) 384 return torch_em.get_data_loader(dataset, batch_size, **loader_kwargs)
Get the Prostate CNB dataloader for semantic segmentation of prostate tissue in whole-slide images.
Arguments:
- path: Filepath to a folder where the data is downloaded for further processing.
- batch_size: The batch size for training.
- patch_shape: The patch shape to use for training.
- n_cases: The number of slides to use, sorted by slide id ('B
- - - - '). By default all 37 slides are used. - level: The pyramid level of the preprocessed data (0 to 5), with a pixel size of 0.243 * 2 ** level micrometer.
- download: Whether to download the data if it is not present.
- label_dtype: The datatype of the labels.
- resize_inputs: Whether to resize the input images.
- kwargs: Additional keyword arguments for
torch_em.default_segmentation_datasetor for the PyTorch DataLoader.
Returns:
The DataLoader.