torch_em.data.datasets.histopathology.precise
The PRECISE dataset contains annotations for the semantic segmentation of prostate tissue and lesions in paired H&E and immunohistochemistry (IHC) whole-slide images of prostate core needle biopsies.
PRECISE (PRostate Expert-annotated Contiguous IHC-H&E Serial sEctions) consists of 54 slide pairs from 25 patients (sub-01 has three sessions, all other patients one). Each pair has a H&E and a HMWCK-AMACR (CKAPM + racemase) IHC whole-slide image, scanned at 0.243 micrometer per pixel with a 3DHISTECH Pannoramic scanner, and a pixel-level mask for each stain, drawn by expert uropathologists (24,387 annotations in total).
The label ids of the masks are:
- 0: background
- 1: tumor
- 2: benign gland
- 3: artifact
- 4: high-grade prostatic intraepithelial neoplasia (HGPIN)
- 5: intraductal carcinoma
- 6: atypical intraductal proliferation
- 7: stroma Not every class occurs in every slide.
The data is located at https://doi.org/10.5281/zenodo.20721779 as a single 55.6 GB zip archive with pyramidal
OME-TIFF images and masks, released under a CC BY 4.0 license. To avoid downloading the whole archive, this module reads
its directory from the server and fetches only the requested slides and masks. Each slide is converted once into a
chunked HDF5 file at a chosen pyramid level (pixel size 0.243 * 2 ** level micrometer, the default level 1 is
0.486 micrometer), the downloaded OME-TIFFs are removed afterwards. Use n_cases to only prepare a subset.
The publication for this dataset was not available at the time of implementation, so please cite the Zenodo record if you use it in your research.
1"""The PRECISE dataset contains annotations for the semantic segmentation of prostate tissue and lesions 2in paired H&E and immunohistochemistry (IHC) whole-slide images of prostate core needle biopsies. 3 4PRECISE (PRostate Expert-annotated Contiguous IHC-H&E Serial sEctions) consists of 54 slide pairs from 25 patients 5(sub-01 has three sessions, all other patients one). Each pair has a H&E and a HMWCK-AMACR (CKAPM + racemase) IHC 6whole-slide image, scanned at 0.243 micrometer per pixel with a 3DHISTECH Pannoramic scanner, and a pixel-level 7mask for each stain, drawn by expert uropathologists (24,387 annotations in total). 8 9The label ids of the masks are: 10- 0: background 11- 1: tumor 12- 2: benign gland 13- 3: artifact 14- 4: high-grade prostatic intraepithelial neoplasia (HGPIN) 15- 5: intraductal carcinoma 16- 6: atypical intraductal proliferation 17- 7: stroma 18Not every class occurs in every slide. 19 20The data is located at https://doi.org/10.5281/zenodo.20721779 as a single 55.6 GB zip archive with pyramidal 21OME-TIFF images and masks, released under a CC BY 4.0 license. To avoid downloading the whole archive, this module reads 22its directory from the server and fetches only the requested slides and masks. Each slide is converted once into a 23chunked HDF5 file at a chosen pyramid level (pixel size 0.243 * 2 ** level micrometer, the default level 1 is 240.486 micrometer), the downloaded OME-TIFFs are removed afterwards. Use `n_cases` to only prepare a subset. 25 26The publication for this dataset was not available at the time of implementation, so please cite the Zenodo record 27if you use it in your research. 28""" 29 30import os 31import re 32import json 33import uuid 34import zlib 35import struct 36from typing import List, Literal, Optional, Tuple, Union 37 38from tqdm import tqdm 39 40import torch 41 42from torch.utils.data import Dataset, DataLoader 43 44import torch_em 45 46from .. import util 47 48 49RECORD_URL = "https://zenodo.org/api/records/20721779/files" 50ZIP_URL = f"{RECORD_URL}/data.zip/content" 51 52SMALL_FILES = { 53 "label_descriptions.json": "31f9687f2611fce5a56e5f983c9a97b6ae0f5d1b6171be65e9bbc961ad2ec11c", 54 "participants.csv": "8b5f18120b83e84eb235eb746eb005f8dbbd0b9d46c30d248c90c0a7b31cdcaf", 55} 56 57STAINS = {"he": "h-e", "ihc": "hmwck-amacr"} 58 59CLASS_NAMES = { 60 0: "background", 61 1: "tumor", 62 2: "benign gland", 63 3: "artifact", 64 4: "high-grade prostatic intraepithelial neoplasia", 65 5: "intraductal carcinoma", 66 6: "atypical intraductal proliferation", 67 7: "stroma", 68} 69 70N_LEVELS = 6 71 72 73def _range_get(url, start, end, **kwargs): 74 import requests 75 76 response = requests.get(url, headers={"Range": f"bytes={start}-{end}"}, **kwargs) 77 response.raise_for_status() 78 return response 79 80 81def _read_zip_entries(path): 82 cache_path = os.path.join(path, "zip_entries.json") 83 if os.path.exists(cache_path): 84 with open(cache_path) as f: 85 return json.load(f) 86 87 import requests 88 89 size = int(requests.head(ZIP_URL, allow_redirects=True).headers["Content-Length"]) 90 tail = _range_get(ZIP_URL, size - 65557, size - 1).content 91 eocd = tail[tail.rfind(b"PK\x05\x06"):] 92 n_entries, cd_size, cd_offset = struct.unpack("<HII", eocd[10:20]) 93 if cd_offset == 0xFFFFFFFF or n_entries == 0xFFFF: 94 locator = tail[tail.rfind(b"PK\x06\x07"):][:20] 95 zip64_offset = struct.unpack("<Q", locator[8:16])[0] 96 record = _range_get(ZIP_URL, zip64_offset, zip64_offset + 55).content 97 n_entries, cd_size, cd_offset = struct.unpack("<QQQ", record[32:56]) 98 99 directory = _range_get(ZIP_URL, cd_offset, cd_offset + cd_size - 1).content 100 entries, pos = {}, 0 101 while pos < len(directory): 102 (signature, _, _, _, _, _, _, crc, comp_size, size, name_len, extra_len, comment_len, _, _, _, 103 header_offset) = struct.unpack("<IHHHHHHIIIHHHHHII", directory[pos:pos + 46]) 104 assert signature == 0x02014B50, "The zip directory could not be parsed." 105 name = directory[pos + 46:pos + 46 + name_len].decode() 106 extra = directory[pos + 46 + name_len:pos + 46 + name_len + extra_len] 107 extra_pos = 0 108 while extra_pos < len(extra): 109 field_id, field_size = struct.unpack("<HH", extra[extra_pos:extra_pos + 4]) 110 if field_id == 1: 111 field, k = extra[extra_pos + 4:extra_pos + 4 + field_size], 0 112 if size == 0xFFFFFFFF: 113 size, k = struct.unpack("<Q", field[k:k + 8])[0], k + 8 114 if comp_size == 0xFFFFFFFF: 115 comp_size, k = struct.unpack("<Q", field[k:k + 8])[0], k + 8 116 if header_offset == 0xFFFFFFFF: 117 header_offset = struct.unpack("<Q", field[k:k + 8])[0] 118 extra_pos += 4 + field_size 119 if size > 0: 120 entries[name] = {"crc": crc, "comp_size": comp_size, "size": size, "offset": header_offset} 121 pos += 46 + name_len + extra_len + comment_len 122 123 tmp_path = f"{cache_path}.{uuid.uuid4().hex}.tmp" 124 with open(tmp_path, "w") as f: 125 json.dump(entries, f) 126 os.replace(tmp_path, cache_path) 127 return entries 128 129 130def _fetch_member(entry, out_path): 131 header = _range_get(ZIP_URL, entry["offset"], entry["offset"] + 29).content 132 name_len, extra_len = struct.unpack("<HH", header[26:30]) 133 start = entry["offset"] + 30 + name_len + extra_len 134 135 decompressor, crc, n_bytes = zlib.decompressobj(-15), 0, 0 136 tmp_path = f"{out_path}.{uuid.uuid4().hex}.tmp" 137 with _range_get(ZIP_URL, start, start + entry["comp_size"] - 1, stream=True) as response, \ 138 open(tmp_path, "wb") as f: 139 for chunk in response.iter_content(1 << 20): 140 data = decompressor.decompress(chunk) 141 crc, n_bytes = zlib.crc32(data, crc), n_bytes + len(data) 142 f.write(data) 143 data = decompressor.flush() 144 crc, n_bytes = zlib.crc32(data, crc), n_bytes + len(data) 145 f.write(data) 146 147 if crc != entry["crc"] or n_bytes != entry["size"]: 148 os.remove(tmp_path) 149 raise RuntimeError(f"The download of {os.path.basename(out_path)} is corrupted, please try again.") 150 os.replace(tmp_path, out_path) 151 152 153def _get_cases(entries, stain): 154 stain_dir = STAINS[stain] 155 stain = re.escape(stain_dir) 156 pattern = re.compile(rf"^data/sub-\d+/ses-\d+/wsi_{stain}/(sub-\d+_ses-\d+)_{stain}\.ome\.tif$") 157 cases = {} 158 for name in entries: 159 match = pattern.match(name) 160 if match: 161 case = match.group(1) 162 mask_name = name.replace(".ome.tif", "_mask.ome.tif") 163 assert mask_name in entries, f"Cannot find the mask for {name}." 164 cases[case] = (name, mask_name) 165 return dict(sorted(cases.items())) 166 167 168def _open_level(series, level_index): 169 import zarr 170 171 # The pyramidal TIFFs are tiled, so the zarr view reads only the requested tiles. 172 array = zarr.open(series.aszarr(), mode="r") 173 return array if hasattr(array, "shape") else array[str(level_index)] 174 175 176def _convert_case(image_path, mask_path, out_path, level, tile=4096): 177 import h5py 178 import tifffile 179 import numpy as np 180 181 image_series = tifffile.TiffFile(image_path).series[0] 182 mask_series = tifffile.TiffFile(mask_path).series[0] 183 image, mask = _open_level(image_series, level), _open_level(mask_series, level) 184 185 # The pyramid levels of the image and mask are rounded differently, so the mask level can be a few pixels off. 186 # It is mapped to the image grid by nearest neighbor sampling, unless the shapes really differ. 187 height, width = image.shape[:2] 188 if any(abs(image.shape[i] - mask.shape[i]) > 1e-3 * image.shape[i] + 1 for i in range(2)): 189 raise RuntimeError(f"The image {image.shape} and mask {mask.shape} shapes of {image_path} do not match.") 190 rows = np.minimum(((np.arange(height) + 0.5) * mask.shape[0] / height).astype(int), mask.shape[0] - 1) 191 cols = np.minimum(((np.arange(width) + 0.5) * mask.shape[1] / width).astype(int), mask.shape[1] - 1) 192 193 tmp_path = f"{out_path}.{uuid.uuid4().hex}.tmp" 194 with h5py.File(tmp_path, "w") as f: 195 raw = f.create_dataset( 196 "raw", shape=(3, height, width), dtype="uint8", compression="gzip", chunks=(3, 512, 512) 197 ) 198 labels = f.create_dataset( 199 "labels", shape=(height, width), dtype="uint8", compression="gzip", chunks=(512, 512) 200 ) 201 for y in range(0, height, tile): 202 for x in range(0, width, tile): 203 bb = (slice(y, min(y + tile, height)), slice(x, min(x + tile, width))) 204 raw[(slice(None),) + bb] = image[bb].transpose(2, 0, 1) 205 r, c = rows[bb[0]], cols[bb[1]] 206 mask_block = mask[r[0]:r[-1] + 1, c[0]:c[-1] + 1] 207 labels[bb] = mask_block[np.ix_(r - r[0], c - c[0])] 208 os.replace(tmp_path, out_path) 209 210 211def get_precise_data( 212 path: Union[os.PathLike, str], 213 stain: Literal["he", "ihc"] = "he", 214 n_cases: Optional[int] = None, 215 level: int = 1, 216 download: bool = False, 217) -> str: 218 """Download and preprocess the PRECISE dataset. 219 220 NOTE: The full archive is 55.6 GB and is never downloaded as a whole, only the requested slides are fetched. 221 Use `n_cases` to only prepare a subset, e.g. a slide takes about 0.5 to 1.4 GB to download. 222 223 Args: 224 path: Filepath to a folder where the data is downloaded for further processing. 225 stain: The choice of stain. Either 'he' (H&E) or 'ihc' (HMWCK-AMACR IHC). 226 n_cases: The number of slides to use, sorted by slide id ('sub-<patient>_ses-<session>'). 227 By default all 54 slides of the stain are used. 228 level: The pyramid level of the preprocessed data (0 to 5), with a pixel size of 0.243 * 2 ** level 229 micrometer. The preprocessed data is stored separately for each level. 230 download: Whether to download the data if it is not present. 231 232 Returns: 233 Filepath to the folder where the preprocessed data is stored. 234 """ 235 if stain not in STAINS: 236 raise ValueError(f"'{stain}' is not a valid stain. Choose one of {list(STAINS)}.") 237 if level not in range(N_LEVELS): 238 raise ValueError(f"'{level}' is not a valid pyramid level. Choose a level from 0 to {N_LEVELS - 1}.") 239 240 os.makedirs(path, exist_ok=True) 241 for filename, checksum in SMALL_FILES.items(): 242 util.download_source( 243 path=os.path.join(path, filename), url=f"{RECORD_URL}/{filename}/content", download=download, 244 checksum=checksum, 245 ) 246 247 if not download and not os.path.exists(os.path.join(path, "zip_entries.json")): 248 raise RuntimeError(f"Cannot find the data at {path}, but download was set to False.") 249 entries = _read_zip_entries(path) 250 cases = _get_cases(entries, stain) 251 case_ids = list(cases)[:n_cases] 252 253 preprocessed_dir = os.path.join(path, "preprocessed", f"level{level}", stain) 254 raw_dir = os.path.join(path, "raw", stain) 255 os.makedirs(preprocessed_dir, exist_ok=True) 256 os.makedirs(raw_dir, exist_ok=True) 257 258 missing = [case for case in case_ids if not os.path.exists(os.path.join(preprocessed_dir, f"{case}.h5"))] 259 if missing and not download: 260 raise RuntimeError(f"Cannot find the data at {path}, but download was set to False.") 261 262 for case in tqdm(missing, desc="Prepare PRECISE slides"): 263 image_name, mask_name = cases[case] 264 image_path, mask_path = (os.path.join(raw_dir, os.path.basename(name)) for name in (image_name, mask_name)) 265 for name, out_path in ((image_name, image_path), (mask_name, mask_path)): 266 if not os.path.exists(out_path): 267 _fetch_member(entries[name], out_path) 268 269 _convert_case(image_path, mask_path, os.path.join(preprocessed_dir, f"{case}.h5"), level) 270 os.remove(image_path) 271 os.remove(mask_path) 272 273 return preprocessed_dir 274 275 276def get_precise_paths( 277 path: Union[os.PathLike, str], 278 stain: Literal["he", "ihc"] = "he", 279 n_cases: Optional[int] = None, 280 level: int = 1, 281 download: bool = False, 282) -> List[str]: 283 """Get paths to the PRECISE data. 284 285 Args: 286 path: Filepath to a folder where the data is downloaded for further processing. 287 stain: The choice of stain. Either 'he' (H&E) or 'ihc' (HMWCK-AMACR IHC). 288 n_cases: The number of slides to use, sorted by slide id ('sub-<patient>_ses-<session>'). 289 By default all 54 slides of the stain are used. 290 level: The pyramid level of the preprocessed data (0 to 5), with a pixel size of 0.243 * 2 ** level 291 micrometer. 292 download: Whether to download the data if it is not present. 293 294 Returns: 295 List of filepaths to the preprocessed HDF5 files, which contain the image data ('raw') and the 296 label data ('labels'). 297 """ 298 preprocessed_dir = get_precise_data(path, stain, n_cases, level, download) 299 case_ids = list(_get_cases(_read_zip_entries(path), stain))[:n_cases] 300 return [os.path.join(preprocessed_dir, f"{case}.h5") for case in case_ids] 301 302 303def get_precise_dataset( 304 path: Union[os.PathLike, str], 305 patch_shape: Tuple[int, int], 306 stain: Literal["he", "ihc"] = "he", 307 n_cases: Optional[int] = None, 308 level: int = 1, 309 download: bool = False, 310 label_dtype: torch.dtype = torch.int64, 311 resize_inputs: bool = False, 312 **kwargs 313) -> Dataset: 314 """Get the PRECISE dataset for semantic segmentation of prostate tissue and lesions in whole-slide images. 315 316 Args: 317 path: Filepath to a folder where the data is downloaded for further processing. 318 patch_shape: The patch shape to use for training. 319 stain: The choice of stain. Either 'he' (H&E) or 'ihc' (HMWCK-AMACR IHC). 320 n_cases: The number of slides to use, sorted by slide id ('sub-<patient>_ses-<session>'). 321 By default all 54 slides of the stain are used. 322 level: The pyramid level of the preprocessed data (0 to 5), with a pixel size of 0.243 * 2 ** level 323 micrometer. 324 download: Whether to download the data if it is not present. 325 label_dtype: The datatype of the labels. 326 resize_inputs: Whether to resize the input images. 327 kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset`. 328 329 Returns: 330 The segmentation dataset. 331 """ 332 volume_paths = get_precise_paths(path, stain, n_cases, level, download) 333 334 if resize_inputs: 335 resize_kwargs = {"patch_shape": patch_shape, "is_rgb": True} 336 kwargs, patch_shape = util.update_kwargs_for_resize_trafo( 337 kwargs=kwargs, patch_shape=patch_shape, resize_inputs=resize_inputs, resize_kwargs=resize_kwargs 338 ) 339 340 return torch_em.default_segmentation_dataset( 341 raw_paths=volume_paths, 342 raw_key="raw", 343 label_paths=volume_paths, 344 label_key="labels", 345 patch_shape=patch_shape, 346 label_dtype=label_dtype, 347 is_seg_dataset=True, 348 with_channels=True, 349 ndim=2, 350 **kwargs 351 ) 352 353 354def get_precise_loader( 355 path: Union[os.PathLike, str], 356 batch_size: int, 357 patch_shape: Tuple[int, int], 358 stain: Literal["he", "ihc"] = "he", 359 n_cases: Optional[int] = None, 360 level: int = 1, 361 download: bool = False, 362 label_dtype: torch.dtype = torch.int64, 363 resize_inputs: bool = False, 364 **kwargs 365) -> DataLoader: 366 """Get the PRECISE dataloader for semantic segmentation of prostate tissue and lesions in whole-slide images. 367 368 Args: 369 path: Filepath to a folder where the data is downloaded for further processing. 370 batch_size: The batch size for training. 371 patch_shape: The patch shape to use for training. 372 stain: The choice of stain. Either 'he' (H&E) or 'ihc' (HMWCK-AMACR IHC). 373 n_cases: The number of slides to use, sorted by slide id ('sub-<patient>_ses-<session>'). 374 By default all 54 slides of the stain are used. 375 level: The pyramid level of the preprocessed data (0 to 5), with a pixel size of 0.243 * 2 ** level 376 micrometer. 377 download: Whether to download the data if it is not present. 378 label_dtype: The datatype of the labels. 379 resize_inputs: Whether to resize the input images. 380 kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset` or for the PyTorch DataLoader. 381 382 Returns: 383 The DataLoader. 384 """ 385 ds_kwargs, loader_kwargs = util.split_kwargs(torch_em.default_segmentation_dataset, **kwargs) 386 dataset = get_precise_dataset( 387 path=path, patch_shape=patch_shape, stain=stain, n_cases=n_cases, level=level, download=download, 388 label_dtype=label_dtype, resize_inputs=resize_inputs, **ds_kwargs 389 ) 390 return torch_em.get_data_loader(dataset, batch_size, **loader_kwargs)
212def get_precise_data( 213 path: Union[os.PathLike, str], 214 stain: Literal["he", "ihc"] = "he", 215 n_cases: Optional[int] = None, 216 level: int = 1, 217 download: bool = False, 218) -> str: 219 """Download and preprocess the PRECISE dataset. 220 221 NOTE: The full archive is 55.6 GB and is never downloaded as a whole, only the requested slides are fetched. 222 Use `n_cases` to only prepare a subset, e.g. a slide takes about 0.5 to 1.4 GB to download. 223 224 Args: 225 path: Filepath to a folder where the data is downloaded for further processing. 226 stain: The choice of stain. Either 'he' (H&E) or 'ihc' (HMWCK-AMACR IHC). 227 n_cases: The number of slides to use, sorted by slide id ('sub-<patient>_ses-<session>'). 228 By default all 54 slides of the stain are used. 229 level: The pyramid level of the preprocessed data (0 to 5), with a pixel size of 0.243 * 2 ** level 230 micrometer. The preprocessed data is stored separately for each level. 231 download: Whether to download the data if it is not present. 232 233 Returns: 234 Filepath to the folder where the preprocessed data is stored. 235 """ 236 if stain not in STAINS: 237 raise ValueError(f"'{stain}' is not a valid stain. Choose one of {list(STAINS)}.") 238 if level not in range(N_LEVELS): 239 raise ValueError(f"'{level}' is not a valid pyramid level. Choose a level from 0 to {N_LEVELS - 1}.") 240 241 os.makedirs(path, exist_ok=True) 242 for filename, checksum in SMALL_FILES.items(): 243 util.download_source( 244 path=os.path.join(path, filename), url=f"{RECORD_URL}/{filename}/content", download=download, 245 checksum=checksum, 246 ) 247 248 if not download and not os.path.exists(os.path.join(path, "zip_entries.json")): 249 raise RuntimeError(f"Cannot find the data at {path}, but download was set to False.") 250 entries = _read_zip_entries(path) 251 cases = _get_cases(entries, stain) 252 case_ids = list(cases)[:n_cases] 253 254 preprocessed_dir = os.path.join(path, "preprocessed", f"level{level}", stain) 255 raw_dir = os.path.join(path, "raw", stain) 256 os.makedirs(preprocessed_dir, exist_ok=True) 257 os.makedirs(raw_dir, exist_ok=True) 258 259 missing = [case for case in case_ids if not os.path.exists(os.path.join(preprocessed_dir, f"{case}.h5"))] 260 if missing and not download: 261 raise RuntimeError(f"Cannot find the data at {path}, but download was set to False.") 262 263 for case in tqdm(missing, desc="Prepare PRECISE slides"): 264 image_name, mask_name = cases[case] 265 image_path, mask_path = (os.path.join(raw_dir, os.path.basename(name)) for name in (image_name, mask_name)) 266 for name, out_path in ((image_name, image_path), (mask_name, mask_path)): 267 if not os.path.exists(out_path): 268 _fetch_member(entries[name], out_path) 269 270 _convert_case(image_path, mask_path, os.path.join(preprocessed_dir, f"{case}.h5"), level) 271 os.remove(image_path) 272 os.remove(mask_path) 273 274 return preprocessed_dir
Download and preprocess the PRECISE dataset.
NOTE: The full archive is 55.6 GB and is never downloaded as a whole, only the requested slides are fetched.
Use n_cases to only prepare a subset, e.g. a slide takes about 0.5 to 1.4 GB to download.
Arguments:
- path: Filepath to a folder where the data is downloaded for further processing.
- stain: The choice of stain. Either 'he' (H&E) or 'ihc' (HMWCK-AMACR IHC).
- n_cases: The number of slides to use, sorted by slide id ('sub-
_ses- '). By default all 54 slides of the stain are used. - level: The pyramid level of the preprocessed data (0 to 5), with a pixel size of 0.243 * 2 ** level micrometer. The preprocessed data is stored separately for each level.
- download: Whether to download the data if it is not present.
Returns:
Filepath to the folder where the preprocessed data is stored.
277def get_precise_paths( 278 path: Union[os.PathLike, str], 279 stain: Literal["he", "ihc"] = "he", 280 n_cases: Optional[int] = None, 281 level: int = 1, 282 download: bool = False, 283) -> List[str]: 284 """Get paths to the PRECISE data. 285 286 Args: 287 path: Filepath to a folder where the data is downloaded for further processing. 288 stain: The choice of stain. Either 'he' (H&E) or 'ihc' (HMWCK-AMACR IHC). 289 n_cases: The number of slides to use, sorted by slide id ('sub-<patient>_ses-<session>'). 290 By default all 54 slides of the stain are used. 291 level: The pyramid level of the preprocessed data (0 to 5), with a pixel size of 0.243 * 2 ** level 292 micrometer. 293 download: Whether to download the data if it is not present. 294 295 Returns: 296 List of filepaths to the preprocessed HDF5 files, which contain the image data ('raw') and the 297 label data ('labels'). 298 """ 299 preprocessed_dir = get_precise_data(path, stain, n_cases, level, download) 300 case_ids = list(_get_cases(_read_zip_entries(path), stain))[:n_cases] 301 return [os.path.join(preprocessed_dir, f"{case}.h5") for case in case_ids]
Get paths to the PRECISE data.
Arguments:
- path: Filepath to a folder where the data is downloaded for further processing.
- stain: The choice of stain. Either 'he' (H&E) or 'ihc' (HMWCK-AMACR IHC).
- n_cases: The number of slides to use, sorted by slide id ('sub-
_ses- '). By default all 54 slides of the stain are used. - level: The pyramid level of the preprocessed data (0 to 5), with a pixel size of 0.243 * 2 ** level micrometer.
- download: Whether to download the data if it is not present.
Returns:
List of filepaths to the preprocessed HDF5 files, which contain the image data ('raw') and the label data ('labels').
304def get_precise_dataset( 305 path: Union[os.PathLike, str], 306 patch_shape: Tuple[int, int], 307 stain: Literal["he", "ihc"] = "he", 308 n_cases: Optional[int] = None, 309 level: int = 1, 310 download: bool = False, 311 label_dtype: torch.dtype = torch.int64, 312 resize_inputs: bool = False, 313 **kwargs 314) -> Dataset: 315 """Get the PRECISE dataset for semantic segmentation of prostate tissue and lesions in whole-slide images. 316 317 Args: 318 path: Filepath to a folder where the data is downloaded for further processing. 319 patch_shape: The patch shape to use for training. 320 stain: The choice of stain. Either 'he' (H&E) or 'ihc' (HMWCK-AMACR IHC). 321 n_cases: The number of slides to use, sorted by slide id ('sub-<patient>_ses-<session>'). 322 By default all 54 slides of the stain are used. 323 level: The pyramid level of the preprocessed data (0 to 5), with a pixel size of 0.243 * 2 ** level 324 micrometer. 325 download: Whether to download the data if it is not present. 326 label_dtype: The datatype of the labels. 327 resize_inputs: Whether to resize the input images. 328 kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset`. 329 330 Returns: 331 The segmentation dataset. 332 """ 333 volume_paths = get_precise_paths(path, stain, n_cases, level, download) 334 335 if resize_inputs: 336 resize_kwargs = {"patch_shape": patch_shape, "is_rgb": True} 337 kwargs, patch_shape = util.update_kwargs_for_resize_trafo( 338 kwargs=kwargs, patch_shape=patch_shape, resize_inputs=resize_inputs, resize_kwargs=resize_kwargs 339 ) 340 341 return torch_em.default_segmentation_dataset( 342 raw_paths=volume_paths, 343 raw_key="raw", 344 label_paths=volume_paths, 345 label_key="labels", 346 patch_shape=patch_shape, 347 label_dtype=label_dtype, 348 is_seg_dataset=True, 349 with_channels=True, 350 ndim=2, 351 **kwargs 352 )
Get the PRECISE dataset for semantic segmentation of prostate tissue and lesions in whole-slide images.
Arguments:
- path: Filepath to a folder where the data is downloaded for further processing.
- patch_shape: The patch shape to use for training.
- stain: The choice of stain. Either 'he' (H&E) or 'ihc' (HMWCK-AMACR IHC).
- n_cases: The number of slides to use, sorted by slide id ('sub-
_ses- '). By default all 54 slides of the stain are used. - level: The pyramid level of the preprocessed data (0 to 5), with a pixel size of 0.243 * 2 ** level micrometer.
- download: Whether to download the data if it is not present.
- label_dtype: The datatype of the labels.
- resize_inputs: Whether to resize the input images.
- kwargs: Additional keyword arguments for
torch_em.default_segmentation_dataset.
Returns:
The segmentation dataset.
355def get_precise_loader( 356 path: Union[os.PathLike, str], 357 batch_size: int, 358 patch_shape: Tuple[int, int], 359 stain: Literal["he", "ihc"] = "he", 360 n_cases: Optional[int] = None, 361 level: int = 1, 362 download: bool = False, 363 label_dtype: torch.dtype = torch.int64, 364 resize_inputs: bool = False, 365 **kwargs 366) -> DataLoader: 367 """Get the PRECISE dataloader for semantic segmentation of prostate tissue and lesions in whole-slide images. 368 369 Args: 370 path: Filepath to a folder where the data is downloaded for further processing. 371 batch_size: The batch size for training. 372 patch_shape: The patch shape to use for training. 373 stain: The choice of stain. Either 'he' (H&E) or 'ihc' (HMWCK-AMACR IHC). 374 n_cases: The number of slides to use, sorted by slide id ('sub-<patient>_ses-<session>'). 375 By default all 54 slides of the stain are used. 376 level: The pyramid level of the preprocessed data (0 to 5), with a pixel size of 0.243 * 2 ** level 377 micrometer. 378 download: Whether to download the data if it is not present. 379 label_dtype: The datatype of the labels. 380 resize_inputs: Whether to resize the input images. 381 kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset` or for the PyTorch DataLoader. 382 383 Returns: 384 The DataLoader. 385 """ 386 ds_kwargs, loader_kwargs = util.split_kwargs(torch_em.default_segmentation_dataset, **kwargs) 387 dataset = get_precise_dataset( 388 path=path, patch_shape=patch_shape, stain=stain, n_cases=n_cases, level=level, download=download, 389 label_dtype=label_dtype, resize_inputs=resize_inputs, **ds_kwargs 390 ) 391 return torch_em.get_data_loader(dataset, batch_size, **loader_kwargs)
Get the PRECISE dataloader for semantic segmentation of prostate tissue and lesions in whole-slide images.
Arguments:
- path: Filepath to a folder where the data is downloaded for further processing.
- batch_size: The batch size for training.
- patch_shape: The patch shape to use for training.
- stain: The choice of stain. Either 'he' (H&E) or 'ihc' (HMWCK-AMACR IHC).
- n_cases: The number of slides to use, sorted by slide id ('sub-
_ses- '). By default all 54 slides of the stain are used. - level: The pyramid level of the preprocessed data (0 to 5), with a pixel size of 0.243 * 2 ** level micrometer.
- download: Whether to download the data if it is not present.
- label_dtype: The datatype of the labels.
- resize_inputs: Whether to resize the input images.
- kwargs: Additional keyword arguments for
torch_em.default_segmentation_datasetor for the PyTorch DataLoader.
Returns:
The DataLoader.