torch_em.data.datasets.light_microscopy.xenium_lung_treg
This dataset contains cell and nucleus segmentation masks for whole-slide fluorescence images of mouse lung tissue, from the BioImage Archive study S-BIAD3880.
The study images FFPE sections of adult female FVB/N mouse lung with the 10x Genomics
Xenium Prime 5K platform, to map the preinvasive regulatory T cell niches that form around
NTCU-induced airway lesions and how the PI3K-delta inhibitor PI-3065 changes them. It holds
8 samples: 2 untreated controls, 3 NTCU-treated mice, and 3 NTCU-treated mice that also got
PI-3065. See TREATMENT for which group each sample belongs to.
Every sample pairs a set of morphology images with the masks that the Xenium Onboard Analysis (XOA) pipeline produced. Channel 0 always holds the DAPI nuclear stain. The other channels hold the boundary and interior stains of the multi-tissue stain mix, which XOA uses to draw the cell outlines.
NOTE: The masks are the output of XOA, not a manual annotation. XOA segments the nuclei on DAPI, then grows each cell from its nucleus using the boundary and interior stains. Where no stain supports a boundary it falls back to expanding the nucleus by 5 micrometer, so a share of the cells are dilated nuclei rather than observed outlines. The 'nuc_expansion_fraction' attribute of each h5 file reports that share. Treat this data as a weak label source for pretraining, not as a benchmark.
NOTE: A cell can hold more than one nucleus, so the two label sets do not match one to one and their ids do not correspond.
NOTE: This dataset is distinct from the Xenium samples in torch_em.data.datasets.light_microscopy.xenium,
which are hosted directly by 10x Genomics and cover human and mouse tissue other than lung.
NOTE: The images are whole slides that hold large empty regions. Pass a sampler such as
torch_em.data.MinInstanceSampler to avoid drawing empty patches.
The data is hosted at https://www.ebi.ac.uk/biostudies/bioimages/studies/S-BIAD3880 under the CC0 1.0 license. Please cite the corresponding study page if you use this data in your research.
1"""This dataset contains cell and nucleus segmentation masks for whole-slide fluorescence 2images of mouse lung tissue, from the BioImage Archive study S-BIAD3880. 3 4The study images FFPE sections of adult female FVB/N mouse lung with the 10x Genomics 5Xenium Prime 5K platform, to map the preinvasive regulatory T cell niches that form around 6NTCU-induced airway lesions and how the PI3K-delta inhibitor PI-3065 changes them. It holds 78 samples: 2 untreated controls, 3 NTCU-treated mice, and 3 NTCU-treated mice that also got 8PI-3065. See `TREATMENT` for which group each sample belongs to. 9 10Every sample pairs a set of morphology images with the masks that the Xenium Onboard 11Analysis (XOA) pipeline produced. Channel 0 always holds the DAPI nuclear stain. The other 12channels hold the boundary and interior stains of the multi-tissue stain mix, which XOA uses 13to draw the cell outlines. 14 15NOTE: The masks are the output of XOA, not a manual annotation. XOA segments the nuclei on 16DAPI, then grows each cell from its nucleus using the boundary and interior stains. Where no 17stain supports a boundary it falls back to expanding the nucleus by 5 micrometer, so a share 18of the cells are dilated nuclei rather than observed outlines. The 'nuc_expansion_fraction' 19attribute of each h5 file reports that share. Treat this data as a weak label source for 20pretraining, not as a benchmark. 21 22NOTE: A cell can hold more than one nucleus, so the two label sets do not match one to one 23and their ids do not correspond. 24 25NOTE: This dataset is distinct from the Xenium samples in `torch_em.data.datasets.light_microscopy.xenium`, 26which are hosted directly by 10x Genomics and cover human and mouse tissue other than lung. 27 28NOTE: The images are whole slides that hold large empty regions. Pass a sampler such as 29`torch_em.data.MinInstanceSampler` to avoid drawing empty patches. 30 31The data is hosted at https://www.ebi.ac.uk/biostudies/bioimages/studies/S-BIAD3880 under the 32CC0 1.0 license. Please cite the corresponding study page if you use this data in your research. 33""" 34 35import os 36import json 37import zlib 38import shutil 39import struct 40import zipfile 41from glob import glob 42from natsort import natsorted 43from contextlib import contextmanager 44from typing import Dict, List, Literal, Optional, Sequence, Tuple, Union 45 46import requests 47 48from torch.utils.data import DataLoader, Dataset 49 50import torch_em 51 52from .. import util 53 54 55BASE_URL = "https://ftp.ebi.ac.uk/pub/databases/biostudies/S-BIAD/880/S-BIAD3880/Files/BIA_upload" 56 57URLS = { 58 "ctrl69": f"{BASE_URL}/ctrl69_xenium_output.zip", 59 "ctrl70": f"{BASE_URL}/ctrl70_xenium_output.zip", 60 "ntcu06": f"{BASE_URL}/ntcu06_xenium_output.zip", 61 "ntcu36": f"{BASE_URL}/ntcu36_xenium_output.zip", 62 "ntcu46": f"{BASE_URL}/ntcu46_xenium_output.zip", 63 "pi42": f"{BASE_URL}/pi42_xenium_output.zip", 64 "pi51": f"{BASE_URL}/pi51_xenium_output.zip", 65 "pi57": f"{BASE_URL}/pi57_xenium_output.zip", 66} 67 68# The treatment group of every sample, for the h5 attributes. 69TREATMENT = { 70 "ctrl69": "control", 71 "ctrl70": "control", 72 "ntcu06": "ntcu", 73 "ntcu36": "ntcu", 74 "ntcu46": "ntcu", 75 "pi42": "ntcu_pi3065", 76 "pi51": "ntcu_pi3065", 77 "pi57": "ntcu_pi3065", 78} 79 80LABEL_CHANNELS = {"nuclei": "0", "cells": "1"} 81 82BLOCK = 4096 83 84 85def _read_range(url: str, start: int, end: int) -> bytes: 86 """Read the bytes from `start` to `end`, both included.""" 87 response = requests.get(url, headers={"Range": f"bytes={start}-{end}"}, stream=True) 88 response.raise_for_status() 89 return response.content 90 91 92def _read_zip64_extra(extra: bytes, uncompressed: int, compressed: int, offset: int) -> Tuple[int, int, int]: 93 """Replace the fields that overflowed the 32 bit form with their zip64 values.""" 94 position = 0 95 while position + 4 <= len(extra): 96 tag, size = struct.unpack("<HH", extra[position:position + 4]) 97 if tag == 0x0001: 98 values = list(struct.unpack(f"<{size // 8}Q", extra[position + 4:position + 4 + size])) 99 if uncompressed == 0xFFFFFFFF: 100 uncompressed = values.pop(0) 101 if compressed == 0xFFFFFFFF: 102 compressed = values.pop(0) 103 if offset == 0xFFFFFFFF: 104 offset = values.pop(0) 105 break 106 position += 4 + size 107 return uncompressed, compressed, offset 108 109 110def _remote_zip_index(url: str) -> Dict[str, Dict[str, int]]: 111 """Read the central directory of a remote zip, which lists every member and where it sits. 112 113 This turns the bundle into a set of files that can be read one by one, so that a download 114 can skip the transcripts, the per-cycle QC images, and the expression matrix, which 115 segmentation never touches. 116 """ 117 response = requests.head(url) 118 response.raise_for_status() 119 size = int(response.headers["Content-Length"]) 120 121 tail = _read_range(url, max(0, size - 262144), size - 1) 122 position = tail.rfind(b"PK\x06\x06") 123 if position == -1: 124 position = tail.rfind(b"PK\x05\x06") 125 if position == -1: 126 raise RuntimeError(f"Could not find the central directory of the zip at {url}.") 127 directory_size, directory_offset = struct.unpack("<II", tail[position + 12:position + 20]) 128 else: 129 directory_size, directory_offset = struct.unpack("<QQ", tail[position + 40:position + 56]) 130 131 directory = _read_range(url, directory_offset, directory_offset + directory_size + 128) 132 index, position = {}, 0 133 while position + 46 <= len(directory) and directory[position:position + 4] == b"PK\x01\x02": 134 method = struct.unpack("<H", directory[position + 10:position + 12])[0] 135 crc = struct.unpack("<I", directory[position + 16:position + 20])[0] 136 compressed, uncompressed = struct.unpack("<II", directory[position + 20:position + 28]) 137 name_length, extra_length, comment_length = struct.unpack("<HHH", directory[position + 28:position + 34]) 138 offset = struct.unpack("<I", directory[position + 42:position + 46])[0] 139 name = directory[position + 46:position + 46 + name_length].decode("utf-8", errors="replace") 140 extra = directory[position + 46 + name_length:position + 46 + name_length + extra_length] 141 uncompressed, compressed, offset = _read_zip64_extra(extra, uncompressed, compressed, offset) 142 index[name] = {"method": method, "crc": crc, "compressed": compressed, "offset": offset} 143 position += 46 + name_length + extra_length + comment_length 144 return index 145 146 147def _find_root(index: Dict[str, Dict[str, int]], sample: str) -> str: 148 """Find the single top-level output folder that every member of the bundle sits under.""" 149 roots = {name.split("/")[0] for name in index if "/" in name} 150 if len(roots) != 1: 151 raise RuntimeError(f"Expected exactly one output folder in the bundle for '{sample}', found {roots}.") 152 return next(iter(roots)) + "/" 153 154 155def _download_member(url: str, name: str, entry: Dict[str, int], output_path: str) -> None: 156 """Read one member out of a remote zip and write it, checking its crc32.""" 157 from tqdm import tqdm 158 159 # The local header repeats the name and carries its own extra field, so the data of the 160 # member starts behind a header whose length only the header itself gives. 161 header = _read_range(url, entry["offset"], entry["offset"] + 29) 162 if header[:4] != b"PK\x03\x04": 163 raise RuntimeError(f"The member '{name}' does not start with a local header.") 164 name_length, extra_length = struct.unpack("<HH", header[26:30]) 165 start = entry["offset"] + 30 + name_length + extra_length 166 167 if entry["method"] not in (0, 8): 168 raise RuntimeError(f"The member '{name}' uses the unsupported compression method {entry['method']}.") 169 decompressor = zlib.decompressobj(-zlib.MAX_WBITS) if entry["method"] == 8 else None 170 171 crc = 0 172 temporary_path = f"{output_path}.incomplete" 173 os.makedirs(os.path.dirname(output_path), exist_ok=True) 174 175 headers = {"Range": f"bytes={start}-{start + entry['compressed'] - 1}"} 176 with requests.get(url, headers=headers, stream=True) as response: 177 response.raise_for_status() 178 description = f"Download {name}" 179 with open(temporary_path, "wb") as f: 180 with tqdm(total=entry["compressed"], unit="B", unit_scale=True, desc=description) as progress: 181 for chunk in response.iter_content(chunk_size=1 << 20): 182 progress.update(len(chunk)) 183 block = decompressor.decompress(chunk) if decompressor is not None else chunk 184 crc = zlib.crc32(block, crc) 185 f.write(block) 186 if decompressor is not None: 187 block = decompressor.flush() 188 crc = zlib.crc32(block, crc) 189 f.write(block) 190 191 if crc != entry["crc"]: 192 os.remove(temporary_path) 193 raise RuntimeError(f"The member '{name}' has the crc32 {crc}, but the zip lists {entry['crc']}.") 194 os.replace(temporary_path, output_path) 195 196 197def _download_bundle(url: str, bundle_dir: str, sample: str, download: bool, with_stains: bool) -> str: 198 """Download the members that this loader reads, and skip the rest of the bundle.""" 199 if os.path.exists(bundle_dir): 200 return bundle_dir 201 if not download: 202 raise RuntimeError(f"Cannot find the data at {bundle_dir}, but download was set to False") 203 204 index = _remote_zip_index(url) 205 root = _find_root(index, sample) 206 channels = natsorted( 207 n for n in index if n.startswith(f"{root}morphology_focus/") and n.endswith(".ome.tif") 208 ) 209 if not channels: 210 raise RuntimeError(f"The bundle for '{sample}' holds no morphology channel.") 211 if not with_stains: 212 channels = channels[:1] 213 214 names = [f"{root}experiment.xenium", f"{root}cells.zarr.zip"] + channels 215 missing = [n for n in names if n not in index] 216 if missing: 217 raise RuntimeError(f"The bundle for '{sample}' is missing {missing}.") 218 219 temporary_dir = f"{bundle_dir}.tmp" 220 if os.path.exists(temporary_dir): 221 shutil.rmtree(temporary_dir) 222 for name in names: 223 relative_name = name[len(root):] 224 _download_member(url, name, index[name], os.path.join(temporary_dir, relative_name)) 225 226 os.replace(temporary_dir, bundle_dir) 227 return bundle_dir 228 229 230def _channel_paths(bundle_dir: str) -> List[str]: 231 """Return the morphology channels, DAPI first.""" 232 paths = natsorted(glob(os.path.join(bundle_dir, "morphology_focus", "*.ome.tif"))) 233 if not paths: 234 raise RuntimeError(f"Could not find any morphology channel in {bundle_dir}.") 235 return paths 236 237 238@contextmanager 239def _open_image(path: str): 240 """Open one morphology channel at full resolution, without reading it into memory. 241 242 The channels form a multi file OME pyramid, which the OME reader of tifffile rejects. 243 Reading the file on its own gives the pyramid of this channel alone. 244 """ 245 import zarr 246 import tifffile 247 248 with tifffile.TiffFile(path, is_ome=False) as tif: 249 store = tif.series[0].levels[0].aszarr() 250 try: 251 yield zarr.open(store, mode="r")["0"] 252 finally: 253 store.close() 254 255 256def _copy_blockwise(source, target, channel: Optional[int] = None) -> int: 257 """Copy a large 2d array block by block, to keep the memory use bounded. 258 259 The target must be indexed in one step, because indexing an h5 dataset twice writes into 260 a copy of the data rather than into the file. 261 262 Returns the largest value that was copied, which gives the instance count of a mask 263 without a second pass over the whole array. 264 """ 265 height, width = source.shape[-2:] 266 largest = 0 267 for y in range(0, height, BLOCK): 268 for x in range(0, width, BLOCK): 269 box = (slice(y, min(y + BLOCK, height)), slice(x, min(x + BLOCK, width))) 270 block = source[box] 271 if channel is None: 272 target[box] = block 273 else: 274 target[(channel,) + box] = block 275 largest = max(largest, int(block.max())) 276 return largest 277 278 279def _create_h5(bundle_dir: str, output_path: str, sample: str, with_stains: bool) -> str: 280 import h5py 281 import zarr 282 from tqdm import tqdm 283 284 if os.path.exists(output_path): 285 return output_path 286 287 with open(os.path.join(bundle_dir, "experiment.xenium")) as f: 288 manifest = json.load(f) 289 290 masks_zip = os.path.join(bundle_dir, "cells.zarr.zip") 291 masks_dir = os.path.join(bundle_dir, "cells.zarr") 292 if not os.path.exists(masks_dir): 293 with zipfile.ZipFile(masks_zip) as archive: 294 archive.extractall(masks_dir) 295 masks = zarr.open(masks_dir, mode="r")["masks"] 296 297 channel_paths = _channel_paths(bundle_dir) 298 with _open_image(channel_paths[0]) as dapi: 299 shape = dapi.shape 300 for name, key in LABEL_CHANNELS.items(): 301 if masks[key].shape != shape: 302 raise RuntimeError( 303 f"The {name} mask of '{sample}' has the shape {masks[key].shape}, " 304 f"but its DAPI image has the shape {shape}." 305 ) 306 307 temporary_path = f"{output_path}.tmp" 308 with h5py.File(temporary_path, "w") as f: 309 f.attrs["sample"] = sample 310 f.attrs["treatment_group"] = TREATMENT[sample] 311 f.attrs["axes"] = "yx" 312 f.attrs["pixel_size"] = manifest["pixel_size"] 313 f.attrs["xoa_version"] = manifest["analysis_sw_version"] 314 f.attrs["segmentation_stain"] = manifest.get("segmentation_stain", "none") 315 f.attrs["nuc_expansion_fraction"] = manifest.get("segmented_cell_nuc_expansion_frac", float("nan")) 316 f.attrs["num_cells"] = manifest["num_cells"] 317 f.attrs["channel_files"] = [os.path.basename(p) for p in channel_paths] 318 319 chunks = (min(1024, shape[0]), min(1024, shape[1])) 320 raw = f.create_dataset("raw/dapi", shape=shape, dtype="uint16", chunks=chunks, compression="gzip") 321 raw.attrs["pixel_size"] = manifest["pixel_size"] 322 with _open_image(channel_paths[0]) as dapi: 323 _copy_blockwise(dapi, raw) 324 325 if with_stains: 326 stack_shape = (len(channel_paths),) + shape 327 stack = f.create_dataset( 328 "raw/stack", shape=stack_shape, dtype="uint16", chunks=(1,) + chunks, compression="gzip" 329 ) 330 stack.attrs["pixel_size"] = manifest["pixel_size"] 331 for channel, channel_path in enumerate(tqdm(channel_paths, desc=f"Copy channels of '{sample}'")): 332 with _open_image(channel_path) as image: 333 if image.shape != shape: 334 raise RuntimeError( 335 f"The channel {os.path.basename(channel_path)} of '{sample}' has the shape " 336 f"{image.shape}, but its DAPI image has the shape {shape}." 337 ) 338 _copy_blockwise(image, stack, channel=channel) 339 340 for name, key in LABEL_CHANNELS.items(): 341 labels = f.create_dataset( 342 f"labels/{name}", shape=shape, dtype="uint32", chunks=chunks, compression="gzip" 343 ) 344 labels.attrs["pixel_size"] = manifest["pixel_size"] 345 labels.attrs["num_instances"] = _copy_blockwise(masks[key], labels) 346 347 os.replace(temporary_path, output_path) 348 shutil.rmtree(masks_dir, ignore_errors=True) 349 return output_path 350 351 352def get_xenium_lung_treg_data( 353 path: Union[os.PathLike, str], 354 sample: str, 355 download: bool = False, 356 with_stains: bool = True, 357) -> str: 358 """Download one sample of the mouse lung Treg Xenium dataset. 359 360 This function reads the central directory of the remote zip that BioImage Archive 361 hosts for the sample, then reads only the members it needs (the manifest, the cell and 362 nucleus masks, and the morphology channels). It does not fetch the transcripts, the 363 per-cycle QC images, or the expression matrix, which cuts the download by a large margin. 364 365 Args: 366 path: Filepath to a folder where the downloaded data will be saved. 367 sample: The sample to use. See `URLS` for the available samples. 368 download: Whether to download the data if it is not present. 369 with_stains: Whether to also store the boundary and interior stains next to DAPI. 370 The cell masks rest on those stains, so a model that predicts cells needs them. 371 372 Returns: 373 The filepath to the h5 file of the sample. 374 """ 375 if sample not in URLS: 376 raise ValueError(f"'{sample}' is not a valid sample. Choose from {list(URLS)}.") 377 378 output_path = os.path.join(path, "preprocessed", f"{sample}.h5") 379 if os.path.exists(output_path): 380 return output_path 381 382 os.makedirs(os.path.join(path, "preprocessed"), exist_ok=True) 383 bundle_dir = _download_bundle(URLS[sample], os.path.join(path, "bundles", sample), sample, download, with_stains) 384 _create_h5(bundle_dir, output_path, sample, with_stains) 385 shutil.rmtree(bundle_dir, ignore_errors=True) 386 387 return output_path 388 389 390def get_xenium_lung_treg_paths( 391 path: Union[os.PathLike, str], 392 sample: Optional[Union[str, Sequence[str]]] = None, 393 download: bool = False, 394 with_stains: bool = True, 395) -> List[str]: 396 """Get paths to the mouse lung Treg Xenium data. 397 398 Args: 399 path: Filepath to a folder where the downloaded data will be saved. 400 sample: The sample or samples to use. Defaults to all of them. 401 download: Whether to download the data if it is not present. 402 with_stains: Whether to also store the boundary and interior stains next to DAPI. 403 404 Returns: 405 List of filepaths for the h5 data. 406 """ 407 if sample is None: 408 samples = list(URLS) 409 else: 410 samples = [sample] if isinstance(sample, str) else list(sample) 411 412 return [get_xenium_lung_treg_data(path, name, download, with_stains) for name in samples] 413 414 415def get_xenium_lung_treg_dataset( 416 path: Union[os.PathLike, str], 417 patch_shape: Tuple[int, int], 418 sample: Optional[Union[str, Sequence[str]]] = None, 419 label_channel: Literal["nuclei", "cells"] = "nuclei", 420 raw_channel: Literal["dapi", "stack"] = "dapi", 421 offsets: Optional[List[List[int]]] = None, 422 boundaries: bool = False, 423 binary: bool = False, 424 download: bool = False, 425 **kwargs, 426) -> Dataset: 427 """Get the mouse lung Treg Xenium dataset for cell and nucleus segmentation. 428 429 Args: 430 path: Filepath to a folder where the downloaded data will be saved. 431 patch_shape: The 2D patch shape to use for training. 432 sample: The sample or samples to use. Defaults to all of them. 433 label_channel: The masks to use as target. Either 'nuclei' or 'cells'. 434 raw_channel: The images to use as input. Either 'dapi' for the nuclear stain alone, 435 or 'stack' for all morphology channels. 436 offsets: Offset values for affinity computation used as target. 437 boundaries: Whether to compute boundaries as the target. 438 binary: Whether to use a binary segmentation target. 439 download: Whether to download the data if it is not present. 440 kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset`. 441 442 Returns: 443 The segmentation dataset. 444 """ 445 if len(patch_shape) != 2: 446 raise ValueError(f"The Xenium patch shape must be two-dimensional, got {patch_shape}.") 447 if label_channel not in LABEL_CHANNELS: 448 raise ValueError(f"'{label_channel}' is not a valid label channel. Choose from {list(LABEL_CHANNELS)}.") 449 if raw_channel not in ("dapi", "stack"): 450 raise ValueError(f"'{raw_channel}' is not a valid raw channel. Choose from ['dapi', 'stack'].") 451 452 volume_paths = get_xenium_lung_treg_paths(path, sample, download, with_stains=raw_channel == "stack") 453 454 kwargs = util.update_kwargs(kwargs, "with_channels", raw_channel == "stack") 455 kwargs, _ = util.add_instance_label_transform( 456 kwargs, add_binary_target=True, offsets=offsets, boundaries=boundaries, binary=binary, 457 ) 458 kwargs = util.ensure_transforms(ndim=2, **kwargs) 459 460 return torch_em.default_segmentation_dataset( 461 raw_paths=volume_paths, 462 raw_key=f"raw/{raw_channel}", 463 label_paths=volume_paths, 464 label_key=f"labels/{label_channel}", 465 patch_shape=patch_shape, 466 ndim=2, 467 **kwargs, 468 ) 469 470 471def get_xenium_lung_treg_loader( 472 path: Union[os.PathLike, str], 473 batch_size: int, 474 patch_shape: Tuple[int, int], 475 sample: Optional[Union[str, Sequence[str]]] = None, 476 label_channel: Literal["nuclei", "cells"] = "nuclei", 477 raw_channel: Literal["dapi", "stack"] = "dapi", 478 offsets: Optional[List[List[int]]] = None, 479 boundaries: bool = False, 480 binary: bool = False, 481 download: bool = False, 482 **kwargs, 483) -> DataLoader: 484 """Get the mouse lung Treg Xenium dataloader for cell and nucleus segmentation. 485 486 Args: 487 path: Filepath to a folder where the downloaded data will be saved. 488 batch_size: The batch size for training. 489 patch_shape: The 2D patch shape to use for training. 490 sample: The sample or samples to use. Defaults to all of them. 491 label_channel: The masks to use as target. Either 'nuclei' or 'cells'. 492 raw_channel: The images to use as input. Either 'dapi' or 'stack'. 493 offsets: Offset values for affinity computation used as target. 494 boundaries: Whether to compute boundaries as the target. 495 binary: Whether to use a binary segmentation target. 496 download: Whether to download the data if it is not present. 497 kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset` or the PyTorch DataLoader. 498 499 Returns: 500 The DataLoader. 501 """ 502 ds_kwargs, loader_kwargs = util.split_kwargs(torch_em.default_segmentation_dataset, **kwargs) 503 dataset = get_xenium_lung_treg_dataset( 504 path=path, 505 patch_shape=patch_shape, 506 sample=sample, 507 label_channel=label_channel, 508 raw_channel=raw_channel, 509 offsets=offsets, 510 boundaries=boundaries, 511 binary=binary, 512 download=download, 513 **ds_kwargs, 514 ) 515 return torch_em.get_data_loader(dataset, batch_size=batch_size, **loader_kwargs)
353def get_xenium_lung_treg_data( 354 path: Union[os.PathLike, str], 355 sample: str, 356 download: bool = False, 357 with_stains: bool = True, 358) -> str: 359 """Download one sample of the mouse lung Treg Xenium dataset. 360 361 This function reads the central directory of the remote zip that BioImage Archive 362 hosts for the sample, then reads only the members it needs (the manifest, the cell and 363 nucleus masks, and the morphology channels). It does not fetch the transcripts, the 364 per-cycle QC images, or the expression matrix, which cuts the download by a large margin. 365 366 Args: 367 path: Filepath to a folder where the downloaded data will be saved. 368 sample: The sample to use. See `URLS` for the available samples. 369 download: Whether to download the data if it is not present. 370 with_stains: Whether to also store the boundary and interior stains next to DAPI. 371 The cell masks rest on those stains, so a model that predicts cells needs them. 372 373 Returns: 374 The filepath to the h5 file of the sample. 375 """ 376 if sample not in URLS: 377 raise ValueError(f"'{sample}' is not a valid sample. Choose from {list(URLS)}.") 378 379 output_path = os.path.join(path, "preprocessed", f"{sample}.h5") 380 if os.path.exists(output_path): 381 return output_path 382 383 os.makedirs(os.path.join(path, "preprocessed"), exist_ok=True) 384 bundle_dir = _download_bundle(URLS[sample], os.path.join(path, "bundles", sample), sample, download, with_stains) 385 _create_h5(bundle_dir, output_path, sample, with_stains) 386 shutil.rmtree(bundle_dir, ignore_errors=True) 387 388 return output_path
Download one sample of the mouse lung Treg Xenium dataset.
This function reads the central directory of the remote zip that BioImage Archive hosts for the sample, then reads only the members it needs (the manifest, the cell and nucleus masks, and the morphology channels). It does not fetch the transcripts, the per-cycle QC images, or the expression matrix, which cuts the download by a large margin.
Arguments:
- path: Filepath to a folder where the downloaded data will be saved.
- sample: The sample to use. See
URLSfor the available samples. - download: Whether to download the data if it is not present.
- with_stains: Whether to also store the boundary and interior stains next to DAPI. The cell masks rest on those stains, so a model that predicts cells needs them.
Returns:
The filepath to the h5 file of the sample.
391def get_xenium_lung_treg_paths( 392 path: Union[os.PathLike, str], 393 sample: Optional[Union[str, Sequence[str]]] = None, 394 download: bool = False, 395 with_stains: bool = True, 396) -> List[str]: 397 """Get paths to the mouse lung Treg Xenium data. 398 399 Args: 400 path: Filepath to a folder where the downloaded data will be saved. 401 sample: The sample or samples to use. Defaults to all of them. 402 download: Whether to download the data if it is not present. 403 with_stains: Whether to also store the boundary and interior stains next to DAPI. 404 405 Returns: 406 List of filepaths for the h5 data. 407 """ 408 if sample is None: 409 samples = list(URLS) 410 else: 411 samples = [sample] if isinstance(sample, str) else list(sample) 412 413 return [get_xenium_lung_treg_data(path, name, download, with_stains) for name in samples]
Get paths to the mouse lung Treg Xenium data.
Arguments:
- path: Filepath to a folder where the downloaded data will be saved.
- sample: The sample or samples to use. Defaults to all of them.
- download: Whether to download the data if it is not present.
- with_stains: Whether to also store the boundary and interior stains next to DAPI.
Returns:
List of filepaths for the h5 data.
416def get_xenium_lung_treg_dataset( 417 path: Union[os.PathLike, str], 418 patch_shape: Tuple[int, int], 419 sample: Optional[Union[str, Sequence[str]]] = None, 420 label_channel: Literal["nuclei", "cells"] = "nuclei", 421 raw_channel: Literal["dapi", "stack"] = "dapi", 422 offsets: Optional[List[List[int]]] = None, 423 boundaries: bool = False, 424 binary: bool = False, 425 download: bool = False, 426 **kwargs, 427) -> Dataset: 428 """Get the mouse lung Treg Xenium dataset for cell and nucleus segmentation. 429 430 Args: 431 path: Filepath to a folder where the downloaded data will be saved. 432 patch_shape: The 2D patch shape to use for training. 433 sample: The sample or samples to use. Defaults to all of them. 434 label_channel: The masks to use as target. Either 'nuclei' or 'cells'. 435 raw_channel: The images to use as input. Either 'dapi' for the nuclear stain alone, 436 or 'stack' for all morphology channels. 437 offsets: Offset values for affinity computation used as target. 438 boundaries: Whether to compute boundaries as the target. 439 binary: Whether to use a binary segmentation target. 440 download: Whether to download the data if it is not present. 441 kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset`. 442 443 Returns: 444 The segmentation dataset. 445 """ 446 if len(patch_shape) != 2: 447 raise ValueError(f"The Xenium patch shape must be two-dimensional, got {patch_shape}.") 448 if label_channel not in LABEL_CHANNELS: 449 raise ValueError(f"'{label_channel}' is not a valid label channel. Choose from {list(LABEL_CHANNELS)}.") 450 if raw_channel not in ("dapi", "stack"): 451 raise ValueError(f"'{raw_channel}' is not a valid raw channel. Choose from ['dapi', 'stack'].") 452 453 volume_paths = get_xenium_lung_treg_paths(path, sample, download, with_stains=raw_channel == "stack") 454 455 kwargs = util.update_kwargs(kwargs, "with_channels", raw_channel == "stack") 456 kwargs, _ = util.add_instance_label_transform( 457 kwargs, add_binary_target=True, offsets=offsets, boundaries=boundaries, binary=binary, 458 ) 459 kwargs = util.ensure_transforms(ndim=2, **kwargs) 460 461 return torch_em.default_segmentation_dataset( 462 raw_paths=volume_paths, 463 raw_key=f"raw/{raw_channel}", 464 label_paths=volume_paths, 465 label_key=f"labels/{label_channel}", 466 patch_shape=patch_shape, 467 ndim=2, 468 **kwargs, 469 )
Get the mouse lung Treg Xenium dataset for cell and nucleus segmentation.
Arguments:
- path: Filepath to a folder where the downloaded data will be saved.
- patch_shape: The 2D patch shape to use for training.
- sample: The sample or samples to use. Defaults to all of them.
- label_channel: The masks to use as target. Either 'nuclei' or 'cells'.
- raw_channel: The images to use as input. Either 'dapi' for the nuclear stain alone, or 'stack' for all morphology channels.
- offsets: Offset values for affinity computation used as target.
- boundaries: Whether to compute boundaries as the target.
- binary: Whether to use a binary segmentation target.
- download: Whether to download the data if it is not present.
- kwargs: Additional keyword arguments for
torch_em.default_segmentation_dataset.
Returns:
The segmentation dataset.
472def get_xenium_lung_treg_loader( 473 path: Union[os.PathLike, str], 474 batch_size: int, 475 patch_shape: Tuple[int, int], 476 sample: Optional[Union[str, Sequence[str]]] = None, 477 label_channel: Literal["nuclei", "cells"] = "nuclei", 478 raw_channel: Literal["dapi", "stack"] = "dapi", 479 offsets: Optional[List[List[int]]] = None, 480 boundaries: bool = False, 481 binary: bool = False, 482 download: bool = False, 483 **kwargs, 484) -> DataLoader: 485 """Get the mouse lung Treg Xenium dataloader for cell and nucleus segmentation. 486 487 Args: 488 path: Filepath to a folder where the downloaded data will be saved. 489 batch_size: The batch size for training. 490 patch_shape: The 2D patch shape to use for training. 491 sample: The sample or samples to use. Defaults to all of them. 492 label_channel: The masks to use as target. Either 'nuclei' or 'cells'. 493 raw_channel: The images to use as input. Either 'dapi' or 'stack'. 494 offsets: Offset values for affinity computation used as target. 495 boundaries: Whether to compute boundaries as the target. 496 binary: Whether to use a binary segmentation target. 497 download: Whether to download the data if it is not present. 498 kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset` or the PyTorch DataLoader. 499 500 Returns: 501 The DataLoader. 502 """ 503 ds_kwargs, loader_kwargs = util.split_kwargs(torch_em.default_segmentation_dataset, **kwargs) 504 dataset = get_xenium_lung_treg_dataset( 505 path=path, 506 patch_shape=patch_shape, 507 sample=sample, 508 label_channel=label_channel, 509 raw_channel=raw_channel, 510 offsets=offsets, 511 boundaries=boundaries, 512 binary=binary, 513 download=download, 514 **ds_kwargs, 515 ) 516 return torch_em.get_data_loader(dataset, batch_size=batch_size, **loader_kwargs)
Get the mouse lung Treg Xenium dataloader for cell and nucleus segmentation.
Arguments:
- path: Filepath to a folder where the downloaded data will be saved.
- batch_size: The batch size for training.
- patch_shape: The 2D patch shape to use for training.
- sample: The sample or samples to use. Defaults to all of them.
- label_channel: The masks to use as target. Either 'nuclei' or 'cells'.
- raw_channel: The images to use as input. Either 'dapi' or 'stack'.
- offsets: Offset values for affinity computation used as target.
- boundaries: Whether to compute boundaries as the target.
- binary: Whether to use a binary segmentation target.
- download: Whether to download the data if it is not present.
- kwargs: Additional keyword arguments for
torch_em.default_segmentation_datasetor the PyTorch DataLoader.
Returns:
The DataLoader.