torch_em.data.datasets.electron_microscopy.openorganelle_nucleus
This dataset provides nucleus segmentation masks for 30 volumes from the public
"OpenOrganelle" janelia-cosem-datasets S3 bucket (the same bucket used by cellmap.py and
janelia_nucleus.py), none of which are covered by either of those existing loaders. All 30 are
released under CC0-1.0, except jrc_mus-thymus-1 and aic_desmosome-3 whose license and paper
could not be independently verified (aic_desmosome-3 is an Allen Institute dataset, so the
Janelia CC0-1.0 convention documented for the rest does not apply to it).
Three label folder naming conventions are used across the bucket, all distinct from
janelia_nucleus.py's:
labels/inference/nucleus_seg/{level}(jrc_ctl-id8-2,jrc_choroid-plexus-2,jrc_dauer-larva).labels/inference/segmentations/nuc/{level}(most of the remaining 27 datasets).labels/inference/segmentations/nucleus/{level}(jrc_hum-airway-14953vconly).labels/inference/segmentations/nuc_mem/{level}(jrc_mus-liver-4/5/6) - these three are nuclear MEMBRANE labels, a distinct annotation target from a nucleus mask/instance segmentation, tracked viaDATASETS[name]["label_target"].
Label and raw data are usually released at an exactly matching multiscale pyramid level (no
resampling, same axis order). Three exceptions are handled: jrc_dauer-larva and
jrc_hela-h89-1/2 have no exact-scale match and are resampled onto the label grid via
scipy.ndimage.zoom; jrc_choroid-plexus-2 and jrc_mus-epididymis-1/2 store one of the two
arrays with a reversed axis order, corrected via np.transpose.
Datasets whose full matched-resolution raw+label pair is small (roughly under 5 GB) are
downloaded whole, following janelia_nucleus.py's pattern. The rest (full volumes ranging from
tens of GB to several TB) are downloaded as one fixed-size representative crop, auto-detected
from a non-empty region of the coarsest available label pyramid level, following cellmap.py's
per-crop pattern - see DATASETS[name]["full_array"].
jrc_mosquito-stylet-6 is the only dataset whose raw and label arrays are stored in zarr v3 with
sharding (1024^3 outer shards, 64^3 inner chunks); _open_remote_zarr uses zarr.storage.
FsspecStore rather than fsspec.get_mapper specifically so that a crop only pulls the individual
inner chunks it needs via ranged reads, not whole ~1 GB shard files.
Please cite the CellMap project (https://www.janelia.org/project-team/cellmap) if you use this
data (except aic_desmosome-3, an Allen Institute for Cell Science dataset).
1"""This dataset provides nucleus segmentation masks for 30 volumes from the public 2"OpenOrganelle" `janelia-cosem-datasets` S3 bucket (the same bucket used by `cellmap.py` and 3`janelia_nucleus.py`), none of which are covered by either of those existing loaders. All 30 are 4released under CC0-1.0, except `jrc_mus-thymus-1` and `aic_desmosome-3` whose license and paper 5could not be independently verified (`aic_desmosome-3` is an Allen Institute dataset, so the 6Janelia CC0-1.0 convention documented for the rest does not apply to it). 7 8Three label folder naming conventions are used across the bucket, all distinct from 9`janelia_nucleus.py`'s: 10 11- `labels/inference/nucleus_seg/{level}` (`jrc_ctl-id8-2`, `jrc_choroid-plexus-2`, `jrc_dauer-larva`). 12- `labels/inference/segmentations/nuc/{level}` (most of the remaining 27 datasets). 13- `labels/inference/segmentations/nucleus/{level}` (`jrc_hum-airway-14953vc` only). 14- `labels/inference/segmentations/nuc_mem/{level}` (`jrc_mus-liver-4/5/6`) - these three are 15 nuclear MEMBRANE labels, a distinct annotation target from a nucleus mask/instance segmentation, 16 tracked via `DATASETS[name]["label_target"]`. 17 18Label and raw data are usually released at an exactly matching multiscale pyramid level (no 19resampling, same axis order). Three exceptions are handled: `jrc_dauer-larva` and 20`jrc_hela-h89-1/2` have no exact-scale match and are resampled onto the label grid via 21`scipy.ndimage.zoom`; `jrc_choroid-plexus-2` and `jrc_mus-epididymis-1/2` store one of the two 22arrays with a reversed axis order, corrected via `np.transpose`. 23 24Datasets whose full matched-resolution raw+label pair is small (roughly under 5 GB) are 25downloaded whole, following `janelia_nucleus.py`'s pattern. The rest (full volumes ranging from 26tens of GB to several TB) are downloaded as one fixed-size representative crop, auto-detected 27from a non-empty region of the coarsest available label pyramid level, following `cellmap.py`'s 28per-crop pattern - see `DATASETS[name]["full_array"]`. 29 30`jrc_mosquito-stylet-6` is the only dataset whose raw and label arrays are stored in zarr v3 with 31sharding (1024^3 outer shards, 64^3 inner chunks); `_open_remote_zarr` uses `zarr.storage. 32FsspecStore` rather than `fsspec.get_mapper` specifically so that a crop only pulls the individual 33inner chunks it needs via ranged reads, not whole ~1 GB shard files. 34 35Please cite the CellMap project (https://www.janelia.org/project-team/cellmap) if you use this 36data (except `aic_desmosome-3`, an Allen Institute for Cell Science dataset). 37""" 38 39import os 40from typing import List, Optional, Tuple, Union 41 42import numpy as np 43 44from torch.utils.data import DataLoader, Dataset 45 46import torch_em 47 48from .. import util 49 50 51BUCKET_URL = "https://janelia-cosem-datasets.s3.amazonaws.com/" 52 53# dataset name -> recon folder, em array name, label subpath (relative to "labels/inference/"), 54# label pyramid level, label kind, label target, axis to reverse ("none", "label", or "raw"), 55# and whether the full matched-resolution pair is downloaded whole (True) or as one fixed-size 56# crop auto-detected from a non-empty region (False). 57# label_kind: "binary" = label array only ever takes values {0, 1}; "instance" = per-nucleus IDs. 58# label_target: "nucleus" (a nucleus mask/instance segmentation) or "nuclear_membrane" (a nuclear 59# envelope surface label, a distinct annotation target). 60DATASETS = { 61 "jrc_ctl-id8-2": { 62 "recon": "recon-1", "em_name": "fibsem-uint8", "label_subpath": "nucleus_seg", "label_level": "s0", 63 "label_kind": "binary", "label_target": "nucleus", "axis_reverse": "none", "full_array": False, 64 "bounding_box": ((3824, 4048), (448, 672), (4064, 4288)), 65 }, 66 "jrc_choroid-plexus-2": { 67 "recon": "recon-1", "em_name": "fibsem-uint8", "label_subpath": "nucleus_seg", "label_level": "s1", 68 "label_kind": "instance", "label_target": "nucleus", "axis_reverse": "label", "full_array": False, 69 "bounding_box": ((780, 1100), (2220, 2540), (620, 940)), 70 }, 71 "jrc_dauer-larva": { 72 "recon": "recon-1", "em_name": "tem-uint8", "label_subpath": "nucleus_seg", "label_level": "s0", 73 "label_kind": "instance", "label_target": "nucleus", "axis_reverse": "none", "full_array": False, 74 "bounding_box": ((130, 330), (1804, 2004), (11724, 11924)), 75 }, 76 "jrc_ccl81-covid-1": { 77 "recon": "recon-1", "em_name": "fibsem-uint16", "label_subpath": "segmentations/nuc", "label_level": "s0", 78 "label_kind": "instance", "label_target": "nucleus", "axis_reverse": "none", "full_array": False, 79 }, 80 "jrc_cos7-11": { 81 "recon": "recon-1", "em_name": "fibsem-uint16", "label_subpath": "segmentations/nuc", "label_level": "s0", 82 "label_kind": "instance", "label_target": "nucleus", "axis_reverse": "none", "full_array": False, 83 }, 84 "jrc_ctl-id8-3": { 85 "recon": "recon-1", "em_name": "fibsem-uint8", "label_subpath": "segmentations/nuc", "label_level": "s0", 86 "label_kind": "instance", "label_target": "nucleus", "axis_reverse": "none", "full_array": False, 87 }, 88 "jrc_ctl-id8-4": { 89 "recon": "recon-1", "em_name": "fibsem-uint8", "label_subpath": "segmentations/nuc", "label_level": "s0", 90 "label_kind": "instance", "label_target": "nucleus", "axis_reverse": "none", "full_array": False, 91 }, 92 "jrc_ctl-id8-5": { 93 "recon": "recon-1", "em_name": "fibsem-uint8", "label_subpath": "segmentations/nuc", "label_level": "s0", 94 "label_kind": "instance", "label_target": "nucleus", "axis_reverse": "none", "full_array": False, 95 }, 96 "jrc_fly-acc-calyx-1": { 97 "recon": "recon-1", "em_name": "fibsem-uint16", "label_subpath": "segmentations/nuc", "label_level": "s0", 98 "label_kind": "instance", "label_target": "nucleus", "axis_reverse": "none", "full_array": True, 99 }, 100 "jrc_fly-fsb-1": { 101 "recon": "recon-1", "em_name": "fibsem-uint16", "label_subpath": "segmentations/nuc", "label_level": "s0", 102 "label_kind": "instance", "label_target": "nucleus", "axis_reverse": "none", "full_array": True, 103 }, 104 "jrc_hela-21": { 105 "recon": "recon-1", "em_name": "fibsem-uint8", "label_subpath": "segmentations/nuc", "label_level": "s0", 106 "label_kind": "binary", "label_target": "nucleus", "axis_reverse": "none", "full_array": True, 107 }, 108 "jrc_hela-22": { 109 "recon": "recon-1", "em_name": "fibsem-uint8", "label_subpath": "segmentations/nuc", "label_level": "s0", 110 "label_kind": "instance", "label_target": "nucleus", "axis_reverse": "none", "full_array": True, 111 }, 112 "jrc_hela-bfa": { 113 "recon": "recon-1", "em_name": "fibsem-uint8", "label_subpath": "segmentations/nuc", "label_level": "s0", 114 "label_kind": "binary", "label_target": "nucleus", "axis_reverse": "none", "full_array": True, 115 }, 116 "jrc_hela-h89-1": { 117 "recon": "recon-1", "em_name": "fibsem-uint16", "label_subpath": "segmentations/nuc", "label_level": "s0", 118 "label_kind": "instance", "label_target": "nucleus", "axis_reverse": "none", "full_array": True, 119 }, 120 "jrc_hela-h89-2": { 121 "recon": "recon-1", "em_name": "fibsem-uint16", "label_subpath": "segmentations/nuc", "label_level": "s0", 122 "label_kind": "instance", "label_target": "nucleus", "axis_reverse": "none", "full_array": True, 123 }, 124 "jrc_hela-nz-1": { 125 "recon": "recon-2", "em_name": "fibsem-int16", "label_subpath": "segmentations/nuc", "label_level": "s0", 126 "label_kind": "instance", "label_target": "nucleus", "axis_reverse": "none", "full_array": True, 127 }, 128 "jrc_hum-airway-14953vc": { 129 "recon": "recon-1", "em_name": "fibsem-uint8", "label_subpath": "segmentations/nucleus", "label_level": "s0", 130 "label_kind": "instance", "label_target": "nucleus", "axis_reverse": "none", "full_array": True, 131 }, 132 # Stored in zarr v3 with sharding (1024^3 outer shards, 64^3 inner chunks); `_open_remote_zarr` 133 # uses `FsspecStore` rather than `fsspec.get_mapper` so a crop only pulls the inner chunks it 134 # actually needs, not whole ~1GB shard files. Its label pyramid has only one level (s0, full 135 # resolution) - there is no genuinely coarse level to cheaply scan for a non-empty region, so 136 # this uses a verified fixed crop instead of auto-detection. 137 "jrc_mosquito-stylet-6": { 138 "recon": "recon-1", "em_name": "fibsem-uint8", "label_subpath": "segmentations/nuc", "label_level": "s0", 139 "label_kind": "instance", "label_target": "nucleus", "axis_reverse": "none", "full_array": False, 140 "bounding_box": ((1088, 1216), (640, 768), (768, 896)), 141 }, 142 "jrc_mus-epididymis-1": { 143 "recon": "recon-1", "em_name": "fibsem-uint8", "label_subpath": "segmentations/nuc", "label_level": "s0", 144 "label_kind": "binary", "label_target": "nucleus", "axis_reverse": "raw", "full_array": True, 145 }, 146 "jrc_mus-epididymis-2": { 147 "recon": "recon-1", "em_name": "fibsem-uint8", "label_subpath": "segmentations/nuc", "label_level": "s0", 148 "label_kind": "binary", "label_target": "nucleus", "axis_reverse": "raw", "full_array": True, 149 }, 150 "jrc_mus-hippocampus-1": { 151 "recon": "recon-1", "em_name": "fibsem-uint8", "label_subpath": "segmentations/nuc", "label_level": "s0", 152 "label_kind": "instance", "label_target": "nucleus", "axis_reverse": "none", "full_array": False, 153 }, 154 # These three membranes are too thin/sparse to survive downsampling to the coarsest label 155 # pyramid level, so auto-crop-detection finds no non-empty region; use a verified fixed crop. 156 "jrc_mus-liver-4": { 157 "recon": "recon-1", "em_name": "fibsem-uint8", "label_subpath": "segmentations/nuc_mem", "label_level": "s0", 158 "label_kind": "binary", "label_target": "nuclear_membrane", "axis_reverse": "none", "full_array": False, 159 "bounding_box": ((1409, 1537), (2000, 2128), (6000, 6128)), 160 }, 161 "jrc_mus-liver-5": { 162 "recon": "recon-1", "em_name": "fibsem-uint8", "label_subpath": "segmentations/nuc_mem", "label_level": "s0", 163 "label_kind": "binary", "label_target": "nuclear_membrane", "axis_reverse": "none", "full_array": False, 164 "bounding_box": ((2961, 3089), (5096, 5224), (4553, 4681)), 165 }, 166 "jrc_mus-liver-6": { 167 "recon": "recon-1", "em_name": "fibsem-uint8", "label_subpath": "segmentations/nuc_mem", "label_level": "s0", 168 "label_kind": "binary", "label_target": "nuclear_membrane", "axis_reverse": "none", "full_array": False, 169 "bounding_box": ((5600, 5792), (3970, 4162), (6610, 6802)), 170 }, 171 "jrc_mus-pancreas-4": { 172 "recon": "recon-1", "em_name": "fibsem-uint8", "label_subpath": "segmentations/nuc", "label_level": "s0", 173 "label_kind": "instance", "label_target": "nucleus", "axis_reverse": "none", "full_array": False, 174 }, 175 "jrc_mus-sc-zp104a": { 176 "recon": "recon-1", "em_name": "fibsem-uint8", "label_subpath": "segmentations/nuc", "label_level": "s0", 177 "label_kind": "instance", "label_target": "nucleus", "axis_reverse": "none", "full_array": False, 178 }, 179 "jrc_mus-sc-zp105a": { 180 "recon": "recon-1", "em_name": "fibsem-uint8", "label_subpath": "segmentations/nuc", "label_level": "s0", 181 "label_kind": "instance", "label_target": "nucleus", "axis_reverse": "none", "full_array": False, 182 }, 183 "jrc_mus-skin-1": { 184 "recon": "recon-1", "em_name": "fibsem-uint8", "label_subpath": "segmentations/nuc", "label_level": "s0", 185 "label_kind": "instance", "label_target": "nucleus", "axis_reverse": "none", "full_array": False, 186 }, 187 "jrc_mus-thymus-1": { 188 "recon": "recon-1", "em_name": "fibsem-uint8", "label_subpath": "segmentations/nuc", "label_level": "s0", 189 "label_kind": "instance", "label_target": "nucleus", "axis_reverse": "none", "full_array": True, 190 }, 191 "aic_desmosome-3": { 192 "recon": "recon-1", "em_name": "fibsem-uint8", "label_subpath": "segmentations/nuc", "label_level": "s0", 193 "label_kind": "binary", "label_target": "nucleus", "axis_reverse": "none", "full_array": True, 194 }, 195} 196 197# Datasets whose data license and paper could not be independently verified. Everything else in 198# `DATASETS` is CC0-1.0 per the OpenOrganelle/CellMap project convention. 199UNVERIFIED_LICENSE = ("jrc_mus-thymus-1", "aic_desmosome-3") 200 201CROP_SHAPE_DEFAULT = (128, 128, 128) 202 203 204def _open_remote_zarr(s3_path): 205 import zarr 206 from zarr.storage import FsspecStore 207 208 # `FsspecStore` (not `fsspec.get_mapper`) is required for `jrc_mosquito-stylet-6`'s zarr v3 209 # sharded array: it performs ranged partial reads of individual inner chunks, whereas 210 # `fsspec.get_mapper` fetches whole shard files (up to ~1GB each) even for a small crop. 211 store = FsspecStore.from_url(s3_path, storage_options={"anon": True}, read_only=True) 212 return zarr.open(store, mode="r") 213 214 215def _get_json(url): 216 import requests 217 218 resp = requests.get(url, timeout=60) 219 resp.raise_for_status() 220 return resp.json() 221 222 223def _multiscale_levels(group_url): 224 """Read multiscale pyramid levels from a zarr group, trying zarr v2 and v3 metadata. 225 226 `jrc_mosquito-stylet-6`'s raw and label arrays are stored in zarr v3 (`zarr.json`), unlike 227 every other dataset here (zarr v2, `.zattrs`); its label metadata nests multiscales under 228 `attributes.ome`, while its raw metadata puts them directly under `attributes`. 229 """ 230 try: 231 multiscales = _get_json(f"{group_url}/.zattrs")["multiscales"] 232 except Exception: 233 attributes = _get_json(f"{group_url}/zarr.json")["attributes"] 234 multiscales = attributes["ome"]["multiscales"] if "ome" in attributes else attributes["multiscales"] 235 return { 236 d["path"]: tuple(d["coordinateTransformations"][0]["scale"]) 237 for d in multiscales[0]["datasets"] 238 } 239 240 241def _find_matching_em_level(dataset_name, recon, em_name, label_scale): 242 """Find the EM pyramid level whose scale exactly matches the label pyramid level, or `None`.""" 243 em_levels = _multiscale_levels(f"{BUCKET_URL}{dataset_name}/{dataset_name}.zarr/{recon}/em/{em_name}") 244 for level, scale in em_levels.items(): 245 if scale == label_scale: 246 return level, scale 247 return None, None 248 249 250def _nearest_em_level(dataset_name, recon, em_name, label_scale): 251 """Find the EM pyramid level whose scale is closest to the label pyramid level. 252 253 Used only for datasets whose label pyramid has no exact scale match. 254 """ 255 em_levels = _multiscale_levels(f"{BUCKET_URL}{dataset_name}/{dataset_name}.zarr/{recon}/em/{em_name}") 256 best_level, best_dist = None, None 257 for level, scale in em_levels.items(): 258 dist = sum((s - lv) ** 2 for s, lv in zip(scale, label_scale)) 259 if best_dist is None or dist < best_dist: 260 best_level, best_dist = level, dist 261 return best_level, em_levels[best_level] 262 263 264def _auto_bounding_box(dataset_name, recon, label_full_path, label_level, label_scale, crop_shape): 265 """Center a crop on a non-empty region, found via the coarsest available label level. 266 267 Downloading the coarsest level (rather than `label_level`, which may be a huge full-resolution 268 array) keeps this cheap even for multi-TB datasets. 269 """ 270 levels = _multiscale_levels(f"{BUCKET_URL}{dataset_name}/{dataset_name}.zarr/{recon}/{label_full_path}") 271 coarsest_level = max(levels, key=lambda lv: levels[lv]) 272 coarse_url = ( 273 f"s3://janelia-cosem-datasets/{dataset_name}/{dataset_name}.zarr/{recon}/" 274 f"{label_full_path}/{coarsest_level}" 275 ) 276 coarse_arr = _open_remote_zarr(coarse_url) 277 coarse_data = np.asarray(coarse_arr[:]) 278 nonzero = np.argwhere(coarse_data > 0) 279 if len(nonzero) == 0: 280 raise RuntimeError(f"No non-empty '{label_full_path}' region found for '{dataset_name}'.") 281 center_coarse = nonzero[len(nonzero) // 2] 282 scale_factor = np.array(levels[coarsest_level]) / np.array(label_scale) 283 center = (center_coarse * scale_factor).astype(int) 284 half = [c // 2 for c in crop_shape] 285 return tuple((max(0, int(c - h)), int(c + h)) for c, h in zip(center, half)) 286 287 288def get_openorganelle_nucleus_data(path: Union[os.PathLike, str], dataset_name: str, download: bool = False) -> str: 289 """Download nucleus segmentation data for one OpenOrganelle dataset. 290 291 Args: 292 path: Filepath to a folder where the cached zarr store will be saved. 293 dataset_name: The name of the dataset, one of the keys in `DATASETS`. 294 download: Whether to download the data if it is not present. 295 296 Returns: 297 The filepath to the cached zarr store. 298 """ 299 import zarr 300 from zarr.codecs import BloscCodec 301 from scipy.ndimage import zoom 302 303 if dataset_name not in DATASETS: 304 raise ValueError(f"'{dataset_name}' is not a valid dataset name. Choose from {sorted(DATASETS.keys())}.") 305 306 info = DATASETS[dataset_name] 307 recon, em_name = info["recon"], info["em_name"] 308 label_subpath, label_level = info["label_subpath"], info["label_level"] 309 axis_reverse, full_array = info["axis_reverse"], info["full_array"] 310 311 os.makedirs(str(path), exist_ok=True) 312 zarr_path = os.path.join(str(path), f"{dataset_name}.zarr") 313 314 root = zarr.open_group(zarr_path, mode="a") 315 if "raw" in root and "labels" in root: 316 return zarr_path 317 318 if not download: 319 raise RuntimeError(f"No cached data found at '{zarr_path}'. Set download=True to stream it from S3.") 320 321 print(f"Streaming nucleus data for '{dataset_name}' from the janelia-cosem-datasets S3 bucket ...") 322 label_full_path = f"labels/inference/{label_subpath}" 323 label_url = ( 324 f"s3://janelia-cosem-datasets/{dataset_name}/{dataset_name}.zarr/{recon}/{label_full_path}/{label_level}" 325 ) 326 label_arr = _open_remote_zarr(label_url) 327 328 label_scale = _multiscale_levels( 329 f"{BUCKET_URL}{dataset_name}/{dataset_name}.zarr/{recon}/{label_full_path}" 330 )[label_level] 331 em_level, em_scale = _find_matching_em_level(dataset_name, recon, em_name, label_scale) 332 resampled = em_level is None 333 334 if full_array: 335 label_data = np.asarray(label_arr[:]) 336 bounding_box = None 337 else: 338 bounding_box = _auto_bounding_box( 339 dataset_name, recon, label_full_path, label_level, label_scale, CROP_SHAPE_DEFAULT, 340 ) if info.get("bounding_box") is None else info["bounding_box"] 341 label_slices = tuple(slice(*bb) for bb in bounding_box) 342 label_data = label_arr[label_slices] 343 344 if axis_reverse == "label": 345 label_data = np.transpose(label_data, tuple(reversed(range(label_data.ndim)))) 346 raw_bbox = tuple(reversed(bounding_box)) if bounding_box is not None else None 347 else: 348 raw_bbox = bounding_box 349 350 if em_level is not None: 351 em_url = f"s3://janelia-cosem-datasets/{dataset_name}/{dataset_name}.zarr/{recon}/em/{em_name}/{em_level}" 352 em_arr = _open_remote_zarr(em_url) 353 if axis_reverse == "raw": 354 raw_data = np.asarray(em_arr[:]) 355 raw_data = np.transpose(raw_data, tuple(reversed(range(raw_data.ndim)))) 356 if raw_bbox is not None: 357 raw_data = raw_data[tuple(slice(*bb) for bb in raw_bbox)] 358 elif raw_bbox is not None: 359 raw_data = em_arr[tuple(slice(*bb) for bb in raw_bbox)] 360 else: 361 raw_data = np.asarray(em_arr[:]) 362 else: 363 em_level, em_scale = _nearest_em_level(dataset_name, recon, em_name, label_scale) 364 em_url = f"s3://janelia-cosem-datasets/{dataset_name}/{dataset_name}.zarr/{recon}/em/{em_name}/{em_level}" 365 em_arr = _open_remote_zarr(em_url) 366 ratio = [ls / es for ls, es in zip(label_scale, em_scale)] 367 if raw_bbox is not None: 368 em_bbox_native = tuple((int(lo * r), int(hi * r)) for (lo, hi), r in zip(raw_bbox, ratio)) 369 raw_native = em_arr[tuple(slice(*bb) for bb in em_bbox_native)] 370 else: 371 raw_native = np.asarray(em_arr[:]) 372 zoom_factors = [t / s for t, s in zip(label_data.shape, raw_native.shape)] 373 raw_data = zoom(raw_native, zoom_factors, order=1).astype(raw_native.dtype) 374 375 # Guard against off-by-a-few-voxel shape mismatches between independently stored pyramids. 376 common_shape = tuple(min(r, lb) for r, lb in zip(raw_data.shape, label_data.shape)) 377 raw_data = raw_data[tuple(slice(0, s) for s in common_shape)] 378 label_data = label_data[tuple(slice(0, s) for s in common_shape)] 379 380 assert raw_data.shape == label_data.shape, ( 381 f"Shape mismatch for '{dataset_name}': raw {raw_data.shape} vs labels {label_data.shape}" 382 ) 383 384 def _make_array(name, data, shuffle): 385 arr = root.create_array( 386 name, shape=data.shape, chunks=tuple(min(128, s) for s in data.shape), dtype=data.dtype, 387 compressors=BloscCodec(cname="zstd", clevel=6, shuffle=shuffle), 388 ) 389 arr[:] = data 390 391 root.attrs["dataset"] = dataset_name 392 root.attrs["recon"] = recon 393 root.attrs["label_kind"] = info["label_kind"] 394 root.attrs["label_target"] = info["label_target"] 395 root.attrs["label_source"] = label_url 396 root.attrs["em_source"] = em_url 397 root.attrs["resampled"] = resampled 398 root.attrs["license_verified"] = dataset_name not in UNVERIFIED_LICENSE 399 400 _make_array("raw", raw_data, shuffle="shuffle") 401 _make_array("labels", label_data, shuffle="bitshuffle") 402 403 print(f"Cached '{dataset_name}' to '{zarr_path}' (shape {raw_data.shape}).") 404 return zarr_path 405 406 407def get_openorganelle_nucleus_paths( 408 path: Union[os.PathLike, str], dataset_names: Optional[List[str]] = None, download: bool = False, 409) -> List[str]: 410 """Get paths to cached OpenOrganelle nucleus zarr stores. 411 412 Args: 413 path: Filepath to a folder where the cached zarr stores will be saved. 414 dataset_names: The names of the datasets to use. Defaults to all datasets in `DATASETS`. 415 download: Whether to download the data if it is not present. 416 417 Returns: 418 List of filepaths to the cached zarr stores. 419 """ 420 if dataset_names is None: 421 dataset_names = list(DATASETS.keys()) 422 return [get_openorganelle_nucleus_data(path, name, download) for name in dataset_names] 423 424 425def get_openorganelle_nucleus_dataset( 426 path: Union[os.PathLike, str], 427 patch_shape: Tuple[int, int, int], 428 dataset_names: Optional[List[str]] = None, 429 download: bool = False, 430 offsets: Optional[List[List[int]]] = None, 431 boundaries: bool = False, 432 **kwargs, 433) -> Dataset: 434 """Get the OpenOrganelle nucleus dataset for nucleus segmentation. 435 436 Args: 437 path: Filepath to a folder where the cached zarr stores will be saved. 438 patch_shape: The patch shape (z, y, x) to use for training. 439 dataset_names: The names of the datasets to use. Defaults to all datasets in `DATASETS`. 440 download: Whether to download the data if it is not present. 441 offsets: Offset values for affinity computation used as target. 442 boundaries: Whether to compute boundaries as the target. 443 kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset`. 444 445 Returns: 446 The segmentation dataset. 447 """ 448 assert len(patch_shape) == 3 449 450 paths = get_openorganelle_nucleus_paths(path, dataset_names, download) 451 452 kwargs = util.update_kwargs(kwargs, "is_seg_dataset", True) 453 kwargs, _ = util.add_instance_label_transform( 454 kwargs, add_binary_target=False, boundaries=boundaries, offsets=offsets 455 ) 456 457 return torch_em.default_segmentation_dataset( 458 raw_paths=paths, 459 raw_key="raw", 460 label_paths=paths, 461 label_key="labels", 462 patch_shape=patch_shape, 463 **kwargs, 464 ) 465 466 467def get_openorganelle_nucleus_loader( 468 path: Union[os.PathLike, str], 469 patch_shape: Tuple[int, int, int], 470 batch_size: int, 471 dataset_names: Optional[List[str]] = None, 472 download: bool = False, 473 offsets: Optional[List[List[int]]] = None, 474 boundaries: bool = False, 475 **kwargs, 476) -> DataLoader: 477 """Get the DataLoader for nucleus segmentation in the OpenOrganelle nucleus dataset. 478 479 Args: 480 path: Filepath to a folder where the cached zarr stores will be saved. 481 patch_shape: The patch shape (z, y, x) to use for training. 482 batch_size: The batch size for training. 483 dataset_names: The names of the datasets to use. Defaults to all datasets in `DATASETS`. 484 download: Whether to download the data if it is not present. 485 offsets: Offset values for affinity computation used as target. 486 boundaries: Whether to compute boundaries as the target. 487 kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset` 488 or for the PyTorch DataLoader. 489 490 Returns: 491 The DataLoader. 492 """ 493 ds_kwargs, loader_kwargs = util.split_kwargs(torch_em.default_segmentation_dataset, **kwargs) 494 dataset = get_openorganelle_nucleus_dataset( 495 path, patch_shape, dataset_names=dataset_names, download=download, 496 offsets=offsets, boundaries=boundaries, **ds_kwargs 497 ) 498 return torch_em.get_data_loader(dataset, batch_size, **loader_kwargs)
289def get_openorganelle_nucleus_data(path: Union[os.PathLike, str], dataset_name: str, download: bool = False) -> str: 290 """Download nucleus segmentation data for one OpenOrganelle dataset. 291 292 Args: 293 path: Filepath to a folder where the cached zarr store will be saved. 294 dataset_name: The name of the dataset, one of the keys in `DATASETS`. 295 download: Whether to download the data if it is not present. 296 297 Returns: 298 The filepath to the cached zarr store. 299 """ 300 import zarr 301 from zarr.codecs import BloscCodec 302 from scipy.ndimage import zoom 303 304 if dataset_name not in DATASETS: 305 raise ValueError(f"'{dataset_name}' is not a valid dataset name. Choose from {sorted(DATASETS.keys())}.") 306 307 info = DATASETS[dataset_name] 308 recon, em_name = info["recon"], info["em_name"] 309 label_subpath, label_level = info["label_subpath"], info["label_level"] 310 axis_reverse, full_array = info["axis_reverse"], info["full_array"] 311 312 os.makedirs(str(path), exist_ok=True) 313 zarr_path = os.path.join(str(path), f"{dataset_name}.zarr") 314 315 root = zarr.open_group(zarr_path, mode="a") 316 if "raw" in root and "labels" in root: 317 return zarr_path 318 319 if not download: 320 raise RuntimeError(f"No cached data found at '{zarr_path}'. Set download=True to stream it from S3.") 321 322 print(f"Streaming nucleus data for '{dataset_name}' from the janelia-cosem-datasets S3 bucket ...") 323 label_full_path = f"labels/inference/{label_subpath}" 324 label_url = ( 325 f"s3://janelia-cosem-datasets/{dataset_name}/{dataset_name}.zarr/{recon}/{label_full_path}/{label_level}" 326 ) 327 label_arr = _open_remote_zarr(label_url) 328 329 label_scale = _multiscale_levels( 330 f"{BUCKET_URL}{dataset_name}/{dataset_name}.zarr/{recon}/{label_full_path}" 331 )[label_level] 332 em_level, em_scale = _find_matching_em_level(dataset_name, recon, em_name, label_scale) 333 resampled = em_level is None 334 335 if full_array: 336 label_data = np.asarray(label_arr[:]) 337 bounding_box = None 338 else: 339 bounding_box = _auto_bounding_box( 340 dataset_name, recon, label_full_path, label_level, label_scale, CROP_SHAPE_DEFAULT, 341 ) if info.get("bounding_box") is None else info["bounding_box"] 342 label_slices = tuple(slice(*bb) for bb in bounding_box) 343 label_data = label_arr[label_slices] 344 345 if axis_reverse == "label": 346 label_data = np.transpose(label_data, tuple(reversed(range(label_data.ndim)))) 347 raw_bbox = tuple(reversed(bounding_box)) if bounding_box is not None else None 348 else: 349 raw_bbox = bounding_box 350 351 if em_level is not None: 352 em_url = f"s3://janelia-cosem-datasets/{dataset_name}/{dataset_name}.zarr/{recon}/em/{em_name}/{em_level}" 353 em_arr = _open_remote_zarr(em_url) 354 if axis_reverse == "raw": 355 raw_data = np.asarray(em_arr[:]) 356 raw_data = np.transpose(raw_data, tuple(reversed(range(raw_data.ndim)))) 357 if raw_bbox is not None: 358 raw_data = raw_data[tuple(slice(*bb) for bb in raw_bbox)] 359 elif raw_bbox is not None: 360 raw_data = em_arr[tuple(slice(*bb) for bb in raw_bbox)] 361 else: 362 raw_data = np.asarray(em_arr[:]) 363 else: 364 em_level, em_scale = _nearest_em_level(dataset_name, recon, em_name, label_scale) 365 em_url = f"s3://janelia-cosem-datasets/{dataset_name}/{dataset_name}.zarr/{recon}/em/{em_name}/{em_level}" 366 em_arr = _open_remote_zarr(em_url) 367 ratio = [ls / es for ls, es in zip(label_scale, em_scale)] 368 if raw_bbox is not None: 369 em_bbox_native = tuple((int(lo * r), int(hi * r)) for (lo, hi), r in zip(raw_bbox, ratio)) 370 raw_native = em_arr[tuple(slice(*bb) for bb in em_bbox_native)] 371 else: 372 raw_native = np.asarray(em_arr[:]) 373 zoom_factors = [t / s for t, s in zip(label_data.shape, raw_native.shape)] 374 raw_data = zoom(raw_native, zoom_factors, order=1).astype(raw_native.dtype) 375 376 # Guard against off-by-a-few-voxel shape mismatches between independently stored pyramids. 377 common_shape = tuple(min(r, lb) for r, lb in zip(raw_data.shape, label_data.shape)) 378 raw_data = raw_data[tuple(slice(0, s) for s in common_shape)] 379 label_data = label_data[tuple(slice(0, s) for s in common_shape)] 380 381 assert raw_data.shape == label_data.shape, ( 382 f"Shape mismatch for '{dataset_name}': raw {raw_data.shape} vs labels {label_data.shape}" 383 ) 384 385 def _make_array(name, data, shuffle): 386 arr = root.create_array( 387 name, shape=data.shape, chunks=tuple(min(128, s) for s in data.shape), dtype=data.dtype, 388 compressors=BloscCodec(cname="zstd", clevel=6, shuffle=shuffle), 389 ) 390 arr[:] = data 391 392 root.attrs["dataset"] = dataset_name 393 root.attrs["recon"] = recon 394 root.attrs["label_kind"] = info["label_kind"] 395 root.attrs["label_target"] = info["label_target"] 396 root.attrs["label_source"] = label_url 397 root.attrs["em_source"] = em_url 398 root.attrs["resampled"] = resampled 399 root.attrs["license_verified"] = dataset_name not in UNVERIFIED_LICENSE 400 401 _make_array("raw", raw_data, shuffle="shuffle") 402 _make_array("labels", label_data, shuffle="bitshuffle") 403 404 print(f"Cached '{dataset_name}' to '{zarr_path}' (shape {raw_data.shape}).") 405 return zarr_path
Download nucleus segmentation data for one OpenOrganelle dataset.
Arguments:
- path: Filepath to a folder where the cached zarr store will be saved.
- dataset_name: The name of the dataset, one of the keys in
DATASETS. - download: Whether to download the data if it is not present.
Returns:
The filepath to the cached zarr store.
408def get_openorganelle_nucleus_paths( 409 path: Union[os.PathLike, str], dataset_names: Optional[List[str]] = None, download: bool = False, 410) -> List[str]: 411 """Get paths to cached OpenOrganelle nucleus zarr stores. 412 413 Args: 414 path: Filepath to a folder where the cached zarr stores will be saved. 415 dataset_names: The names of the datasets to use. Defaults to all datasets in `DATASETS`. 416 download: Whether to download the data if it is not present. 417 418 Returns: 419 List of filepaths to the cached zarr stores. 420 """ 421 if dataset_names is None: 422 dataset_names = list(DATASETS.keys()) 423 return [get_openorganelle_nucleus_data(path, name, download) for name in dataset_names]
Get paths to cached OpenOrganelle nucleus zarr stores.
Arguments:
- path: Filepath to a folder where the cached zarr stores will be saved.
- dataset_names: The names of the datasets to use. Defaults to all datasets in
DATASETS. - download: Whether to download the data if it is not present.
Returns:
List of filepaths to the cached zarr stores.
426def get_openorganelle_nucleus_dataset( 427 path: Union[os.PathLike, str], 428 patch_shape: Tuple[int, int, int], 429 dataset_names: Optional[List[str]] = None, 430 download: bool = False, 431 offsets: Optional[List[List[int]]] = None, 432 boundaries: bool = False, 433 **kwargs, 434) -> Dataset: 435 """Get the OpenOrganelle nucleus dataset for nucleus segmentation. 436 437 Args: 438 path: Filepath to a folder where the cached zarr stores will be saved. 439 patch_shape: The patch shape (z, y, x) to use for training. 440 dataset_names: The names of the datasets to use. Defaults to all datasets in `DATASETS`. 441 download: Whether to download the data if it is not present. 442 offsets: Offset values for affinity computation used as target. 443 boundaries: Whether to compute boundaries as the target. 444 kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset`. 445 446 Returns: 447 The segmentation dataset. 448 """ 449 assert len(patch_shape) == 3 450 451 paths = get_openorganelle_nucleus_paths(path, dataset_names, download) 452 453 kwargs = util.update_kwargs(kwargs, "is_seg_dataset", True) 454 kwargs, _ = util.add_instance_label_transform( 455 kwargs, add_binary_target=False, boundaries=boundaries, offsets=offsets 456 ) 457 458 return torch_em.default_segmentation_dataset( 459 raw_paths=paths, 460 raw_key="raw", 461 label_paths=paths, 462 label_key="labels", 463 patch_shape=patch_shape, 464 **kwargs, 465 )
Get the OpenOrganelle nucleus dataset for nucleus segmentation.
Arguments:
- path: Filepath to a folder where the cached zarr stores will be saved.
- patch_shape: The patch shape (z, y, x) to use for training.
- dataset_names: The names of the datasets to use. Defaults to all datasets in
DATASETS. - download: Whether to download the data if it is not present.
- offsets: Offset values for affinity computation used as target.
- boundaries: Whether to compute boundaries as the target.
- kwargs: Additional keyword arguments for
torch_em.default_segmentation_dataset.
Returns:
The segmentation dataset.
468def get_openorganelle_nucleus_loader( 469 path: Union[os.PathLike, str], 470 patch_shape: Tuple[int, int, int], 471 batch_size: int, 472 dataset_names: Optional[List[str]] = None, 473 download: bool = False, 474 offsets: Optional[List[List[int]]] = None, 475 boundaries: bool = False, 476 **kwargs, 477) -> DataLoader: 478 """Get the DataLoader for nucleus segmentation in the OpenOrganelle nucleus dataset. 479 480 Args: 481 path: Filepath to a folder where the cached zarr stores will be saved. 482 patch_shape: The patch shape (z, y, x) to use for training. 483 batch_size: The batch size for training. 484 dataset_names: The names of the datasets to use. Defaults to all datasets in `DATASETS`. 485 download: Whether to download the data if it is not present. 486 offsets: Offset values for affinity computation used as target. 487 boundaries: Whether to compute boundaries as the target. 488 kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset` 489 or for the PyTorch DataLoader. 490 491 Returns: 492 The DataLoader. 493 """ 494 ds_kwargs, loader_kwargs = util.split_kwargs(torch_em.default_segmentation_dataset, **kwargs) 495 dataset = get_openorganelle_nucleus_dataset( 496 path, patch_shape, dataset_names=dataset_names, download=download, 497 offsets=offsets, boundaries=boundaries, **ds_kwargs 498 ) 499 return torch_em.get_data_loader(dataset, batch_size, **loader_kwargs)
Get the DataLoader for nucleus segmentation in the OpenOrganelle nucleus dataset.
Arguments:
- path: Filepath to a folder where the cached zarr stores will be saved.
- patch_shape: The patch shape (z, y, x) to use for training.
- batch_size: The batch size for training.
- dataset_names: The names of the datasets to use. Defaults to all datasets in
DATASETS. - download: Whether to download the data if it is not present.
- offsets: Offset values for affinity computation used as target.
- boundaries: Whether to compute boundaries as the target.
- kwargs: Additional keyword arguments for
torch_em.default_segmentation_datasetor for the PyTorch DataLoader.
Returns:
The DataLoader.