torch_em.data.datasets.electron_microscopy.openorganelle_lipid_droplet
This dataset provides lipid droplet segmentation masks for representative crops from four
mouse liver volumes on the public "OpenOrganelle" janelia-cosem-datasets S3 bucket (the same
bucket used by cellmap.py, janelia_nucleus.py, and openorganelle_nucleus.py).
Three of the four datasets (jrc_mus-liver-4/5/6) are already covered by
openorganelle_nucleus.py, but only for their nuclear-membrane (nuc_mem) label; their separate
lipid droplet label, at labels/inference/segmentations/ld/{level}, is a distinct annotation not
provided by that module. jrc_mus-liver-7 is not covered by any existing torch-em loader.
The lipid droplet label is a binary mask (only values 0/255 seen) that exactly matches the raw
FIB-SEM data's multiscale pyramid at every level checked, so no resampling or axis correction is
needed. Lipid droplets in these fatty-liver volumes can be very large (macrovesicular steatosis),
so jrc_mus-liver-5 and jrc_mus-liver-6 use a coarser matched pyramid level (s3, 64nm) for
their fixed crop, chosen to show a real droplet boundary rather than a single, deep-interior,
fully-saturated block; jrc_mus-liver-4 and jrc_mus-liver-7 use the finest level (s0, 8nm).
The full matched-resolution raw+label pairs are far too large to download whole (hundreds of GB
to over 1 TB); this module downloads one fixed, manually verified crop per dataset, matching the
crop that was visually reviewed in napari before this loader was added. Data is CC0-1.0, per the
bucket's established convention, except for jrc_mus-liver-7, whose license could not be
independently confirmed (no dedicated landing page was found for this specific dataset).
Please cite the CellMap project (https://www.janelia.org/project-team/cellmap) if you use this data.
1"""This dataset provides lipid droplet segmentation masks for representative crops from four 2mouse liver volumes on the public "OpenOrganelle" `janelia-cosem-datasets` S3 bucket (the same 3bucket used by `cellmap.py`, `janelia_nucleus.py`, and `openorganelle_nucleus.py`). 4 5Three of the four datasets (`jrc_mus-liver-4/5/6`) are already covered by 6`openorganelle_nucleus.py`, but only for their nuclear-membrane (`nuc_mem`) label; their separate 7lipid droplet label, at `labels/inference/segmentations/ld/{level}`, is a distinct annotation not 8provided by that module. `jrc_mus-liver-7` is not covered by any existing torch-em loader. 9 10The lipid droplet label is a binary mask (only values 0/255 seen) that exactly matches the raw 11FIB-SEM data's multiscale pyramid at every level checked, so no resampling or axis correction is 12needed. Lipid droplets in these fatty-liver volumes can be very large (macrovesicular steatosis), 13so `jrc_mus-liver-5` and `jrc_mus-liver-6` use a coarser matched pyramid level (`s3`, 64nm) for 14their fixed crop, chosen to show a real droplet boundary rather than a single, deep-interior, 15fully-saturated block; `jrc_mus-liver-4` and `jrc_mus-liver-7` use the finest level (`s0`, 8nm). 16 17The full matched-resolution raw+label pairs are far too large to download whole (hundreds of GB 18to over 1 TB); this module downloads one fixed, manually verified crop per dataset, matching the 19crop that was visually reviewed in napari before this loader was added. Data is CC0-1.0, per the 20bucket's established convention, except for `jrc_mus-liver-7`, whose license could not be 21independently confirmed (no dedicated landing page was found for this specific dataset). 22 23Please cite the CellMap project (https://www.janelia.org/project-team/cellmap) if you use this 24data. 25""" 26 27import os 28from typing import List, Optional, Tuple, Union 29 30from torch.utils.data import DataLoader, Dataset 31 32import torch_em 33 34from .. import util 35 36 37BUCKET_URL = "https://janelia-cosem-datasets.s3.amazonaws.com/" 38LABEL_SUBPATH = "labels/inference/segmentations/ld" 39 40# dataset name -> recon folder, EM array name, ld pyramid level to use, and the bounding box 41# (in that level's voxel coordinates) of the fixed, visually reviewed crop. 42DATASETS = { 43 "jrc_mus-liver-4": { 44 "recon": "recon-1", "em_name": "fibsem-uint8", "label_level": "s0", 45 "bounding_box": ((704, 832), (1984, 2112), (4672, 4800)), 46 }, 47 "jrc_mus-liver-5": { 48 "recon": "recon-1", "em_name": "fibsem-uint8", "label_level": "s3", 49 "bounding_box": ((448, 640), (352, 544), (0, 128)), 50 }, 51 "jrc_mus-liver-6": { 52 "recon": "recon-1", "em_name": "fibsem-uint8", "label_level": "s3", 53 "bounding_box": ((288, 480), (0, 96), (736, 928)), 54 }, 55 "jrc_mus-liver-7": { 56 "recon": "recon-1", "em_name": "fibsem-uint8", "label_level": "s0", 57 "bounding_box": ((960, 1088), (5056, 5184), (8640, 8768)), 58 }, 59} 60 61# `jrc_mus-liver-7` has no dedicated landing page confirming its license; the others follow the 62# bucket's established CC0-1.0 convention (already verified for their existing torch-em entries). 63UNVERIFIED_LICENSE = ("jrc_mus-liver-7",) 64 65 66def _open_remote_zarr(s3_path): 67 import zarr 68 from zarr.storage import FsspecStore 69 70 # `FsspecStore` performs ranged partial reads, so a bounded crop only pulls the chunks it 71 # actually needs, unlike `fsspec.get_mapper` which can fetch whole shard/chunk files. 72 store = FsspecStore.from_url(s3_path, storage_options={"anon": True}, read_only=True) 73 return zarr.open(store, mode="r") 74 75 76def _get_json(url): 77 import requests 78 79 resp = requests.get(url, timeout=60) 80 resp.raise_for_status() 81 return resp.json() 82 83 84def _multiscale_levels(group_url): 85 attrs = _get_json(f"{group_url}/.zattrs") 86 return { 87 d["path"]: tuple(d["coordinateTransformations"][0]["scale"]) 88 for d in attrs["multiscales"][0]["datasets"] 89 } 90 91 92def _find_matching_em_level(dataset_name, recon, em_name, label_scale): 93 em_levels = _multiscale_levels(f"{BUCKET_URL}{dataset_name}/{dataset_name}.zarr/{recon}/em/{em_name}") 94 for level, scale in em_levels.items(): 95 if scale == label_scale: 96 return level 97 raise RuntimeError(f"No EM pyramid level of '{dataset_name}' matches the ld label scale {label_scale}.") 98 99 100def get_openorganelle_lipid_droplet_data( 101 path: Union[os.PathLike, str], dataset_name: str, download: bool = False 102) -> str: 103 """Download the reviewed lipid droplet crop for one OpenOrganelle dataset. 104 105 Args: 106 path: Filepath to a folder where the cached zarr store will be saved. 107 dataset_name: The name of the dataset, one of the keys in `DATASETS`. 108 download: Whether to download the data if it is not present. 109 110 Returns: 111 The filepath to the cached zarr store. 112 """ 113 import zarr 114 from zarr.codecs import BloscCodec 115 116 if dataset_name not in DATASETS: 117 raise ValueError(f"'{dataset_name}' is not a valid dataset name. Choose from {sorted(DATASETS.keys())}.") 118 119 info = DATASETS[dataset_name] 120 recon, em_name, label_level = info["recon"], info["em_name"], info["label_level"] 121 bounding_box = info["bounding_box"] 122 123 os.makedirs(str(path), exist_ok=True) 124 zarr_path = os.path.join(str(path), f"{dataset_name}.zarr") 125 126 root = zarr.open_group(zarr_path, mode="a") 127 if "raw" in root and "labels" in root: 128 return zarr_path 129 130 if not download: 131 raise RuntimeError(f"No cached crop found at '{zarr_path}'. Set download=True to stream it from S3.") 132 133 print(f"Streaming a lipid droplet crop for '{dataset_name}' from the janelia-cosem-datasets S3 bucket ...") 134 label_url = f"s3://janelia-cosem-datasets/{dataset_name}/{dataset_name}.zarr/{recon}/{LABEL_SUBPATH}/{label_level}" 135 label_arr = _open_remote_zarr(label_url) 136 label_slices = tuple(slice(*bb) for bb in bounding_box) 137 label_data = label_arr[label_slices] 138 139 label_scale = _multiscale_levels( 140 f"{BUCKET_URL}{dataset_name}/{dataset_name}.zarr/{recon}/{LABEL_SUBPATH}" 141 )[label_level] 142 em_level = _find_matching_em_level(dataset_name, recon, em_name, label_scale) 143 em_url = f"s3://janelia-cosem-datasets/{dataset_name}/{dataset_name}.zarr/{recon}/em/{em_name}/{em_level}" 144 em_arr = _open_remote_zarr(em_url) 145 raw_data = em_arr[label_slices] 146 147 common_shape = tuple(min(r, lb) for r, lb in zip(raw_data.shape, label_data.shape)) 148 raw_data = raw_data[tuple(slice(0, s) for s in common_shape)] 149 label_data = label_data[tuple(slice(0, s) for s in common_shape)] 150 151 assert raw_data.shape == label_data.shape, ( 152 f"Shape mismatch for '{dataset_name}': raw {raw_data.shape} vs labels {label_data.shape}" 153 ) 154 155 def _make_array(name, data, shuffle): 156 arr = root.create_array( 157 name, shape=data.shape, chunks=tuple(min(128, s) for s in data.shape), dtype=data.dtype, 158 compressors=BloscCodec(cname="zstd", clevel=6, shuffle=shuffle), 159 ) 160 arr[:] = data 161 162 root.attrs["dataset"] = dataset_name 163 root.attrs["recon"] = recon 164 root.attrs["label_target"] = "lipid_droplet" 165 root.attrs["label_source"] = label_url 166 root.attrs["em_source"] = em_url 167 root.attrs["license_verified"] = dataset_name not in UNVERIFIED_LICENSE 168 169 _make_array("raw", raw_data, shuffle="shuffle") 170 _make_array("labels", label_data, shuffle="bitshuffle") 171 172 print(f"Cached '{dataset_name}' to '{zarr_path}' (shape {raw_data.shape}).") 173 return zarr_path 174 175 176def get_openorganelle_lipid_droplet_paths( 177 path: Union[os.PathLike, str], dataset_names: Optional[List[str]] = None, download: bool = False, 178) -> List[str]: 179 """Get paths to cached OpenOrganelle lipid droplet zarr stores. 180 181 Args: 182 path: Filepath to a folder where the cached zarr stores will be saved. 183 dataset_names: The names of the datasets to use. Defaults to all datasets in `DATASETS`. 184 download: Whether to download the data if it is not present. 185 186 Returns: 187 List of filepaths to the cached zarr stores. 188 """ 189 if dataset_names is None: 190 dataset_names = list(DATASETS.keys()) 191 return [get_openorganelle_lipid_droplet_data(path, name, download) for name in dataset_names] 192 193 194def get_openorganelle_lipid_droplet_dataset( 195 path: Union[os.PathLike, str], 196 patch_shape: Tuple[int, int, int], 197 dataset_names: Optional[List[str]] = None, 198 download: bool = False, 199 **kwargs, 200) -> Dataset: 201 """Get the OpenOrganelle lipid droplet dataset for lipid droplet segmentation. 202 203 Args: 204 path: Filepath to a folder where the cached zarr stores will be saved. 205 patch_shape: The patch shape (z, y, x) to use for training. 206 dataset_names: The names of the datasets to use. Defaults to all datasets in `DATASETS`. 207 download: Whether to download the data if it is not present. 208 kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset`. 209 210 Returns: 211 The segmentation dataset. 212 """ 213 assert len(patch_shape) == 3 214 215 paths = get_openorganelle_lipid_droplet_paths(path, dataset_names, download) 216 217 kwargs = util.update_kwargs(kwargs, "is_seg_dataset", True) 218 219 return torch_em.default_segmentation_dataset( 220 raw_paths=paths, 221 raw_key="raw", 222 label_paths=paths, 223 label_key="labels", 224 patch_shape=patch_shape, 225 **kwargs, 226 ) 227 228 229def get_openorganelle_lipid_droplet_loader( 230 path: Union[os.PathLike, str], 231 patch_shape: Tuple[int, int, int], 232 batch_size: int, 233 dataset_names: Optional[List[str]] = None, 234 download: bool = False, 235 **kwargs, 236) -> DataLoader: 237 """Get the DataLoader for lipid droplet segmentation in the OpenOrganelle lipid droplet dataset. 238 239 Args: 240 path: Filepath to a folder where the cached zarr stores will be saved. 241 patch_shape: The patch shape (z, y, x) to use for training. 242 batch_size: The batch size for training. 243 dataset_names: The names of the datasets to use. Defaults to all datasets in `DATASETS`. 244 download: Whether to download the data if it is not present. 245 kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset` 246 or for the PyTorch DataLoader. 247 248 Returns: 249 The DataLoader. 250 """ 251 ds_kwargs, loader_kwargs = util.split_kwargs(torch_em.default_segmentation_dataset, **kwargs) 252 dataset = get_openorganelle_lipid_droplet_dataset( 253 path, patch_shape, dataset_names=dataset_names, download=download, **ds_kwargs 254 ) 255 return torch_em.get_data_loader(dataset, batch_size, **loader_kwargs)
101def get_openorganelle_lipid_droplet_data( 102 path: Union[os.PathLike, str], dataset_name: str, download: bool = False 103) -> str: 104 """Download the reviewed lipid droplet crop for one OpenOrganelle dataset. 105 106 Args: 107 path: Filepath to a folder where the cached zarr store will be saved. 108 dataset_name: The name of the dataset, one of the keys in `DATASETS`. 109 download: Whether to download the data if it is not present. 110 111 Returns: 112 The filepath to the cached zarr store. 113 """ 114 import zarr 115 from zarr.codecs import BloscCodec 116 117 if dataset_name not in DATASETS: 118 raise ValueError(f"'{dataset_name}' is not a valid dataset name. Choose from {sorted(DATASETS.keys())}.") 119 120 info = DATASETS[dataset_name] 121 recon, em_name, label_level = info["recon"], info["em_name"], info["label_level"] 122 bounding_box = info["bounding_box"] 123 124 os.makedirs(str(path), exist_ok=True) 125 zarr_path = os.path.join(str(path), f"{dataset_name}.zarr") 126 127 root = zarr.open_group(zarr_path, mode="a") 128 if "raw" in root and "labels" in root: 129 return zarr_path 130 131 if not download: 132 raise RuntimeError(f"No cached crop found at '{zarr_path}'. Set download=True to stream it from S3.") 133 134 print(f"Streaming a lipid droplet crop for '{dataset_name}' from the janelia-cosem-datasets S3 bucket ...") 135 label_url = f"s3://janelia-cosem-datasets/{dataset_name}/{dataset_name}.zarr/{recon}/{LABEL_SUBPATH}/{label_level}" 136 label_arr = _open_remote_zarr(label_url) 137 label_slices = tuple(slice(*bb) for bb in bounding_box) 138 label_data = label_arr[label_slices] 139 140 label_scale = _multiscale_levels( 141 f"{BUCKET_URL}{dataset_name}/{dataset_name}.zarr/{recon}/{LABEL_SUBPATH}" 142 )[label_level] 143 em_level = _find_matching_em_level(dataset_name, recon, em_name, label_scale) 144 em_url = f"s3://janelia-cosem-datasets/{dataset_name}/{dataset_name}.zarr/{recon}/em/{em_name}/{em_level}" 145 em_arr = _open_remote_zarr(em_url) 146 raw_data = em_arr[label_slices] 147 148 common_shape = tuple(min(r, lb) for r, lb in zip(raw_data.shape, label_data.shape)) 149 raw_data = raw_data[tuple(slice(0, s) for s in common_shape)] 150 label_data = label_data[tuple(slice(0, s) for s in common_shape)] 151 152 assert raw_data.shape == label_data.shape, ( 153 f"Shape mismatch for '{dataset_name}': raw {raw_data.shape} vs labels {label_data.shape}" 154 ) 155 156 def _make_array(name, data, shuffle): 157 arr = root.create_array( 158 name, shape=data.shape, chunks=tuple(min(128, s) for s in data.shape), dtype=data.dtype, 159 compressors=BloscCodec(cname="zstd", clevel=6, shuffle=shuffle), 160 ) 161 arr[:] = data 162 163 root.attrs["dataset"] = dataset_name 164 root.attrs["recon"] = recon 165 root.attrs["label_target"] = "lipid_droplet" 166 root.attrs["label_source"] = label_url 167 root.attrs["em_source"] = em_url 168 root.attrs["license_verified"] = dataset_name not in UNVERIFIED_LICENSE 169 170 _make_array("raw", raw_data, shuffle="shuffle") 171 _make_array("labels", label_data, shuffle="bitshuffle") 172 173 print(f"Cached '{dataset_name}' to '{zarr_path}' (shape {raw_data.shape}).") 174 return zarr_path
Download the reviewed lipid droplet crop for one OpenOrganelle dataset.
Arguments:
- path: Filepath to a folder where the cached zarr store will be saved.
- dataset_name: The name of the dataset, one of the keys in
DATASETS. - download: Whether to download the data if it is not present.
Returns:
The filepath to the cached zarr store.
177def get_openorganelle_lipid_droplet_paths( 178 path: Union[os.PathLike, str], dataset_names: Optional[List[str]] = None, download: bool = False, 179) -> List[str]: 180 """Get paths to cached OpenOrganelle lipid droplet zarr stores. 181 182 Args: 183 path: Filepath to a folder where the cached zarr stores will be saved. 184 dataset_names: The names of the datasets to use. Defaults to all datasets in `DATASETS`. 185 download: Whether to download the data if it is not present. 186 187 Returns: 188 List of filepaths to the cached zarr stores. 189 """ 190 if dataset_names is None: 191 dataset_names = list(DATASETS.keys()) 192 return [get_openorganelle_lipid_droplet_data(path, name, download) for name in dataset_names]
Get paths to cached OpenOrganelle lipid droplet zarr stores.
Arguments:
- path: Filepath to a folder where the cached zarr stores will be saved.
- dataset_names: The names of the datasets to use. Defaults to all datasets in
DATASETS. - download: Whether to download the data if it is not present.
Returns:
List of filepaths to the cached zarr stores.
195def get_openorganelle_lipid_droplet_dataset( 196 path: Union[os.PathLike, str], 197 patch_shape: Tuple[int, int, int], 198 dataset_names: Optional[List[str]] = None, 199 download: bool = False, 200 **kwargs, 201) -> Dataset: 202 """Get the OpenOrganelle lipid droplet dataset for lipid droplet segmentation. 203 204 Args: 205 path: Filepath to a folder where the cached zarr stores will be saved. 206 patch_shape: The patch shape (z, y, x) to use for training. 207 dataset_names: The names of the datasets to use. Defaults to all datasets in `DATASETS`. 208 download: Whether to download the data if it is not present. 209 kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset`. 210 211 Returns: 212 The segmentation dataset. 213 """ 214 assert len(patch_shape) == 3 215 216 paths = get_openorganelle_lipid_droplet_paths(path, dataset_names, download) 217 218 kwargs = util.update_kwargs(kwargs, "is_seg_dataset", True) 219 220 return torch_em.default_segmentation_dataset( 221 raw_paths=paths, 222 raw_key="raw", 223 label_paths=paths, 224 label_key="labels", 225 patch_shape=patch_shape, 226 **kwargs, 227 )
Get the OpenOrganelle lipid droplet dataset for lipid droplet segmentation.
Arguments:
- path: Filepath to a folder where the cached zarr stores will be saved.
- patch_shape: The patch shape (z, y, x) to use for training.
- dataset_names: The names of the datasets to use. Defaults to all datasets in
DATASETS. - download: Whether to download the data if it is not present.
- kwargs: Additional keyword arguments for
torch_em.default_segmentation_dataset.
Returns:
The segmentation dataset.
230def get_openorganelle_lipid_droplet_loader( 231 path: Union[os.PathLike, str], 232 patch_shape: Tuple[int, int, int], 233 batch_size: int, 234 dataset_names: Optional[List[str]] = None, 235 download: bool = False, 236 **kwargs, 237) -> DataLoader: 238 """Get the DataLoader for lipid droplet segmentation in the OpenOrganelle lipid droplet dataset. 239 240 Args: 241 path: Filepath to a folder where the cached zarr stores will be saved. 242 patch_shape: The patch shape (z, y, x) to use for training. 243 batch_size: The batch size for training. 244 dataset_names: The names of the datasets to use. Defaults to all datasets in `DATASETS`. 245 download: Whether to download the data if it is not present. 246 kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset` 247 or for the PyTorch DataLoader. 248 249 Returns: 250 The DataLoader. 251 """ 252 ds_kwargs, loader_kwargs = util.split_kwargs(torch_em.default_segmentation_dataset, **kwargs) 253 dataset = get_openorganelle_lipid_droplet_dataset( 254 path, patch_shape, dataset_names=dataset_names, download=download, **ds_kwargs 255 ) 256 return torch_em.get_data_loader(dataset, batch_size, **loader_kwargs)
Get the DataLoader for lipid droplet segmentation in the OpenOrganelle lipid droplet dataset.
Arguments:
- path: Filepath to a folder where the cached zarr stores will be saved.
- patch_shape: The patch shape (z, y, x) to use for training.
- batch_size: The batch size for training.
- dataset_names: The names of the datasets to use. Defaults to all datasets in
DATASETS. - download: Whether to download the data if it is not present.
- kwargs: Additional keyword arguments for
torch_em.default_segmentation_datasetor for the PyTorch DataLoader.
Returns:
The DataLoader.