torch_em.data.datasets.electron_microscopy.janelia_nucleus
The Janelia CellMap nucleus dataset provides nucleus instance segmentation masks for FIB-SEM
volumes from the public janelia-cosem-datasets S3 bucket (the same "OpenOrganelle" bucket used
by other CellMap data), covering several mouse tissues and one Drosophila tissue.
For 13 of the 16 volumes, nuclei were segmented automatically with Cellpose 3.0.9 and only
cursorily corrected by hand, following the protocol
"Generating Nuclei Segmentations for vEM datasets using Cellpose"
(https://dx.doi.org/10.17504/protocols.io.n2bvjnq8pgk5/v1). Treat these labels the same way as
the MitoNet auto-labels in mitonet_predicted_kidney.py: useful pseudo-labels, not verified
ground truth.
The 3 jrc_mus-nacc-* volumes (mouse nucleus accumbens) are instead densely, manually segmented
in Amira-Avizo, following
"Using Amira to manually segment organelles in vEM for machine learning V.3"
(https://dx.doi.org/10.17504/protocols.io.bp2l61rb5vqe/v3). These can be treated as reliable
ground truth.
Each volume is released as its own Figshare deposit (e.g. https://doi.org/10.6084/m9.figshare.26506555),
under CC-BY-4.0. The Figshare entries are metadata-only pointers: their one "file" is a link-only
stub whose actual content lives on the S3 bucket, at
s3://janelia-cosem-datasets/{dataset}/{dataset}.zarr/{recon}/em/... (raw) and
.../{recon}/labels/inference/segmentations/nuc/... (nucleus instance labels), where recon is
either "recon-1" or "recon-2" depending on the deposit.
The raw EM data is stored as a multiscale OME-Zarr pyramid at up to 4 nm isotropic resolution,
while the nucleus labels are only released at one, coarser resolution (their pyramid level "s0").
This module finds the raw pyramid level whose resolution matches the label's "s0" level (their
shapes then match exactly, so no rescaling is needed, unlike the mito/tissue labels in
hemibrain.py which need explicit upsampling) and downloads only that matched pair, in full,
for each requested dataset name.
Please cite the Cellpose or Amira-Avizo protocol above (as appropriate for the chosen dataset) and the CellMap project (https://www.janelia.org/project-team/cellmap) if you use this data.
1"""The Janelia CellMap nucleus dataset provides nucleus instance segmentation masks for FIB-SEM 2volumes from the public `janelia-cosem-datasets` S3 bucket (the same "OpenOrganelle" bucket used 3by other CellMap data), covering several mouse tissues and one Drosophila tissue. 4 5For 13 of the 16 volumes, nuclei were segmented automatically with Cellpose 3.0.9 and only 6cursorily corrected by hand, following the protocol 7"Generating Nuclei Segmentations for vEM datasets using Cellpose" 8(https://dx.doi.org/10.17504/protocols.io.n2bvjnq8pgk5/v1). Treat these labels the same way as 9the MitoNet auto-labels in `mitonet_predicted_kidney.py`: useful pseudo-labels, not verified 10ground truth. 11 12The 3 `jrc_mus-nacc-*` volumes (mouse nucleus accumbens) are instead densely, manually segmented 13in Amira-Avizo, following 14"Using Amira to manually segment organelles in vEM for machine learning V.3" 15(https://dx.doi.org/10.17504/protocols.io.bp2l61rb5vqe/v3). These can be treated as reliable 16ground truth. 17 18Each volume is released as its own Figshare deposit (e.g. https://doi.org/10.6084/m9.figshare.26506555), 19under CC-BY-4.0. The Figshare entries are metadata-only pointers: their one "file" is a link-only 20stub whose actual content lives on the S3 bucket, at 21`s3://janelia-cosem-datasets/{dataset}/{dataset}.zarr/{recon}/em/...` (raw) and 22`.../{recon}/labels/inference/segmentations/nuc/...` (nucleus instance labels), where `recon` is 23either "recon-1" or "recon-2" depending on the deposit. 24 25The raw EM data is stored as a multiscale OME-Zarr pyramid at up to 4 nm isotropic resolution, 26while the nucleus labels are only released at one, coarser resolution (their pyramid level "s0"). 27This module finds the raw pyramid level whose resolution matches the label's "s0" level (their 28shapes then match exactly, so no rescaling is needed, unlike the mito/tissue labels in 29`hemibrain.py` which need explicit upsampling) and downloads only that matched pair, in full, 30for each requested dataset name. 31 32Please cite the Cellpose or Amira-Avizo protocol above (as appropriate for the chosen dataset) 33and the CellMap project (https://www.janelia.org/project-team/cellmap) if you use this data. 34""" 35 36import os 37import re 38from typing import List, Optional, Tuple, Union 39 40import numpy as np 41 42from torch.utils.data import DataLoader, Dataset 43 44import torch_em 45 46from .. import util 47 48 49BUCKET_URL = "https://janelia-cosem-datasets.s3.amazonaws.com/" 50 51# dataset name -> (figshare article id, recon folder, annotation type) 52# "cellpose" = automatic Cellpose 3.0.9 prediction with cursory manual correction. 53# "manual" = dense manual segmentation in Amira-Avizo. 54DATASETS = { 55 "jrc_mus-pancreas-1": (26506552, "recon-1", "cellpose"), 56 "jrc_mus-pancreas-2": (26506555, "recon-1", "cellpose"), 57 "jrc_mus-pancreas-3": (26506561, "recon-1", "cellpose"), 58 "jrc_mus-granule-neurons-1": (26506510, "recon-2", "cellpose"), 59 "jrc_mus-granule-neurons-2": (26506513, "recon-2", "cellpose"), 60 "jrc_mus-granule-neurons-3": (26506516, "recon-2", "cellpose"), 61 "jrc_mus-guard-hair-follicle": (26506519, "recon-1", "cellpose"), 62 "jrc_mus-kidney": (26506522, "recon-1", "cellpose"), 63 "jrc_mus-kidney-2": (26506534, "recon-1", "cellpose"), 64 "jrc_mus-liver-2": (26506537, "recon-1", "cellpose"), 65 "jrc_mus-meissner-corpuscle-1": (26506540, "recon-1", "cellpose"), 66 "jrc_mus-meissner-corpuscle-2": (26506543, "recon-1", "cellpose"), 67 "jrc_fly-mb-z0419-20": (26506507, "recon-1", "cellpose"), 68 "jrc_mus-nacc-2": (26513209, "recon-2", "manual"), 69 "jrc_mus-nacc-3": (26513212, "recon-2", "manual"), 70 "jrc_mus-nacc-4": (26513215, "recon-2", "manual"), 71} 72 73 74def _get_json(url): 75 import requests 76 resp = requests.get(url, timeout=60) 77 resp.raise_for_status() 78 return resp.json() 79 80 81def _list_prefixes(prefix): 82 import requests 83 resp = requests.get(BUCKET_URL, params={"prefix": prefix, "delimiter": "/"}, timeout=60) 84 resp.raise_for_status() 85 return re.findall(r"<Prefix>([^<]*)</Prefix>", resp.text) 86 87 88def _find_em_array_name(dataset_name, recon): 89 prefix = f"{dataset_name}/{dataset_name}.zarr/{recon}/em/" 90 names = [p[len(prefix):].rstrip("/") for p in _list_prefixes(prefix) if p != prefix] 91 if not names: 92 raise RuntimeError(f"Could not find an 'em' array for '{dataset_name}' under '{prefix}'.") 93 return names[0] 94 95 96def _find_matching_em_level(dataset_name, recon, em_name): 97 em_attrs = _get_json(f"{BUCKET_URL}{dataset_name}/{dataset_name}.zarr/{recon}/em/{em_name}/.zattrs") 98 nuc_attrs = _get_json( 99 f"{BUCKET_URL}{dataset_name}/{dataset_name}.zarr/{recon}/labels/inference/segmentations/nuc/.zattrs" 100 ) 101 nuc_scale = nuc_attrs["multiscales"][0]["datasets"][0]["coordinateTransformations"][0]["scale"] 102 for level in em_attrs["multiscales"][0]["datasets"]: 103 scale = level["coordinateTransformations"][0]["scale"] 104 if scale == nuc_scale: 105 return level["path"] 106 raise RuntimeError( 107 f"No EM resolution level of '{dataset_name}' matches the nucleus label resolution {nuc_scale}." 108 ) 109 110 111def _read_zarr_array(s3_path): 112 import zarr 113 import fsspec 114 115 store = fsspec.get_mapper(s3_path, anon=True) 116 array = zarr.open(store, mode="r") 117 return np.asarray(array[:]) 118 119 120def get_janelia_nucleus_data(path: Union[os.PathLike, str], dataset_name: str, download: bool = False) -> str: 121 """Download nucleus instance segmentation data for one Janelia CellMap dataset. 122 123 Args: 124 path: Filepath to a folder where the cached zarr store will be saved. 125 dataset_name: The name of the dataset, one of the keys in `DATASETS`. 126 download: Whether to download the data if it is not present. 127 128 Returns: 129 The filepath to the cached zarr store. 130 """ 131 import zarr 132 from zarr.codecs import BloscCodec 133 134 if dataset_name not in DATASETS: 135 raise ValueError(f"'{dataset_name}' is not a valid dataset name. Choose from {sorted(DATASETS.keys())}.") 136 137 os.makedirs(str(path), exist_ok=True) 138 zarr_path = os.path.join(str(path), f"{dataset_name}.zarr") 139 140 root = zarr.open_group(zarr_path, mode="a") 141 if "raw" in root and "labels" in root: 142 return zarr_path 143 144 if not download: 145 raise RuntimeError(f"No cached data found at '{zarr_path}'. Set download=True to stream it from S3.") 146 147 _, recon, annotation = DATASETS[dataset_name] 148 em_name = _find_em_array_name(dataset_name, recon) 149 em_level = _find_matching_em_level(dataset_name, recon, em_name) 150 151 em_s3 = f"s3://janelia-cosem-datasets/{dataset_name}/{dataset_name}.zarr/{recon}/em/{em_name}/{em_level}" 152 nuc_s3 = ( 153 f"s3://janelia-cosem-datasets/{dataset_name}/{dataset_name}.zarr/{recon}/" 154 "labels/inference/segmentations/nuc/s0" 155 ) 156 157 print(f"Streaming '{dataset_name}' ({recon}, {annotation} labels) from the janelia-cosem-datasets S3 bucket ...") 158 raw = _read_zarr_array(em_s3) 159 labels = _read_zarr_array(nuc_s3) 160 assert raw.shape == labels.shape, f"Shape mismatch for '{dataset_name}': raw {raw.shape} vs labels {labels.shape}" 161 162 def _make_array(name, data, shuffle): 163 arr = root.create_array( 164 name, shape=data.shape, chunks=tuple(min(128, s) for s in data.shape), dtype=data.dtype, 165 compressors=BloscCodec(cname="zstd", clevel=6, shuffle=shuffle), 166 ) 167 arr[:] = data 168 169 root.attrs["dataset"] = dataset_name 170 root.attrs["recon"] = recon 171 root.attrs["annotation_type"] = annotation 172 root.attrs["labels_are_manual"] = annotation == "manual" 173 root.attrs["em_source"] = em_s3 174 root.attrs["label_source"] = nuc_s3 175 176 _make_array("raw", raw, shuffle="shuffle") 177 _make_array("labels", labels, shuffle="bitshuffle") 178 179 print(f"Cached '{dataset_name}' to '{zarr_path}' (shape {raw.shape}).") 180 return zarr_path 181 182 183def get_janelia_nucleus_paths( 184 path: Union[os.PathLike, str], dataset_names: Optional[List[str]] = None, download: bool = False, 185) -> List[str]: 186 """Get paths to cached Janelia CellMap nucleus zarr stores. 187 188 Args: 189 path: Filepath to a folder where the cached zarr stores will be saved. 190 dataset_names: The names of the datasets to use. Defaults to all datasets in `DATASETS`. 191 download: Whether to download the data if it is not present. 192 193 Returns: 194 List of filepaths to the cached zarr stores. 195 """ 196 if dataset_names is None: 197 dataset_names = list(DATASETS.keys()) 198 return [get_janelia_nucleus_data(path, name, download) for name in dataset_names] 199 200 201def get_janelia_nucleus_dataset( 202 path: Union[os.PathLike, str], 203 patch_shape: Tuple[int, int, int], 204 dataset_names: Optional[List[str]] = None, 205 download: bool = False, 206 offsets: Optional[List[List[int]]] = None, 207 boundaries: bool = False, 208 **kwargs, 209) -> Dataset: 210 """Get the Janelia CellMap nucleus dataset for nucleus instance segmentation. 211 212 Args: 213 path: Filepath to a folder where the cached zarr stores will be saved. 214 patch_shape: The patch shape (z, y, x) to use for training. 215 dataset_names: The names of the datasets to use. Defaults to all datasets in `DATASETS`. 216 download: Whether to download the data if it is not present. 217 offsets: Offset values for affinity computation used as target. 218 boundaries: Whether to compute boundaries as the target. 219 kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset`. 220 221 Returns: 222 The segmentation dataset. 223 """ 224 assert len(patch_shape) == 3 225 226 paths = get_janelia_nucleus_paths(path, dataset_names, download) 227 228 kwargs = util.update_kwargs(kwargs, "is_seg_dataset", True) 229 kwargs, _ = util.add_instance_label_transform( 230 kwargs, add_binary_target=False, boundaries=boundaries, offsets=offsets 231 ) 232 233 return torch_em.default_segmentation_dataset( 234 raw_paths=paths, 235 raw_key="raw", 236 label_paths=paths, 237 label_key="labels", 238 patch_shape=patch_shape, 239 **kwargs, 240 ) 241 242 243def get_janelia_nucleus_loader( 244 path: Union[os.PathLike, str], 245 patch_shape: Tuple[int, int, int], 246 batch_size: int, 247 dataset_names: Optional[List[str]] = None, 248 download: bool = False, 249 offsets: Optional[List[List[int]]] = None, 250 boundaries: bool = False, 251 **kwargs, 252) -> DataLoader: 253 """Get the DataLoader for nucleus instance segmentation in the Janelia CellMap nucleus dataset. 254 255 Args: 256 path: Filepath to a folder where the cached zarr stores will be saved. 257 patch_shape: The patch shape (z, y, x) to use for training. 258 batch_size: The batch size for training. 259 dataset_names: The names of the datasets to use. Defaults to all datasets in `DATASETS`. 260 download: Whether to download the data if it is not present. 261 offsets: Offset values for affinity computation used as target. 262 boundaries: Whether to compute boundaries as the target. 263 kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset` 264 or for the PyTorch DataLoader. 265 266 Returns: 267 The DataLoader. 268 """ 269 ds_kwargs, loader_kwargs = util.split_kwargs(torch_em.default_segmentation_dataset, **kwargs) 270 dataset = get_janelia_nucleus_dataset( 271 path, patch_shape, dataset_names=dataset_names, download=download, 272 offsets=offsets, boundaries=boundaries, **ds_kwargs 273 ) 274 return torch_em.get_data_loader(dataset, batch_size, **loader_kwargs)
121def get_janelia_nucleus_data(path: Union[os.PathLike, str], dataset_name: str, download: bool = False) -> str: 122 """Download nucleus instance segmentation data for one Janelia CellMap dataset. 123 124 Args: 125 path: Filepath to a folder where the cached zarr store will be saved. 126 dataset_name: The name of the dataset, one of the keys in `DATASETS`. 127 download: Whether to download the data if it is not present. 128 129 Returns: 130 The filepath to the cached zarr store. 131 """ 132 import zarr 133 from zarr.codecs import BloscCodec 134 135 if dataset_name not in DATASETS: 136 raise ValueError(f"'{dataset_name}' is not a valid dataset name. Choose from {sorted(DATASETS.keys())}.") 137 138 os.makedirs(str(path), exist_ok=True) 139 zarr_path = os.path.join(str(path), f"{dataset_name}.zarr") 140 141 root = zarr.open_group(zarr_path, mode="a") 142 if "raw" in root and "labels" in root: 143 return zarr_path 144 145 if not download: 146 raise RuntimeError(f"No cached data found at '{zarr_path}'. Set download=True to stream it from S3.") 147 148 _, recon, annotation = DATASETS[dataset_name] 149 em_name = _find_em_array_name(dataset_name, recon) 150 em_level = _find_matching_em_level(dataset_name, recon, em_name) 151 152 em_s3 = f"s3://janelia-cosem-datasets/{dataset_name}/{dataset_name}.zarr/{recon}/em/{em_name}/{em_level}" 153 nuc_s3 = ( 154 f"s3://janelia-cosem-datasets/{dataset_name}/{dataset_name}.zarr/{recon}/" 155 "labels/inference/segmentations/nuc/s0" 156 ) 157 158 print(f"Streaming '{dataset_name}' ({recon}, {annotation} labels) from the janelia-cosem-datasets S3 bucket ...") 159 raw = _read_zarr_array(em_s3) 160 labels = _read_zarr_array(nuc_s3) 161 assert raw.shape == labels.shape, f"Shape mismatch for '{dataset_name}': raw {raw.shape} vs labels {labels.shape}" 162 163 def _make_array(name, data, shuffle): 164 arr = root.create_array( 165 name, shape=data.shape, chunks=tuple(min(128, s) for s in data.shape), dtype=data.dtype, 166 compressors=BloscCodec(cname="zstd", clevel=6, shuffle=shuffle), 167 ) 168 arr[:] = data 169 170 root.attrs["dataset"] = dataset_name 171 root.attrs["recon"] = recon 172 root.attrs["annotation_type"] = annotation 173 root.attrs["labels_are_manual"] = annotation == "manual" 174 root.attrs["em_source"] = em_s3 175 root.attrs["label_source"] = nuc_s3 176 177 _make_array("raw", raw, shuffle="shuffle") 178 _make_array("labels", labels, shuffle="bitshuffle") 179 180 print(f"Cached '{dataset_name}' to '{zarr_path}' (shape {raw.shape}).") 181 return zarr_path
Download nucleus instance segmentation data for one Janelia CellMap dataset.
Arguments:
- path: Filepath to a folder where the cached zarr store will be saved.
- dataset_name: The name of the dataset, one of the keys in
DATASETS. - download: Whether to download the data if it is not present.
Returns:
The filepath to the cached zarr store.
184def get_janelia_nucleus_paths( 185 path: Union[os.PathLike, str], dataset_names: Optional[List[str]] = None, download: bool = False, 186) -> List[str]: 187 """Get paths to cached Janelia CellMap nucleus zarr stores. 188 189 Args: 190 path: Filepath to a folder where the cached zarr stores will be saved. 191 dataset_names: The names of the datasets to use. Defaults to all datasets in `DATASETS`. 192 download: Whether to download the data if it is not present. 193 194 Returns: 195 List of filepaths to the cached zarr stores. 196 """ 197 if dataset_names is None: 198 dataset_names = list(DATASETS.keys()) 199 return [get_janelia_nucleus_data(path, name, download) for name in dataset_names]
Get paths to cached Janelia CellMap nucleus zarr stores.
Arguments:
- path: Filepath to a folder where the cached zarr stores will be saved.
- dataset_names: The names of the datasets to use. Defaults to all datasets in
DATASETS. - download: Whether to download the data if it is not present.
Returns:
List of filepaths to the cached zarr stores.
202def get_janelia_nucleus_dataset( 203 path: Union[os.PathLike, str], 204 patch_shape: Tuple[int, int, int], 205 dataset_names: Optional[List[str]] = None, 206 download: bool = False, 207 offsets: Optional[List[List[int]]] = None, 208 boundaries: bool = False, 209 **kwargs, 210) -> Dataset: 211 """Get the Janelia CellMap nucleus dataset for nucleus instance segmentation. 212 213 Args: 214 path: Filepath to a folder where the cached zarr stores will be saved. 215 patch_shape: The patch shape (z, y, x) to use for training. 216 dataset_names: The names of the datasets to use. Defaults to all datasets in `DATASETS`. 217 download: Whether to download the data if it is not present. 218 offsets: Offset values for affinity computation used as target. 219 boundaries: Whether to compute boundaries as the target. 220 kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset`. 221 222 Returns: 223 The segmentation dataset. 224 """ 225 assert len(patch_shape) == 3 226 227 paths = get_janelia_nucleus_paths(path, dataset_names, download) 228 229 kwargs = util.update_kwargs(kwargs, "is_seg_dataset", True) 230 kwargs, _ = util.add_instance_label_transform( 231 kwargs, add_binary_target=False, boundaries=boundaries, offsets=offsets 232 ) 233 234 return torch_em.default_segmentation_dataset( 235 raw_paths=paths, 236 raw_key="raw", 237 label_paths=paths, 238 label_key="labels", 239 patch_shape=patch_shape, 240 **kwargs, 241 )
Get the Janelia CellMap nucleus dataset for nucleus instance segmentation.
Arguments:
- path: Filepath to a folder where the cached zarr stores will be saved.
- patch_shape: The patch shape (z, y, x) to use for training.
- dataset_names: The names of the datasets to use. Defaults to all datasets in
DATASETS. - download: Whether to download the data if it is not present.
- offsets: Offset values for affinity computation used as target.
- boundaries: Whether to compute boundaries as the target.
- kwargs: Additional keyword arguments for
torch_em.default_segmentation_dataset.
Returns:
The segmentation dataset.
244def get_janelia_nucleus_loader( 245 path: Union[os.PathLike, str], 246 patch_shape: Tuple[int, int, int], 247 batch_size: int, 248 dataset_names: Optional[List[str]] = None, 249 download: bool = False, 250 offsets: Optional[List[List[int]]] = None, 251 boundaries: bool = False, 252 **kwargs, 253) -> DataLoader: 254 """Get the DataLoader for nucleus instance segmentation in the Janelia CellMap nucleus dataset. 255 256 Args: 257 path: Filepath to a folder where the cached zarr stores will be saved. 258 patch_shape: The patch shape (z, y, x) to use for training. 259 batch_size: The batch size for training. 260 dataset_names: The names of the datasets to use. Defaults to all datasets in `DATASETS`. 261 download: Whether to download the data if it is not present. 262 offsets: Offset values for affinity computation used as target. 263 boundaries: Whether to compute boundaries as the target. 264 kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset` 265 or for the PyTorch DataLoader. 266 267 Returns: 268 The DataLoader. 269 """ 270 ds_kwargs, loader_kwargs = util.split_kwargs(torch_em.default_segmentation_dataset, **kwargs) 271 dataset = get_janelia_nucleus_dataset( 272 path, patch_shape, dataset_names=dataset_names, download=download, 273 offsets=offsets, boundaries=boundaries, **ds_kwargs 274 ) 275 return torch_em.get_data_loader(dataset, batch_size, **loader_kwargs)
Get the DataLoader for nucleus instance segmentation in the Janelia CellMap nucleus dataset.
Arguments:
- path: Filepath to a folder where the cached zarr stores will be saved.
- patch_shape: The patch shape (z, y, x) to use for training.
- batch_size: The batch size for training.
- dataset_names: The names of the datasets to use. Defaults to all datasets in
DATASETS. - download: Whether to download the data if it is not present.
- offsets: Offset values for affinity computation used as target.
- boundaries: Whether to compute boundaries as the target.
- kwargs: Additional keyword arguments for
torch_em.default_segmentation_datasetor for the PyTorch DataLoader.
Returns:
The DataLoader.