torch_em.data.datasets.electron_microscopy.cortex_connectomics
The CortexConnectomics dataset contains dense electron microscopy connectomics volumes of cortex from mouse, macaque, and human, with proofread neuron instance segmentation, acquired to compare synaptic circuit architecture across species.
Nine of the twelve datasets published with the paper include a proofread segmentation layer; the remaining three (two human MultiSEM volumes and one extended mouse posterior parietal cortex volume) provide only raw EM data and are not included here.
Segmentation is dense but not uniformly proofread throughout each volume: some regions still show block-stitching seams from the automated reconstruction, visible as brief instance ID mismatches at internal block boundaries.
The data is available at https://demo.wk1.connectomics.hpccloud.mpg.de (organization 'Helmstaedter Lab', Max Planck Institute for Brain Research). No license is specified by the data provider. The dataset was published in https://doi.org/10.1126/science.abo0924. Please cite this publication if you use the dataset in your research.
1"""The CortexConnectomics dataset contains dense electron microscopy connectomics volumes of 2cortex from mouse, macaque, and human, with proofread neuron instance segmentation, acquired to 3compare synaptic circuit architecture across species. 4 5Nine of the twelve datasets published with the paper include a proofread segmentation layer; the 6remaining three (two human MultiSEM volumes and one extended mouse posterior parietal cortex 7volume) provide only raw EM data and are not included here. 8 9Segmentation is dense but not uniformly proofread throughout each volume: some regions still show 10block-stitching seams from the automated reconstruction, visible as brief instance ID mismatches 11at internal block boundaries. 12 13The data is available at https://demo.wk1.connectomics.hpccloud.mpg.de (organization 14'Helmstaedter Lab', Max Planck Institute for Brain Research). No license is specified by the data 15provider. 16The dataset was published in https://doi.org/10.1126/science.abo0924. 17Please cite this publication if you use the dataset in your research. 18""" 19 20import os 21from typing import Dict, List, Literal, Tuple, Union 22 23import numpy as np 24 25from torch.utils.data import DataLoader, Dataset 26 27import torch_em 28 29from .. import util 30 31 32BASE_URL = "https://demo.wk1.connectomics.hpccloud.mpg.de/data/zarr" 33 34SAMPLES: Dict[str, dict] = { 35 "mouse_s1": { 36 "dataset_id": "62b19325010000860033e7c9", "shape_zyx": (7630, 5030, 7874), 37 "resolution_nm": (28, 11.24, 11.24), 38 }, 39 "mouse_ppc": { 40 "dataset_id": "5c1cd439010000e66f8c2786", "shape_zyx": (4864, 12032, 12032), "resolution_nm": (30, 12, 12), 41 }, 42 "mouse_acc": { 43 "dataset_id": "5c1cd43801000026708c2783", "shape_zyx": (3328, 14336, 9216), "resolution_nm": (30, 12, 12), 44 }, 45 "mouse_v2": { 46 "dataset_id": "5c1cd4380100004f6f8c2785", "shape_zyx": (5120, 10240, 7168), "resolution_nm": (30, 12, 12), 47 }, 48 "mouse_a2": { 49 "dataset_id": "62b17edd010000f20075d7b9", "shape_zyx": (3642, 16102, 10394), 50 "resolution_nm": (30, 11.24, 11.24), 51 }, 52 "macaque_s1": { 53 "dataset_id": "62b17f19010000aa0075d7bb", "shape_zyx": (3359, 19544, 14942), 54 "resolution_nm": (30, 11.24, 11.24), 55 }, 56 "macaque_stg": { 57 "dataset_id": "62b17f55010000eb0075d7be", "shape_zyx": (3599, 20324, 15907), 58 "resolution_nm": (30, 11.24, 11.24), 59 }, 60 "human_stg": { 61 "dataset_id": "62b17f19010000aa0075d7bc", "shape_zyx": (3759, 19260, 14845), 62 "resolution_nm": (30, 11.24, 11.24), 63 }, 64 "human_ifg": { 65 "dataset_id": "62b17fcd010000eb0075d7c0", "shape_zyx": (2626, 19197, 15110), 66 "resolution_nm": (30, 11.24, 11.24), 67 }, 68} 69 70 71def _bbox_to_str(bounding_box): 72 import hashlib 73 return hashlib.md5("_".join(str(v) for v in bounding_box).encode()).hexdigest()[:12] 74 75 76def _open_remote_array(dataset_id, layer): 77 import fsspec 78 import zarr 79 return zarr.open(fsspec.get_mapper(f"{BASE_URL}/{dataset_id}/{layer}/1-1-1"), mode="r", zarr_format=2) 80 81 82def get_cortex_connectomics_data( 83 path: Union[os.PathLike, str], 84 bounding_box: Tuple[int, int, int, int, int, int], 85 sample: Literal[ 86 "mouse_s1", "mouse_ppc", "mouse_acc", "mouse_v2", "mouse_a2", "macaque_s1", "macaque_stg", 87 "human_stg", "human_ifg", 88 ] = "mouse_s1", 89 download: bool = False, 90) -> str: 91 """Stream a subvolume of the CortexConnectomics dataset and cache it as a zarr v3 store. 92 93 Args: 94 path: Filepath to a folder where the cached zarr store will be saved. 95 bounding_box: The region to fetch as (z_min, z_max, y_min, y_max, x_min, x_max) in voxel 96 coordinates, at the dataset's native resolution (see `SAMPLES`). 97 sample: Which species / cortical region to use. See `SAMPLES` for the available keys. 98 download: Whether to stream and cache the data if it is not present. 99 100 Returns: 101 The filepath to the cached zarr store. 102 """ 103 import zarr 104 from zarr.codecs import BloscCodec 105 106 if sample not in SAMPLES: 107 raise ValueError(f"sample must be one of {list(SAMPLES)}, got {sample!r}") 108 109 os.makedirs(str(path), exist_ok=True) 110 zarr_path = os.path.join(str(path), f"{sample}_{_bbox_to_str(bounding_box)}.zarr") 111 112 root = zarr.open_group(zarr_path, mode="a") 113 if "raw" in root and "labels" in root: 114 return zarr_path 115 116 if not download: 117 raise RuntimeError(f"No cached data found at '{zarr_path}'. Set download=True to stream it.") 118 119 z_min, z_max, y_min, y_max, x_min, x_max = bounding_box 120 info = SAMPLES[sample] 121 shape_z, shape_y, shape_x = info["shape_zyx"] 122 assert z_max <= shape_z and y_max <= shape_y and x_max <= shape_x, \ 123 f"Bounding box exceeds the {sample} volume {info['shape_zyx']}" 124 125 color = _open_remote_array(info["dataset_id"], "color") 126 seg = _open_remote_array(info["dataset_id"], "segmentation") 127 128 # the remote arrays are indexed as (channel, x, y, z); torch_em volumes use (z, y, x). 129 raw = np.transpose(color[0, x_min:x_max, y_min:y_max, z_min:z_max], (2, 1, 0)) 130 labels = np.transpose(seg[0, x_min:x_max, y_min:y_max, z_min:z_max], (2, 1, 0)) 131 132 def _make_array(name, data, shuffle): 133 array = root.create_array( 134 name, shape=data.shape, chunks=(32, 512, 512), dtype=data.dtype, 135 compressors=BloscCodec(cname="zstd", clevel=6, shuffle=shuffle), 136 ) 137 array[:] = data 138 139 root.attrs["bounding_box"] = list(bounding_box) 140 root.attrs["sample"] = sample 141 root.attrs["resolution_nm"] = list(info["resolution_nm"]) 142 root.attrs["labels_are_exhaustive"] = False 143 144 _make_array("raw", raw, shuffle="shuffle") 145 _make_array("labels", labels, shuffle="bitshuffle") 146 147 return zarr_path 148 149 150def get_cortex_connectomics_paths( 151 path: Union[os.PathLike, str], 152 bounding_boxes: List[Tuple[int, int, int, int, int, int]], 153 sample: Literal[ 154 "mouse_s1", "mouse_ppc", "mouse_acc", "mouse_v2", "mouse_a2", "macaque_s1", "macaque_stg", 155 "human_stg", "human_ifg", 156 ] = "mouse_s1", 157 download: bool = False, 158) -> List[str]: 159 """Get paths to cached CortexConnectomics zarr stores. 160 161 Args: 162 path: Filepath to a folder where the cached zarr stores will be saved. 163 bounding_boxes: List of regions to fetch, each as (z_min, z_max, y_min, y_max, x_min, x_max). 164 sample: Which species / cortical region to use. See `SAMPLES` for the available keys. 165 download: Whether to stream and cache the data if it is not present. 166 167 Returns: 168 List of filepaths to the cached zarr stores. 169 """ 170 return [get_cortex_connectomics_data(path, bbox, sample, download) for bbox in bounding_boxes] 171 172 173def get_cortex_connectomics_dataset( 174 path: Union[os.PathLike, str], 175 patch_shape: Tuple[int, int, int], 176 bounding_boxes: List[Tuple[int, int, int, int, int, int]], 177 sample: Literal[ 178 "mouse_s1", "mouse_ppc", "mouse_acc", "mouse_v2", "mouse_a2", "macaque_s1", "macaque_stg", 179 "human_stg", "human_ifg", 180 ] = "mouse_s1", 181 download: bool = False, 182 **kwargs, 183) -> Dataset: 184 """Get the CortexConnectomics dataset for neuron instance segmentation. 185 186 Labels are dense within the proofread region of each volume, but that region can be smaller 187 than the full raw volume; choose bounding boxes accordingly. 188 189 Args: 190 path: Filepath to a folder where the cached zarr stores will be saved. 191 patch_shape: The patch shape (z, y, x) to use for training. 192 bounding_boxes: List of subvolumes to use, each as (z_min, z_max, y_min, y_max, x_min, x_max). 193 sample: Which species / cortical region to use. See `SAMPLES` for the available keys. 194 download: Whether to stream and cache data if not already present. 195 kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset`. 196 197 Returns: 198 The segmentation dataset. 199 """ 200 assert len(patch_shape) == 3 201 202 paths = get_cortex_connectomics_paths(path, bounding_boxes, sample, download) 203 kwargs = util.update_kwargs(kwargs, "is_seg_dataset", True) 204 205 return torch_em.default_segmentation_dataset( 206 raw_paths=paths, 207 raw_key="raw", 208 label_paths=paths, 209 label_key="labels", 210 patch_shape=patch_shape, 211 **kwargs, 212 ) 213 214 215def get_cortex_connectomics_loader( 216 path: Union[os.PathLike, str], 217 patch_shape: Tuple[int, int, int], 218 batch_size: int, 219 bounding_boxes: List[Tuple[int, int, int, int, int, int]], 220 sample: Literal[ 221 "mouse_s1", "mouse_ppc", "mouse_acc", "mouse_v2", "mouse_a2", "macaque_s1", "macaque_stg", 222 "human_stg", "human_ifg", 223 ] = "mouse_s1", 224 download: bool = False, 225 **kwargs, 226) -> DataLoader: 227 """Get the DataLoader for neuron instance segmentation in the CortexConnectomics dataset. 228 229 Labels are dense within the proofread region of each volume, but that region can be smaller 230 than the full raw volume; choose bounding boxes accordingly. 231 232 Args: 233 path: Filepath to a folder where the cached zarr stores will be saved. 234 patch_shape: The patch shape (z, y, x) to use for training. 235 batch_size: The batch size for training. 236 bounding_boxes: List of subvolumes to use, each as (z_min, z_max, y_min, y_max, x_min, x_max). 237 sample: Which species / cortical region to use. See `SAMPLES` for the available keys. 238 download: Whether to stream and cache data if not already present. 239 kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset` or for the 240 PyTorch DataLoader. 241 242 Returns: 243 The DataLoader. 244 """ 245 ds_kwargs, loader_kwargs = util.split_kwargs(torch_em.default_segmentation_dataset, **kwargs) 246 dataset = get_cortex_connectomics_dataset(path, patch_shape, bounding_boxes, sample, download, **ds_kwargs) 247 return torch_em.get_data_loader(dataset, batch_size, **loader_kwargs)
83def get_cortex_connectomics_data( 84 path: Union[os.PathLike, str], 85 bounding_box: Tuple[int, int, int, int, int, int], 86 sample: Literal[ 87 "mouse_s1", "mouse_ppc", "mouse_acc", "mouse_v2", "mouse_a2", "macaque_s1", "macaque_stg", 88 "human_stg", "human_ifg", 89 ] = "mouse_s1", 90 download: bool = False, 91) -> str: 92 """Stream a subvolume of the CortexConnectomics dataset and cache it as a zarr v3 store. 93 94 Args: 95 path: Filepath to a folder where the cached zarr store will be saved. 96 bounding_box: The region to fetch as (z_min, z_max, y_min, y_max, x_min, x_max) in voxel 97 coordinates, at the dataset's native resolution (see `SAMPLES`). 98 sample: Which species / cortical region to use. See `SAMPLES` for the available keys. 99 download: Whether to stream and cache the data if it is not present. 100 101 Returns: 102 The filepath to the cached zarr store. 103 """ 104 import zarr 105 from zarr.codecs import BloscCodec 106 107 if sample not in SAMPLES: 108 raise ValueError(f"sample must be one of {list(SAMPLES)}, got {sample!r}") 109 110 os.makedirs(str(path), exist_ok=True) 111 zarr_path = os.path.join(str(path), f"{sample}_{_bbox_to_str(bounding_box)}.zarr") 112 113 root = zarr.open_group(zarr_path, mode="a") 114 if "raw" in root and "labels" in root: 115 return zarr_path 116 117 if not download: 118 raise RuntimeError(f"No cached data found at '{zarr_path}'. Set download=True to stream it.") 119 120 z_min, z_max, y_min, y_max, x_min, x_max = bounding_box 121 info = SAMPLES[sample] 122 shape_z, shape_y, shape_x = info["shape_zyx"] 123 assert z_max <= shape_z and y_max <= shape_y and x_max <= shape_x, \ 124 f"Bounding box exceeds the {sample} volume {info['shape_zyx']}" 125 126 color = _open_remote_array(info["dataset_id"], "color") 127 seg = _open_remote_array(info["dataset_id"], "segmentation") 128 129 # the remote arrays are indexed as (channel, x, y, z); torch_em volumes use (z, y, x). 130 raw = np.transpose(color[0, x_min:x_max, y_min:y_max, z_min:z_max], (2, 1, 0)) 131 labels = np.transpose(seg[0, x_min:x_max, y_min:y_max, z_min:z_max], (2, 1, 0)) 132 133 def _make_array(name, data, shuffle): 134 array = root.create_array( 135 name, shape=data.shape, chunks=(32, 512, 512), dtype=data.dtype, 136 compressors=BloscCodec(cname="zstd", clevel=6, shuffle=shuffle), 137 ) 138 array[:] = data 139 140 root.attrs["bounding_box"] = list(bounding_box) 141 root.attrs["sample"] = sample 142 root.attrs["resolution_nm"] = list(info["resolution_nm"]) 143 root.attrs["labels_are_exhaustive"] = False 144 145 _make_array("raw", raw, shuffle="shuffle") 146 _make_array("labels", labels, shuffle="bitshuffle") 147 148 return zarr_path
Stream a subvolume of the CortexConnectomics dataset and cache it as a zarr v3 store.
Arguments:
- path: Filepath to a folder where the cached zarr store will be saved.
- bounding_box: The region to fetch as (z_min, z_max, y_min, y_max, x_min, x_max) in voxel
coordinates, at the dataset's native resolution (see
SAMPLES). - sample: Which species / cortical region to use. See
SAMPLESfor the available keys. - download: Whether to stream and cache the data if it is not present.
Returns:
The filepath to the cached zarr store.
151def get_cortex_connectomics_paths( 152 path: Union[os.PathLike, str], 153 bounding_boxes: List[Tuple[int, int, int, int, int, int]], 154 sample: Literal[ 155 "mouse_s1", "mouse_ppc", "mouse_acc", "mouse_v2", "mouse_a2", "macaque_s1", "macaque_stg", 156 "human_stg", "human_ifg", 157 ] = "mouse_s1", 158 download: bool = False, 159) -> List[str]: 160 """Get paths to cached CortexConnectomics zarr stores. 161 162 Args: 163 path: Filepath to a folder where the cached zarr stores will be saved. 164 bounding_boxes: List of regions to fetch, each as (z_min, z_max, y_min, y_max, x_min, x_max). 165 sample: Which species / cortical region to use. See `SAMPLES` for the available keys. 166 download: Whether to stream and cache the data if it is not present. 167 168 Returns: 169 List of filepaths to the cached zarr stores. 170 """ 171 return [get_cortex_connectomics_data(path, bbox, sample, download) for bbox in bounding_boxes]
Get paths to cached CortexConnectomics zarr stores.
Arguments:
- path: Filepath to a folder where the cached zarr stores will be saved.
- bounding_boxes: List of regions to fetch, each as (z_min, z_max, y_min, y_max, x_min, x_max).
- sample: Which species / cortical region to use. See
SAMPLESfor the available keys. - download: Whether to stream and cache the data if it is not present.
Returns:
List of filepaths to the cached zarr stores.
174def get_cortex_connectomics_dataset( 175 path: Union[os.PathLike, str], 176 patch_shape: Tuple[int, int, int], 177 bounding_boxes: List[Tuple[int, int, int, int, int, int]], 178 sample: Literal[ 179 "mouse_s1", "mouse_ppc", "mouse_acc", "mouse_v2", "mouse_a2", "macaque_s1", "macaque_stg", 180 "human_stg", "human_ifg", 181 ] = "mouse_s1", 182 download: bool = False, 183 **kwargs, 184) -> Dataset: 185 """Get the CortexConnectomics dataset for neuron instance segmentation. 186 187 Labels are dense within the proofread region of each volume, but that region can be smaller 188 than the full raw volume; choose bounding boxes accordingly. 189 190 Args: 191 path: Filepath to a folder where the cached zarr stores will be saved. 192 patch_shape: The patch shape (z, y, x) to use for training. 193 bounding_boxes: List of subvolumes to use, each as (z_min, z_max, y_min, y_max, x_min, x_max). 194 sample: Which species / cortical region to use. See `SAMPLES` for the available keys. 195 download: Whether to stream and cache data if not already present. 196 kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset`. 197 198 Returns: 199 The segmentation dataset. 200 """ 201 assert len(patch_shape) == 3 202 203 paths = get_cortex_connectomics_paths(path, bounding_boxes, sample, download) 204 kwargs = util.update_kwargs(kwargs, "is_seg_dataset", True) 205 206 return torch_em.default_segmentation_dataset( 207 raw_paths=paths, 208 raw_key="raw", 209 label_paths=paths, 210 label_key="labels", 211 patch_shape=patch_shape, 212 **kwargs, 213 )
Get the CortexConnectomics dataset for neuron instance segmentation.
Labels are dense within the proofread region of each volume, but that region can be smaller than the full raw volume; choose bounding boxes accordingly.
Arguments:
- path: Filepath to a folder where the cached zarr stores will be saved.
- patch_shape: The patch shape (z, y, x) to use for training.
- bounding_boxes: List of subvolumes to use, each as (z_min, z_max, y_min, y_max, x_min, x_max).
- sample: Which species / cortical region to use. See
SAMPLESfor the available keys. - download: Whether to stream and cache data if not already present.
- kwargs: Additional keyword arguments for
torch_em.default_segmentation_dataset.
Returns:
The segmentation dataset.
216def get_cortex_connectomics_loader( 217 path: Union[os.PathLike, str], 218 patch_shape: Tuple[int, int, int], 219 batch_size: int, 220 bounding_boxes: List[Tuple[int, int, int, int, int, int]], 221 sample: Literal[ 222 "mouse_s1", "mouse_ppc", "mouse_acc", "mouse_v2", "mouse_a2", "macaque_s1", "macaque_stg", 223 "human_stg", "human_ifg", 224 ] = "mouse_s1", 225 download: bool = False, 226 **kwargs, 227) -> DataLoader: 228 """Get the DataLoader for neuron instance segmentation in the CortexConnectomics dataset. 229 230 Labels are dense within the proofread region of each volume, but that region can be smaller 231 than the full raw volume; choose bounding boxes accordingly. 232 233 Args: 234 path: Filepath to a folder where the cached zarr stores will be saved. 235 patch_shape: The patch shape (z, y, x) to use for training. 236 batch_size: The batch size for training. 237 bounding_boxes: List of subvolumes to use, each as (z_min, z_max, y_min, y_max, x_min, x_max). 238 sample: Which species / cortical region to use. See `SAMPLES` for the available keys. 239 download: Whether to stream and cache data if not already present. 240 kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset` or for the 241 PyTorch DataLoader. 242 243 Returns: 244 The DataLoader. 245 """ 246 ds_kwargs, loader_kwargs = util.split_kwargs(torch_em.default_segmentation_dataset, **kwargs) 247 dataset = get_cortex_connectomics_dataset(path, patch_shape, bounding_boxes, sample, download, **ds_kwargs) 248 return torch_em.get_data_loader(dataset, batch_size, **loader_kwargs)
Get the DataLoader for neuron instance segmentation in the CortexConnectomics dataset.
Labels are dense within the proofread region of each volume, but that region can be smaller than the full raw volume; choose bounding boxes accordingly.
Arguments:
- path: Filepath to a folder where the cached zarr stores will be saved.
- patch_shape: The patch shape (z, y, x) to use for training.
- batch_size: The batch size for training.
- bounding_boxes: List of subvolumes to use, each as (z_min, z_max, y_min, y_max, x_min, x_max).
- sample: Which species / cortical region to use. See
SAMPLESfor the available keys. - download: Whether to stream and cache data if not already present.
- kwargs: Additional keyword arguments for
torch_em.default_segmentation_datasetor for the PyTorch DataLoader.
Returns:
The DataLoader.