torch_em.data.datasets.light_microscopy.wing_disc_timelapse
The Wing Disc Timelapse dataset contains annotations for cell instance segmentation in E-cadherin-GFP spinning disc confocal timelapse movies of cultured Drosophila wing disc epithelia.
The data consists of five movies (biological replicates) of the pouch region of ex vivo cultured wing discs, with
one 2D frame of the apical surface every 5 minutes: 'Ecd_20141010_P2', 'Ecd_20150310_P2', 'Ecd_20150418_P1',
'Ecd_20150608_P1' and 'Ecd_20161213_F0' (see MOVIES). Every frame is segmented with TissueAnalyzer and the
segmentation was corrected manually, and the cells were tracked over the movies. Each frame has about 1,800 to 3,000
cells. The movies differ in image size, and 808 frames with segmentation are available in total. Adjacent frames of a
movie are highly correlated, so training and validation data should be split by movie (see the movies argument).
The loader exposes the raw E-cadherin frame together with a cell instance label image, where the pixels of the cell bonds (cell boundaries) are background. The instance ids are consecutive per frame. The label images are derived from the released 'tracked_cells_resized' images, in which every cell has a unique color, and the 'handCorrection' bond images. The cell ids are consistent over time in the release, but this module does not keep them, and it does not use the cell divisions and the tracking database either.
This dataset is different from torch_em.data.datasets.light_microscopy.wing_disc (3D nucleus segmentation in
confocal images of wing discs) and torch_em.data.datasets.light_microscopy.flywing (128 x 128 patches from an
older epithelial benchmark).
NOTE: The data is stored in one zip archive per movie (2.9 GB in total), and this module converts the frames once into small tif files. The archives are kept next to the converted data.
The data is located at https://doi.org/10.5281/zenodo.22760592 and is released under a CC-BY-4.0 license. This dataset is from the publications https://doi.org/10.7554/eLife.57964 and https://doi.org/10.1242/dev.155069. Please cite them if you use this dataset for your research.
1"""The Wing Disc Timelapse dataset contains annotations for cell instance segmentation in E-cadherin-GFP 2spinning disc confocal timelapse movies of cultured Drosophila wing disc epithelia. 3 4The data consists of five movies (biological replicates) of the pouch region of ex vivo cultured wing discs, with 5one 2D frame of the apical surface every 5 minutes: 'Ecd_20141010_P2', 'Ecd_20150310_P2', 'Ecd_20150418_P1', 6'Ecd_20150608_P1' and 'Ecd_20161213_F0' (see `MOVIES`). Every frame is segmented with TissueAnalyzer and the 7segmentation was corrected manually, and the cells were tracked over the movies. Each frame has about 1,800 to 3,000 8cells. The movies differ in image size, and 808 frames with segmentation are available in total. Adjacent frames of a 9movie are highly correlated, so training and validation data should be split by movie (see the `movies` argument). 10 11The loader exposes the raw E-cadherin frame together with a cell instance label image, where the pixels of the cell 12bonds (cell boundaries) are background. The instance ids are consecutive per frame. The label images are derived 13from the released 'tracked_cells_resized' images, in which every cell has a unique color, and the 'handCorrection' 14bond images. The cell ids are consistent over time in the release, but this module does not keep them, and it does 15not use the cell divisions and the tracking database either. 16 17This dataset is different from `torch_em.data.datasets.light_microscopy.wing_disc` (3D nucleus segmentation in 18confocal images of wing discs) and `torch_em.data.datasets.light_microscopy.flywing` (128 x 128 patches from an 19older epithelial benchmark). 20 21NOTE: The data is stored in one zip archive per movie (2.9 GB in total), and this module converts the frames once 22into small tif files. The archives are kept next to the converted data. 23 24The data is located at https://doi.org/10.5281/zenodo.22760592 and is released under a CC-BY-4.0 license. 25This dataset is from the publications https://doi.org/10.7554/eLife.57964 and https://doi.org/10.1242/dev.155069. 26Please cite them if you use this dataset for your research. 27""" 28 29import os 30import re 31import uuid 32from glob import glob 33from natsort import natsorted 34from concurrent import futures 35from typing import Union, Tuple, Optional, Sequence, List 36 37from torch.utils.data import Dataset, DataLoader 38 39import torch_em 40 41from .. import util 42 43 44URL_BASE = "https://zenodo.org/records/22760592/files" 45 46MOVIES = { 47 "Ecd_20141010_P2": "0eb41b1e0fbfb23b1af3950c14ad06834c40c6debf639ab572c2ecff006fb478", 48 "Ecd_20150310_P2": "07b74f7594a99bd3243e7fde4f1c23967f7256db7db91cccc4a4eb0dd0de4a0c", 49 "Ecd_20150418_P1": "dcfc9f79b387b0657723731bcf79fab0d600600161f57577cb54748239a1ef69", 50 "Ecd_20150608_P1": "79c52b04b414fc0f4dcb1a73053c2e96f5db02b39b9dff3e0eb25a54e2d25f2d", 51 "Ecd_20161213_F0": "8a3ee532c89950bc0f4880ff243593e0c7917f080e352bd41ac6f215ac9da6c0", 52} 53 54COMPLETE_MARKER = "complete" 55FRAME_PATTERN = re.compile(r".*/Segmentation/[^/]+_(\d+)/(original|tracked_cells_resized|handCorrection)\.(png|tif)$") 56 57 58def _write_atomic(path, array): 59 import tifffile 60 61 tmp_path = f"{path}.{uuid.uuid4().hex}.incomplete.tif" 62 tifffile.imwrite(tmp_path, array, compression="zlib") 63 os.replace(tmp_path, path) 64 65 66def _read_member(archive, name): 67 from io import BytesIO 68 import imageio.v3 as imageio 69 70 return imageio.imread(BytesIO(archive.read(name)), extension=os.path.splitext(name)[1]) 71 72 73def _convert_frame(archive, members, image_path, label_path): 74 import numpy as np 75 76 raw = _read_member(archive, members["original"]) 77 raw = raw[..., 0] if raw.ndim == 3 else raw 78 79 bonds = _read_member(archive, members["handCorrection"]) 80 bonds = bonds.max(axis=-1) > 0 if bonds.ndim == 3 else bonds > 0 81 82 colors = _read_member(archive, members["tracked_cells_resized"]).astype("int64") 83 codes = colors[..., 0] * 65536 + colors[..., 1] * 256 + colors[..., 2] 84 codes[bonds] = 0 85 86 foreground = codes > 0 87 _, inverse = np.unique(codes[foreground], return_inverse=True) 88 labels = np.zeros(codes.shape, dtype="uint16") 89 labels[foreground] = inverse + 1 90 assert labels.max() < 65535 and raw.shape == labels.shape, f"Unexpected frame '{members['original']}'." 91 92 _write_atomic(image_path, raw.astype("uint8")) 93 _write_atomic(label_path, labels) 94 95 96def _convert_movie(zip_path, out_dir): 97 import zipfile 98 99 os.makedirs(os.path.join(out_dir, "images"), exist_ok=True) 100 os.makedirs(os.path.join(out_dir, "labels"), exist_ok=True) 101 102 with zipfile.ZipFile(zip_path) as archive: 103 frames = {} 104 for name in archive.namelist(): 105 match = FRAME_PATTERN.match(name) 106 if match is not None: 107 frames.setdefault(int(match.group(1)), {})[match.group(2)] = name 108 109 for frame, members in sorted(frames.items()): 110 if len(members) < 3: 111 continue 112 _convert_frame( 113 archive, members, os.path.join(out_dir, "images", f"{frame:03d}.tif"), 114 os.path.join(out_dir, "labels", f"{frame:03d}.tif"), 115 ) 116 117 with open(os.path.join(out_dir, COMPLETE_MARKER), "w"): 118 pass 119 120 121def _validate(movies): 122 if movies is None: 123 return list(MOVIES) 124 invalid = [m for m in movies if m not in MOVIES] 125 if invalid: 126 raise ValueError(f"{invalid} are not valid movies. Choose from {list(MOVIES)}.") 127 return list(movies) 128 129 130def get_wing_disc_timelapse_data( 131 path: Union[os.PathLike, str], movies: Optional[Sequence[str]] = None, download: bool = False 132) -> str: 133 """Download the Wing Disc Timelapse dataset. 134 135 Args: 136 path: Filepath to a folder where the data is downloaded for further processing. 137 movies: The movies to use. By default all five movies are used. See `MOVIES` for the valid names. 138 download: Whether to download the data if it is not present. 139 140 Returns: 141 Filepath to the folder with the converted frames. 142 """ 143 movies = _validate(movies) 144 converted_dir = os.path.join(path, "converted") 145 146 pending = [m for m in movies if not os.path.exists(os.path.join(converted_dir, m, COMPLETE_MARKER))] 147 if not pending: 148 return converted_dir 149 150 os.makedirs(path, exist_ok=True) 151 for movie in pending: 152 util.download_source( 153 path=os.path.join(path, f"{movie}.zip"), url=f"{URL_BASE}/{movie}.zip", download=download, 154 checksum=MOVIES[movie], 155 ) 156 157 with futures.ProcessPoolExecutor(min(len(pending), os.cpu_count() or 1)) as pool: 158 tasks = [ 159 pool.submit(_convert_movie, os.path.join(path, f"{movie}.zip"), os.path.join(converted_dir, movie)) 160 for movie in pending 161 ] 162 for task in tasks: 163 task.result() 164 165 return converted_dir 166 167 168def get_wing_disc_timelapse_paths( 169 path: Union[os.PathLike, str], movies: Optional[Sequence[str]] = None, download: bool = False, 170) -> Tuple[List[str], List[str]]: 171 """Get paths to the Wing Disc Timelapse data. 172 173 Args: 174 path: Filepath to a folder where the data is downloaded for further processing. 175 movies: The movies to use. By default all five movies are used. See `MOVIES` for the valid names. 176 download: Whether to download the data if it is not present. 177 178 Returns: 179 List of filepaths for the image data. 180 List of filepaths for the label data. 181 """ 182 movies = _validate(movies) 183 converted_dir = get_wing_disc_timelapse_data(path, movies, download) 184 185 raw_paths, label_paths = [], [] 186 for movie in movies: 187 movie_raw_paths = natsorted(glob(os.path.join(converted_dir, movie, "images", "*.tif"))) 188 raw_paths.extend(movie_raw_paths) 189 label_paths.extend( 190 os.path.join(converted_dir, movie, "labels", os.path.basename(p)) for p in movie_raw_paths 191 ) 192 193 assert len(raw_paths) == len(label_paths) and len(raw_paths) > 0 194 assert all(os.path.exists(p) for p in label_paths) 195 196 return raw_paths, label_paths 197 198 199def get_wing_disc_timelapse_dataset( 200 path: Union[os.PathLike, str], 201 patch_shape: Tuple[int, int], 202 movies: Optional[Sequence[str]] = None, 203 resize_inputs: bool = False, 204 download: bool = False, 205 **kwargs 206) -> Dataset: 207 """Get the Wing Disc Timelapse dataset for cell instance segmentation in epithelial timelapse microscopy. 208 209 Args: 210 path: Filepath to a folder where the data is downloaded for further processing. 211 patch_shape: The patch shape to use for training. 212 movies: The movies to use. By default all five movies are used. See `MOVIES` for the valid names. 213 resize_inputs: Whether to resize the inputs to the patch shape. 214 download: Whether to download the data if it is not present. 215 kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset`. 216 217 Returns: 218 The segmentation dataset. 219 """ 220 raw_paths, label_paths = get_wing_disc_timelapse_paths(path, movies, download) 221 222 if resize_inputs: 223 resize_kwargs = {"patch_shape": patch_shape, "is_rgb": False} 224 kwargs, patch_shape = util.update_kwargs_for_resize_trafo( 225 kwargs=kwargs, patch_shape=patch_shape, resize_inputs=resize_inputs, resize_kwargs=resize_kwargs 226 ) 227 228 return torch_em.default_segmentation_dataset( 229 raw_paths=raw_paths, 230 raw_key=None, 231 label_paths=label_paths, 232 label_key=None, 233 is_seg_dataset=False, 234 patch_shape=patch_shape, 235 **kwargs 236 ) 237 238 239def get_wing_disc_timelapse_loader( 240 path: Union[os.PathLike, str], 241 batch_size: int, 242 patch_shape: Tuple[int, int], 243 movies: Optional[Sequence[str]] = None, 244 resize_inputs: bool = False, 245 download: bool = False, 246 **kwargs 247) -> DataLoader: 248 """Get the Wing Disc Timelapse dataloader for cell instance segmentation in epithelial timelapse microscopy. 249 250 Args: 251 path: Filepath to a folder where the data is downloaded for further processing. 252 batch_size: The batch size for training. 253 patch_shape: The patch shape to use for training. 254 movies: The movies to use. By default all five movies are used. See `MOVIES` for the valid names. 255 resize_inputs: Whether to resize the inputs to the patch shape. 256 download: Whether to download the data if it is not present. 257 kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset` or for the PyTorch DataLoader. 258 259 Returns: 260 The DataLoader. 261 """ 262 ds_kwargs, loader_kwargs = util.split_kwargs(torch_em.default_segmentation_dataset, **kwargs) 263 dataset = get_wing_disc_timelapse_dataset(path, patch_shape, movies, resize_inputs, download, **ds_kwargs) 264 return torch_em.get_data_loader(dataset, batch_size, **loader_kwargs)
131def get_wing_disc_timelapse_data( 132 path: Union[os.PathLike, str], movies: Optional[Sequence[str]] = None, download: bool = False 133) -> str: 134 """Download the Wing Disc Timelapse dataset. 135 136 Args: 137 path: Filepath to a folder where the data is downloaded for further processing. 138 movies: The movies to use. By default all five movies are used. See `MOVIES` for the valid names. 139 download: Whether to download the data if it is not present. 140 141 Returns: 142 Filepath to the folder with the converted frames. 143 """ 144 movies = _validate(movies) 145 converted_dir = os.path.join(path, "converted") 146 147 pending = [m for m in movies if not os.path.exists(os.path.join(converted_dir, m, COMPLETE_MARKER))] 148 if not pending: 149 return converted_dir 150 151 os.makedirs(path, exist_ok=True) 152 for movie in pending: 153 util.download_source( 154 path=os.path.join(path, f"{movie}.zip"), url=f"{URL_BASE}/{movie}.zip", download=download, 155 checksum=MOVIES[movie], 156 ) 157 158 with futures.ProcessPoolExecutor(min(len(pending), os.cpu_count() or 1)) as pool: 159 tasks = [ 160 pool.submit(_convert_movie, os.path.join(path, f"{movie}.zip"), os.path.join(converted_dir, movie)) 161 for movie in pending 162 ] 163 for task in tasks: 164 task.result() 165 166 return converted_dir
Download the Wing Disc Timelapse dataset.
Arguments:
- path: Filepath to a folder where the data is downloaded for further processing.
- movies: The movies to use. By default all five movies are used. See
MOVIESfor the valid names. - download: Whether to download the data if it is not present.
Returns:
Filepath to the folder with the converted frames.
169def get_wing_disc_timelapse_paths( 170 path: Union[os.PathLike, str], movies: Optional[Sequence[str]] = None, download: bool = False, 171) -> Tuple[List[str], List[str]]: 172 """Get paths to the Wing Disc Timelapse data. 173 174 Args: 175 path: Filepath to a folder where the data is downloaded for further processing. 176 movies: The movies to use. By default all five movies are used. See `MOVIES` for the valid names. 177 download: Whether to download the data if it is not present. 178 179 Returns: 180 List of filepaths for the image data. 181 List of filepaths for the label data. 182 """ 183 movies = _validate(movies) 184 converted_dir = get_wing_disc_timelapse_data(path, movies, download) 185 186 raw_paths, label_paths = [], [] 187 for movie in movies: 188 movie_raw_paths = natsorted(glob(os.path.join(converted_dir, movie, "images", "*.tif"))) 189 raw_paths.extend(movie_raw_paths) 190 label_paths.extend( 191 os.path.join(converted_dir, movie, "labels", os.path.basename(p)) for p in movie_raw_paths 192 ) 193 194 assert len(raw_paths) == len(label_paths) and len(raw_paths) > 0 195 assert all(os.path.exists(p) for p in label_paths) 196 197 return raw_paths, label_paths
Get paths to the Wing Disc Timelapse data.
Arguments:
- path: Filepath to a folder where the data is downloaded for further processing.
- movies: The movies to use. By default all five movies are used. See
MOVIESfor the valid names. - download: Whether to download the data if it is not present.
Returns:
List of filepaths for the image data. List of filepaths for the label data.
200def get_wing_disc_timelapse_dataset( 201 path: Union[os.PathLike, str], 202 patch_shape: Tuple[int, int], 203 movies: Optional[Sequence[str]] = None, 204 resize_inputs: bool = False, 205 download: bool = False, 206 **kwargs 207) -> Dataset: 208 """Get the Wing Disc Timelapse dataset for cell instance segmentation in epithelial timelapse microscopy. 209 210 Args: 211 path: Filepath to a folder where the data is downloaded for further processing. 212 patch_shape: The patch shape to use for training. 213 movies: The movies to use. By default all five movies are used. See `MOVIES` for the valid names. 214 resize_inputs: Whether to resize the inputs to the patch shape. 215 download: Whether to download the data if it is not present. 216 kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset`. 217 218 Returns: 219 The segmentation dataset. 220 """ 221 raw_paths, label_paths = get_wing_disc_timelapse_paths(path, movies, download) 222 223 if resize_inputs: 224 resize_kwargs = {"patch_shape": patch_shape, "is_rgb": False} 225 kwargs, patch_shape = util.update_kwargs_for_resize_trafo( 226 kwargs=kwargs, patch_shape=patch_shape, resize_inputs=resize_inputs, resize_kwargs=resize_kwargs 227 ) 228 229 return torch_em.default_segmentation_dataset( 230 raw_paths=raw_paths, 231 raw_key=None, 232 label_paths=label_paths, 233 label_key=None, 234 is_seg_dataset=False, 235 patch_shape=patch_shape, 236 **kwargs 237 )
Get the Wing Disc Timelapse dataset for cell instance segmentation in epithelial timelapse microscopy.
Arguments:
- path: Filepath to a folder where the data is downloaded for further processing.
- patch_shape: The patch shape to use for training.
- movies: The movies to use. By default all five movies are used. See
MOVIESfor the valid names. - resize_inputs: Whether to resize the inputs to the patch shape.
- download: Whether to download the data if it is not present.
- kwargs: Additional keyword arguments for
torch_em.default_segmentation_dataset.
Returns:
The segmentation dataset.
240def get_wing_disc_timelapse_loader( 241 path: Union[os.PathLike, str], 242 batch_size: int, 243 patch_shape: Tuple[int, int], 244 movies: Optional[Sequence[str]] = None, 245 resize_inputs: bool = False, 246 download: bool = False, 247 **kwargs 248) -> DataLoader: 249 """Get the Wing Disc Timelapse dataloader for cell instance segmentation in epithelial timelapse microscopy. 250 251 Args: 252 path: Filepath to a folder where the data is downloaded for further processing. 253 batch_size: The batch size for training. 254 patch_shape: The patch shape to use for training. 255 movies: The movies to use. By default all five movies are used. See `MOVIES` for the valid names. 256 resize_inputs: Whether to resize the inputs to the patch shape. 257 download: Whether to download the data if it is not present. 258 kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset` or for the PyTorch DataLoader. 259 260 Returns: 261 The DataLoader. 262 """ 263 ds_kwargs, loader_kwargs = util.split_kwargs(torch_em.default_segmentation_dataset, **kwargs) 264 dataset = get_wing_disc_timelapse_dataset(path, patch_shape, movies, resize_inputs, download, **ds_kwargs) 265 return torch_em.get_data_loader(dataset, batch_size, **loader_kwargs)
Get the Wing Disc Timelapse dataloader for cell instance segmentation in epithelial timelapse microscopy.
Arguments:
- path: Filepath to a folder where the data is downloaded for further processing.
- batch_size: The batch size for training.
- patch_shape: The patch shape to use for training.
- movies: The movies to use. By default all five movies are used. See
MOVIESfor the valid names. - resize_inputs: Whether to resize the inputs to the patch shape.
- download: Whether to download the data if it is not present.
- kwargs: Additional keyword arguments for
torch_em.default_segmentation_datasetor for the PyTorch DataLoader.
Returns:
The DataLoader.