torch_em.data.datasets.light_microscopy.wing_disc_timelapse

The Wing Disc Timelapse dataset contains annotations for cell instance segmentation in E-cadherin-GFP spinning disc confocal timelapse movies of cultured Drosophila wing disc epithelia.

The data consists of five movies (biological replicates) of the pouch region of ex vivo cultured wing discs, with one 2D frame of the apical surface every 5 minutes: 'Ecd_20141010_P2', 'Ecd_20150310_P2', 'Ecd_20150418_P1', 'Ecd_20150608_P1' and 'Ecd_20161213_F0' (see MOVIES). Every frame is segmented with TissueAnalyzer and the segmentation was corrected manually, and the cells were tracked over the movies. Each frame has about 1,800 to 3,000 cells. The movies differ in image size, and 808 frames with segmentation are available in total. Adjacent frames of a movie are highly correlated, so training and validation data should be split by movie (see the movies argument).

The loader exposes the raw E-cadherin frame together with a cell instance label image, where the pixels of the cell bonds (cell boundaries) are background. The instance ids are consecutive per frame. The label images are derived from the released 'tracked_cells_resized' images, in which every cell has a unique color, and the 'handCorrection' bond images. The cell ids are consistent over time in the release, but this module does not keep them, and it does not use the cell divisions and the tracking database either.

This dataset is different from torch_em.data.datasets.light_microscopy.wing_disc (3D nucleus segmentation in confocal images of wing discs) and torch_em.data.datasets.light_microscopy.flywing (128 x 128 patches from an older epithelial benchmark).

NOTE: The data is stored in one zip archive per movie (2.9 GB in total), and this module converts the frames once into small tif files. The archives are kept next to the converted data.

The data is located at https://doi.org/10.5281/zenodo.22760592 and is released under a CC-BY-4.0 license. This dataset is from the publications https://doi.org/10.7554/eLife.57964 and https://doi.org/10.1242/dev.155069. Please cite them if you use this dataset for your research.

  1"""The Wing Disc Timelapse dataset contains annotations for cell instance segmentation in E-cadherin-GFP
  2spinning disc confocal timelapse movies of cultured Drosophila wing disc epithelia.
  3
  4The data consists of five movies (biological replicates) of the pouch region of ex vivo cultured wing discs, with
  5one 2D frame of the apical surface every 5 minutes: 'Ecd_20141010_P2', 'Ecd_20150310_P2', 'Ecd_20150418_P1',
  6'Ecd_20150608_P1' and 'Ecd_20161213_F0' (see `MOVIES`). Every frame is segmented with TissueAnalyzer and the
  7segmentation was corrected manually, and the cells were tracked over the movies. Each frame has about 1,800 to 3,000
  8cells. The movies differ in image size, and 808 frames with segmentation are available in total. Adjacent frames of a
  9movie are highly correlated, so training and validation data should be split by movie (see the `movies` argument).
 10
 11The loader exposes the raw E-cadherin frame together with a cell instance label image, where the pixels of the cell
 12bonds (cell boundaries) are background. The instance ids are consecutive per frame. The label images are derived
 13from the released 'tracked_cells_resized' images, in which every cell has a unique color, and the 'handCorrection'
 14bond images. The cell ids are consistent over time in the release, but this module does not keep them, and it does
 15not use the cell divisions and the tracking database either.
 16
 17This dataset is different from `torch_em.data.datasets.light_microscopy.wing_disc` (3D nucleus segmentation in
 18confocal images of wing discs) and `torch_em.data.datasets.light_microscopy.flywing` (128 x 128 patches from an
 19older epithelial benchmark).
 20
 21NOTE: The data is stored in one zip archive per movie (2.9 GB in total), and this module converts the frames once
 22into small tif files. The archives are kept next to the converted data.
 23
 24The data is located at https://doi.org/10.5281/zenodo.22760592 and is released under a CC-BY-4.0 license.
 25This dataset is from the publications https://doi.org/10.7554/eLife.57964 and https://doi.org/10.1242/dev.155069.
 26Please cite them if you use this dataset for your research.
 27"""
 28
 29import os
 30import re
 31import uuid
 32from glob import glob
 33from natsort import natsorted
 34from concurrent import futures
 35from typing import Union, Tuple, Optional, Sequence, List
 36
 37from torch.utils.data import Dataset, DataLoader
 38
 39import torch_em
 40
 41from .. import util
 42
 43
 44URL_BASE = "https://zenodo.org/records/22760592/files"
 45
 46MOVIES = {
 47    "Ecd_20141010_P2": "0eb41b1e0fbfb23b1af3950c14ad06834c40c6debf639ab572c2ecff006fb478",
 48    "Ecd_20150310_P2": "07b74f7594a99bd3243e7fde4f1c23967f7256db7db91cccc4a4eb0dd0de4a0c",
 49    "Ecd_20150418_P1": "dcfc9f79b387b0657723731bcf79fab0d600600161f57577cb54748239a1ef69",
 50    "Ecd_20150608_P1": "79c52b04b414fc0f4dcb1a73053c2e96f5db02b39b9dff3e0eb25a54e2d25f2d",
 51    "Ecd_20161213_F0": "8a3ee532c89950bc0f4880ff243593e0c7917f080e352bd41ac6f215ac9da6c0",
 52}
 53
 54COMPLETE_MARKER = "complete"
 55FRAME_PATTERN = re.compile(r".*/Segmentation/[^/]+_(\d+)/(original|tracked_cells_resized|handCorrection)\.(png|tif)$")
 56
 57
 58def _write_atomic(path, array):
 59    import tifffile
 60
 61    tmp_path = f"{path}.{uuid.uuid4().hex}.incomplete.tif"
 62    tifffile.imwrite(tmp_path, array, compression="zlib")
 63    os.replace(tmp_path, path)
 64
 65
 66def _read_member(archive, name):
 67    from io import BytesIO
 68    import imageio.v3 as imageio
 69
 70    return imageio.imread(BytesIO(archive.read(name)), extension=os.path.splitext(name)[1])
 71
 72
 73def _convert_frame(archive, members, image_path, label_path):
 74    import numpy as np
 75
 76    raw = _read_member(archive, members["original"])
 77    raw = raw[..., 0] if raw.ndim == 3 else raw
 78
 79    bonds = _read_member(archive, members["handCorrection"])
 80    bonds = bonds.max(axis=-1) > 0 if bonds.ndim == 3 else bonds > 0
 81
 82    colors = _read_member(archive, members["tracked_cells_resized"]).astype("int64")
 83    codes = colors[..., 0] * 65536 + colors[..., 1] * 256 + colors[..., 2]
 84    codes[bonds] = 0
 85
 86    foreground = codes > 0
 87    _, inverse = np.unique(codes[foreground], return_inverse=True)
 88    labels = np.zeros(codes.shape, dtype="uint16")
 89    labels[foreground] = inverse + 1
 90    assert labels.max() < 65535 and raw.shape == labels.shape, f"Unexpected frame '{members['original']}'."
 91
 92    _write_atomic(image_path, raw.astype("uint8"))
 93    _write_atomic(label_path, labels)
 94
 95
 96def _convert_movie(zip_path, out_dir):
 97    import zipfile
 98
 99    os.makedirs(os.path.join(out_dir, "images"), exist_ok=True)
100    os.makedirs(os.path.join(out_dir, "labels"), exist_ok=True)
101
102    with zipfile.ZipFile(zip_path) as archive:
103        frames = {}
104        for name in archive.namelist():
105            match = FRAME_PATTERN.match(name)
106            if match is not None:
107                frames.setdefault(int(match.group(1)), {})[match.group(2)] = name
108
109        for frame, members in sorted(frames.items()):
110            if len(members) < 3:
111                continue
112            _convert_frame(
113                archive, members, os.path.join(out_dir, "images", f"{frame:03d}.tif"),
114                os.path.join(out_dir, "labels", f"{frame:03d}.tif"),
115            )
116
117    with open(os.path.join(out_dir, COMPLETE_MARKER), "w"):
118        pass
119
120
121def _validate(movies):
122    if movies is None:
123        return list(MOVIES)
124    invalid = [m for m in movies if m not in MOVIES]
125    if invalid:
126        raise ValueError(f"{invalid} are not valid movies. Choose from {list(MOVIES)}.")
127    return list(movies)
128
129
130def get_wing_disc_timelapse_data(
131    path: Union[os.PathLike, str], movies: Optional[Sequence[str]] = None, download: bool = False
132) -> str:
133    """Download the Wing Disc Timelapse dataset.
134
135    Args:
136        path: Filepath to a folder where the data is downloaded for further processing.
137        movies: The movies to use. By default all five movies are used. See `MOVIES` for the valid names.
138        download: Whether to download the data if it is not present.
139
140    Returns:
141        Filepath to the folder with the converted frames.
142    """
143    movies = _validate(movies)
144    converted_dir = os.path.join(path, "converted")
145
146    pending = [m for m in movies if not os.path.exists(os.path.join(converted_dir, m, COMPLETE_MARKER))]
147    if not pending:
148        return converted_dir
149
150    os.makedirs(path, exist_ok=True)
151    for movie in pending:
152        util.download_source(
153            path=os.path.join(path, f"{movie}.zip"), url=f"{URL_BASE}/{movie}.zip", download=download,
154            checksum=MOVIES[movie],
155        )
156
157    with futures.ProcessPoolExecutor(min(len(pending), os.cpu_count() or 1)) as pool:
158        tasks = [
159            pool.submit(_convert_movie, os.path.join(path, f"{movie}.zip"), os.path.join(converted_dir, movie))
160            for movie in pending
161        ]
162        for task in tasks:
163            task.result()
164
165    return converted_dir
166
167
168def get_wing_disc_timelapse_paths(
169    path: Union[os.PathLike, str], movies: Optional[Sequence[str]] = None, download: bool = False,
170) -> Tuple[List[str], List[str]]:
171    """Get paths to the Wing Disc Timelapse data.
172
173    Args:
174        path: Filepath to a folder where the data is downloaded for further processing.
175        movies: The movies to use. By default all five movies are used. See `MOVIES` for the valid names.
176        download: Whether to download the data if it is not present.
177
178    Returns:
179        List of filepaths for the image data.
180        List of filepaths for the label data.
181    """
182    movies = _validate(movies)
183    converted_dir = get_wing_disc_timelapse_data(path, movies, download)
184
185    raw_paths, label_paths = [], []
186    for movie in movies:
187        movie_raw_paths = natsorted(glob(os.path.join(converted_dir, movie, "images", "*.tif")))
188        raw_paths.extend(movie_raw_paths)
189        label_paths.extend(
190            os.path.join(converted_dir, movie, "labels", os.path.basename(p)) for p in movie_raw_paths
191        )
192
193    assert len(raw_paths) == len(label_paths) and len(raw_paths) > 0
194    assert all(os.path.exists(p) for p in label_paths)
195
196    return raw_paths, label_paths
197
198
199def get_wing_disc_timelapse_dataset(
200    path: Union[os.PathLike, str],
201    patch_shape: Tuple[int, int],
202    movies: Optional[Sequence[str]] = None,
203    resize_inputs: bool = False,
204    download: bool = False,
205    **kwargs
206) -> Dataset:
207    """Get the Wing Disc Timelapse dataset for cell instance segmentation in epithelial timelapse microscopy.
208
209    Args:
210        path: Filepath to a folder where the data is downloaded for further processing.
211        patch_shape: The patch shape to use for training.
212        movies: The movies to use. By default all five movies are used. See `MOVIES` for the valid names.
213        resize_inputs: Whether to resize the inputs to the patch shape.
214        download: Whether to download the data if it is not present.
215        kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset`.
216
217    Returns:
218        The segmentation dataset.
219    """
220    raw_paths, label_paths = get_wing_disc_timelapse_paths(path, movies, download)
221
222    if resize_inputs:
223        resize_kwargs = {"patch_shape": patch_shape, "is_rgb": False}
224        kwargs, patch_shape = util.update_kwargs_for_resize_trafo(
225            kwargs=kwargs, patch_shape=patch_shape, resize_inputs=resize_inputs, resize_kwargs=resize_kwargs
226        )
227
228    return torch_em.default_segmentation_dataset(
229        raw_paths=raw_paths,
230        raw_key=None,
231        label_paths=label_paths,
232        label_key=None,
233        is_seg_dataset=False,
234        patch_shape=patch_shape,
235        **kwargs
236    )
237
238
239def get_wing_disc_timelapse_loader(
240    path: Union[os.PathLike, str],
241    batch_size: int,
242    patch_shape: Tuple[int, int],
243    movies: Optional[Sequence[str]] = None,
244    resize_inputs: bool = False,
245    download: bool = False,
246    **kwargs
247) -> DataLoader:
248    """Get the Wing Disc Timelapse dataloader for cell instance segmentation in epithelial timelapse microscopy.
249
250    Args:
251        path: Filepath to a folder where the data is downloaded for further processing.
252        batch_size: The batch size for training.
253        patch_shape: The patch shape to use for training.
254        movies: The movies to use. By default all five movies are used. See `MOVIES` for the valid names.
255        resize_inputs: Whether to resize the inputs to the patch shape.
256        download: Whether to download the data if it is not present.
257        kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset` or for the PyTorch DataLoader.
258
259    Returns:
260        The DataLoader.
261    """
262    ds_kwargs, loader_kwargs = util.split_kwargs(torch_em.default_segmentation_dataset, **kwargs)
263    dataset = get_wing_disc_timelapse_dataset(path, patch_shape, movies, resize_inputs, download, **ds_kwargs)
264    return torch_em.get_data_loader(dataset, batch_size, **loader_kwargs)
URL_BASE = 'https://zenodo.org/records/22760592/files'
MOVIES = {'Ecd_20141010_P2': '0eb41b1e0fbfb23b1af3950c14ad06834c40c6debf639ab572c2ecff006fb478', 'Ecd_20150310_P2': '07b74f7594a99bd3243e7fde4f1c23967f7256db7db91cccc4a4eb0dd0de4a0c', 'Ecd_20150418_P1': 'dcfc9f79b387b0657723731bcf79fab0d600600161f57577cb54748239a1ef69', 'Ecd_20150608_P1': '79c52b04b414fc0f4dcb1a73053c2e96f5db02b39b9dff3e0eb25a54e2d25f2d', 'Ecd_20161213_F0': '8a3ee532c89950bc0f4880ff243593e0c7917f080e352bd41ac6f215ac9da6c0'}
COMPLETE_MARKER = 'complete'
FRAME_PATTERN = re.compile('.*/Segmentation/[^/]+_(\\d+)/(original|tracked_cells_resized|handCorrection)\\.(png|tif)$')
def get_wing_disc_timelapse_data( path: Union[os.PathLike, str], movies: Optional[Sequence[str]] = None, download: bool = False) -> str:
131def get_wing_disc_timelapse_data(
132    path: Union[os.PathLike, str], movies: Optional[Sequence[str]] = None, download: bool = False
133) -> str:
134    """Download the Wing Disc Timelapse dataset.
135
136    Args:
137        path: Filepath to a folder where the data is downloaded for further processing.
138        movies: The movies to use. By default all five movies are used. See `MOVIES` for the valid names.
139        download: Whether to download the data if it is not present.
140
141    Returns:
142        Filepath to the folder with the converted frames.
143    """
144    movies = _validate(movies)
145    converted_dir = os.path.join(path, "converted")
146
147    pending = [m for m in movies if not os.path.exists(os.path.join(converted_dir, m, COMPLETE_MARKER))]
148    if not pending:
149        return converted_dir
150
151    os.makedirs(path, exist_ok=True)
152    for movie in pending:
153        util.download_source(
154            path=os.path.join(path, f"{movie}.zip"), url=f"{URL_BASE}/{movie}.zip", download=download,
155            checksum=MOVIES[movie],
156        )
157
158    with futures.ProcessPoolExecutor(min(len(pending), os.cpu_count() or 1)) as pool:
159        tasks = [
160            pool.submit(_convert_movie, os.path.join(path, f"{movie}.zip"), os.path.join(converted_dir, movie))
161            for movie in pending
162        ]
163        for task in tasks:
164            task.result()
165
166    return converted_dir

Download the Wing Disc Timelapse dataset.

Arguments:
  • path: Filepath to a folder where the data is downloaded for further processing.
  • movies: The movies to use. By default all five movies are used. See MOVIES for the valid names.
  • download: Whether to download the data if it is not present.
Returns:

Filepath to the folder with the converted frames.

def get_wing_disc_timelapse_paths( path: Union[os.PathLike, str], movies: Optional[Sequence[str]] = None, download: bool = False) -> Tuple[List[str], List[str]]:
169def get_wing_disc_timelapse_paths(
170    path: Union[os.PathLike, str], movies: Optional[Sequence[str]] = None, download: bool = False,
171) -> Tuple[List[str], List[str]]:
172    """Get paths to the Wing Disc Timelapse data.
173
174    Args:
175        path: Filepath to a folder where the data is downloaded for further processing.
176        movies: The movies to use. By default all five movies are used. See `MOVIES` for the valid names.
177        download: Whether to download the data if it is not present.
178
179    Returns:
180        List of filepaths for the image data.
181        List of filepaths for the label data.
182    """
183    movies = _validate(movies)
184    converted_dir = get_wing_disc_timelapse_data(path, movies, download)
185
186    raw_paths, label_paths = [], []
187    for movie in movies:
188        movie_raw_paths = natsorted(glob(os.path.join(converted_dir, movie, "images", "*.tif")))
189        raw_paths.extend(movie_raw_paths)
190        label_paths.extend(
191            os.path.join(converted_dir, movie, "labels", os.path.basename(p)) for p in movie_raw_paths
192        )
193
194    assert len(raw_paths) == len(label_paths) and len(raw_paths) > 0
195    assert all(os.path.exists(p) for p in label_paths)
196
197    return raw_paths, label_paths

Get paths to the Wing Disc Timelapse data.

Arguments:
  • path: Filepath to a folder where the data is downloaded for further processing.
  • movies: The movies to use. By default all five movies are used. See MOVIES for the valid names.
  • download: Whether to download the data if it is not present.
Returns:

List of filepaths for the image data. List of filepaths for the label data.

def get_wing_disc_timelapse_dataset( path: Union[os.PathLike, str], patch_shape: Tuple[int, int], movies: Optional[Sequence[str]] = None, resize_inputs: bool = False, download: bool = False, **kwargs) -> torch.utils.data.dataset.Dataset:
200def get_wing_disc_timelapse_dataset(
201    path: Union[os.PathLike, str],
202    patch_shape: Tuple[int, int],
203    movies: Optional[Sequence[str]] = None,
204    resize_inputs: bool = False,
205    download: bool = False,
206    **kwargs
207) -> Dataset:
208    """Get the Wing Disc Timelapse dataset for cell instance segmentation in epithelial timelapse microscopy.
209
210    Args:
211        path: Filepath to a folder where the data is downloaded for further processing.
212        patch_shape: The patch shape to use for training.
213        movies: The movies to use. By default all five movies are used. See `MOVIES` for the valid names.
214        resize_inputs: Whether to resize the inputs to the patch shape.
215        download: Whether to download the data if it is not present.
216        kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset`.
217
218    Returns:
219        The segmentation dataset.
220    """
221    raw_paths, label_paths = get_wing_disc_timelapse_paths(path, movies, download)
222
223    if resize_inputs:
224        resize_kwargs = {"patch_shape": patch_shape, "is_rgb": False}
225        kwargs, patch_shape = util.update_kwargs_for_resize_trafo(
226            kwargs=kwargs, patch_shape=patch_shape, resize_inputs=resize_inputs, resize_kwargs=resize_kwargs
227        )
228
229    return torch_em.default_segmentation_dataset(
230        raw_paths=raw_paths,
231        raw_key=None,
232        label_paths=label_paths,
233        label_key=None,
234        is_seg_dataset=False,
235        patch_shape=patch_shape,
236        **kwargs
237    )

Get the Wing Disc Timelapse dataset for cell instance segmentation in epithelial timelapse microscopy.

Arguments:
  • path: Filepath to a folder where the data is downloaded for further processing.
  • patch_shape: The patch shape to use for training.
  • movies: The movies to use. By default all five movies are used. See MOVIES for the valid names.
  • resize_inputs: Whether to resize the inputs to the patch shape.
  • download: Whether to download the data if it is not present.
  • kwargs: Additional keyword arguments for torch_em.default_segmentation_dataset.
Returns:

The segmentation dataset.

def get_wing_disc_timelapse_loader( path: Union[os.PathLike, str], batch_size: int, patch_shape: Tuple[int, int], movies: Optional[Sequence[str]] = None, resize_inputs: bool = False, download: bool = False, **kwargs) -> torch.utils.data.dataloader.DataLoader:
240def get_wing_disc_timelapse_loader(
241    path: Union[os.PathLike, str],
242    batch_size: int,
243    patch_shape: Tuple[int, int],
244    movies: Optional[Sequence[str]] = None,
245    resize_inputs: bool = False,
246    download: bool = False,
247    **kwargs
248) -> DataLoader:
249    """Get the Wing Disc Timelapse dataloader for cell instance segmentation in epithelial timelapse microscopy.
250
251    Args:
252        path: Filepath to a folder where the data is downloaded for further processing.
253        batch_size: The batch size for training.
254        patch_shape: The patch shape to use for training.
255        movies: The movies to use. By default all five movies are used. See `MOVIES` for the valid names.
256        resize_inputs: Whether to resize the inputs to the patch shape.
257        download: Whether to download the data if it is not present.
258        kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset` or for the PyTorch DataLoader.
259
260    Returns:
261        The DataLoader.
262    """
263    ds_kwargs, loader_kwargs = util.split_kwargs(torch_em.default_segmentation_dataset, **kwargs)
264    dataset = get_wing_disc_timelapse_dataset(path, patch_shape, movies, resize_inputs, download, **ds_kwargs)
265    return torch_em.get_data_loader(dataset, batch_size, **loader_kwargs)

Get the Wing Disc Timelapse dataloader for cell instance segmentation in epithelial timelapse microscopy.

Arguments:
  • path: Filepath to a folder where the data is downloaded for further processing.
  • batch_size: The batch size for training.
  • patch_shape: The patch shape to use for training.
  • movies: The movies to use. By default all five movies are used. See MOVIES for the valid names.
  • resize_inputs: Whether to resize the inputs to the patch shape.
  • download: Whether to download the data if it is not present.
  • kwargs: Additional keyword arguments for torch_em.default_segmentation_dataset or for the PyTorch DataLoader.
Returns:

The DataLoader.