torch_em.data.datasets.medical.echonet_pediatric

EchoNet-Pediatric contains annotations for left ventricle segmentation in pediatric echocardiography videos, split into apical-4-chamber (A4C) and parasternal-short-axis (PSAX) views.

The dataset comprises 7,643 videos from 1,958 patients aged 0-18 years, collected at Lucile Packard Children's Hospital Stanford between 2014 and 2021. Each video only has expert tracings for the two labeled frames (end-diastole and end-systole) used to compute the left ventricular ejection fraction, not dense per-frame masks. The dataset is located at https://echonet.github.io/pediatric/ and distributed through the Stanford AIMI Center Shared Datasets Portal under a non-commercial research use agreement: registration as an individual user is required, re-distribution (including sharing the download link) is forbidden, and re-identification attempts are prohibited. This module cannot download the dataset automatically; see get_echonet_pediatric_data for the manual steps.

This dataset is from the publication https://doi.org/10.1016/j.echo.2023.01.015 (cite the DOI https://doi.org/10.71718/d05h-gy43 for the data itself). Please cite them if you use this dataset in your research.

NOTE: Reading the videos requires 'opencv-python' ('cv2'), and the tracing rasterization requires 'scikit-image'.

  1"""EchoNet-Pediatric contains annotations for left ventricle segmentation in pediatric
  2echocardiography videos, split into apical-4-chamber (A4C) and parasternal-short-axis (PSAX) views.
  3
  4The dataset comprises 7,643 videos from 1,958 patients aged 0-18 years, collected at Lucile Packard
  5Children's Hospital Stanford between 2014 and 2021. Each video only has expert tracings for the two
  6labeled frames (end-diastole and end-systole) used to compute the left ventricular ejection fraction,
  7not dense per-frame masks. The dataset is located at https://echonet.github.io/pediatric/ and
  8distributed through the Stanford AIMI Center Shared Datasets Portal under a non-commercial research
  9use agreement: registration as an individual user is required, re-distribution (including sharing the
 10download link) is forbidden, and re-identification attempts are prohibited. This module cannot
 11download the dataset automatically; see `get_echonet_pediatric_data` for the manual steps.
 12
 13This dataset is from the publication https://doi.org/10.1016/j.echo.2023.01.015 (cite the DOI
 14https://doi.org/10.71718/d05h-gy43 for the data itself). Please cite them if you use this dataset in
 15your research.
 16
 17NOTE: Reading the videos requires 'opencv-python' ('cv2'), and the tracing rasterization requires
 18'scikit-image'.
 19"""
 20
 21import os
 22from glob import glob
 23from tqdm import tqdm
 24from natsort import natsorted
 25from typing import Union, Tuple, Literal, List
 26
 27import numpy as np
 28import pandas as pd
 29import imageio.v3 as imageio
 30
 31from torch.utils.data import Dataset, DataLoader
 32
 33import torch_em
 34
 35from .. import util
 36
 37
 38VIEWS = ("A4C", "PSAX")
 39SPLITS = ("TRAIN", "VAL", "TEST")
 40
 41
 42def _trace_to_mask(x1, y1, x2, y2, shape):
 43    from skimage.draw import polygon
 44
 45    # Follows the rasterization used by the EchoNet dataset releases: the first coordinate pair is
 46    # the long axis of the left ventricle, and the remaining pairs are the perpendicular short-axis
 47    # distances. Walking the two sides of the short-axis points (one side forwards, the other
 48    # backwards) traces the outline of the traced region, which is then filled in.
 49    x = np.concatenate((x1[1:], x2[1:][::-1]))
 50    y = np.concatenate((y1[1:], y2[1:][::-1]))
 51
 52    rows, cols = polygon(np.round(y).astype(int), np.round(x).astype(int), shape=shape)
 53    mask = np.zeros(shape, dtype="uint8")
 54    mask[rows, cols] = 1
 55    return mask
 56
 57
 58def _preprocess_view(data_dir, view, preprocessed_dir):
 59    import cv2
 60
 61    file_list = pd.read_csv(os.path.join(data_dir, view, "FileList.csv"))
 62    tracings = pd.read_csv(os.path.join(data_dir, view, "VolumeTracings.csv"))
 63
 64    videos_dir = os.path.join(data_dir, view, "Videos")
 65    view_out_dir = os.path.join(preprocessed_dir, view)
 66    os.makedirs(view_out_dir, exist_ok=True)
 67
 68    image_paths, gt_paths = [], []
 69    for _, row in tqdm(file_list.iterrows(), total=len(file_list), desc=f"Preprocess EchoNet-Pediatric {view}"):
 70        file_name = row["FileName"]
 71        stem = os.path.splitext(file_name)[0]
 72        split = str(row["Split"]).upper()
 73
 74        video_tracings = tracings[tracings["FileName"] == file_name]
 75        if video_tracings.empty:
 76            video_tracings = tracings[tracings["FileName"] == stem + ".avi"]
 77        if video_tracings.empty:
 78            continue
 79
 80        video_path = os.path.join(videos_dir, file_name)
 81        if not os.path.exists(video_path):
 82            video_path = os.path.join(videos_dir, stem + ".avi")
 83
 84        capture = cv2.VideoCapture(video_path)
 85
 86        for frame_idx, frame_tracings in video_tracings.groupby("Frame"):
 87            image_path = os.path.join(view_out_dir, f"{stem}_{split}_{frame_idx}.tif")
 88            mask_path = os.path.join(view_out_dir, f"{stem}_{split}_{frame_idx}_mask.tif")
 89
 90            if os.path.exists(image_path) and os.path.exists(mask_path):
 91                image_paths.append(image_path)
 92                gt_paths.append(mask_path)
 93                continue
 94
 95            capture.set(cv2.CAP_PROP_POS_FRAMES, int(frame_idx))
 96            success, frame = capture.read()
 97            if not success:
 98                continue
 99
100            frame = cv2.cvtColor(frame, cv2.COLOR_BGR2GRAY)
101            mask = _trace_to_mask(
102                frame_tracings["X1"].to_numpy(),
103                frame_tracings["Y1"].to_numpy(),
104                frame_tracings["X2"].to_numpy(),
105                frame_tracings["Y2"].to_numpy(),
106                frame.shape,
107            )
108
109            imageio.imwrite(image_path, frame)
110            imageio.imwrite(mask_path, mask)
111
112            image_paths.append(image_path)
113            gt_paths.append(mask_path)
114
115        capture.release()
116
117    return image_paths, gt_paths
118
119
120def get_echonet_pediatric_data(
121    path: Union[os.PathLike, str], view: Literal["A4C", "PSAX"], download: bool = False
122) -> str:
123    """Obtain the EchoNet-Pediatric dataset.
124
125    NOTE: 'torch_em' cannot download this dataset, as it requires individual registration and
126    agreement to the EchoNet-Pediatric Research Use Agreement. Please follow these steps:
127    - Visit https://echonet.github.io/pediatric/ and click on the download / access instructions.
128    - Register (individually, per user) with the Stanford AIMI Center Shared Datasets Portal and
129      agree to the Research Use Agreement (non-commercial research use only, no re-distribution).
130    - Download the dataset and place it at `path`, so that it has the following structure, with one
131      folder per view:
132      `path/A4C/Videos`, `path/A4C/FileList.csv`, `path/A4C/VolumeTracings.csv`, and analogously for
133      `path/PSAX`.
134
135    Args:
136        path: Filepath to a folder where the dataset is stored.
137        view: The choice of view. Either 'A4C' or 'PSAX'.
138        download: Whether to download the data if it is not present. This is not supported for this
139            dataset, and setting it to True raises an error explaining the manual steps above.
140
141    Returns:
142        Filepath to the folder where the dataset is stored.
143    """
144    assert view in VIEWS, f"'{view}' is not a valid view choice for the EchoNet-Pediatric dataset."
145
146    if download:
147        raise NotImplementedError(
148            "Download is set to True, but 'torch_em' cannot download the EchoNet-Pediatric dataset. "
149            "See `get_echonet_pediatric_data` for the manual registration and download steps."
150        )
151
152    view_dir = os.path.join(path, view)
153    if not (
154        os.path.exists(os.path.join(view_dir, "FileList.csv"))
155        and os.path.exists(os.path.join(view_dir, "VolumeTracings.csv"))
156        and os.path.exists(os.path.join(view_dir, "Videos"))
157    ):
158        raise RuntimeError(
159            f"Cannot find the EchoNet-Pediatric '{view}' data at '{view_dir}'. "
160            "This dataset requires manual download, see `get_echonet_pediatric_data` for the steps."
161        )
162
163    return path
164
165
166def get_echonet_pediatric_paths(
167    path: Union[os.PathLike, str],
168    view: Literal["A4C", "PSAX"],
169    split: Literal["TRAIN", "VAL", "TEST", None] = None,
170    download: bool = False,
171) -> Tuple[List[str], List[str]]:
172    """Get paths to the EchoNet-Pediatric data.
173
174    Args:
175        path: Filepath to a folder where the dataset is stored.
176        view: The choice of view. Either 'A4C' or 'PSAX'.
177        split: The choice of data split. Either 'TRAIN', 'VAL' or 'TEST'. By default, all splits are used.
178        download: Whether to download the data if it is not present. See `get_echonet_pediatric_data`.
179
180    Returns:
181        List of filepaths for the image data.
182        List of filepaths for the label data.
183    """
184    data_dir = get_echonet_pediatric_data(path, view, download)
185
186    preprocessed_dir = os.path.join(path, "preprocessed")
187    view_out_dir = os.path.join(preprocessed_dir, view)
188
189    image_paths = natsorted(glob(os.path.join(view_out_dir, "*.tif")))
190    image_paths = [p for p in image_paths if not p.endswith("_mask.tif")]
191    if not image_paths:
192        image_paths, _ = _preprocess_view(data_dir, view, preprocessed_dir)
193        image_paths = natsorted(image_paths)
194
195    gt_paths = [p.replace(".tif", "_mask.tif") for p in image_paths]
196    assert len(image_paths) == len(gt_paths) and len(image_paths) > 0
197
198    if split is not None:
199        assert split in SPLITS, f"'{split}' is not a valid split choice for the EchoNet-Pediatric dataset."
200        image_paths = [p for p in image_paths if f"_{split}_" in os.path.basename(p)]
201        gt_paths = [p for p in gt_paths if f"_{split}_" in os.path.basename(p)]
202
203    return image_paths, gt_paths
204
205
206def get_echonet_pediatric_dataset(
207    path: Union[os.PathLike, str],
208    patch_shape: Tuple[int, int],
209    view: Literal["A4C", "PSAX"],
210    split: Literal["TRAIN", "VAL", "TEST", None] = None,
211    resize_inputs: bool = False,
212    download: bool = False,
213    **kwargs
214) -> Dataset:
215    """Get the EchoNet-Pediatric dataset for left ventricle segmentation in echocardiography videos.
216
217    Args:
218        path: Filepath to a folder where the dataset is stored.
219        patch_shape: The patch shape to use for training.
220        view: The choice of view. Either 'A4C' or 'PSAX'.
221        split: The choice of data split. Either 'TRAIN', 'VAL' or 'TEST'. By default, all splits are used.
222        resize_inputs: Whether to resize inputs to the desired patch shape.
223        download: Whether to download the data if it is not present. See `get_echonet_pediatric_data`.
224        kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset`.
225
226    Returns:
227        The segmentation dataset.
228    """
229    image_paths, gt_paths = get_echonet_pediatric_paths(path, view, split, download)
230
231    if resize_inputs:
232        resize_kwargs = {"patch_shape": patch_shape, "is_rgb": False}
233        kwargs, patch_shape = util.update_kwargs_for_resize_trafo(
234            kwargs=kwargs, patch_shape=patch_shape, resize_inputs=resize_inputs, resize_kwargs=resize_kwargs
235        )
236
237    return torch_em.default_segmentation_dataset(
238        raw_paths=image_paths,
239        raw_key=None,
240        label_paths=gt_paths,
241        label_key=None,
242        patch_shape=patch_shape,
243        is_seg_dataset=False,
244        **kwargs
245    )
246
247
248def get_echonet_pediatric_loader(
249    path: Union[os.PathLike, str],
250    batch_size: int,
251    patch_shape: Tuple[int, int],
252    view: Literal["A4C", "PSAX"],
253    split: Literal["TRAIN", "VAL", "TEST", None] = None,
254    resize_inputs: bool = False,
255    download: bool = False,
256    **kwargs
257) -> DataLoader:
258    """Get the EchoNet-Pediatric dataloader for left ventricle segmentation in echocardiography videos.
259
260    Args:
261        path: Filepath to a folder where the dataset is stored.
262        batch_size: The batch size for training.
263        patch_shape: The patch shape to use for training.
264        view: The choice of view. Either 'A4C' or 'PSAX'.
265        split: The choice of data split. Either 'TRAIN', 'VAL' or 'TEST'. By default, all splits are used.
266        resize_inputs: Whether to resize inputs to the desired patch shape.
267        download: Whether to download the data if it is not present. See `get_echonet_pediatric_data`.
268        kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset` or for the PyTorch DataLoader.
269
270    Returns:
271        The DataLoader.
272    """
273    ds_kwargs, loader_kwargs = util.split_kwargs(torch_em.default_segmentation_dataset, **kwargs)
274    dataset = get_echonet_pediatric_dataset(path, patch_shape, view, split, resize_inputs, download, **ds_kwargs)
275    return torch_em.get_data_loader(dataset, batch_size, **loader_kwargs)
VIEWS = ('A4C', 'PSAX')
SPLITS = ('TRAIN', 'VAL', 'TEST')
def get_echonet_pediatric_data( path: Union[os.PathLike, str], view: Literal['A4C', 'PSAX'], download: bool = False) -> str:
121def get_echonet_pediatric_data(
122    path: Union[os.PathLike, str], view: Literal["A4C", "PSAX"], download: bool = False
123) -> str:
124    """Obtain the EchoNet-Pediatric dataset.
125
126    NOTE: 'torch_em' cannot download this dataset, as it requires individual registration and
127    agreement to the EchoNet-Pediatric Research Use Agreement. Please follow these steps:
128    - Visit https://echonet.github.io/pediatric/ and click on the download / access instructions.
129    - Register (individually, per user) with the Stanford AIMI Center Shared Datasets Portal and
130      agree to the Research Use Agreement (non-commercial research use only, no re-distribution).
131    - Download the dataset and place it at `path`, so that it has the following structure, with one
132      folder per view:
133      `path/A4C/Videos`, `path/A4C/FileList.csv`, `path/A4C/VolumeTracings.csv`, and analogously for
134      `path/PSAX`.
135
136    Args:
137        path: Filepath to a folder where the dataset is stored.
138        view: The choice of view. Either 'A4C' or 'PSAX'.
139        download: Whether to download the data if it is not present. This is not supported for this
140            dataset, and setting it to True raises an error explaining the manual steps above.
141
142    Returns:
143        Filepath to the folder where the dataset is stored.
144    """
145    assert view in VIEWS, f"'{view}' is not a valid view choice for the EchoNet-Pediatric dataset."
146
147    if download:
148        raise NotImplementedError(
149            "Download is set to True, but 'torch_em' cannot download the EchoNet-Pediatric dataset. "
150            "See `get_echonet_pediatric_data` for the manual registration and download steps."
151        )
152
153    view_dir = os.path.join(path, view)
154    if not (
155        os.path.exists(os.path.join(view_dir, "FileList.csv"))
156        and os.path.exists(os.path.join(view_dir, "VolumeTracings.csv"))
157        and os.path.exists(os.path.join(view_dir, "Videos"))
158    ):
159        raise RuntimeError(
160            f"Cannot find the EchoNet-Pediatric '{view}' data at '{view_dir}'. "
161            "This dataset requires manual download, see `get_echonet_pediatric_data` for the steps."
162        )
163
164    return path

Obtain the EchoNet-Pediatric dataset.

NOTE: 'torch_em' cannot download this dataset, as it requires individual registration and agreement to the EchoNet-Pediatric Research Use Agreement. Please follow these steps:

  • Visit https://echonet.github.io/pediatric/ and click on the download / access instructions.
  • Register (individually, per user) with the Stanford AIMI Center Shared Datasets Portal and agree to the Research Use Agreement (non-commercial research use only, no re-distribution).
  • Download the dataset and place it at path, so that it has the following structure, with one folder per view: path/A4C/Videos, path/A4C/FileList.csv, path/A4C/VolumeTracings.csv, and analogously for path/PSAX.
Arguments:
  • path: Filepath to a folder where the dataset is stored.
  • view: The choice of view. Either 'A4C' or 'PSAX'.
  • download: Whether to download the data if it is not present. This is not supported for this dataset, and setting it to True raises an error explaining the manual steps above.
Returns:

Filepath to the folder where the dataset is stored.

def get_echonet_pediatric_paths( path: Union[os.PathLike, str], view: Literal['A4C', 'PSAX'], split: Literal['TRAIN', 'VAL', 'TEST', None] = None, download: bool = False) -> Tuple[List[str], List[str]]:
167def get_echonet_pediatric_paths(
168    path: Union[os.PathLike, str],
169    view: Literal["A4C", "PSAX"],
170    split: Literal["TRAIN", "VAL", "TEST", None] = None,
171    download: bool = False,
172) -> Tuple[List[str], List[str]]:
173    """Get paths to the EchoNet-Pediatric data.
174
175    Args:
176        path: Filepath to a folder where the dataset is stored.
177        view: The choice of view. Either 'A4C' or 'PSAX'.
178        split: The choice of data split. Either 'TRAIN', 'VAL' or 'TEST'. By default, all splits are used.
179        download: Whether to download the data if it is not present. See `get_echonet_pediatric_data`.
180
181    Returns:
182        List of filepaths for the image data.
183        List of filepaths for the label data.
184    """
185    data_dir = get_echonet_pediatric_data(path, view, download)
186
187    preprocessed_dir = os.path.join(path, "preprocessed")
188    view_out_dir = os.path.join(preprocessed_dir, view)
189
190    image_paths = natsorted(glob(os.path.join(view_out_dir, "*.tif")))
191    image_paths = [p for p in image_paths if not p.endswith("_mask.tif")]
192    if not image_paths:
193        image_paths, _ = _preprocess_view(data_dir, view, preprocessed_dir)
194        image_paths = natsorted(image_paths)
195
196    gt_paths = [p.replace(".tif", "_mask.tif") for p in image_paths]
197    assert len(image_paths) == len(gt_paths) and len(image_paths) > 0
198
199    if split is not None:
200        assert split in SPLITS, f"'{split}' is not a valid split choice for the EchoNet-Pediatric dataset."
201        image_paths = [p for p in image_paths if f"_{split}_" in os.path.basename(p)]
202        gt_paths = [p for p in gt_paths if f"_{split}_" in os.path.basename(p)]
203
204    return image_paths, gt_paths

Get paths to the EchoNet-Pediatric data.

Arguments:
  • path: Filepath to a folder where the dataset is stored.
  • view: The choice of view. Either 'A4C' or 'PSAX'.
  • split: The choice of data split. Either 'TRAIN', 'VAL' or 'TEST'. By default, all splits are used.
  • download: Whether to download the data if it is not present. See get_echonet_pediatric_data.
Returns:

List of filepaths for the image data. List of filepaths for the label data.

def get_echonet_pediatric_dataset( path: Union[os.PathLike, str], patch_shape: Tuple[int, int], view: Literal['A4C', 'PSAX'], split: Literal['TRAIN', 'VAL', 'TEST', None] = None, resize_inputs: bool = False, download: bool = False, **kwargs) -> torch.utils.data.dataset.Dataset:
207def get_echonet_pediatric_dataset(
208    path: Union[os.PathLike, str],
209    patch_shape: Tuple[int, int],
210    view: Literal["A4C", "PSAX"],
211    split: Literal["TRAIN", "VAL", "TEST", None] = None,
212    resize_inputs: bool = False,
213    download: bool = False,
214    **kwargs
215) -> Dataset:
216    """Get the EchoNet-Pediatric dataset for left ventricle segmentation in echocardiography videos.
217
218    Args:
219        path: Filepath to a folder where the dataset is stored.
220        patch_shape: The patch shape to use for training.
221        view: The choice of view. Either 'A4C' or 'PSAX'.
222        split: The choice of data split. Either 'TRAIN', 'VAL' or 'TEST'. By default, all splits are used.
223        resize_inputs: Whether to resize inputs to the desired patch shape.
224        download: Whether to download the data if it is not present. See `get_echonet_pediatric_data`.
225        kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset`.
226
227    Returns:
228        The segmentation dataset.
229    """
230    image_paths, gt_paths = get_echonet_pediatric_paths(path, view, split, download)
231
232    if resize_inputs:
233        resize_kwargs = {"patch_shape": patch_shape, "is_rgb": False}
234        kwargs, patch_shape = util.update_kwargs_for_resize_trafo(
235            kwargs=kwargs, patch_shape=patch_shape, resize_inputs=resize_inputs, resize_kwargs=resize_kwargs
236        )
237
238    return torch_em.default_segmentation_dataset(
239        raw_paths=image_paths,
240        raw_key=None,
241        label_paths=gt_paths,
242        label_key=None,
243        patch_shape=patch_shape,
244        is_seg_dataset=False,
245        **kwargs
246    )

Get the EchoNet-Pediatric dataset for left ventricle segmentation in echocardiography videos.

Arguments:
  • path: Filepath to a folder where the dataset is stored.
  • patch_shape: The patch shape to use for training.
  • view: The choice of view. Either 'A4C' or 'PSAX'.
  • split: The choice of data split. Either 'TRAIN', 'VAL' or 'TEST'. By default, all splits are used.
  • resize_inputs: Whether to resize inputs to the desired patch shape.
  • download: Whether to download the data if it is not present. See get_echonet_pediatric_data.
  • kwargs: Additional keyword arguments for torch_em.default_segmentation_dataset.
Returns:

The segmentation dataset.

def get_echonet_pediatric_loader( path: Union[os.PathLike, str], batch_size: int, patch_shape: Tuple[int, int], view: Literal['A4C', 'PSAX'], split: Literal['TRAIN', 'VAL', 'TEST', None] = None, resize_inputs: bool = False, download: bool = False, **kwargs) -> torch.utils.data.dataloader.DataLoader:
249def get_echonet_pediatric_loader(
250    path: Union[os.PathLike, str],
251    batch_size: int,
252    patch_shape: Tuple[int, int],
253    view: Literal["A4C", "PSAX"],
254    split: Literal["TRAIN", "VAL", "TEST", None] = None,
255    resize_inputs: bool = False,
256    download: bool = False,
257    **kwargs
258) -> DataLoader:
259    """Get the EchoNet-Pediatric dataloader for left ventricle segmentation in echocardiography videos.
260
261    Args:
262        path: Filepath to a folder where the dataset is stored.
263        batch_size: The batch size for training.
264        patch_shape: The patch shape to use for training.
265        view: The choice of view. Either 'A4C' or 'PSAX'.
266        split: The choice of data split. Either 'TRAIN', 'VAL' or 'TEST'. By default, all splits are used.
267        resize_inputs: Whether to resize inputs to the desired patch shape.
268        download: Whether to download the data if it is not present. See `get_echonet_pediatric_data`.
269        kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset` or for the PyTorch DataLoader.
270
271    Returns:
272        The DataLoader.
273    """
274    ds_kwargs, loader_kwargs = util.split_kwargs(torch_em.default_segmentation_dataset, **kwargs)
275    dataset = get_echonet_pediatric_dataset(path, patch_shape, view, split, resize_inputs, download, **ds_kwargs)
276    return torch_em.get_data_loader(dataset, batch_size, **loader_kwargs)

Get the EchoNet-Pediatric dataloader for left ventricle segmentation in echocardiography videos.

Arguments:
  • path: Filepath to a folder where the dataset is stored.
  • batch_size: The batch size for training.
  • patch_shape: The patch shape to use for training.
  • view: The choice of view. Either 'A4C' or 'PSAX'.
  • split: The choice of data split. Either 'TRAIN', 'VAL' or 'TEST'. By default, all splits are used.
  • resize_inputs: Whether to resize inputs to the desired patch shape.
  • download: Whether to download the data if it is not present. See get_echonet_pediatric_data.
  • kwargs: Additional keyword arguments for torch_em.default_segmentation_dataset or for the PyTorch DataLoader.
Returns:

The DataLoader.