torch_em.data.datasets.medical.echonet_pediatric
EchoNet-Pediatric contains annotations for left ventricle segmentation in pediatric echocardiography videos, split into apical-4-chamber (A4C) and parasternal-short-axis (PSAX) views.
The dataset comprises 7,643 videos from 1,958 patients aged 0-18 years, collected at Lucile Packard
Children's Hospital Stanford between 2014 and 2021. Each video only has expert tracings for the two
labeled frames (end-diastole and end-systole) used to compute the left ventricular ejection fraction,
not dense per-frame masks. The dataset is located at https://echonet.github.io/pediatric/ and
distributed through the Stanford AIMI Center Shared Datasets Portal under a non-commercial research
use agreement: registration as an individual user is required, re-distribution (including sharing the
download link) is forbidden, and re-identification attempts are prohibited. This module cannot
download the dataset automatically; see get_echonet_pediatric_data for the manual steps.
This dataset is from the publication https://doi.org/10.1016/j.echo.2023.01.015 (cite the DOI https://doi.org/10.71718/d05h-gy43 for the data itself). Please cite them if you use this dataset in your research.
NOTE: Reading the videos requires 'opencv-python' ('cv2'), and the tracing rasterization requires 'scikit-image'.
1"""EchoNet-Pediatric contains annotations for left ventricle segmentation in pediatric 2echocardiography videos, split into apical-4-chamber (A4C) and parasternal-short-axis (PSAX) views. 3 4The dataset comprises 7,643 videos from 1,958 patients aged 0-18 years, collected at Lucile Packard 5Children's Hospital Stanford between 2014 and 2021. Each video only has expert tracings for the two 6labeled frames (end-diastole and end-systole) used to compute the left ventricular ejection fraction, 7not dense per-frame masks. The dataset is located at https://echonet.github.io/pediatric/ and 8distributed through the Stanford AIMI Center Shared Datasets Portal under a non-commercial research 9use agreement: registration as an individual user is required, re-distribution (including sharing the 10download link) is forbidden, and re-identification attempts are prohibited. This module cannot 11download the dataset automatically; see `get_echonet_pediatric_data` for the manual steps. 12 13This dataset is from the publication https://doi.org/10.1016/j.echo.2023.01.015 (cite the DOI 14https://doi.org/10.71718/d05h-gy43 for the data itself). Please cite them if you use this dataset in 15your research. 16 17NOTE: Reading the videos requires 'opencv-python' ('cv2'), and the tracing rasterization requires 18'scikit-image'. 19""" 20 21import os 22from glob import glob 23from tqdm import tqdm 24from natsort import natsorted 25from typing import Union, Tuple, Literal, List 26 27import numpy as np 28import pandas as pd 29import imageio.v3 as imageio 30 31from torch.utils.data import Dataset, DataLoader 32 33import torch_em 34 35from .. import util 36 37 38VIEWS = ("A4C", "PSAX") 39SPLITS = ("TRAIN", "VAL", "TEST") 40 41 42def _trace_to_mask(x1, y1, x2, y2, shape): 43 from skimage.draw import polygon 44 45 # Follows the rasterization used by the EchoNet dataset releases: the first coordinate pair is 46 # the long axis of the left ventricle, and the remaining pairs are the perpendicular short-axis 47 # distances. Walking the two sides of the short-axis points (one side forwards, the other 48 # backwards) traces the outline of the traced region, which is then filled in. 49 x = np.concatenate((x1[1:], x2[1:][::-1])) 50 y = np.concatenate((y1[1:], y2[1:][::-1])) 51 52 rows, cols = polygon(np.round(y).astype(int), np.round(x).astype(int), shape=shape) 53 mask = np.zeros(shape, dtype="uint8") 54 mask[rows, cols] = 1 55 return mask 56 57 58def _preprocess_view(data_dir, view, preprocessed_dir): 59 import cv2 60 61 file_list = pd.read_csv(os.path.join(data_dir, view, "FileList.csv")) 62 tracings = pd.read_csv(os.path.join(data_dir, view, "VolumeTracings.csv")) 63 64 videos_dir = os.path.join(data_dir, view, "Videos") 65 view_out_dir = os.path.join(preprocessed_dir, view) 66 os.makedirs(view_out_dir, exist_ok=True) 67 68 image_paths, gt_paths = [], [] 69 for _, row in tqdm(file_list.iterrows(), total=len(file_list), desc=f"Preprocess EchoNet-Pediatric {view}"): 70 file_name = row["FileName"] 71 stem = os.path.splitext(file_name)[0] 72 split = str(row["Split"]).upper() 73 74 video_tracings = tracings[tracings["FileName"] == file_name] 75 if video_tracings.empty: 76 video_tracings = tracings[tracings["FileName"] == stem + ".avi"] 77 if video_tracings.empty: 78 continue 79 80 video_path = os.path.join(videos_dir, file_name) 81 if not os.path.exists(video_path): 82 video_path = os.path.join(videos_dir, stem + ".avi") 83 84 capture = cv2.VideoCapture(video_path) 85 86 for frame_idx, frame_tracings in video_tracings.groupby("Frame"): 87 image_path = os.path.join(view_out_dir, f"{stem}_{split}_{frame_idx}.tif") 88 mask_path = os.path.join(view_out_dir, f"{stem}_{split}_{frame_idx}_mask.tif") 89 90 if os.path.exists(image_path) and os.path.exists(mask_path): 91 image_paths.append(image_path) 92 gt_paths.append(mask_path) 93 continue 94 95 capture.set(cv2.CAP_PROP_POS_FRAMES, int(frame_idx)) 96 success, frame = capture.read() 97 if not success: 98 continue 99 100 frame = cv2.cvtColor(frame, cv2.COLOR_BGR2GRAY) 101 mask = _trace_to_mask( 102 frame_tracings["X1"].to_numpy(), 103 frame_tracings["Y1"].to_numpy(), 104 frame_tracings["X2"].to_numpy(), 105 frame_tracings["Y2"].to_numpy(), 106 frame.shape, 107 ) 108 109 imageio.imwrite(image_path, frame) 110 imageio.imwrite(mask_path, mask) 111 112 image_paths.append(image_path) 113 gt_paths.append(mask_path) 114 115 capture.release() 116 117 return image_paths, gt_paths 118 119 120def get_echonet_pediatric_data( 121 path: Union[os.PathLike, str], view: Literal["A4C", "PSAX"], download: bool = False 122) -> str: 123 """Obtain the EchoNet-Pediatric dataset. 124 125 NOTE: 'torch_em' cannot download this dataset, as it requires individual registration and 126 agreement to the EchoNet-Pediatric Research Use Agreement. Please follow these steps: 127 - Visit https://echonet.github.io/pediatric/ and click on the download / access instructions. 128 - Register (individually, per user) with the Stanford AIMI Center Shared Datasets Portal and 129 agree to the Research Use Agreement (non-commercial research use only, no re-distribution). 130 - Download the dataset and place it at `path`, so that it has the following structure, with one 131 folder per view: 132 `path/A4C/Videos`, `path/A4C/FileList.csv`, `path/A4C/VolumeTracings.csv`, and analogously for 133 `path/PSAX`. 134 135 Args: 136 path: Filepath to a folder where the dataset is stored. 137 view: The choice of view. Either 'A4C' or 'PSAX'. 138 download: Whether to download the data if it is not present. This is not supported for this 139 dataset, and setting it to True raises an error explaining the manual steps above. 140 141 Returns: 142 Filepath to the folder where the dataset is stored. 143 """ 144 assert view in VIEWS, f"'{view}' is not a valid view choice for the EchoNet-Pediatric dataset." 145 146 if download: 147 raise NotImplementedError( 148 "Download is set to True, but 'torch_em' cannot download the EchoNet-Pediatric dataset. " 149 "See `get_echonet_pediatric_data` for the manual registration and download steps." 150 ) 151 152 view_dir = os.path.join(path, view) 153 if not ( 154 os.path.exists(os.path.join(view_dir, "FileList.csv")) 155 and os.path.exists(os.path.join(view_dir, "VolumeTracings.csv")) 156 and os.path.exists(os.path.join(view_dir, "Videos")) 157 ): 158 raise RuntimeError( 159 f"Cannot find the EchoNet-Pediatric '{view}' data at '{view_dir}'. " 160 "This dataset requires manual download, see `get_echonet_pediatric_data` for the steps." 161 ) 162 163 return path 164 165 166def get_echonet_pediatric_paths( 167 path: Union[os.PathLike, str], 168 view: Literal["A4C", "PSAX"], 169 split: Literal["TRAIN", "VAL", "TEST", None] = None, 170 download: bool = False, 171) -> Tuple[List[str], List[str]]: 172 """Get paths to the EchoNet-Pediatric data. 173 174 Args: 175 path: Filepath to a folder where the dataset is stored. 176 view: The choice of view. Either 'A4C' or 'PSAX'. 177 split: The choice of data split. Either 'TRAIN', 'VAL' or 'TEST'. By default, all splits are used. 178 download: Whether to download the data if it is not present. See `get_echonet_pediatric_data`. 179 180 Returns: 181 List of filepaths for the image data. 182 List of filepaths for the label data. 183 """ 184 data_dir = get_echonet_pediatric_data(path, view, download) 185 186 preprocessed_dir = os.path.join(path, "preprocessed") 187 view_out_dir = os.path.join(preprocessed_dir, view) 188 189 image_paths = natsorted(glob(os.path.join(view_out_dir, "*.tif"))) 190 image_paths = [p for p in image_paths if not p.endswith("_mask.tif")] 191 if not image_paths: 192 image_paths, _ = _preprocess_view(data_dir, view, preprocessed_dir) 193 image_paths = natsorted(image_paths) 194 195 gt_paths = [p.replace(".tif", "_mask.tif") for p in image_paths] 196 assert len(image_paths) == len(gt_paths) and len(image_paths) > 0 197 198 if split is not None: 199 assert split in SPLITS, f"'{split}' is not a valid split choice for the EchoNet-Pediatric dataset." 200 image_paths = [p for p in image_paths if f"_{split}_" in os.path.basename(p)] 201 gt_paths = [p for p in gt_paths if f"_{split}_" in os.path.basename(p)] 202 203 return image_paths, gt_paths 204 205 206def get_echonet_pediatric_dataset( 207 path: Union[os.PathLike, str], 208 patch_shape: Tuple[int, int], 209 view: Literal["A4C", "PSAX"], 210 split: Literal["TRAIN", "VAL", "TEST", None] = None, 211 resize_inputs: bool = False, 212 download: bool = False, 213 **kwargs 214) -> Dataset: 215 """Get the EchoNet-Pediatric dataset for left ventricle segmentation in echocardiography videos. 216 217 Args: 218 path: Filepath to a folder where the dataset is stored. 219 patch_shape: The patch shape to use for training. 220 view: The choice of view. Either 'A4C' or 'PSAX'. 221 split: The choice of data split. Either 'TRAIN', 'VAL' or 'TEST'. By default, all splits are used. 222 resize_inputs: Whether to resize inputs to the desired patch shape. 223 download: Whether to download the data if it is not present. See `get_echonet_pediatric_data`. 224 kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset`. 225 226 Returns: 227 The segmentation dataset. 228 """ 229 image_paths, gt_paths = get_echonet_pediatric_paths(path, view, split, download) 230 231 if resize_inputs: 232 resize_kwargs = {"patch_shape": patch_shape, "is_rgb": False} 233 kwargs, patch_shape = util.update_kwargs_for_resize_trafo( 234 kwargs=kwargs, patch_shape=patch_shape, resize_inputs=resize_inputs, resize_kwargs=resize_kwargs 235 ) 236 237 return torch_em.default_segmentation_dataset( 238 raw_paths=image_paths, 239 raw_key=None, 240 label_paths=gt_paths, 241 label_key=None, 242 patch_shape=patch_shape, 243 is_seg_dataset=False, 244 **kwargs 245 ) 246 247 248def get_echonet_pediatric_loader( 249 path: Union[os.PathLike, str], 250 batch_size: int, 251 patch_shape: Tuple[int, int], 252 view: Literal["A4C", "PSAX"], 253 split: Literal["TRAIN", "VAL", "TEST", None] = None, 254 resize_inputs: bool = False, 255 download: bool = False, 256 **kwargs 257) -> DataLoader: 258 """Get the EchoNet-Pediatric dataloader for left ventricle segmentation in echocardiography videos. 259 260 Args: 261 path: Filepath to a folder where the dataset is stored. 262 batch_size: The batch size for training. 263 patch_shape: The patch shape to use for training. 264 view: The choice of view. Either 'A4C' or 'PSAX'. 265 split: The choice of data split. Either 'TRAIN', 'VAL' or 'TEST'. By default, all splits are used. 266 resize_inputs: Whether to resize inputs to the desired patch shape. 267 download: Whether to download the data if it is not present. See `get_echonet_pediatric_data`. 268 kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset` or for the PyTorch DataLoader. 269 270 Returns: 271 The DataLoader. 272 """ 273 ds_kwargs, loader_kwargs = util.split_kwargs(torch_em.default_segmentation_dataset, **kwargs) 274 dataset = get_echonet_pediatric_dataset(path, patch_shape, view, split, resize_inputs, download, **ds_kwargs) 275 return torch_em.get_data_loader(dataset, batch_size, **loader_kwargs)
121def get_echonet_pediatric_data( 122 path: Union[os.PathLike, str], view: Literal["A4C", "PSAX"], download: bool = False 123) -> str: 124 """Obtain the EchoNet-Pediatric dataset. 125 126 NOTE: 'torch_em' cannot download this dataset, as it requires individual registration and 127 agreement to the EchoNet-Pediatric Research Use Agreement. Please follow these steps: 128 - Visit https://echonet.github.io/pediatric/ and click on the download / access instructions. 129 - Register (individually, per user) with the Stanford AIMI Center Shared Datasets Portal and 130 agree to the Research Use Agreement (non-commercial research use only, no re-distribution). 131 - Download the dataset and place it at `path`, so that it has the following structure, with one 132 folder per view: 133 `path/A4C/Videos`, `path/A4C/FileList.csv`, `path/A4C/VolumeTracings.csv`, and analogously for 134 `path/PSAX`. 135 136 Args: 137 path: Filepath to a folder where the dataset is stored. 138 view: The choice of view. Either 'A4C' or 'PSAX'. 139 download: Whether to download the data if it is not present. This is not supported for this 140 dataset, and setting it to True raises an error explaining the manual steps above. 141 142 Returns: 143 Filepath to the folder where the dataset is stored. 144 """ 145 assert view in VIEWS, f"'{view}' is not a valid view choice for the EchoNet-Pediatric dataset." 146 147 if download: 148 raise NotImplementedError( 149 "Download is set to True, but 'torch_em' cannot download the EchoNet-Pediatric dataset. " 150 "See `get_echonet_pediatric_data` for the manual registration and download steps." 151 ) 152 153 view_dir = os.path.join(path, view) 154 if not ( 155 os.path.exists(os.path.join(view_dir, "FileList.csv")) 156 and os.path.exists(os.path.join(view_dir, "VolumeTracings.csv")) 157 and os.path.exists(os.path.join(view_dir, "Videos")) 158 ): 159 raise RuntimeError( 160 f"Cannot find the EchoNet-Pediatric '{view}' data at '{view_dir}'. " 161 "This dataset requires manual download, see `get_echonet_pediatric_data` for the steps." 162 ) 163 164 return path
Obtain the EchoNet-Pediatric dataset.
NOTE: 'torch_em' cannot download this dataset, as it requires individual registration and agreement to the EchoNet-Pediatric Research Use Agreement. Please follow these steps:
- Visit https://echonet.github.io/pediatric/ and click on the download / access instructions.
- Register (individually, per user) with the Stanford AIMI Center Shared Datasets Portal and agree to the Research Use Agreement (non-commercial research use only, no re-distribution).
- Download the dataset and place it at
path, so that it has the following structure, with one folder per view:path/A4C/Videos,path/A4C/FileList.csv,path/A4C/VolumeTracings.csv, and analogously forpath/PSAX.
Arguments:
- path: Filepath to a folder where the dataset is stored.
- view: The choice of view. Either 'A4C' or 'PSAX'.
- download: Whether to download the data if it is not present. This is not supported for this dataset, and setting it to True raises an error explaining the manual steps above.
Returns:
Filepath to the folder where the dataset is stored.
167def get_echonet_pediatric_paths( 168 path: Union[os.PathLike, str], 169 view: Literal["A4C", "PSAX"], 170 split: Literal["TRAIN", "VAL", "TEST", None] = None, 171 download: bool = False, 172) -> Tuple[List[str], List[str]]: 173 """Get paths to the EchoNet-Pediatric data. 174 175 Args: 176 path: Filepath to a folder where the dataset is stored. 177 view: The choice of view. Either 'A4C' or 'PSAX'. 178 split: The choice of data split. Either 'TRAIN', 'VAL' or 'TEST'. By default, all splits are used. 179 download: Whether to download the data if it is not present. See `get_echonet_pediatric_data`. 180 181 Returns: 182 List of filepaths for the image data. 183 List of filepaths for the label data. 184 """ 185 data_dir = get_echonet_pediatric_data(path, view, download) 186 187 preprocessed_dir = os.path.join(path, "preprocessed") 188 view_out_dir = os.path.join(preprocessed_dir, view) 189 190 image_paths = natsorted(glob(os.path.join(view_out_dir, "*.tif"))) 191 image_paths = [p for p in image_paths if not p.endswith("_mask.tif")] 192 if not image_paths: 193 image_paths, _ = _preprocess_view(data_dir, view, preprocessed_dir) 194 image_paths = natsorted(image_paths) 195 196 gt_paths = [p.replace(".tif", "_mask.tif") for p in image_paths] 197 assert len(image_paths) == len(gt_paths) and len(image_paths) > 0 198 199 if split is not None: 200 assert split in SPLITS, f"'{split}' is not a valid split choice for the EchoNet-Pediatric dataset." 201 image_paths = [p for p in image_paths if f"_{split}_" in os.path.basename(p)] 202 gt_paths = [p for p in gt_paths if f"_{split}_" in os.path.basename(p)] 203 204 return image_paths, gt_paths
Get paths to the EchoNet-Pediatric data.
Arguments:
- path: Filepath to a folder where the dataset is stored.
- view: The choice of view. Either 'A4C' or 'PSAX'.
- split: The choice of data split. Either 'TRAIN', 'VAL' or 'TEST'. By default, all splits are used.
- download: Whether to download the data if it is not present. See
get_echonet_pediatric_data.
Returns:
List of filepaths for the image data. List of filepaths for the label data.
207def get_echonet_pediatric_dataset( 208 path: Union[os.PathLike, str], 209 patch_shape: Tuple[int, int], 210 view: Literal["A4C", "PSAX"], 211 split: Literal["TRAIN", "VAL", "TEST", None] = None, 212 resize_inputs: bool = False, 213 download: bool = False, 214 **kwargs 215) -> Dataset: 216 """Get the EchoNet-Pediatric dataset for left ventricle segmentation in echocardiography videos. 217 218 Args: 219 path: Filepath to a folder where the dataset is stored. 220 patch_shape: The patch shape to use for training. 221 view: The choice of view. Either 'A4C' or 'PSAX'. 222 split: The choice of data split. Either 'TRAIN', 'VAL' or 'TEST'. By default, all splits are used. 223 resize_inputs: Whether to resize inputs to the desired patch shape. 224 download: Whether to download the data if it is not present. See `get_echonet_pediatric_data`. 225 kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset`. 226 227 Returns: 228 The segmentation dataset. 229 """ 230 image_paths, gt_paths = get_echonet_pediatric_paths(path, view, split, download) 231 232 if resize_inputs: 233 resize_kwargs = {"patch_shape": patch_shape, "is_rgb": False} 234 kwargs, patch_shape = util.update_kwargs_for_resize_trafo( 235 kwargs=kwargs, patch_shape=patch_shape, resize_inputs=resize_inputs, resize_kwargs=resize_kwargs 236 ) 237 238 return torch_em.default_segmentation_dataset( 239 raw_paths=image_paths, 240 raw_key=None, 241 label_paths=gt_paths, 242 label_key=None, 243 patch_shape=patch_shape, 244 is_seg_dataset=False, 245 **kwargs 246 )
Get the EchoNet-Pediatric dataset for left ventricle segmentation in echocardiography videos.
Arguments:
- path: Filepath to a folder where the dataset is stored.
- patch_shape: The patch shape to use for training.
- view: The choice of view. Either 'A4C' or 'PSAX'.
- split: The choice of data split. Either 'TRAIN', 'VAL' or 'TEST'. By default, all splits are used.
- resize_inputs: Whether to resize inputs to the desired patch shape.
- download: Whether to download the data if it is not present. See
get_echonet_pediatric_data. - kwargs: Additional keyword arguments for
torch_em.default_segmentation_dataset.
Returns:
The segmentation dataset.
249def get_echonet_pediatric_loader( 250 path: Union[os.PathLike, str], 251 batch_size: int, 252 patch_shape: Tuple[int, int], 253 view: Literal["A4C", "PSAX"], 254 split: Literal["TRAIN", "VAL", "TEST", None] = None, 255 resize_inputs: bool = False, 256 download: bool = False, 257 **kwargs 258) -> DataLoader: 259 """Get the EchoNet-Pediatric dataloader for left ventricle segmentation in echocardiography videos. 260 261 Args: 262 path: Filepath to a folder where the dataset is stored. 263 batch_size: The batch size for training. 264 patch_shape: The patch shape to use for training. 265 view: The choice of view. Either 'A4C' or 'PSAX'. 266 split: The choice of data split. Either 'TRAIN', 'VAL' or 'TEST'. By default, all splits are used. 267 resize_inputs: Whether to resize inputs to the desired patch shape. 268 download: Whether to download the data if it is not present. See `get_echonet_pediatric_data`. 269 kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset` or for the PyTorch DataLoader. 270 271 Returns: 272 The DataLoader. 273 """ 274 ds_kwargs, loader_kwargs = util.split_kwargs(torch_em.default_segmentation_dataset, **kwargs) 275 dataset = get_echonet_pediatric_dataset(path, patch_shape, view, split, resize_inputs, download, **ds_kwargs) 276 return torch_em.get_data_loader(dataset, batch_size, **loader_kwargs)
Get the EchoNet-Pediatric dataloader for left ventricle segmentation in echocardiography videos.
Arguments:
- path: Filepath to a folder where the dataset is stored.
- batch_size: The batch size for training.
- patch_shape: The patch shape to use for training.
- view: The choice of view. Either 'A4C' or 'PSAX'.
- split: The choice of data split. Either 'TRAIN', 'VAL' or 'TEST'. By default, all splits are used.
- resize_inputs: Whether to resize inputs to the desired patch shape.
- download: Whether to download the data if it is not present. See
get_echonet_pediatric_data. - kwargs: Additional keyword arguments for
torch_em.default_segmentation_datasetor for the PyTorch DataLoader.
Returns:
The DataLoader.