torch_em.data.datasets.medical.proteas

The PROTEAS dataset contains annotations for brain metastasis segmentation in longitudinal MRI.

The dataset consists of 40 patients with metastatic brain cancer (45 archives, as a few patients are split into 'a' and 'b' courses of treatment) and 185 imaging studies (MRI and CT) over the course of radiotherapy and follow-up. It provides manual segmentations of 65 brain metastases, for each MRI study of a patient (the 'baseline' and the follow-ups 'fu1', 'fu2', ...), together with the radiotherapy plan and radiomics tables. The MRI studies come as skull-stripped and co-registered T1 ('t1'), contrast-enhanced T1 ('t1c'), T2 ('t2') and FLAIR ('fla') volumes of shape (240, 240, 155), on the same voxel grid as the segmentation masks. This loader pairs a chosen MRI sequence of each study with its mask. The raw DICOM series, the planning CT and the dose maps are not used and not kept.

The masks distinguish three tumor regions, with the label ids 1 = necrotic core, 2 = enhancing tumor and 3 = edema (see LABEL_IDS). NOTE: The dataset record does not document the label ids and they do not follow the BraTS numbering. They were inferred from the data: label 1 is enclosed by label 2, label 2 is bright in the contrast-enhanced T1 volumes and label 3 is the outermost region, bright in the FLAIR volumes.

NOTE: This is not the same as the BEAMSTER dataset in torch_em.data.datasets.medical.beamster, which has binary metastasis masks on contrast-enhanced T1 scans from a single time point, without longitudinal follow-up.

The data is located at https://doi.org/10.5281/zenodo.17253793 (v1 of the record, released in October 2025 under a CC BY 4.0 license). The latest version of the same record, https://doi.org/10.5281/zenodo.20025432, is access restricted, so this loader uses the open v1 record, whose availability may change.

The dataset is from the publication https://doi.org/10.1038/s41597-025-06131-0. Please cite it if you use this dataset for your research.

  1"""The PROTEAS dataset contains annotations for brain metastasis segmentation in longitudinal MRI.
  2
  3The dataset consists of 40 patients with metastatic brain cancer (45 archives, as a few patients are split into
  4'a' and 'b' courses of treatment) and 185 imaging studies (MRI and CT) over the course of radiotherapy and follow-up.
  5It provides manual segmentations of 65 brain metastases, for each MRI study of a patient (the 'baseline' and the
  6follow-ups 'fu1', 'fu2', ...), together with the radiotherapy plan and radiomics tables. The MRI studies come as
  7skull-stripped and co-registered T1 ('t1'), contrast-enhanced T1 ('t1c'), T2 ('t2') and FLAIR ('fla') volumes of shape
  8(240, 240, 155), on the same voxel grid as the segmentation masks. This loader pairs a chosen MRI sequence of each
  9study with its mask. The raw DICOM series, the planning CT and the dose maps are not used and not kept.
 10
 11The masks distinguish three tumor regions, with the label ids 1 = necrotic core, 2 = enhancing tumor and 3 = edema
 12(see `LABEL_IDS`). NOTE: The dataset record does not document the label ids and they do not follow the BraTS
 13numbering. They were inferred from the data: label 1 is enclosed by label 2, label 2 is bright in the
 14contrast-enhanced T1 volumes and label 3 is the outermost region, bright in the FLAIR volumes.
 15
 16NOTE: This is not the same as the BEAMSTER dataset in `torch_em.data.datasets.medical.beamster`, which has binary
 17metastasis masks on contrast-enhanced T1 scans from a single time point, without longitudinal follow-up.
 18
 19The data is located at https://doi.org/10.5281/zenodo.17253793 (v1 of the record, released in October 2025 under a
 20CC BY 4.0 license). The latest version of the same record, https://doi.org/10.5281/zenodo.20025432, is access
 21restricted, so this loader uses the open v1 record, whose availability may change.
 22
 23The dataset is from the publication https://doi.org/10.1038/s41597-025-06131-0.
 24Please cite it if you use this dataset for your research.
 25"""
 26
 27import os
 28import uuid
 29import shutil
 30import zipfile
 31from glob import glob
 32from concurrent import futures
 33from natsort import natsorted
 34from typing import Union, Tuple, List, Optional, Literal
 35
 36from tqdm import tqdm
 37from torch.utils.data import Dataset, DataLoader
 38
 39import torch_em
 40
 41from .. import util
 42
 43
 44URL = "https://zenodo.org/api/records/17253793/files/{patient_id}.zip/content"
 45
 46CHECKSUMS = {
 47    "P01": "9355b7f50d3e690737c7a5e4d22046d8df206212d4b09fa31ec1ad9f469ba54b",
 48    "P02": "0b4978e3036e1e5c3c8f0f55168c74b612152686ce3d7f8681e933b89e7098d8",
 49    "P03": "6f0a8e9d66016c06ec7e186ad3dac72f620c56a74345de66c75b70f4d3a73f9a",
 50    "P04a": "ec12923914188ef81eed02a7f6f4aed1ae839a40b1d6e63a47af90ebff540487",
 51    "P04b": "83c13be378f3efe79388bf3ad909f6426b6793b54ea3333c6a57955934218a35",
 52    "P05": "282d3d8f2ce3fb8a79cffa69ae6e28f414770cbf2e404211597474f58e11130a",
 53    "P06": "0a5d2f7f9998454abcf9143ec9e29309cf97d68946c8b41eaad6e7c90fdefa14",
 54    "P07a": "9698954c9f277036f710affa2590cbfb5f4aff0a17041307bd4411e9d9afe36e",
 55    "P07b": "9ea5e9ffcfe4ff80246dca983d3ea5da3309cc96c20c51b0ac00f7b500854d35",
 56    "P08": "6e6f559840766dd10af8ce6f20a94747f20319537ee3f568a68eba578e093e7b",
 57    "P09": "14bd3371ddb19f4fb486f469f27b0db42f78d19ac4c975abf2b30b0859ef35d0",
 58    "P10": "b5dd8535efa50a6102e1387338750273cec39ac8beee3b38bdffd7edc5c605a9",
 59    "P11": "2b2ab945002116de468dd1caf50914e272c0ee473b6708b37d497d548dffd2e0",
 60    "P12": "c08da88ed4b56e64d3db161b0b318d0e1b98baa6c28266fd9b1181979337a2e1",
 61    "P13": "ab0842a95e31e69ca0c64f4cfbbc495af481a3f8ab798a5cd3f6b49c01dae2c8",
 62    "P14": "ed2d56706456a2aca8e26cee365349326126d0271efa97141a16e86ea95afeaa",
 63    "P15": "b4abed11f5b9a9eedf7df659afe0e4ca4b57f135454963af71e610e7e6a19469",
 64    "P16": "beb26fb1c01f77c80f22eee591c73c5657b2cc3426d19d3bd4e6cceec525be79",
 65    "P17a": "1a776b528c60a7a49924c7bf1c59eedc93c37e6e0083c0d8f16ddb2932daf0f4",
 66    "P17b": "13c396baa500ee3b2c82b72c89e15f9c044850a02c7aa3615f2cced9ab6b39b7",
 67    "P18": "f064cfbb7303d5c8ca6597b60ca9550152c8c6c85de7bf900269d2708dfa930d",
 68    "P19": "56902860f712f4bce08e20c642217d4f18c45f3297a6d8376284f121a6c2cf18",
 69    "P20a": "24d04ba76de099dad365951f74788f436f086e87bafcbdd4704c884fc5d82a97",
 70    "P20b": "0072f1d4158727aa4088b3fb32294512a78aa7772da0f7a98063ee051e7660ba",
 71    "P21": "124192750c623382fc4911a112e9b372e90952e75a4e9da3fa602e467bccb3d3",
 72    "P22": "1011e6b69a5f864ba43032ded01ff6bbde6a02fc72354d2c63f7307c96a890af",
 73    "P23a": "818425160a1514fb2d3f6b5eb86b439cb4cc4f56b93c1e2c3b5690fb53e2ab99",
 74    "P23b": "2241d9c24d03cd69583e709bf498ab6abc86e791664048cd59d3de8453dd98dc",
 75    "P24": "b660ea009c3604cf5f4891ce6abb7fabf25f428f34a515353df55aa67b0944b5",
 76    "P25": "ebfd97a5efdb798f7c9a53c0536d25125d3fc14d85d90b7ef2b7d42949ec1503",
 77    "P26": "b526e4a2ec8f4a77097ccf925667ee6a4a527a6d213722d9f853f19a4546b31a",
 78    "P27": "76017fe710787309d4aa0ceb6a590c58fc5130c0e9e08a63119bf4b9b65ec5d7",
 79    "P28": "6e0c811e7e6da46296b9dc6e73050e65d8abada81f1900713c3fddb728a5b989",
 80    "P29": "6432a22a555c15d0086b2e7ada50a14f1927af1009fd6ec4fa7bf05d5efe4f18",
 81    "P30": "90c8beef41f82004181d1b7cd7db390f81723e3f57040b5ffa1f55205eb5734b",
 82    "P31": "553a410ebe4bc0d159b93f048e38021c22cd7ee4b506d9aa72ffb0241a463628",
 83    "P32": "05192b28f86022c40b82b1d82192428bd8387fe3378436bb0cc4333cc062e824",
 84    "P33": "64ee5d2d83fa2aebc4838603b8ada018d2082af71dd7e1e680585344377f1e63",
 85    "P34": "60063b57524b0978942b28eec70ac8cf28af053623454e1c0fdf446168663f5d",
 86    "P35": "73a23c05f6a1700384eb4d7b9ccf96ff745f6e7f1627e27a678d859521006dd0",
 87    "P36": "f4f897589a16831fe2ff2dd5d3185d39a98ce23313e03837d9399297e0a89ace",
 88    "P37": "d645a3910d394d82620588a6249fa6a9ece61a0f06f042acb9c69e7e296bb947",
 89    "P38": "5c8c12811462024ad51ccba928bc8ce405fb397632e4b675dc072f3c4a5dbe06",
 90    "P39": "bb3c06bcda4125597c353eaba8949bce420efcea831bb91cc047af14eb5b1296",
 91    "P40": "6a09cb9a8bbdd383760efa490255f18d0f43c348630bffc59dbf5d3ea985a569",
 92}
 93
 94SEQUENCES = ["t1", "t1c", "t2", "fla"]
 95
 96LABEL_IDS = {"necrotic_core": 1, "enhancing_tumor": 2, "edema": 3}
 97
 98
 99def _extract_patient(zip_path, patient_id, patient_dir):
100    tmp_dir = f"{patient_dir}.{uuid.uuid4().hex}.incomplete"
101    with zipfile.ZipFile(zip_path) as zf:
102        members = [
103            name for name in zf.namelist()
104            if name.endswith(".nii.gz") and ("/BraTS/" in name or "/tumor_segmentation/" in name)
105        ]
106        zf.extractall(tmp_dir, members)
107
108    os.replace(os.path.join(tmp_dir, patient_id), patient_dir)
109    shutil.rmtree(tmp_dir)
110
111
112def _download_patient(patient_id, path):
113    zip_path = os.path.join(path, f"{patient_id}.zip")
114    util.download_source(
115        path=zip_path, url=URL.format(patient_id=patient_id), download=True, checksum=CHECKSUMS[patient_id],
116    )
117    _extract_patient(zip_path, patient_id, os.path.join(path, "patients", patient_id))
118    os.remove(zip_path)
119
120
121def get_proteas_data(
122    path: Union[os.PathLike, str], n_patients: Optional[int] = None, download: bool = False,
123) -> str:
124    """Download the PROTEAS dataset.
125
126    NOTE: The archives of all patients are about 15 GB, as they contain the DICOM series next to the NIfTI volumes
127    used here. Use `n_patients` to only download a subset for a quick start.
128
129    Args:
130        path: Filepath to a folder where the data is downloaded for further processing.
131        n_patients: The number of patients (archives) to download, sorted by patient id. By default all 45 are
132            downloaded.
133        download: Whether to download the data if it is not present.
134
135    Returns:
136        Filepath where the extracted data is stored.
137    """
138    patient_dir = os.path.join(path, "patients")
139    patient_ids = sorted(CHECKSUMS)[:n_patients]
140    missing = [pid for pid in patient_ids if not os.path.exists(os.path.join(patient_dir, pid))]
141    if not missing:
142        return patient_dir
143
144    if not download:
145        raise RuntimeError(f"Cannot find the data at {path}, but download was set to False.")
146
147    os.makedirs(patient_dir, exist_ok=True)
148    with futures.ThreadPoolExecutor(4) as pool:
149        tasks = [pool.submit(_download_patient, pid, path) for pid in missing]
150        for task in tqdm(futures.as_completed(tasks), total=len(tasks), desc="Download PROTEAS patients"):
151            task.result()
152
153    return patient_dir
154
155
156def get_proteas_paths(
157    path: Union[os.PathLike, str],
158    sequence: Literal["t1", "t1c", "t2", "fla"] = "t1c",
159    n_patients: Optional[int] = None,
160    download: bool = False,
161) -> Tuple[List[str], List[str]]:
162    """Get paths to the PROTEAS data.
163
164    Args:
165        path: Filepath to a folder where the data is downloaded for further processing.
166        sequence: The choice of MRI sequence. One of 't1', 't1c' (contrast-enhanced T1), 't2' or 'fla' (FLAIR).
167        n_patients: The number of patients to use, sorted by patient id. By default all 45 are used.
168        download: Whether to download the data if it is not present.
169
170    Returns:
171        List of filepaths for the image data, one per annotated MRI study.
172        List of filepaths for the label data.
173    """
174    if sequence not in SEQUENCES:
175        raise ValueError(f"'{sequence}' is not a valid sequence. Choose one of {SEQUENCES}.")
176
177    patient_dir = get_proteas_data(path, n_patients, download)
178
179    raw_paths, label_paths = [], []
180    for patient_id in sorted(CHECKSUMS)[:n_patients]:
181        for label_path in natsorted(glob(os.path.join(patient_dir, patient_id, "tumor_segmentation", "*.nii.gz"))):
182            study = os.path.basename(label_path)[:-len(".nii.gz")].split("_tumor_mask_")[1]
183            raw_path = os.path.join(patient_dir, patient_id, "BraTS", study, f"{sequence}.nii.gz")
184            if os.path.exists(raw_path):
185                raw_paths.append(raw_path)
186                label_paths.append(label_path)
187
188    assert len(raw_paths) == len(label_paths) and len(raw_paths) > 0
189
190    return raw_paths, label_paths
191
192
193def get_proteas_dataset(
194    path: Union[os.PathLike, str],
195    patch_shape: Tuple[int, int, int],
196    sequence: Literal["t1", "t1c", "t2", "fla"] = "t1c",
197    n_patients: Optional[int] = None,
198    resize_inputs: bool = False,
199    download: bool = False,
200    **kwargs
201) -> Dataset:
202    """Get the PROTEAS dataset for brain metastasis segmentation in longitudinal MRI.
203
204    Args:
205        path: Filepath to a folder where the data is downloaded for further processing.
206        patch_shape: The patch shape to use for training.
207        sequence: The choice of MRI sequence. One of 't1', 't1c' (contrast-enhanced T1), 't2' or 'fla' (FLAIR).
208        n_patients: The number of patients to use, sorted by patient id. By default all 45 are used.
209        resize_inputs: Whether to resize the inputs to the patch shape.
210        download: Whether to download the data if it is not present.
211        kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset`.
212
213    Returns:
214        The segmentation dataset.
215    """
216    raw_paths, label_paths = get_proteas_paths(path, sequence, n_patients, download)
217
218    if resize_inputs:
219        resize_kwargs = {"patch_shape": patch_shape, "is_rgb": False}
220        kwargs, patch_shape = util.update_kwargs_for_resize_trafo(
221            kwargs=kwargs, patch_shape=patch_shape, resize_inputs=resize_inputs, resize_kwargs=resize_kwargs
222        )
223
224    return torch_em.default_segmentation_dataset(
225        raw_paths=raw_paths,
226        raw_key="data",
227        label_paths=label_paths,
228        label_key="data",
229        is_seg_dataset=True,
230        patch_shape=patch_shape,
231        ndim=3,
232        **kwargs
233    )
234
235
236def get_proteas_loader(
237    path: Union[os.PathLike, str],
238    batch_size: int,
239    patch_shape: Tuple[int, int, int],
240    sequence: Literal["t1", "t1c", "t2", "fla"] = "t1c",
241    n_patients: Optional[int] = None,
242    resize_inputs: bool = False,
243    download: bool = False,
244    **kwargs
245) -> DataLoader:
246    """Get the PROTEAS dataloader for brain metastasis segmentation in longitudinal MRI.
247
248    Args:
249        path: Filepath to a folder where the data is downloaded for further processing.
250        batch_size: The batch size for training.
251        patch_shape: The patch shape to use for training.
252        sequence: The choice of MRI sequence. One of 't1', 't1c' (contrast-enhanced T1), 't2' or 'fla' (FLAIR).
253        n_patients: The number of patients to use, sorted by patient id. By default all 45 are used.
254        resize_inputs: Whether to resize the inputs to the patch shape.
255        download: Whether to download the data if it is not present.
256        kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset` or for the PyTorch DataLoader.
257
258    Returns:
259        The DataLoader.
260    """
261    ds_kwargs, loader_kwargs = util.split_kwargs(torch_em.default_segmentation_dataset, **kwargs)
262    dataset = get_proteas_dataset(
263        path, patch_shape, sequence, n_patients, resize_inputs, download, **ds_kwargs
264    )
265    return torch_em.get_data_loader(dataset, batch_size, **loader_kwargs)
URL = 'https://zenodo.org/api/records/17253793/files/{patient_id}.zip/content'
CHECKSUMS = {'P01': '9355b7f50d3e690737c7a5e4d22046d8df206212d4b09fa31ec1ad9f469ba54b', 'P02': '0b4978e3036e1e5c3c8f0f55168c74b612152686ce3d7f8681e933b89e7098d8', 'P03': '6f0a8e9d66016c06ec7e186ad3dac72f620c56a74345de66c75b70f4d3a73f9a', 'P04a': 'ec12923914188ef81eed02a7f6f4aed1ae839a40b1d6e63a47af90ebff540487', 'P04b': '83c13be378f3efe79388bf3ad909f6426b6793b54ea3333c6a57955934218a35', 'P05': '282d3d8f2ce3fb8a79cffa69ae6e28f414770cbf2e404211597474f58e11130a', 'P06': '0a5d2f7f9998454abcf9143ec9e29309cf97d68946c8b41eaad6e7c90fdefa14', 'P07a': '9698954c9f277036f710affa2590cbfb5f4aff0a17041307bd4411e9d9afe36e', 'P07b': '9ea5e9ffcfe4ff80246dca983d3ea5da3309cc96c20c51b0ac00f7b500854d35', 'P08': '6e6f559840766dd10af8ce6f20a94747f20319537ee3f568a68eba578e093e7b', 'P09': '14bd3371ddb19f4fb486f469f27b0db42f78d19ac4c975abf2b30b0859ef35d0', 'P10': 'b5dd8535efa50a6102e1387338750273cec39ac8beee3b38bdffd7edc5c605a9', 'P11': '2b2ab945002116de468dd1caf50914e272c0ee473b6708b37d497d548dffd2e0', 'P12': 'c08da88ed4b56e64d3db161b0b318d0e1b98baa6c28266fd9b1181979337a2e1', 'P13': 'ab0842a95e31e69ca0c64f4cfbbc495af481a3f8ab798a5cd3f6b49c01dae2c8', 'P14': 'ed2d56706456a2aca8e26cee365349326126d0271efa97141a16e86ea95afeaa', 'P15': 'b4abed11f5b9a9eedf7df659afe0e4ca4b57f135454963af71e610e7e6a19469', 'P16': 'beb26fb1c01f77c80f22eee591c73c5657b2cc3426d19d3bd4e6cceec525be79', 'P17a': '1a776b528c60a7a49924c7bf1c59eedc93c37e6e0083c0d8f16ddb2932daf0f4', 'P17b': '13c396baa500ee3b2c82b72c89e15f9c044850a02c7aa3615f2cced9ab6b39b7', 'P18': 'f064cfbb7303d5c8ca6597b60ca9550152c8c6c85de7bf900269d2708dfa930d', 'P19': '56902860f712f4bce08e20c642217d4f18c45f3297a6d8376284f121a6c2cf18', 'P20a': '24d04ba76de099dad365951f74788f436f086e87bafcbdd4704c884fc5d82a97', 'P20b': '0072f1d4158727aa4088b3fb32294512a78aa7772da0f7a98063ee051e7660ba', 'P21': '124192750c623382fc4911a112e9b372e90952e75a4e9da3fa602e467bccb3d3', 'P22': '1011e6b69a5f864ba43032ded01ff6bbde6a02fc72354d2c63f7307c96a890af', 'P23a': '818425160a1514fb2d3f6b5eb86b439cb4cc4f56b93c1e2c3b5690fb53e2ab99', 'P23b': '2241d9c24d03cd69583e709bf498ab6abc86e791664048cd59d3de8453dd98dc', 'P24': 'b660ea009c3604cf5f4891ce6abb7fabf25f428f34a515353df55aa67b0944b5', 'P25': 'ebfd97a5efdb798f7c9a53c0536d25125d3fc14d85d90b7ef2b7d42949ec1503', 'P26': 'b526e4a2ec8f4a77097ccf925667ee6a4a527a6d213722d9f853f19a4546b31a', 'P27': '76017fe710787309d4aa0ceb6a590c58fc5130c0e9e08a63119bf4b9b65ec5d7', 'P28': '6e0c811e7e6da46296b9dc6e73050e65d8abada81f1900713c3fddb728a5b989', 'P29': '6432a22a555c15d0086b2e7ada50a14f1927af1009fd6ec4fa7bf05d5efe4f18', 'P30': '90c8beef41f82004181d1b7cd7db390f81723e3f57040b5ffa1f55205eb5734b', 'P31': '553a410ebe4bc0d159b93f048e38021c22cd7ee4b506d9aa72ffb0241a463628', 'P32': '05192b28f86022c40b82b1d82192428bd8387fe3378436bb0cc4333cc062e824', 'P33': '64ee5d2d83fa2aebc4838603b8ada018d2082af71dd7e1e680585344377f1e63', 'P34': '60063b57524b0978942b28eec70ac8cf28af053623454e1c0fdf446168663f5d', 'P35': '73a23c05f6a1700384eb4d7b9ccf96ff745f6e7f1627e27a678d859521006dd0', 'P36': 'f4f897589a16831fe2ff2dd5d3185d39a98ce23313e03837d9399297e0a89ace', 'P37': 'd645a3910d394d82620588a6249fa6a9ece61a0f06f042acb9c69e7e296bb947', 'P38': '5c8c12811462024ad51ccba928bc8ce405fb397632e4b675dc072f3c4a5dbe06', 'P39': 'bb3c06bcda4125597c353eaba8949bce420efcea831bb91cc047af14eb5b1296', 'P40': '6a09cb9a8bbdd383760efa490255f18d0f43c348630bffc59dbf5d3ea985a569'}
SEQUENCES = ['t1', 't1c', 't2', 'fla']
LABEL_IDS = {'necrotic_core': 1, 'enhancing_tumor': 2, 'edema': 3}
def get_proteas_data( path: Union[os.PathLike, str], n_patients: Optional[int] = None, download: bool = False) -> str:
122def get_proteas_data(
123    path: Union[os.PathLike, str], n_patients: Optional[int] = None, download: bool = False,
124) -> str:
125    """Download the PROTEAS dataset.
126
127    NOTE: The archives of all patients are about 15 GB, as they contain the DICOM series next to the NIfTI volumes
128    used here. Use `n_patients` to only download a subset for a quick start.
129
130    Args:
131        path: Filepath to a folder where the data is downloaded for further processing.
132        n_patients: The number of patients (archives) to download, sorted by patient id. By default all 45 are
133            downloaded.
134        download: Whether to download the data if it is not present.
135
136    Returns:
137        Filepath where the extracted data is stored.
138    """
139    patient_dir = os.path.join(path, "patients")
140    patient_ids = sorted(CHECKSUMS)[:n_patients]
141    missing = [pid for pid in patient_ids if not os.path.exists(os.path.join(patient_dir, pid))]
142    if not missing:
143        return patient_dir
144
145    if not download:
146        raise RuntimeError(f"Cannot find the data at {path}, but download was set to False.")
147
148    os.makedirs(patient_dir, exist_ok=True)
149    with futures.ThreadPoolExecutor(4) as pool:
150        tasks = [pool.submit(_download_patient, pid, path) for pid in missing]
151        for task in tqdm(futures.as_completed(tasks), total=len(tasks), desc="Download PROTEAS patients"):
152            task.result()
153
154    return patient_dir

Download the PROTEAS dataset.

NOTE: The archives of all patients are about 15 GB, as they contain the DICOM series next to the NIfTI volumes used here. Use n_patients to only download a subset for a quick start.

Arguments:
  • path: Filepath to a folder where the data is downloaded for further processing.
  • n_patients: The number of patients (archives) to download, sorted by patient id. By default all 45 are downloaded.
  • download: Whether to download the data if it is not present.
Returns:

Filepath where the extracted data is stored.

def get_proteas_paths( path: Union[os.PathLike, str], sequence: Literal['t1', 't1c', 't2', 'fla'] = 't1c', n_patients: Optional[int] = None, download: bool = False) -> Tuple[List[str], List[str]]:
157def get_proteas_paths(
158    path: Union[os.PathLike, str],
159    sequence: Literal["t1", "t1c", "t2", "fla"] = "t1c",
160    n_patients: Optional[int] = None,
161    download: bool = False,
162) -> Tuple[List[str], List[str]]:
163    """Get paths to the PROTEAS data.
164
165    Args:
166        path: Filepath to a folder where the data is downloaded for further processing.
167        sequence: The choice of MRI sequence. One of 't1', 't1c' (contrast-enhanced T1), 't2' or 'fla' (FLAIR).
168        n_patients: The number of patients to use, sorted by patient id. By default all 45 are used.
169        download: Whether to download the data if it is not present.
170
171    Returns:
172        List of filepaths for the image data, one per annotated MRI study.
173        List of filepaths for the label data.
174    """
175    if sequence not in SEQUENCES:
176        raise ValueError(f"'{sequence}' is not a valid sequence. Choose one of {SEQUENCES}.")
177
178    patient_dir = get_proteas_data(path, n_patients, download)
179
180    raw_paths, label_paths = [], []
181    for patient_id in sorted(CHECKSUMS)[:n_patients]:
182        for label_path in natsorted(glob(os.path.join(patient_dir, patient_id, "tumor_segmentation", "*.nii.gz"))):
183            study = os.path.basename(label_path)[:-len(".nii.gz")].split("_tumor_mask_")[1]
184            raw_path = os.path.join(patient_dir, patient_id, "BraTS", study, f"{sequence}.nii.gz")
185            if os.path.exists(raw_path):
186                raw_paths.append(raw_path)
187                label_paths.append(label_path)
188
189    assert len(raw_paths) == len(label_paths) and len(raw_paths) > 0
190
191    return raw_paths, label_paths

Get paths to the PROTEAS data.

Arguments:
  • path: Filepath to a folder where the data is downloaded for further processing.
  • sequence: The choice of MRI sequence. One of 't1', 't1c' (contrast-enhanced T1), 't2' or 'fla' (FLAIR).
  • n_patients: The number of patients to use, sorted by patient id. By default all 45 are used.
  • download: Whether to download the data if it is not present.
Returns:

List of filepaths for the image data, one per annotated MRI study. List of filepaths for the label data.

def get_proteas_dataset( path: Union[os.PathLike, str], patch_shape: Tuple[int, int, int], sequence: Literal['t1', 't1c', 't2', 'fla'] = 't1c', n_patients: Optional[int] = None, resize_inputs: bool = False, download: bool = False, **kwargs) -> torch.utils.data.dataset.Dataset:
194def get_proteas_dataset(
195    path: Union[os.PathLike, str],
196    patch_shape: Tuple[int, int, int],
197    sequence: Literal["t1", "t1c", "t2", "fla"] = "t1c",
198    n_patients: Optional[int] = None,
199    resize_inputs: bool = False,
200    download: bool = False,
201    **kwargs
202) -> Dataset:
203    """Get the PROTEAS dataset for brain metastasis segmentation in longitudinal MRI.
204
205    Args:
206        path: Filepath to a folder where the data is downloaded for further processing.
207        patch_shape: The patch shape to use for training.
208        sequence: The choice of MRI sequence. One of 't1', 't1c' (contrast-enhanced T1), 't2' or 'fla' (FLAIR).
209        n_patients: The number of patients to use, sorted by patient id. By default all 45 are used.
210        resize_inputs: Whether to resize the inputs to the patch shape.
211        download: Whether to download the data if it is not present.
212        kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset`.
213
214    Returns:
215        The segmentation dataset.
216    """
217    raw_paths, label_paths = get_proteas_paths(path, sequence, n_patients, download)
218
219    if resize_inputs:
220        resize_kwargs = {"patch_shape": patch_shape, "is_rgb": False}
221        kwargs, patch_shape = util.update_kwargs_for_resize_trafo(
222            kwargs=kwargs, patch_shape=patch_shape, resize_inputs=resize_inputs, resize_kwargs=resize_kwargs
223        )
224
225    return torch_em.default_segmentation_dataset(
226        raw_paths=raw_paths,
227        raw_key="data",
228        label_paths=label_paths,
229        label_key="data",
230        is_seg_dataset=True,
231        patch_shape=patch_shape,
232        ndim=3,
233        **kwargs
234    )

Get the PROTEAS dataset for brain metastasis segmentation in longitudinal MRI.

Arguments:
  • path: Filepath to a folder where the data is downloaded for further processing.
  • patch_shape: The patch shape to use for training.
  • sequence: The choice of MRI sequence. One of 't1', 't1c' (contrast-enhanced T1), 't2' or 'fla' (FLAIR).
  • n_patients: The number of patients to use, sorted by patient id. By default all 45 are used.
  • resize_inputs: Whether to resize the inputs to the patch shape.
  • download: Whether to download the data if it is not present.
  • kwargs: Additional keyword arguments for torch_em.default_segmentation_dataset.
Returns:

The segmentation dataset.

def get_proteas_loader( path: Union[os.PathLike, str], batch_size: int, patch_shape: Tuple[int, int, int], sequence: Literal['t1', 't1c', 't2', 'fla'] = 't1c', n_patients: Optional[int] = None, resize_inputs: bool = False, download: bool = False, **kwargs) -> torch.utils.data.dataloader.DataLoader:
237def get_proteas_loader(
238    path: Union[os.PathLike, str],
239    batch_size: int,
240    patch_shape: Tuple[int, int, int],
241    sequence: Literal["t1", "t1c", "t2", "fla"] = "t1c",
242    n_patients: Optional[int] = None,
243    resize_inputs: bool = False,
244    download: bool = False,
245    **kwargs
246) -> DataLoader:
247    """Get the PROTEAS dataloader for brain metastasis segmentation in longitudinal MRI.
248
249    Args:
250        path: Filepath to a folder where the data is downloaded for further processing.
251        batch_size: The batch size for training.
252        patch_shape: The patch shape to use for training.
253        sequence: The choice of MRI sequence. One of 't1', 't1c' (contrast-enhanced T1), 't2' or 'fla' (FLAIR).
254        n_patients: The number of patients to use, sorted by patient id. By default all 45 are used.
255        resize_inputs: Whether to resize the inputs to the patch shape.
256        download: Whether to download the data if it is not present.
257        kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset` or for the PyTorch DataLoader.
258
259    Returns:
260        The DataLoader.
261    """
262    ds_kwargs, loader_kwargs = util.split_kwargs(torch_em.default_segmentation_dataset, **kwargs)
263    dataset = get_proteas_dataset(
264        path, patch_shape, sequence, n_patients, resize_inputs, download, **ds_kwargs
265    )
266    return torch_em.get_data_loader(dataset, batch_size, **loader_kwargs)

Get the PROTEAS dataloader for brain metastasis segmentation in longitudinal MRI.

Arguments:
  • path: Filepath to a folder where the data is downloaded for further processing.
  • batch_size: The batch size for training.
  • patch_shape: The patch shape to use for training.
  • sequence: The choice of MRI sequence. One of 't1', 't1c' (contrast-enhanced T1), 't2' or 'fla' (FLAIR).
  • n_patients: The number of patients to use, sorted by patient id. By default all 45 are used.
  • resize_inputs: Whether to resize the inputs to the patch shape.
  • download: Whether to download the data if it is not present.
  • kwargs: Additional keyword arguments for torch_em.default_segmentation_dataset or for the PyTorch DataLoader.
Returns:

The DataLoader.