torch_em.data.datasets.medical.acrin_6698
The ACRIN-6698 dataset contains annotations for functional tumor volume segmentation in dynamic contrast-enhanced (DCE) breast MRI, acquired as part of the ACRIN 6698 / I-SPY2 TRIAL.
The collection ships 1110 "VOLSER" (volume of signal enhancement ratio) DICOM-SEG objects: a binary mask
of the functional tumor volume, obtained by thresholding a per-voxel signal enhancement ratio (SER) map
computed from the pre- and post-contrast DCE-MRI phases. The mask is stored either "uni-lateral cropped"
(covering one breast) or "bi-lateral" (covering both), cropped in-plane to the analyzed region. The
collection also ships a second, unrelated family of 1103 "DWI SEG" objects (whole-tumor masks drawn on
diffusion-weighted imaging derivative series); those are covered separately by medical.acrin_6698_dwi.
Unlike most DICOM-SEG collections handled elsewhere in this package, these SEG objects carry neither a
ReferencedSeriesSequence nor a per-frame DerivationImageSequence: only a StudyInstanceUID and a
FrameOfReferenceUID. Matching a mask to its source series therefore requires querying the TCIA REST API
for all series of the mask's study and picking the DCE-MRI series with matching geometry, rather than the
usual reference-UID lookup used in eg. nsclc_radiogenomics.py or adrenal_acc.py.
Inspecting real downloaded data (a random sample of 100 of the 1110 VOLSER masks, across both uni-lateral
and bi-lateral cases) shows that every study contains exactly one MR series whose SeriesDescription is the
mask's own description with "Analysis Mask" replaced by "SER" (eg. "ISPY2: VOLSER: uni-lateral cropped:
Analysis Mask" pairs with "ISPY2: VOLSER: uni-lateral cropped: SER"). This SER series is the actual signal
enhancement ratio map the mask was thresholded from: it shares the mask's FrameOfReferenceUID, in-plane
PixelSpacing, ImageOrientationPatient and - since it is cropped to the same region - the exact same set
of per-slice ImagePositionPatient values, unlike the full-field-of-view source DCE-MRI acquisition (eg.
"ISPY2: Ax Vibrant PRE/POST"), which shares the frame of reference and in-plane geometry but not the
cropped extent. This module therefore uses the description-matched SER series as the raw image paired with
each mask. As a safety net (in case the description-based match is ever wrong or absent for some study),
the matched candidate's geometry is verified by resampling the mask onto the candidate's voxel grid (see
_resample_labels in adrenal_acc.py) and checking that (nearly) all mask foreground voxels land inside
it; if this fails, every other MR series of the study with the same FrameOfReferenceUID is tried in turn,
and the case is skipped (with a warning) if none of them pass the geometry check either. In the validated
sample of 100 masks, the description-based match succeeded and passed geometry verification in all 100
cases (0% failure rate); a larger production run may still see occasional skips for studies with unusual or
missing SER series, which are handled gracefully rather than raised as errors.
Two further data quirks were found by inspecting the actual downloaded pixel data of 5 validated cases and had to be handled explicitly:
- The SEG objects use
SegmentationType = FRACTIONALwithSegmentationFractionalType = OCCUPANCY, notBINARY, even thoughSegmentSequencelists only a single segment, and itsContent Descriptionreads "Bit-map mask of SER tumor segmentation". This is NOT a 0-255 fractional occupancy value: it is an inverse, bit-encoded exclusion mask, as documented in the "Analysis masks from FTV processing" support document linked from the TCIA collection page (https://www.cancerimagingarchive.net/collection/acrin-6698/, filename 'Analysis-mask-files-description.v20211020.docx', by the I-SPY imaging group, UCSF - fetched and parsed directly, not guessed). Per that document: "The masks are INVERSE masks, in that a mask value of 0 indicates that a voxel was included in the measured functional tumor volume (FTV)". Each of up to 5 masking/exclusion steps that removed a voxel from the FTV sets one bit of the byte value: 1 = failed the PE (percent enhancement) threshold, 2 = failed the 3D minimum-neighbor-count (MNC) connectivity filter, 32 = outside the manually-drawn rectangular VOI, 64 = manually OMITted by a trained observer, and a background-intensity-threshold exclusion bit (documented as 8, but empirically 16 in the ACRIN-6698 SEG files actually inspected - see below). A voxel excluded by several steps has the bitwise OR of their values; only{0, 1, 2, 16, 17, 32, 33, 34, 48, 49, ...}-style combinations are therefore valid, and this matches byte-for-byte what is observed in the 5 validated cases (exactly{0, 1, 2, 17, 32, 33, 34, 49}, ie. every combination of {PE, background, VOI} that occurred in the sample, decomposing cleanly as 17 = 16 (background) + 1 (PE), 33 = 32 (VOI) + 1 (PE), 34 = 32 (VOI) + 2 (MNC), 49 = 32 + 16 + 1). This also explains the sharp-edged rectangular sub-regions seen in an earlier (incorrect) version of this module that treated every non-background-mode voxel as foreground: those rectangles are the boundary of the manually-drawn VOI itself (voxels just inside vs. just outside it differ only in whether the 32 bit is set), not a real anatomical structure. This module now instead uses the documented semantics directly: foreground (the FTV) is exactly the voxels with value 0. Verified for the 5 validated cases: the resulting mask is a small, spatially compact, contiguous 3D blob (or, in 1/5 cases, empty - a legitimate "no residual enhancing tumor" result, since ACRIN-6698 includes post-treatment timepoints) whose per-slice area traces a smooth, single-peaked profile, and which spatially coincides exactly with the locations of elevated signal in the matched SER map (see the second quirk below andcheck_acrin_6698.py) - unlike the rectangle-contaminated mask from treating "not background" as foreground. - The matched SER (and other VOLSER-derived) MR series store their quantitative ratio values with a very
small
RescaleSlope(0.001 in the validated cases, ie. the raw stored int16 pixel value is the SER ratio times 1000). Applying the slope/intercept and rounding to an integer volume, asadrenal_acc.pydoes for CT Hounsfield units, collapses nearly the entire dynamic range to {0, 1, 2, 3} and destroys the image. This module therefore keeps the raw stored pixel values as-is for the source volume, without applying the rescale slope/intercept. Separately, and not a bug: the SER maps themselves are extremely sparse (only 0.02-1.9% of voxels nonzero in the validated cases) - by design, per the same support document, the background/non-enhancing/non-tissue voxels are zeroed out by the upstream processing pipeline before this derived series is exported, so most of the cropped volume being exactly 0 is a genuine property of the data, not a display, windowing or download issue.
Because the collection is large (1110 relevant SEG series across 385 patients), use max_cases or
case_ids to only download and preprocess a subset.
NOTE: This requires the pydicom python package.
The dataset is located at https://www.cancerimagingarchive.net/collection/acrin-6698/ and is distributed under the CC BY 4.0 license.
This dataset is from the publication https://doi.org/10.1148/radiol.2018180273. The data was released at https://doi.org/10.7937/tcia.kk02-6d95. Please cite it if you use this dataset in your research.
1"""The ACRIN-6698 dataset contains annotations for functional tumor volume segmentation in dynamic 2contrast-enhanced (DCE) breast MRI, acquired as part of the ACRIN 6698 / I-SPY2 TRIAL. 3 4The collection ships 1110 "VOLSER" (volume of signal enhancement ratio) DICOM-SEG objects: a binary mask 5of the functional tumor volume, obtained by thresholding a per-voxel signal enhancement ratio (SER) map 6computed from the pre- and post-contrast DCE-MRI phases. The mask is stored either "uni-lateral cropped" 7(covering one breast) or "bi-lateral" (covering both), cropped in-plane to the analyzed region. The 8collection also ships a second, unrelated family of 1103 "DWI SEG" objects (whole-tumor masks drawn on 9diffusion-weighted imaging derivative series); those are covered separately by `medical.acrin_6698_dwi`. 10 11Unlike most DICOM-SEG collections handled elsewhere in this package, these SEG objects carry neither a 12`ReferencedSeriesSequence` nor a per-frame `DerivationImageSequence`: only a `StudyInstanceUID` and a 13`FrameOfReferenceUID`. Matching a mask to its source series therefore requires querying the TCIA REST API 14for all series of the mask's study and picking the DCE-MRI series with matching geometry, rather than the 15usual reference-UID lookup used in eg. `nsclc_radiogenomics.py` or `adrenal_acc.py`. 16 17Inspecting real downloaded data (a random sample of 100 of the 1110 VOLSER masks, across both uni-lateral 18and bi-lateral cases) shows that every study contains exactly one MR series whose `SeriesDescription` is the 19mask's own description with "Analysis Mask" replaced by "SER" (eg. "ISPY2: VOLSER: uni-lateral cropped: 20Analysis Mask" pairs with "ISPY2: VOLSER: uni-lateral cropped: SER"). This SER series is the actual signal 21enhancement ratio map the mask was thresholded from: it shares the mask's `FrameOfReferenceUID`, in-plane 22`PixelSpacing`, `ImageOrientationPatient` and - since it is cropped to the same region - the exact same set 23of per-slice `ImagePositionPatient` values, unlike the full-field-of-view source DCE-MRI acquisition (eg. 24"ISPY2: Ax Vibrant PRE/POST"), which shares the frame of reference and in-plane geometry but not the 25cropped extent. This module therefore uses the description-matched SER series as the raw image paired with 26each mask. As a safety net (in case the description-based match is ever wrong or absent for some study), 27the matched candidate's geometry is verified by resampling the mask onto the candidate's voxel grid (see 28`_resample_labels` in `adrenal_acc.py`) and checking that (nearly) all mask foreground voxels land inside 29it; if this fails, every other MR series of the study with the same `FrameOfReferenceUID` is tried in turn, 30and the case is skipped (with a warning) if none of them pass the geometry check either. In the validated 31sample of 100 masks, the description-based match succeeded and passed geometry verification in all 100 32cases (0% failure rate); a larger production run may still see occasional skips for studies with unusual or 33missing SER series, which are handled gracefully rather than raised as errors. 34 35Two further data quirks were found by inspecting the actual downloaded pixel data of 5 validated cases and 36had to be handled explicitly: 37- The SEG objects use `SegmentationType = FRACTIONAL` with `SegmentationFractionalType = OCCUPANCY`, not 38 `BINARY`, even though `SegmentSequence` lists only a single segment, and its `Content Description` reads 39 "Bit-map mask of SER tumor segmentation". This is NOT a 0-255 fractional occupancy value: it is an 40 *inverse*, bit-encoded exclusion mask, as documented in the "Analysis masks from FTV processing" support 41 document linked from the TCIA collection page (https://www.cancerimagingarchive.net/collection/acrin-6698/, 42 filename 'Analysis-mask-files-description.v20211020.docx', by the I-SPY imaging group, UCSF - fetched and 43 parsed directly, not guessed). Per that document: "The masks are INVERSE masks, in that a mask value of 0 44 indicates that a voxel was included in the measured functional tumor volume (FTV)". Each of up to 5 45 masking/exclusion steps that removed a voxel from the FTV sets one bit of the byte value: 1 = failed the 46 PE (percent enhancement) threshold, 2 = failed the 3D minimum-neighbor-count (MNC) connectivity filter, 47 32 = outside the manually-drawn rectangular VOI, 64 = manually OMITted by a trained observer, and a 48 background-intensity-threshold exclusion bit (documented as 8, but empirically 16 in the ACRIN-6698 SEG 49 files actually inspected - see below). A voxel excluded by several steps has the bitwise OR of their 50 values; only `{0, 1, 2, 16, 17, 32, 33, 34, 48, 49, ...}`-style combinations are therefore valid, and this 51 matches byte-for-byte what is observed in the 5 validated cases (exactly `{0, 1, 2, 17, 32, 33, 34, 49}`, 52 ie. every combination of {PE, background, VOI} that occurred in the sample, decomposing cleanly as 53 17 = 16 (background) + 1 (PE), 33 = 32 (VOI) + 1 (PE), 34 = 32 (VOI) + 2 (MNC), 49 = 32 + 16 + 1). This 54 also explains the sharp-edged rectangular sub-regions seen in an earlier (incorrect) version of this 55 module that treated every non-background-mode voxel as foreground: those rectangles are the boundary of 56 the manually-drawn VOI itself (voxels just inside vs. just outside it differ only in whether the 32 bit is 57 set), not a real anatomical structure. This module now instead uses the documented semantics directly: 58 foreground (the FTV) is exactly the voxels with value 0. Verified for the 5 validated cases: the resulting 59 mask is a small, spatially compact, contiguous 3D blob (or, in 1/5 cases, empty - a legitimate "no residual 60 enhancing tumor" result, since ACRIN-6698 includes post-treatment timepoints) whose per-slice area traces 61 a smooth, single-peaked profile, and which spatially coincides exactly with the locations of elevated 62 signal in the matched SER map (see the second quirk below and `check_acrin_6698.py`) - unlike the 63 rectangle-contaminated mask from treating "not background" as foreground. 64- The matched SER (and other VOLSER-derived) MR series store their quantitative ratio values with a very 65 small `RescaleSlope` (0.001 in the validated cases, ie. the raw stored int16 pixel value is the SER ratio 66 times 1000). Applying the slope/intercept and rounding to an integer volume, as `adrenal_acc.py` does for 67 CT Hounsfield units, collapses nearly the entire dynamic range to {0, 1, 2, 3} and destroys the image. This 68 module therefore keeps the raw stored pixel values as-is for the source volume, without applying the 69 rescale slope/intercept. Separately, and not a bug: the SER maps themselves are extremely sparse (only 70 0.02-1.9% of voxels nonzero in the validated cases) - by design, per the same support document, the 71 background/non-enhancing/non-tissue voxels are zeroed out by the upstream processing pipeline before this 72 derived series is exported, so most of the cropped volume being exactly 0 is a genuine property of the 73 data, not a display, windowing or download issue. 74 75Because the collection is large (1110 relevant SEG series across 385 patients), use `max_cases` or 76`case_ids` to only download and preprocess a subset. 77 78NOTE: This requires the pydicom python package. 79 80The dataset is located at https://www.cancerimagingarchive.net/collection/acrin-6698/ and is distributed 81under the CC BY 4.0 license. 82 83This dataset is from the publication https://doi.org/10.1148/radiol.2018180273. 84The data was released at https://doi.org/10.7937/tcia.kk02-6d95. 85Please cite it if you use this dataset in your research. 86""" 87 88import os 89from glob import glob 90from tqdm import tqdm 91from warnings import warn 92from natsort import natsorted 93from typing import Union, Tuple, List, Optional 94 95import numpy as np 96import requests 97 98from torch.utils.data import Dataset, DataLoader 99 100import torch_em 101 102from .adrenal_acc import _resample_labels 103from .. import util 104 105 106COLLECTION = "ACRIN-6698" 107 108LABEL_IDS = {"functional_tumor_volume": 1} 109 110# The minimum fraction of the mask's foreground voxels that must land inside a candidate source series 111# (after resampling onto its voxel grid) for that candidate to be accepted as the match. 112MIN_GEOMETRY_OVERLAP = 0.9 113 114 115def _load_dicom_volume(series_dir): 116 """Stack a DICOM series into a volume with axes (z, y, x) and slices sorted along the slice normal. 117 118 Unlike `adrenal_acc._load_dicom_volume`, the rescale slope/intercept is NOT applied here: see the 119 module docstring for why this matters for the ACRIN-6698 VOLSER-derived MR series (SER, PE2, PE6, ...). 120 121 Returns the volume and the affine matrix that maps voxel indices (z, y, x) to DICOM patient coordinates. 122 """ 123 import pydicom 124 125 slices = [pydicom.dcmread(dcm_path) for dcm_path in natsorted(glob(os.path.join(series_dir, "*.dcm")))] 126 orientation = np.array([float(v) for v in slices[0].ImageOrientationPatient]) 127 row_dir, col_dir = orientation[:3], orientation[3:] 128 normal = np.cross(row_dir, col_dir) 129 slices.sort(key=lambda dcm: np.dot([float(v) for v in dcm.ImagePositionPatient], normal)) 130 131 volume = np.stack([dcm.pixel_array for dcm in slices]) 132 133 positions = np.array([[float(v) for v in dcm.ImagePositionPatient] for dcm in slices]) 134 spacing = [float(v) for v in slices[0].PixelSpacing] # The spacing between rows and between columns. 135 affine = np.eye(4) 136 affine[:3, 0] = (positions[-1] - positions[0]) / (len(slices) - 1) 137 affine[:3, 1] = col_dir * spacing[0] 138 affine[:3, 2] = row_dir * spacing[1] 139 affine[:3, 3] = positions[0] 140 return volume, affine 141 142 143def _load_dicom_seg(seg_path): 144 """Load a VOLSER analysis mask DICOM-SEG object as a label volume with axes (z, y, x). 145 146 Unlike `adrenal_acc._load_dicom_seg`, foreground is not determined by a nonzero pixel value: the 147 ACRIN-6698 VOLSER masks are documented (see the module docstring) as INVERSE, bit-encoded exclusion 148 masks, where a voxel value of exactly 0 means it was included in the functional tumor volume (FTV) and 149 any nonzero value means it was excluded by one or more processing steps. 150 151 Returns the label volume and the affine matrix that maps its voxel indices to DICOM patient coordinates. 152 """ 153 import pydicom 154 155 seg = pydicom.dcmread(seg_path) 156 frames = seg.pixel_array 157 if frames.ndim == 2: # A segmentation with a single frame. 158 frames = frames[None] 159 160 shared_group = seg.SharedFunctionalGroupsSequence[0] 161 orientation = np.array([float(v) for v in shared_group.PlaneOrientationSequence[0].ImageOrientationPatient]) 162 row_dir, col_dir = orientation[:3], orientation[3:] 163 normal = np.cross(row_dir, col_dir) 164 pixel_measures = shared_group.PixelMeasuresSequence[0] 165 spacing = [float(v) for v in pixel_measures.PixelSpacing] # The spacing between rows and between columns. 166 167 frame_groups = seg.PerFrameFunctionalGroupsSequence 168 positions = np.array([[float(v) for v in g.PlanePositionSequence[0].ImagePositionPatient] for g in frame_groups]) 169 projections = positions @ normal 170 if "SpacingBetweenSlices" in pixel_measures: 171 slice_spacing = float(pixel_measures.SpacingBetweenSlices) 172 elif len(projections) > 1: 173 slice_spacing = np.min(np.diff(np.unique(np.round(projections, 3)))) 174 else: 175 slice_spacing = float(pixel_measures.SliceThickness) 176 slice_ids = np.round((projections - projections.min()) / slice_spacing).astype("int") 177 178 # A value of 0 means the voxel was included in the FTV, see the module docstring; any nonzero value 179 # means it was excluded (the bitwise OR of one or more masking-step codes). 180 segment_number = int(seg.SegmentSequence[0].SegmentNumber) 181 182 labels = np.zeros((slice_ids.max() + 1, seg.Rows, seg.Columns), dtype="uint8") 183 for frame, slice_id in zip(frames, slice_ids): 184 labels[slice_id][frame == 0] = segment_number 185 186 affine = np.eye(4) 187 affine[:3, 0] = normal * slice_spacing 188 affine[:3, 1] = col_dir * spacing[0] 189 affine[:3, 2] = row_dir * spacing[1] 190 affine[:3, 3] = positions[np.argmin(projections)] 191 return labels, affine 192 193 194def _get_seg_metadata(path: str, download: bool) -> List[dict]: 195 """Query the TCIA REST API for the metadata of all SEG series of the collection, without downloading images.""" 196 import json 197 198 metadata_path = os.path.join(path, "acrin_6698_seg_series.json") 199 if not os.path.exists(metadata_path): 200 if not download: 201 raise RuntimeError(f"Cannot find the data at {path}, but download was set to False.") 202 response = requests.get(util.NBIA_API_URL + "getSeries", params={"Collection": COLLECTION, "Modality": "SEG"}) 203 response.raise_for_status() 204 with open(metadata_path, "w") as f: 205 json.dump(response.json(), f, indent=2) 206 207 with open(metadata_path, "r") as f: 208 return json.load(f) 209 210 211def _is_volser_mask(series_description: str) -> bool: 212 return "VOLSER" in series_description and "Analysis Mask" in series_description 213 214 215def _get_study_series(study_uid: str, cache: dict) -> List[dict]: 216 """Query the TCIA REST API for the metadata of all series of a study, cached per study.""" 217 if study_uid not in cache: 218 response = requests.get(util.NBIA_API_URL + "getSeries", params={"StudyInstanceUID": study_uid}) 219 response.raise_for_status() 220 cache[study_uid] = response.json() 221 return cache[study_uid] 222 223 224def _match_by_description(seg_description: str, study_series: List[dict]) -> Optional[str]: 225 """Find the SER series whose description matches the mask's description (see the module docstring).""" 226 prefix = seg_description.replace("Analysis Mask", "SER").strip() 227 matches = [ 228 series["SeriesInstanceUID"] for series in study_series 229 if series["Modality"] == "MR" and series["SeriesDescription"].strip() == prefix 230 ] 231 return matches[0] if len(matches) == 1 else None 232 233 234def _candidate_mr_uids(seg_description: str, study_series: List[dict]) -> List[str]: 235 """List the candidate source series of a study, the description match (if any) tried first.""" 236 primary = _match_by_description(seg_description, study_series) 237 others = [ 238 series["SeriesInstanceUID"] for series in study_series 239 if series["Modality"] == "MR" and series["SeriesInstanceUID"] != primary 240 ] 241 return ([primary] if primary is not None else []) + others 242 243 244def _resolve_source_volume(seg_path, seg_labels, seg_affine, frame_of_reference, mr_uids, dicom_dir, download): 245 """Find the source MR series matching a mask's geometry and return its volume, aligned labels and affine. 246 247 Tries the candidate series in order (the description-matched one first, see `_candidate_mr_uids`), 248 downloading each on demand, and accepts the first one whose frame of reference matches and whose 249 resampled mask overlap exceeds `MIN_GEOMETRY_OVERLAP`. A mask can legitimately have no foreground voxels 250 at all (an empty functional tumor volume, eg. a complete response at a later trial timepoint, see the 251 module docstring), in which case the overlap fraction is undefined; the first candidate with a matching 252 frame of reference and pixel grid (`Rows`/`Columns`) is accepted instead, without downloading and trying 253 every other series of the study. Returns (None, None, None) if none match. 254 """ 255 import pydicom 256 257 n_foreground = int((seg_labels > 0).sum()) 258 seg_dcm = pydicom.dcmread(seg_path, stop_before_pixels=True) 259 seg_shape = (int(seg_dcm.Rows), int(seg_dcm.Columns)) 260 261 for mr_uid in mr_uids: 262 mr_dir = os.path.join(dicom_dir, mr_uid) 263 if not glob(os.path.join(mr_dir, "*.dcm")): 264 if not download: 265 continue 266 util.download_tcia_series([mr_uid], dst=dicom_dir, csv_filename=os.path.join(dicom_dir, "..", "acrin_6698_mr_extra")) # noqa 267 dcm_paths = glob(os.path.join(mr_dir, "*.dcm")) 268 if not dcm_paths: 269 continue 270 271 first_dcm = pydicom.dcmread(dcm_paths[0], stop_before_pixels=True) 272 if str(getattr(first_dcm, "FrameOfReferenceUID", None)) != frame_of_reference: 273 continue 274 275 if n_foreground == 0: 276 if (int(first_dcm.Rows), int(first_dcm.Columns)) != seg_shape: 277 continue 278 volume, mr_affine = _load_dicom_volume(mr_dir) 279 labels = np.zeros(volume.shape, dtype=seg_labels.dtype) 280 return volume, labels, mr_affine 281 282 volume, mr_affine = _load_dicom_volume(mr_dir) 283 labels = _resample_labels(seg_labels, seg_affine, volume.shape, mr_affine) 284 overlap = int((labels > 0).sum()) / n_foreground 285 if overlap >= MIN_GEOMETRY_OVERLAP: 286 return volume, labels, mr_affine 287 288 return None, None, None 289 290 291def _preprocess_acrin_6698(seg_series: List[dict], dicom_dir: str, preprocessed_dir: str, download: bool) -> None: 292 import h5py 293 import pydicom 294 295 os.makedirs(preprocessed_dir, exist_ok=True) 296 study_series_cache = {} 297 for series in tqdm(seg_series, desc="Preprocess ACRIN-6698"): 298 seg_uid = series["SeriesInstanceUID"] 299 out_path = os.path.join(preprocessed_dir, f"{seg_uid}.h5") 300 if os.path.exists(out_path): 301 continue 302 303 seg_paths = glob(os.path.join(dicom_dir, seg_uid, "*.dcm")) 304 if not seg_paths: 305 warn(f"Skipping {seg_uid}, whose SEG DICOM data could not be found.") 306 continue 307 308 seg_dcm = pydicom.dcmread(seg_paths[0], stop_before_pixels=True) 309 frame_of_reference = str(seg_dcm.FrameOfReferenceUID) 310 311 study_series = _get_study_series(str(seg_dcm.StudyInstanceUID), study_series_cache) 312 mr_uids = _candidate_mr_uids(series["SeriesDescription"], study_series) 313 if not mr_uids: 314 warn(f"Skipping {seg_uid}, whose study has no MR series to match against.") 315 continue 316 317 seg_labels, seg_affine = _load_dicom_seg(seg_paths[0]) 318 volume, labels, _ = _resolve_source_volume( 319 seg_paths[0], seg_labels, seg_affine, frame_of_reference, mr_uids, dicom_dir, download 320 ) 321 if volume is None: 322 warn(f"Skipping {seg_uid}, for which no MR series with matching geometry could be found.") 323 continue 324 325 labels = (labels > 0).astype("uint8") * LABEL_IDS["functional_tumor_volume"] 326 with h5py.File(out_path, "w") as f: 327 f.create_dataset("raw", data=volume, compression="gzip") 328 f.create_dataset("labels", data=labels, compression="gzip") 329 330 331def get_acrin_6698_data( 332 path: Union[os.PathLike, str], 333 max_cases: Optional[int] = None, 334 case_ids: Optional[List[str]] = None, 335 download: bool = False, 336) -> str: 337 """Download the ACRIN-6698 dataset. 338 339 Args: 340 path: Filepath to a folder where the data is downloaded for further processing. 341 max_cases: The maximum number of VOLSER masks to download and preprocess, taken in deterministic 342 order of their series instance UID. Only the SEG series and their matched source MR series are 343 downloaded. Mutually exclusive with `case_ids`. 344 case_ids: Explicit list of VOLSER mask series instance UIDs to download and preprocess (see the 345 'acrin_6698_seg_series.json' metadata file written to `path` for the available UIDs and their 346 descriptions). Mutually exclusive with `max_cases`. 347 download: Whether to download the data if it is not present. 348 349 Returns: 350 Filepath where the preprocessed data is stored. 351 """ 352 assert max_cases is None or case_ids is None, "'max_cases' and 'case_ids' are mutually exclusive." 353 354 # NOTE: The preprocessing below skips volumes that were converted already, so an interrupted run resumes. 355 preprocessed_dir = os.path.join(path, "preprocessed") 356 os.makedirs(path, exist_ok=True) 357 358 seg_series = _get_seg_metadata(path, download) 359 seg_series = [series for series in seg_series if _is_volser_mask(series["SeriesDescription"])] 360 seg_series = natsorted(seg_series, key=lambda series: series["SeriesInstanceUID"]) 361 362 if max_cases is not None: 363 seg_series = seg_series[:max_cases] 364 elif case_ids is not None: 365 seg_series = [series for series in seg_series if series["SeriesInstanceUID"] in case_ids] 366 367 dicom_dir = os.path.join(path, "dicom") 368 if download: 369 seg_uids = [series["SeriesInstanceUID"] for series in seg_series] 370 util.download_tcia_series(seg_uids, dst=dicom_dir, csv_filename=os.path.join(path, "acrin_6698_seg")) 371 372 # The primary (description-matched) candidate source series of each mask's study is downloaded here 373 # already, so that the preprocessing loop below only has to download extra candidates for the rare 374 # case where the primary candidate's geometry does not match (see the module docstring). 375 study_series_cache = {} 376 primary_mr_uids = set() 377 for series in seg_series: 378 seg_paths = glob(os.path.join(dicom_dir, series["SeriesInstanceUID"], "*.dcm")) 379 if not seg_paths: 380 continue 381 import pydicom 382 study_uid = str(pydicom.dcmread(seg_paths[0], stop_before_pixels=True).StudyInstanceUID) 383 study_series = _get_study_series(study_uid, study_series_cache) 384 primary_uid = _match_by_description(series["SeriesDescription"], study_series) 385 if primary_uid is not None: 386 primary_mr_uids.add(primary_uid) 387 util.download_tcia_series( 388 sorted(primary_mr_uids), dst=dicom_dir, csv_filename=os.path.join(path, "acrin_6698_mr") 389 ) 390 elif not glob(os.path.join(dicom_dir, "*", "*.dcm")): 391 raise RuntimeError(f"Cannot find the data at {path}, but download was set to False.") 392 393 _preprocess_acrin_6698(seg_series, dicom_dir, preprocessed_dir, download) 394 return preprocessed_dir 395 396 397def get_acrin_6698_paths( 398 path: Union[os.PathLike, str], 399 max_cases: Optional[int] = None, 400 case_ids: Optional[List[str]] = None, 401 download: bool = False, 402) -> List[str]: 403 """Get paths to the ACRIN-6698 data. 404 405 Args: 406 path: Filepath to a folder where the data is downloaded for further processing. 407 max_cases: The maximum number of cases to use. See `get_acrin_6698_data` for details. 408 case_ids: Explicit list of case ids to use. See `get_acrin_6698_data` for details. 409 download: Whether to download the data if it is not present. 410 411 Returns: 412 List of filepaths for the hdf5 files, which contain the image data ('raw') and the label data ('labels'). 413 """ 414 preprocessed_dir = get_acrin_6698_data(path, max_cases, case_ids, download) 415 volume_paths = natsorted(glob(os.path.join(preprocessed_dir, "*.h5"))) 416 assert len(volume_paths) > 0, f"Could not find any preprocessed volumes in '{preprocessed_dir}'." 417 return volume_paths 418 419 420def get_acrin_6698_dataset( 421 path: Union[os.PathLike, str], 422 patch_shape: Tuple[int, ...], 423 resize_inputs: bool = False, 424 max_cases: Optional[int] = None, 425 case_ids: Optional[List[str]] = None, 426 download: bool = False, 427 **kwargs 428) -> Dataset: 429 """Get the ACRIN-6698 dataset for functional tumor volume segmentation in breast DCE-MRI. 430 431 Args: 432 path: Filepath to a folder where the data is downloaded for further processing. 433 patch_shape: The patch shape to use for training. 434 resize_inputs: Whether to resize inputs to the desired patch shape. 435 max_cases: The maximum number of cases to use. See `get_acrin_6698_data` for details. 436 case_ids: Explicit list of case ids to use. See `get_acrin_6698_data` for details. 437 download: Whether to download the data if it is not present. 438 kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset`. 439 440 Returns: 441 The segmentation dataset. 442 """ 443 volume_paths = get_acrin_6698_paths(path, max_cases, case_ids, download) 444 445 if resize_inputs: 446 resize_kwargs = {"patch_shape": patch_shape, "is_rgb": False} 447 kwargs, patch_shape = util.update_kwargs_for_resize_trafo( 448 kwargs=kwargs, patch_shape=patch_shape, resize_inputs=resize_inputs, resize_kwargs=resize_kwargs 449 ) 450 451 return torch_em.default_segmentation_dataset( 452 raw_paths=volume_paths, 453 raw_key="raw", 454 label_paths=volume_paths, 455 label_key="labels", 456 patch_shape=patch_shape, 457 is_seg_dataset=True, 458 **kwargs 459 ) 460 461 462def get_acrin_6698_loader( 463 path: Union[os.PathLike, str], 464 batch_size: int, 465 patch_shape: Tuple[int, ...], 466 resize_inputs: bool = False, 467 max_cases: Optional[int] = None, 468 case_ids: Optional[List[str]] = None, 469 download: bool = False, 470 **kwargs 471) -> DataLoader: 472 """Get the ACRIN-6698 dataloader for functional tumor volume segmentation in breast DCE-MRI. 473 474 Args: 475 path: Filepath to a folder where the data is downloaded for further processing. 476 batch_size: The batch size for training. 477 patch_shape: The patch shape to use for training. 478 resize_inputs: Whether to resize inputs to the desired patch shape. 479 max_cases: The maximum number of cases to use. See `get_acrin_6698_data` for details. 480 case_ids: Explicit list of case ids to use. See `get_acrin_6698_data` for details. 481 download: Whether to download the data if it is not present. 482 kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset` or for the PyTorch DataLoader. 483 484 Returns: 485 The DataLoader. 486 """ 487 ds_kwargs, loader_kwargs = util.split_kwargs(torch_em.default_segmentation_dataset, **kwargs) 488 dataset = get_acrin_6698_dataset(path, patch_shape, resize_inputs, max_cases, case_ids, download, **ds_kwargs) 489 return torch_em.get_data_loader(dataset, batch_size, **loader_kwargs)
332def get_acrin_6698_data( 333 path: Union[os.PathLike, str], 334 max_cases: Optional[int] = None, 335 case_ids: Optional[List[str]] = None, 336 download: bool = False, 337) -> str: 338 """Download the ACRIN-6698 dataset. 339 340 Args: 341 path: Filepath to a folder where the data is downloaded for further processing. 342 max_cases: The maximum number of VOLSER masks to download and preprocess, taken in deterministic 343 order of their series instance UID. Only the SEG series and their matched source MR series are 344 downloaded. Mutually exclusive with `case_ids`. 345 case_ids: Explicit list of VOLSER mask series instance UIDs to download and preprocess (see the 346 'acrin_6698_seg_series.json' metadata file written to `path` for the available UIDs and their 347 descriptions). Mutually exclusive with `max_cases`. 348 download: Whether to download the data if it is not present. 349 350 Returns: 351 Filepath where the preprocessed data is stored. 352 """ 353 assert max_cases is None or case_ids is None, "'max_cases' and 'case_ids' are mutually exclusive." 354 355 # NOTE: The preprocessing below skips volumes that were converted already, so an interrupted run resumes. 356 preprocessed_dir = os.path.join(path, "preprocessed") 357 os.makedirs(path, exist_ok=True) 358 359 seg_series = _get_seg_metadata(path, download) 360 seg_series = [series for series in seg_series if _is_volser_mask(series["SeriesDescription"])] 361 seg_series = natsorted(seg_series, key=lambda series: series["SeriesInstanceUID"]) 362 363 if max_cases is not None: 364 seg_series = seg_series[:max_cases] 365 elif case_ids is not None: 366 seg_series = [series for series in seg_series if series["SeriesInstanceUID"] in case_ids] 367 368 dicom_dir = os.path.join(path, "dicom") 369 if download: 370 seg_uids = [series["SeriesInstanceUID"] for series in seg_series] 371 util.download_tcia_series(seg_uids, dst=dicom_dir, csv_filename=os.path.join(path, "acrin_6698_seg")) 372 373 # The primary (description-matched) candidate source series of each mask's study is downloaded here 374 # already, so that the preprocessing loop below only has to download extra candidates for the rare 375 # case where the primary candidate's geometry does not match (see the module docstring). 376 study_series_cache = {} 377 primary_mr_uids = set() 378 for series in seg_series: 379 seg_paths = glob(os.path.join(dicom_dir, series["SeriesInstanceUID"], "*.dcm")) 380 if not seg_paths: 381 continue 382 import pydicom 383 study_uid = str(pydicom.dcmread(seg_paths[0], stop_before_pixels=True).StudyInstanceUID) 384 study_series = _get_study_series(study_uid, study_series_cache) 385 primary_uid = _match_by_description(series["SeriesDescription"], study_series) 386 if primary_uid is not None: 387 primary_mr_uids.add(primary_uid) 388 util.download_tcia_series( 389 sorted(primary_mr_uids), dst=dicom_dir, csv_filename=os.path.join(path, "acrin_6698_mr") 390 ) 391 elif not glob(os.path.join(dicom_dir, "*", "*.dcm")): 392 raise RuntimeError(f"Cannot find the data at {path}, but download was set to False.") 393 394 _preprocess_acrin_6698(seg_series, dicom_dir, preprocessed_dir, download) 395 return preprocessed_dir
Download the ACRIN-6698 dataset.
Arguments:
- path: Filepath to a folder where the data is downloaded for further processing.
- max_cases: The maximum number of VOLSER masks to download and preprocess, taken in deterministic
order of their series instance UID. Only the SEG series and their matched source MR series are
downloaded. Mutually exclusive with
case_ids. - case_ids: Explicit list of VOLSER mask series instance UIDs to download and preprocess (see the
'acrin_6698_seg_series.json' metadata file written to
pathfor the available UIDs and their descriptions). Mutually exclusive withmax_cases. - download: Whether to download the data if it is not present.
Returns:
Filepath where the preprocessed data is stored.
398def get_acrin_6698_paths( 399 path: Union[os.PathLike, str], 400 max_cases: Optional[int] = None, 401 case_ids: Optional[List[str]] = None, 402 download: bool = False, 403) -> List[str]: 404 """Get paths to the ACRIN-6698 data. 405 406 Args: 407 path: Filepath to a folder where the data is downloaded for further processing. 408 max_cases: The maximum number of cases to use. See `get_acrin_6698_data` for details. 409 case_ids: Explicit list of case ids to use. See `get_acrin_6698_data` for details. 410 download: Whether to download the data if it is not present. 411 412 Returns: 413 List of filepaths for the hdf5 files, which contain the image data ('raw') and the label data ('labels'). 414 """ 415 preprocessed_dir = get_acrin_6698_data(path, max_cases, case_ids, download) 416 volume_paths = natsorted(glob(os.path.join(preprocessed_dir, "*.h5"))) 417 assert len(volume_paths) > 0, f"Could not find any preprocessed volumes in '{preprocessed_dir}'." 418 return volume_paths
Get paths to the ACRIN-6698 data.
Arguments:
- path: Filepath to a folder where the data is downloaded for further processing.
- max_cases: The maximum number of cases to use. See
get_acrin_6698_datafor details. - case_ids: Explicit list of case ids to use. See
get_acrin_6698_datafor details. - download: Whether to download the data if it is not present.
Returns:
List of filepaths for the hdf5 files, which contain the image data ('raw') and the label data ('labels').
421def get_acrin_6698_dataset( 422 path: Union[os.PathLike, str], 423 patch_shape: Tuple[int, ...], 424 resize_inputs: bool = False, 425 max_cases: Optional[int] = None, 426 case_ids: Optional[List[str]] = None, 427 download: bool = False, 428 **kwargs 429) -> Dataset: 430 """Get the ACRIN-6698 dataset for functional tumor volume segmentation in breast DCE-MRI. 431 432 Args: 433 path: Filepath to a folder where the data is downloaded for further processing. 434 patch_shape: The patch shape to use for training. 435 resize_inputs: Whether to resize inputs to the desired patch shape. 436 max_cases: The maximum number of cases to use. See `get_acrin_6698_data` for details. 437 case_ids: Explicit list of case ids to use. See `get_acrin_6698_data` for details. 438 download: Whether to download the data if it is not present. 439 kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset`. 440 441 Returns: 442 The segmentation dataset. 443 """ 444 volume_paths = get_acrin_6698_paths(path, max_cases, case_ids, download) 445 446 if resize_inputs: 447 resize_kwargs = {"patch_shape": patch_shape, "is_rgb": False} 448 kwargs, patch_shape = util.update_kwargs_for_resize_trafo( 449 kwargs=kwargs, patch_shape=patch_shape, resize_inputs=resize_inputs, resize_kwargs=resize_kwargs 450 ) 451 452 return torch_em.default_segmentation_dataset( 453 raw_paths=volume_paths, 454 raw_key="raw", 455 label_paths=volume_paths, 456 label_key="labels", 457 patch_shape=patch_shape, 458 is_seg_dataset=True, 459 **kwargs 460 )
Get the ACRIN-6698 dataset for functional tumor volume segmentation in breast DCE-MRI.
Arguments:
- path: Filepath to a folder where the data is downloaded for further processing.
- patch_shape: The patch shape to use for training.
- resize_inputs: Whether to resize inputs to the desired patch shape.
- max_cases: The maximum number of cases to use. See
get_acrin_6698_datafor details. - case_ids: Explicit list of case ids to use. See
get_acrin_6698_datafor details. - download: Whether to download the data if it is not present.
- kwargs: Additional keyword arguments for
torch_em.default_segmentation_dataset.
Returns:
The segmentation dataset.
463def get_acrin_6698_loader( 464 path: Union[os.PathLike, str], 465 batch_size: int, 466 patch_shape: Tuple[int, ...], 467 resize_inputs: bool = False, 468 max_cases: Optional[int] = None, 469 case_ids: Optional[List[str]] = None, 470 download: bool = False, 471 **kwargs 472) -> DataLoader: 473 """Get the ACRIN-6698 dataloader for functional tumor volume segmentation in breast DCE-MRI. 474 475 Args: 476 path: Filepath to a folder where the data is downloaded for further processing. 477 batch_size: The batch size for training. 478 patch_shape: The patch shape to use for training. 479 resize_inputs: Whether to resize inputs to the desired patch shape. 480 max_cases: The maximum number of cases to use. See `get_acrin_6698_data` for details. 481 case_ids: Explicit list of case ids to use. See `get_acrin_6698_data` for details. 482 download: Whether to download the data if it is not present. 483 kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset` or for the PyTorch DataLoader. 484 485 Returns: 486 The DataLoader. 487 """ 488 ds_kwargs, loader_kwargs = util.split_kwargs(torch_em.default_segmentation_dataset, **kwargs) 489 dataset = get_acrin_6698_dataset(path, patch_shape, resize_inputs, max_cases, case_ids, download, **ds_kwargs) 490 return torch_em.get_data_loader(dataset, batch_size, **loader_kwargs)
Get the ACRIN-6698 dataloader for functional tumor volume segmentation in breast DCE-MRI.
Arguments:
- path: Filepath to a folder where the data is downloaded for further processing.
- batch_size: The batch size for training.
- patch_shape: The patch shape to use for training.
- resize_inputs: Whether to resize inputs to the desired patch shape.
- max_cases: The maximum number of cases to use. See
get_acrin_6698_datafor details. - case_ids: Explicit list of case ids to use. See
get_acrin_6698_datafor details. - download: Whether to download the data if it is not present.
- kwargs: Additional keyword arguments for
torch_em.default_segmentation_datasetor for the PyTorch DataLoader.
Returns:
The DataLoader.