torch_em.data.datasets.histopathology.sthelar
The STHELAR dataset contains annotations for nucleus instance segmentation and cell type classification in H&E stained histopathology images of 13 human tissue types (breast, cervix, colon, heart, kidney, liver, lung, lymph node, ovary, pancreas, prostate, skin and tonsil).
The dataset links Xenium spatial transcriptomics with H&E whole slide images. The nuclei masks and the cell types
are derived from the spatial transcriptomics data of the slides (cell types via Tangram alignment to single-cell
reference atlases, followed by clustering and marker gene based refinement) and are NOT manual annotations,
see the publication for the quality control. The slides are cut into patches of 256x256 pixels with an overlap
of 64 pixels: the Hugging Face release used here contains 154,814 patches at 20x and 587,555 patches at 40x
magnification, extracted from 27 slides (see SLIDES).
Each patch is stored in a separate hdf5 file, which contains the RGB image ('image', channel first) and the labels
'labels/instances' (nucleus instance ids, unique per patch, 0 is background) and 'labels/semantic' (cell type of
each nucleus pixel). The cell types are the harmonized 'cells_final_label_group' of the per-slide metadata, with
the label ids given by CLASS_NAMES (1-based, 0 is background). Nuclei can be cut at the patch border, and
neighbouring patches overlap, so the same cell can appear in several patches.
NOTE: The data is stored in large parquet files, each holding the patches of one or several slides, and this
module converts them into the patch files once. The parquet files are kept next to the converted data. All slides
need 18 GB (20x) or 54 GB (40x) to download plus the converted patches. Use slides to only download and convert a
subset. This requires the pyarrow python package.
The data is located at https://huggingface.co/datasets/FelicieGS/STHELAR_20x and https://huggingface.co/datasets/FelicieGS/STHELAR_40x, released under a CC-BY-4.0 license. The full data, including the per-cell analyses, is available at https://doi.org/10.6019/S-BIAD2146.
This dataset is from the publication https://doi.org/10.1038/s41597-026-06937-6. Please cite it if you use this dataset for your research.
1"""The STHELAR dataset contains annotations for nucleus instance segmentation and cell type classification 2in H&E stained histopathology images of 13 human tissue types (breast, cervix, colon, heart, kidney, liver, 3lung, lymph node, ovary, pancreas, prostate, skin and tonsil). 4 5The dataset links Xenium spatial transcriptomics with H&E whole slide images. The nuclei masks and the cell types 6are derived from the spatial transcriptomics data of the slides (cell types via Tangram alignment to single-cell 7reference atlases, followed by clustering and marker gene based refinement) and are NOT manual annotations, 8see the publication for the quality control. The slides are cut into patches of 256x256 pixels with an overlap 9of 64 pixels: the Hugging Face release used here contains 154,814 patches at 20x and 587,555 patches at 40x 10magnification, extracted from 27 slides (see `SLIDES`). 11 12Each patch is stored in a separate hdf5 file, which contains the RGB image ('image', channel first) and the labels 13'labels/instances' (nucleus instance ids, unique per patch, 0 is background) and 'labels/semantic' (cell type of 14each nucleus pixel). The cell types are the harmonized 'cells_final_label_group' of the per-slide metadata, with 15the label ids given by `CLASS_NAMES` (1-based, 0 is background). Nuclei can be cut at the patch border, and 16neighbouring patches overlap, so the same cell can appear in several patches. 17 18NOTE: The data is stored in large parquet files, each holding the patches of one or several slides, and this 19module converts them into the patch files once. The parquet files are kept next to the converted data. All slides 20need 18 GB (20x) or 54 GB (40x) to download plus the converted patches. Use `slides` to only download and convert a 21subset. This requires the pyarrow python package. 22 23The data is located at https://huggingface.co/datasets/FelicieGS/STHELAR_20x and 24https://huggingface.co/datasets/FelicieGS/STHELAR_40x, released under a CC-BY-4.0 license. The full data, 25including the per-cell analyses, is available at https://doi.org/10.6019/S-BIAD2146. 26 27This dataset is from the publication https://doi.org/10.1038/s41597-026-06937-6. 28Please cite it if you use this dataset for your research. 29""" 30 31import os 32import uuid 33from io import BytesIO 34from glob import glob 35from tqdm import tqdm 36from concurrent import futures 37from natsort import natsorted 38from typing import Union, Tuple, List, Literal, Optional, Sequence 39 40import numpy as np 41import imageio.v3 as imageio 42from scipy.sparse import load_npz 43 44from torch.utils.data import Dataset, DataLoader 45 46import torch_em 47 48from .. import util 49 50 51CLASS_NAMES = [ 52 "Epithelial", "Blood_vessel", "Fibroblast_Myofibroblast", "Myeloid", "B_Plasma", "T_NK", "Melanocyte", 53 "Specialized", "Other", 54] 55CLASS_IDS = {name: i for i, name in enumerate(CLASS_NAMES, start=1)} 56 57MAGNIFICATIONS = ["20x", "40x"] 58COMPLETE_MARKER = ".complete" 59 60SLIDES = [ 61 "breast_s0", "breast_s1", "breast_s3", "breast_s6", 62 "cervix_s0", "colon_s1", "colon_s2", "heart_s0", 63 "kidney_s0", "kidney_s1", "liver_s0", "liver_s1", 64 "lung_s1", "lung_s3", "lymph_node_s0", "ovary_s0", 65 "ovary_s1", "pancreatic_s0", "pancreatic_s1", "pancreatic_s2", 66 "prostate_s0", "skin_s1", "skin_s2", "skin_s3", 67 "skin_s4", "tonsil_s0", "tonsil_s1", 68] 69 70REPOS = {"20x": "FelicieGS/STHELAR_20x", "40x": "FelicieGS/STHELAR_40x"} 71REVISIONS = { 72 "20x": "a01e22cdb7368edf595edf6150d0ac57deb5e548", 73 "40x": "e32a8cdd50eff2d38e237f3729e9ac85bbb5203b", 74} 75 76SLIDE_SHARDS = { 77 "20x": { 78 "breast_s0": [0, 1], 79 "breast_s1": [1, 2, 3, 4], 80 "breast_s3": [4, 5, 6], 81 "breast_s6": [6, 7], 82 "cervix_s0": [7, 8], 83 "colon_s1": [8, 9], 84 "colon_s2": [9], 85 "heart_s0": [9], 86 "kidney_s0": [9], 87 "kidney_s1": [9, 10], 88 "liver_s0": [10], 89 "liver_s1": [10, 11], 90 "lung_s1": [11], 91 "lung_s3": [11, 12], 92 "lymph_node_s0": [12], 93 "ovary_s0": [12], 94 "ovary_s1": [12, 13], 95 "pancreatic_s0": [13, 14], 96 "pancreatic_s1": [14], 97 "pancreatic_s2": [14], 98 "prostate_s0": [14, 15], 99 "skin_s1": [15], 100 "skin_s2": [15, 16], 101 "skin_s3": [16], 102 "skin_s4": [16], 103 "tonsil_s0": [16, 17], 104 "tonsil_s1": [17], 105 }, 106 "40x": { 107 "breast_s0": [0, 1, 2, 3, 4], 108 "breast_s1": [4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15], 109 "breast_s3": [15, 16, 17, 18], 110 "breast_s6": [18, 19, 20, 21, 22], 111 "cervix_s0": [22, 23, 24, 25, 26], 112 "colon_s1": [26, 27], 113 "colon_s2": [27, 28, 29], 114 "heart_s0": [29], 115 "kidney_s0": [29, 30], 116 "kidney_s1": [30], 117 "liver_s0": [30, 31, 32], 118 "liver_s1": [32, 33], 119 "lung_s1": [33, 34], 120 "lung_s3": [34, 35, 36], 121 "lymph_node_s0": [36, 37], 122 "ovary_s0": [37, 38], 123 "ovary_s1": [38, 39, 40, 41], 124 "pancreatic_s0": [41, 42, 43], 125 "pancreatic_s1": [43, 44], 126 "pancreatic_s2": [44, 45], 127 "prostate_s0": [45, 46], 128 "skin_s1": [46, 47], 129 "skin_s2": [47, 48], 130 "skin_s3": [48, 49], 131 "skin_s4": [49], 132 "tonsil_s0": [49, 50, 51, 52], 133 "tonsil_s1": [52, 53], 134 }, 135} 136 137SHARD_CHECKSUMS = { 138 "20x": [ 139 "2beb26d148a5bd164289708794ebbcdb79b974135732509db8eb5d8a99e61840", 140 "8c4a010b9049a64f93eb3bc819789a713640b815be0b7c7d84b0aff50093a86a", 141 "f1bbcc6b5d3792554e376f5353b9ddc4168167ccb773823dd55bb3967d36cdf9", 142 "e4642280e2a51c3175e2555f2110f88c3c3c448a4c1c0877f20f7b1780c5ba54", 143 "8254d611a92ecd81e8e4f985b352f41767177dc6c941ff4f61e3f9499bce87bb", 144 "64cf474fab05b96ced0cd0e14b429503975ee4e9af472f7d28d07c66c2f43390", 145 "c45fff34cb07e2e28565a8a283b32c032c3c49a6d7b16f819d4970e405e4bcf6", 146 "3201ca4493a73aaef48113127ad33f022c0e3aa878e3ad2935d21ba1f593e86b", 147 "2bdfe7443d7247c89801ea6ad51f2a948b46fd022e1bf33614a8bfc3e9f6c657", 148 "fd0be2332dcda2e95718a245e00943a2289d6fee940a70032d2b3a04bc2110cf", 149 "88cdba926f9487d89e88b09926b2fa1a808bb9f839ef675080980ca262554dfe", 150 "b3de85309f6b2f044b49ada57bbf2e9589293b1da81e436a0c615a945351759e", 151 "cfd75a5137dba283bc2d14cc47f56b1055f30f84a8fc056714f6714fdde1137d", 152 "2dd980ca2cb79ce746fafafdd6ff905c48014734d6bbca0750591a9d22a6a7f5", 153 "3d6b02ceba429b91f122423505f79bcc2c4b7a81ea22dc6b0f63f6b195a2ab63", 154 "051691b8a73b811c929c19cbcfd0611e36c3cfee585d8c837201e6651be9f5c9", 155 "99d9d589bc21788516237414a136c76758a76099713f3eea08d9b6df034bedd8", 156 "d17ba1032ff76fae8d8a9c8ba529facd073519b81423e686790ad6fba773c0eb", 157 ], 158 "40x": [ 159 "fded6f9036cdba4c0550e995d1dc2cd46f68ba14276419378d49c523a2ace915", 160 "9090da789118603f02cc451ca4e28a4e0643c447e0a6f61e1f79925022a3a875", 161 "c9ed431b0d63631e299c9fb909812b4a0f51160d4a780cb5f9e5611a3e1265b5", 162 "3e548fa7f80d8ee8455fe399f15eb554954fa987de0348764c0c5f9dac97373c", 163 "2cb12c3eab945479624e93658fc54e6ba441d3eaf279401f440fc2943291f319", 164 "7c198b860fac69fc952fcc42a03ced84e414a4c6618a8322e4c4092ae93f2185", 165 "742ff210dd58adc96e9759037811ddeefd865f042b01e8206e784045dc1868f6", 166 "8263ebaee38fdce69d6a052c9488043b952136e93a8abe7c3852171c382c3cf2", 167 "5b66f363aa6dd86d98eaed356ca16b53df0dcb6cb4ae6d9dc2c72f15ba38865c", 168 "7d39f3cde0a3000b4ab6b0bfb9d845ba14cdb80574fd1f17cb64e0cb8e81a5e9", 169 "90ede649a250ea7877d97b400bbe64f01506b96de5b6e420eb9a5ab08cebd279", 170 "5e79e72a0fa406e0f070c65b47b07a9b390c9c712a055185e0ab34059fa2ed84", 171 "df044c5bfc34be9e3d930d286fd7665e19810ed3d8a42dde7233b0deb8e5b8b3", 172 "57baa4ee89d96e0b4712e874a680a7cd3ce4b92b38385576d9593aadcfa7bf7f", 173 "db5d9801b45ef358d1c8b90865ebcef8660f4628ab5195a9d0651a8373dffada", 174 "1cdd7d08ec094f4f950825af322518558928523a1e58c44e59b6e9d44f8ec6e1", 175 "463c18d8cacd82e05738a00b6bd09bbfc1f53eab93c04abf78df0935afb6c4df", 176 "b2a265d3a65d1a3c1461763fde7a088b4f6d3ed94f5ec6bb04e3debcb84201aa", 177 "b62961040f0ba1e5ac56dc3a328a1b478001bee668e03b03fde4a75cc2cf7266", 178 "354f00172234bd2e788b7874c31bab201eff84fafaf19ad425e2d8705b79b8d6", 179 "90a20a79276eb571d8840ee3391923eca32d075c940926a674a51f8fa86a5cf8", 180 "38601ae47fa7ac50a6327e466395acb82f20adb1a18314d578f11458f2d55574", 181 "96fbf1edb41689439d7e327dab80a24a6b720f143054b35de1193d21c49cad14", 182 "889e24ea440f9bbc0e086f91086e4f30a6a409c0aa4ab4004ec8887a7897e3ea", 183 "38dc9cf1a2bd50c6b21e32f4fc2dfc630d0691acee6e17159ff2f741bcd36379", 184 "19f8bd33957c0fa13e54f858a8ea422095ee4ad30708c9f872c66f81ec7f135a", 185 "a478b13ae480f546f3f74df8af6e5e7be78a83d82574814d3e4994205e317ebb", 186 "b7c6a826ebfa182c421c652fe850c2efc0d9e23090ea6378be09925297197862", 187 "042eef94be71a40268f0fd12a88bded6d1d99ec08aace18233728823a61523de", 188 "c1308467499672cd49110a7ddea1bb0ce0d009faba744da67a6da0e305d93bb5", 189 "7ba700b6d595e7e6f72103392ebc16cd1cdf94c79174ae79765cf96768e8bddc", 190 "ad928308055f7e86813ab613a05aa0c6bd65f78cfe7bab8892c38f90264fbba9", 191 "780810e23adac20be6e5c8fd503e4261c5cf78dfd53ce3e36056294288a6dd44", 192 "89f577a75a07026fb3293fcac5a9d451958232a0960866908c5282f9eaae5817", 193 "8df8f7310cbba63e047e8ae9d30987e09f4697ae927cc7c7b7e08d22bcf093f6", 194 "aaf8357b174e2a53532bcb213d3ba21e79dfa4c82cf39ebf338bb6351f1f1de6", 195 "6301ba19059423d481e11617237c058f9b8382f4993401baa14ffc9732d6e3c6", 196 "289e3dbbaaf77e06a57e4aa711f61f5eb4d7fb6b3f97247967874400db64ea57", 197 "5d67d1601a292dc9f371f2ea3349036fde6bbcb18eaa035349c3600ec0dbeb31", 198 "ec7b89e7acd33c9dfbc6b9e12737bcb6da887fd2ca8634cb330b32455c4c5833", 199 "cacb6d40d9dd4e9fe0441c18c0d2ed687da738fc7c12c1767feaa96ad6a52f3d", 200 "571a72273be1fa6f3c14e99087ba130f2f83b098f9674291e1f705f0e4e2752d", 201 "1470a4745b0fe58253d7c484e4d6304e9fc3c008158f66fe403f6fb00e698ee4", 202 "f33e597a2aba5eff981144c5eaa89d51844de76a1a9dcf58b22da7e9a6a301a4", 203 "42e617f6f15ce018d2d11a5ba20527c97414aaea14debf28600240542cf55f9c", 204 "000ecc178dab26dcca21a282d410b51b2b3bec8edb800edc60baa40fd34124b7", 205 "2240e03cfd020c8a9e53524e335f17d1edb4b28319584726efd77429587e9547", 206 "487f7a86152e7acfd623a27fade0fc0c88ad89367f764971bba1f619bff5e140", 207 "af125205e4ff222db33611a2de2ef398e485f11ad2e492632ecb6524de0a0b07", 208 "67ea31dc49c0022d803e143c41e79a61feca2bfb98d0e001c45523b33741d4eb", 209 "d0883ae9c68b7eafad1ae9b8b75a3befcc1aaa86bb1120d022b90854e9d73ed8", 210 "10285d36a9024b2a036b3f019f294dcc6a38d41b348a8176702a7ce87c587f02", 211 "d28f98f0bf468f2fc14f95f1c721175513269ee9a18a13d4314c5e5a848d2e23", 212 "94f2cea4dfd344933c942d14243744b882cc2c10d42628f95ff51c7253c28df7", 213 ], 214} 215 216METADATA_CHECKSUMS = { 217 "20x": { 218 "breast_s0": "1cdb4bba93afb1c3d3eeca8f2caacb6a5d10b22f07b51acce274f8f015588bf5", 219 "breast_s1": "abf1ebba18841e7a86f3114516a07574221f7d6978989c232dedcc7a258169e6", 220 "breast_s3": "435e1a77a0acb2842d08276ad7e3cc67916a6e3949a56aca2260b831b5e0e26e", 221 "breast_s6": "02b60f4bcb26646bc9273acd0dfc6ee98f6a6811266c7b862a625c2f5d87c15d", 222 "cervix_s0": "3a198703ca5e228c81ea03166082316f857cc42416da5b303090206ca24abb4e", 223 "colon_s1": "0f11618dae9209537855ce9568139181cc7d6223ea98810907ad299e77a0067e", 224 "colon_s2": "823b0be5277fabf4930eae5d4402f4fabafd8876dc674c19e641b203a20a96b9", 225 "heart_s0": "4f1a38ab08f83603faac675a3fba49fcc3528fda91a5c288b4efd5d2f34a878a", 226 "kidney_s0": "a4db038e375a8f47532f689fd457516e761acf9c23cb90e128ac5d134b1b4f20", 227 "kidney_s1": "a9bf51936b349456b836c80de6ba131852e91e3e040f20eda3802d837a27533c", 228 "liver_s0": "17d9a987a88460824dabde15a55e5e864cf606aed57e2d3f2f8b4505b9c2d2eb", 229 "liver_s1": "1e50bcda0c31cdd25a84bed02d6c548881a3f3e8a20b57d3daa4982dad505bc1", 230 "lung_s1": "83b059126e8a1bac93eee91cd36ffda90e02aac406f241dc17b09963bf488463", 231 "lung_s3": "639ded2b7c195a44ad7708fef31e4314505fdced5a8097e0e95a3ab5624cd361", 232 "lymph_node_s0": "040c24a0bcd156ecbdfb13d476bae127bf4339cd189dae460ee42f7c4bb0242a", 233 "ovary_s0": "b77a2d78eaff3a69bc4796ab9310c1f708b0368ceaaeb7a3196490e4bbc5160f", 234 "ovary_s1": "2e66104150da680f19789a1e6a97b06f0e5710c83a2803a6754278838c8c59ad", 235 "pancreatic_s0": "5b242492f5ddd965cfa8e6fa991ceea51b9d549fccfb56e91c121e46baba7a6f", 236 "pancreatic_s1": "a1f37d4b57028f814be2f55ff31f74c87f1e82f482558071086ae79078d8dd0e", 237 "pancreatic_s2": "f006c40cbf16a5a2ccf0782fe872f7150d42c5247fec5e93452ae74e132cac56", 238 "prostate_s0": "e719328ca0239dc0ec5f252f2dea88e5c0e94595b7911f4759c1df935bfbd114", 239 "skin_s1": "60354b20b2824abea0dfec2c70e4df2555329d6c2672091ce47e07a88f038af5", 240 "skin_s2": "ef4cd89a2df0ef04f777a184667661f1266a0e4cb6a2fc633e57f84b1783f95c", 241 "skin_s3": "150412b1a9ff4489756567f970a9663b5c894fd50093241316f5e72ef7008e49", 242 "skin_s4": "df44e89d5efbdbfb5d3702e92921ebc38060ee5df02bbb8a87604467780013ea", 243 "tonsil_s0": "8edd68ce0858c303386ed017264b89d7ffa90ba61dec9cc937ed849668394493", 244 "tonsil_s1": "3dc4adb307b85ac7e53bcf65a8d77ce02b62f058d58fcab8543618e804f815e9", 245 }, 246 "40x": { 247 "breast_s0": "f8ddf4696e0a0dd927a8b0c1e24242d3f4575190c7a26fe80f78d0932bb23227", 248 "breast_s1": "5fea33e2716a92054f02f3d8d75ca96a37f7e75acf91821451f6bd771d2fbc00", 249 "breast_s3": "457a3fdabaad46f8e007958419a5828a6d50e438f236b1bb0064b8d604712730", 250 "breast_s6": "64036da240ccf08273156939dbe1a3545951d07d1f62d97db7d3a3f7c8758689", 251 "cervix_s0": "dc980ba27a43a6fe826e02cf529cf17e35f98a22781a530cd5c2b298ebb89d1c", 252 "colon_s1": "27a31ef22fe716285f63dab79bf45aa9cdf6e8e860b8ca725910a239b800050c", 253 "colon_s2": "ad1d3e2cfa20c39da3c4db03c4760732a0e4d851b3ef7f8b45f30cb6ba74756f", 254 "heart_s0": "cfa6690e8be28714583edb866cb3fadf575733cdd084b6d3e8cc68e49b3917d8", 255 "kidney_s0": "70ea019dc4c89647fe4ea34f65061b1386cac46cf69f78e036b1d3f14b6c1766", 256 "kidney_s1": "88fd5484af6f5d1734a7f0f7a8fe6e50151a39c73c48a11892baf79ffd1f724e", 257 "liver_s0": "b6f74e1ecfaa59bbb75921ba6855877f5692487e6890e7a396cd0684d7cb467d", 258 "liver_s1": "a6058b2e2ed771e7cf82dfc1f22c3590fe99d003abdd22d18cbe9266f44bc03d", 259 "lung_s1": "041fbcae0f18ae626c2b5daa37231020bad1edf589ea5149418ca9b727eae2c8", 260 "lung_s3": "4c2ad9337adb0683604f0c8683288cb8eb902f095dc09eeb8764c58ab6c75386", 261 "lymph_node_s0": "e1b4659526ee64a22eada40cec267a7a1b75fc3079d84f6123bfe547bf2716b2", 262 "ovary_s0": "9e561f24a1cc9b6a3dfd9e2c2320c5d11a9ebae9d54bc5ddd6d9db3c446b5375", 263 "ovary_s1": "79f100d35cedee540d5f5319549a381faf8ebde736601e329a99f815b3978980", 264 "pancreatic_s0": "f154d79606c8d1acff94c032bfc975e8518913f488da2bc9224d8ae32162c05b", 265 "pancreatic_s1": "e46c59592eea273450012d599dc09e87a970dab35eae1e9f00f431ecdd1201ed", 266 "pancreatic_s2": "efccb9598ad581bcc0da9fd7c4b2b585b7c405774f0359e2380dd2ae51f56a54", 267 "prostate_s0": "5397a52b30862f60780a8d105cacfa5a5ef7db66896edc88def2a6aa7b55cb3e", 268 "skin_s1": "4b4faf5248e8555fd68395f16d7a6ca653c02f7b45171da224f77026e5246f8c", 269 "skin_s2": "18b149fae149ac1b35c09701d8815b5818c90b5670643269a465d6aa34597b62", 270 "skin_s3": "2cd2e0c9d49c56192cb2eec65d2f4bec3feb0fef3a986472a2c50113e61e5bd5", 271 "skin_s4": "e5879ec0348b02eb1c6cb8ca04b7c33f88d50328e95008309453ef24e21e1702", 272 "tonsil_s0": "0f95581fdc179715ac190df5e6a02df03e48baf156c4c700fc011101ba9315ff", 273 "tonsil_s1": "5f40a96f6c11cdfba8f80f4eb0bdf14c0b34000166c5d116a3714bf25e478222", 274 }, 275} 276 277 278def _validate(magnification, slides): 279 if magnification not in MAGNIFICATIONS: 280 raise ValueError(f"'{magnification}' is not a valid magnification. Choose one of {MAGNIFICATIONS}.") 281 if slides is None: 282 return list(SLIDES) 283 unknown = [s for s in slides if s not in SLIDES] 284 if unknown: 285 raise ValueError(f"{unknown} are not valid slides. Choose from {SLIDES}.") 286 return [s for s in SLIDES if s in slides] 287 288 289def _shard_name(magnification, shard_id): 290 return f"train-{shard_id:05d}-of-{len(SHARD_CHECKSUMS[magnification]):05d}.parquet" 291 292 293def _resolve_url(magnification, relative_path): 294 return f"https://huggingface.co/datasets/{REPOS[magnification]}/resolve/{REVISIONS[magnification]}/{relative_path}" 295 296 297def _load_lookup(metadata_path): 298 import pandas as pd 299 300 df = pd.read_parquet(metadata_path, columns=["cell_id_int", "cells_final_label_group"]) 301 codes = df["cells_final_label_group"].map(CLASS_IDS) 302 assert not codes.isna().any(), f"Unknown cell types in {metadata_path}." 303 ids = df["cell_id_int"].to_numpy().astype("int64") 304 order = np.argsort(ids) 305 return ids[order], codes.to_numpy().astype("uint8")[order] 306 307 308def _write_patch(out_path, image_bytes, cell_id_bytes, lookup): 309 if os.path.exists(out_path): 310 return 311 312 import h5py 313 314 image = imageio.imread(image_bytes, extension=".png") 315 assert image.ndim == 3 and image.shape[-1] == 3 316 cell_ids = load_npz(BytesIO(cell_id_bytes)).toarray().astype("int64") 317 assert cell_ids.shape == image.shape[:2] 318 319 unique_ids, inverse = np.unique(cell_ids, return_inverse=True) 320 instances = inverse.reshape(cell_ids.shape) 321 if unique_ids[0] == 0: 322 nucleus_ids = unique_ids[1:] 323 else: 324 nucleus_ids = unique_ids 325 instances = instances + 1 326 327 meta_ids, meta_codes = lookup 328 positions = np.minimum(np.searchsorted(meta_ids, nucleus_ids), len(meta_ids) - 1) 329 assert (meta_ids[positions] == nucleus_ids).all(), "Found cell ids without metadata." 330 semantic = np.concatenate([[0], meta_codes[positions]]).astype("uint8")[instances] 331 332 tmp_path = f"{out_path}.{uuid.uuid4().hex}.incomplete" 333 with h5py.File(tmp_path, "w") as f: 334 f.create_dataset("image", data=image.transpose((2, 0, 1)), compression="gzip") 335 f.create_dataset("labels/instances", data=instances.astype("int32"), compression="gzip") 336 f.create_dataset("labels/semantic", data=semantic, compression="gzip") 337 os.replace(tmp_path, out_path) 338 339 340def _convert_shard(shard_path, slides, lookups, preprocessed_dir, n_workers): 341 import pyarrow.parquet as pq 342 343 parquet_file = pq.ParquetFile(shard_path) 344 with futures.ThreadPoolExecutor(n_workers) as pool: 345 for row_group in tqdm(range(parquet_file.num_row_groups), desc=f"Convert {os.path.basename(shard_path)}"): 346 table = parquet_file.read_row_group(row_group, columns=["file_name", "slide_id", "image", "cell_id_map"]) 347 slide_ids = table.column("slide_id").to_pylist() 348 rows = [i for i, slide_id in enumerate(slide_ids) if slide_id in slides] 349 if not rows: 350 continue 351 352 file_names = table.column("file_name").to_pylist() 353 images = table.column("image").to_pylist() 354 cell_id_maps = table.column("cell_id_map").to_pylist() 355 tasks = [] 356 for i in rows: 357 out_path = os.path.join(preprocessed_dir, slide_ids[i], f"{os.path.splitext(file_names[i])[0]}.h5") 358 tasks.append( 359 pool.submit(_write_patch, out_path, images[i]["bytes"], cell_id_maps[i], lookups[slide_ids[i]]) 360 ) 361 for task in tasks: 362 task.result() 363 364 365def get_sthelar_data( 366 path: Union[os.PathLike, str], 367 magnification: Literal["20x", "40x"] = "20x", 368 slides: Optional[Sequence[str]] = None, 369 download: bool = False, 370) -> str: 371 """Download the STHELAR dataset and convert its patches to hdf5 files. 372 373 NOTE: All slides need 18 GB (20x) or 54 GB (40x) for the download. Use `slides` for a subset. The parquet files 374 of the slides are kept, since a parquet file can contain the patches of several slides. 375 376 Args: 377 path: Filepath to a folder where the data is downloaded for further processing. 378 magnification: The choice of magnification. Either '20x' or '40x'. 379 slides: The slides to use, see `SLIDES` for the valid choices. By default all slides are used. 380 download: Whether to download the data if it is not present. 381 382 Returns: 383 Filepath where the converted patches are stored, in one sub-folder per slide. 384 """ 385 slides = _validate(magnification, slides) 386 data_dir = os.path.join(path, magnification) 387 preprocessed_dir = os.path.join(data_dir, "preprocessed") 388 389 pending = [s for s in slides if not os.path.exists(os.path.join(preprocessed_dir, s, COMPLETE_MARKER))] 390 if not pending: 391 return preprocessed_dir 392 393 os.makedirs(os.path.join(data_dir, "data"), exist_ok=True) 394 os.makedirs(os.path.join(data_dir, "cell_metadata"), exist_ok=True) 395 396 lookups = {} 397 for slide in pending: 398 relative_path = f"cell_metadata/{slide}_cell_metadata.parquet" 399 metadata_path = os.path.join(data_dir, relative_path) 400 util.download_source( 401 path=metadata_path, url=_resolve_url(magnification, relative_path), download=download, 402 checksum=METADATA_CHECKSUMS[magnification][slide], 403 ) 404 lookups[slide] = _load_lookup(metadata_path) 405 os.makedirs(os.path.join(preprocessed_dir, slide), exist_ok=True) 406 407 n_cpus = len(os.sched_getaffinity(0)) if hasattr(os, "sched_getaffinity") else (os.cpu_count() or 1) 408 n_workers = min(8, n_cpus) 409 410 shard_ids = sorted({i for slide in pending for i in SLIDE_SHARDS[magnification][slide]}) 411 for shard_id in shard_ids: 412 relative_path = f"data/{_shard_name(magnification, shard_id)}" 413 shard_path = os.path.join(data_dir, relative_path) 414 util.download_source( 415 path=shard_path, url=_resolve_url(magnification, relative_path), download=download, 416 checksum=SHARD_CHECKSUMS[magnification][shard_id], 417 ) 418 _convert_shard(shard_path, set(pending), lookups, preprocessed_dir, n_workers) 419 420 for slide in pending: 421 with open(os.path.join(preprocessed_dir, slide, COMPLETE_MARKER), "w"): 422 pass 423 424 return preprocessed_dir 425 426 427def get_sthelar_paths( 428 path: Union[os.PathLike, str], 429 magnification: Literal["20x", "40x"] = "20x", 430 slides: Optional[Sequence[str]] = None, 431 download: bool = False, 432) -> List[str]: 433 """Get paths to the STHELAR data. 434 435 Args: 436 path: Filepath to a folder where the data is downloaded for further processing. 437 magnification: The choice of magnification. Either '20x' or '40x'. 438 slides: The slides to use, see `SLIDES` for the valid choices. By default all slides are used. 439 download: Whether to download the data if it is not present. 440 441 Returns: 442 List of filepaths for the hdf5 files, which contain the image ('image') and the labels 443 ('labels/instances' and 'labels/semantic') of a patch. 444 """ 445 slides = _validate(magnification, slides) 446 preprocessed_dir = get_sthelar_data(path, magnification, slides, download) 447 448 data_paths = [] 449 for slide in slides: 450 data_paths.extend(natsorted(glob(os.path.join(preprocessed_dir, slide, "*.h5")))) 451 452 assert len(data_paths) > 0 453 return data_paths 454 455 456def get_sthelar_dataset( 457 path: Union[os.PathLike, str], 458 patch_shape: Tuple[int, int], 459 magnification: Literal["20x", "40x"] = "20x", 460 slides: Optional[Sequence[str]] = None, 461 label_type: Literal["instances", "semantic"] = "instances", 462 resize_inputs: bool = False, 463 download: bool = False, 464 **kwargs 465) -> Dataset: 466 """Get the STHELAR dataset for nucleus segmentation and cell type classification. 467 468 Args: 469 path: Filepath to a folder where the data is downloaded for further processing. 470 patch_shape: The patch shape to use for training. 471 magnification: The choice of magnification. Either '20x' or '40x'. 472 slides: The slides to use, see `SLIDES` for the valid choices. By default all slides are used. 473 label_type: The choice of labels. Either 'instances' for the nucleus instances or 'semantic' for the 474 cell types of the nuclei, with the ids given by `CLASS_NAMES`. 475 resize_inputs: Whether to resize the inputs to the patch shape. 476 download: Whether to download the data if it is not present. 477 kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset`. 478 479 Returns: 480 The segmentation dataset. 481 """ 482 if label_type not in ("instances", "semantic"): 483 raise ValueError(f"'{label_type}' is not a valid label type. Choose either 'instances' or 'semantic'.") 484 485 data_paths = get_sthelar_paths(path, magnification, slides, download) 486 487 if resize_inputs: 488 resize_kwargs = {"patch_shape": patch_shape, "is_rgb": True} 489 kwargs, patch_shape = util.update_kwargs_for_resize_trafo( 490 kwargs=kwargs, patch_shape=patch_shape, resize_inputs=resize_inputs, resize_kwargs=resize_kwargs 491 ) 492 493 return torch_em.default_segmentation_dataset( 494 raw_paths=data_paths, 495 raw_key="image", 496 label_paths=data_paths, 497 label_key=f"labels/{label_type}", 498 patch_shape=patch_shape, 499 ndim=2, 500 with_channels=True, 501 **kwargs 502 ) 503 504 505def get_sthelar_loader( 506 path: Union[os.PathLike, str], 507 batch_size: int, 508 patch_shape: Tuple[int, int], 509 magnification: Literal["20x", "40x"] = "20x", 510 slides: Optional[Sequence[str]] = None, 511 label_type: Literal["instances", "semantic"] = "instances", 512 resize_inputs: bool = False, 513 download: bool = False, 514 **kwargs 515) -> DataLoader: 516 """Get the STHELAR dataloader for nucleus segmentation and cell type classification. 517 518 Args: 519 path: Filepath to a folder where the data is downloaded for further processing. 520 batch_size: The batch size for training. 521 patch_shape: The patch shape to use for training. 522 magnification: The choice of magnification. Either '20x' or '40x'. 523 slides: The slides to use, see `SLIDES` for the valid choices. By default all slides are used. 524 label_type: The choice of labels. Either 'instances' for the nucleus instances or 'semantic' for the 525 cell types of the nuclei, with the ids given by `CLASS_NAMES`. 526 resize_inputs: Whether to resize the inputs to the patch shape. 527 download: Whether to download the data if it is not present. 528 kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset` or for the PyTorch DataLoader. 529 530 Returns: 531 The DataLoader. 532 """ 533 ds_kwargs, loader_kwargs = util.split_kwargs(torch_em.default_segmentation_dataset, **kwargs) 534 dataset = get_sthelar_dataset( 535 path, patch_shape, magnification, slides, label_type, resize_inputs, download, **ds_kwargs 536 ) 537 return torch_em.get_data_loader(dataset, batch_size, **loader_kwargs)
366def get_sthelar_data( 367 path: Union[os.PathLike, str], 368 magnification: Literal["20x", "40x"] = "20x", 369 slides: Optional[Sequence[str]] = None, 370 download: bool = False, 371) -> str: 372 """Download the STHELAR dataset and convert its patches to hdf5 files. 373 374 NOTE: All slides need 18 GB (20x) or 54 GB (40x) for the download. Use `slides` for a subset. The parquet files 375 of the slides are kept, since a parquet file can contain the patches of several slides. 376 377 Args: 378 path: Filepath to a folder where the data is downloaded for further processing. 379 magnification: The choice of magnification. Either '20x' or '40x'. 380 slides: The slides to use, see `SLIDES` for the valid choices. By default all slides are used. 381 download: Whether to download the data if it is not present. 382 383 Returns: 384 Filepath where the converted patches are stored, in one sub-folder per slide. 385 """ 386 slides = _validate(magnification, slides) 387 data_dir = os.path.join(path, magnification) 388 preprocessed_dir = os.path.join(data_dir, "preprocessed") 389 390 pending = [s for s in slides if not os.path.exists(os.path.join(preprocessed_dir, s, COMPLETE_MARKER))] 391 if not pending: 392 return preprocessed_dir 393 394 os.makedirs(os.path.join(data_dir, "data"), exist_ok=True) 395 os.makedirs(os.path.join(data_dir, "cell_metadata"), exist_ok=True) 396 397 lookups = {} 398 for slide in pending: 399 relative_path = f"cell_metadata/{slide}_cell_metadata.parquet" 400 metadata_path = os.path.join(data_dir, relative_path) 401 util.download_source( 402 path=metadata_path, url=_resolve_url(magnification, relative_path), download=download, 403 checksum=METADATA_CHECKSUMS[magnification][slide], 404 ) 405 lookups[slide] = _load_lookup(metadata_path) 406 os.makedirs(os.path.join(preprocessed_dir, slide), exist_ok=True) 407 408 n_cpus = len(os.sched_getaffinity(0)) if hasattr(os, "sched_getaffinity") else (os.cpu_count() or 1) 409 n_workers = min(8, n_cpus) 410 411 shard_ids = sorted({i for slide in pending for i in SLIDE_SHARDS[magnification][slide]}) 412 for shard_id in shard_ids: 413 relative_path = f"data/{_shard_name(magnification, shard_id)}" 414 shard_path = os.path.join(data_dir, relative_path) 415 util.download_source( 416 path=shard_path, url=_resolve_url(magnification, relative_path), download=download, 417 checksum=SHARD_CHECKSUMS[magnification][shard_id], 418 ) 419 _convert_shard(shard_path, set(pending), lookups, preprocessed_dir, n_workers) 420 421 for slide in pending: 422 with open(os.path.join(preprocessed_dir, slide, COMPLETE_MARKER), "w"): 423 pass 424 425 return preprocessed_dir
Download the STHELAR dataset and convert its patches to hdf5 files.
NOTE: All slides need 18 GB (20x) or 54 GB (40x) for the download. Use slides for a subset. The parquet files
of the slides are kept, since a parquet file can contain the patches of several slides.
Arguments:
- path: Filepath to a folder where the data is downloaded for further processing.
- magnification: The choice of magnification. Either '20x' or '40x'.
- slides: The slides to use, see
SLIDESfor the valid choices. By default all slides are used. - download: Whether to download the data if it is not present.
Returns:
Filepath where the converted patches are stored, in one sub-folder per slide.
428def get_sthelar_paths( 429 path: Union[os.PathLike, str], 430 magnification: Literal["20x", "40x"] = "20x", 431 slides: Optional[Sequence[str]] = None, 432 download: bool = False, 433) -> List[str]: 434 """Get paths to the STHELAR data. 435 436 Args: 437 path: Filepath to a folder where the data is downloaded for further processing. 438 magnification: The choice of magnification. Either '20x' or '40x'. 439 slides: The slides to use, see `SLIDES` for the valid choices. By default all slides are used. 440 download: Whether to download the data if it is not present. 441 442 Returns: 443 List of filepaths for the hdf5 files, which contain the image ('image') and the labels 444 ('labels/instances' and 'labels/semantic') of a patch. 445 """ 446 slides = _validate(magnification, slides) 447 preprocessed_dir = get_sthelar_data(path, magnification, slides, download) 448 449 data_paths = [] 450 for slide in slides: 451 data_paths.extend(natsorted(glob(os.path.join(preprocessed_dir, slide, "*.h5")))) 452 453 assert len(data_paths) > 0 454 return data_paths
Get paths to the STHELAR data.
Arguments:
- path: Filepath to a folder where the data is downloaded for further processing.
- magnification: The choice of magnification. Either '20x' or '40x'.
- slides: The slides to use, see
SLIDESfor the valid choices. By default all slides are used. - download: Whether to download the data if it is not present.
Returns:
List of filepaths for the hdf5 files, which contain the image ('image') and the labels ('labels/instances' and 'labels/semantic') of a patch.
457def get_sthelar_dataset( 458 path: Union[os.PathLike, str], 459 patch_shape: Tuple[int, int], 460 magnification: Literal["20x", "40x"] = "20x", 461 slides: Optional[Sequence[str]] = None, 462 label_type: Literal["instances", "semantic"] = "instances", 463 resize_inputs: bool = False, 464 download: bool = False, 465 **kwargs 466) -> Dataset: 467 """Get the STHELAR dataset for nucleus segmentation and cell type classification. 468 469 Args: 470 path: Filepath to a folder where the data is downloaded for further processing. 471 patch_shape: The patch shape to use for training. 472 magnification: The choice of magnification. Either '20x' or '40x'. 473 slides: The slides to use, see `SLIDES` for the valid choices. By default all slides are used. 474 label_type: The choice of labels. Either 'instances' for the nucleus instances or 'semantic' for the 475 cell types of the nuclei, with the ids given by `CLASS_NAMES`. 476 resize_inputs: Whether to resize the inputs to the patch shape. 477 download: Whether to download the data if it is not present. 478 kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset`. 479 480 Returns: 481 The segmentation dataset. 482 """ 483 if label_type not in ("instances", "semantic"): 484 raise ValueError(f"'{label_type}' is not a valid label type. Choose either 'instances' or 'semantic'.") 485 486 data_paths = get_sthelar_paths(path, magnification, slides, download) 487 488 if resize_inputs: 489 resize_kwargs = {"patch_shape": patch_shape, "is_rgb": True} 490 kwargs, patch_shape = util.update_kwargs_for_resize_trafo( 491 kwargs=kwargs, patch_shape=patch_shape, resize_inputs=resize_inputs, resize_kwargs=resize_kwargs 492 ) 493 494 return torch_em.default_segmentation_dataset( 495 raw_paths=data_paths, 496 raw_key="image", 497 label_paths=data_paths, 498 label_key=f"labels/{label_type}", 499 patch_shape=patch_shape, 500 ndim=2, 501 with_channels=True, 502 **kwargs 503 )
Get the STHELAR dataset for nucleus segmentation and cell type classification.
Arguments:
- path: Filepath to a folder where the data is downloaded for further processing.
- patch_shape: The patch shape to use for training.
- magnification: The choice of magnification. Either '20x' or '40x'.
- slides: The slides to use, see
SLIDESfor the valid choices. By default all slides are used. - label_type: The choice of labels. Either 'instances' for the nucleus instances or 'semantic' for the
cell types of the nuclei, with the ids given by
CLASS_NAMES. - resize_inputs: Whether to resize the inputs to the patch shape.
- download: Whether to download the data if it is not present.
- kwargs: Additional keyword arguments for
torch_em.default_segmentation_dataset.
Returns:
The segmentation dataset.
506def get_sthelar_loader( 507 path: Union[os.PathLike, str], 508 batch_size: int, 509 patch_shape: Tuple[int, int], 510 magnification: Literal["20x", "40x"] = "20x", 511 slides: Optional[Sequence[str]] = None, 512 label_type: Literal["instances", "semantic"] = "instances", 513 resize_inputs: bool = False, 514 download: bool = False, 515 **kwargs 516) -> DataLoader: 517 """Get the STHELAR dataloader for nucleus segmentation and cell type classification. 518 519 Args: 520 path: Filepath to a folder where the data is downloaded for further processing. 521 batch_size: The batch size for training. 522 patch_shape: The patch shape to use for training. 523 magnification: The choice of magnification. Either '20x' or '40x'. 524 slides: The slides to use, see `SLIDES` for the valid choices. By default all slides are used. 525 label_type: The choice of labels. Either 'instances' for the nucleus instances or 'semantic' for the 526 cell types of the nuclei, with the ids given by `CLASS_NAMES`. 527 resize_inputs: Whether to resize the inputs to the patch shape. 528 download: Whether to download the data if it is not present. 529 kwargs: Additional keyword arguments for `torch_em.default_segmentation_dataset` or for the PyTorch DataLoader. 530 531 Returns: 532 The DataLoader. 533 """ 534 ds_kwargs, loader_kwargs = util.split_kwargs(torch_em.default_segmentation_dataset, **kwargs) 535 dataset = get_sthelar_dataset( 536 path, patch_shape, magnification, slides, label_type, resize_inputs, download, **ds_kwargs 537 ) 538 return torch_em.get_data_loader(dataset, batch_size, **loader_kwargs)
Get the STHELAR dataloader for nucleus segmentation and cell type classification.
Arguments:
- path: Filepath to a folder where the data is downloaded for further processing.
- batch_size: The batch size for training.
- patch_shape: The patch shape to use for training.
- magnification: The choice of magnification. Either '20x' or '40x'.
- slides: The slides to use, see
SLIDESfor the valid choices. By default all slides are used. - label_type: The choice of labels. Either 'instances' for the nucleus instances or 'semantic' for the
cell types of the nuclei, with the ids given by
CLASS_NAMES. - resize_inputs: Whether to resize the inputs to the patch shape.
- download: Whether to download the data if it is not present.
- kwargs: Additional keyword arguments for
torch_em.default_segmentation_datasetor for the PyTorch DataLoader.
Returns:
The DataLoader.