Bioinformatics analysis of the genes involved in the extension of prostate cancer to adjacent lymph nodes by supervised and unsupervised machine learning methods: The role of SPAG1 and PLEKHF2 https://www.sciencedirect.com/science/article/pii/S0888754320300963?via%3Dihub
@article{shamsara2020genomics,
author = {Shamsara, Elham and Shamsara, Jamal},
title = {Bioinformatics analysis of the genes involved in the extension of prostate cancer to adjacent lymph nodes by supervised and unsupervised machine learning methods: The role of {SPAG1} and {PLEKHF2}},
journal = {Genomics},
year = {2020},
volume = {112},
number = {6},
pages = {3871--3882},
doi = {10.1016/j.ygeno.2020.06.035}
}Built on mlmolprop, pinned to the
version this notebook was verified against: v1.1.0, on which Prad.ipynb
runs end-to-end with no errors. Later releases have not been re-verified, so
the pin is what reproduces the published results.
pip install "mlmolprop @ git+https://github.com/shamsaraj/mlmolprop.git@v1.1.0"
pip install seabornRequires Python >= 3.12. seaborn is used directly in the notebook for the
composition histogram and the clustermap.
data_Prad/ has the real data behind the analysis, from the paper's
Research Data
supplementary file on ScienceDirect (489 TCGA-PRAD patients, clinical
fields plus 45,606 expression/mutation columns):
data_curated3-only_Mtrans.csv: clinical columns (PATIENT_ID,PATH_N_STAGE,PATH_T_STAGE,PRIOR_DX,RADIATION_THERAPY) plus the ~20,000M_-prefixed expression columns -- the input the notebook's (commented-out)data_prep()call was built from.M_1000_T_noscale.data/M_30_T_noscale.data: the pickled, already-processed train/test splits thatPrad.ipynbactually loads viafile2object(mlmolprop.data_prep, mean-imputed,PATH_N_STAGE == 3rows dropped, top 1000/30 features by ANOVA F-test,TARGET=PATH_N_STAGE,rs=10). Regenerate them by uncommenting thedata_prep()call in the second code cell.
The supplementary file's other two tables -- the full combined table
(data_curated_trans.csv, 127 MB, over GitHub's per-file limit) and its
non-M_-prefixed complement -- aren't used by this notebook and aren't
included here; get them directly from the Research Data link above if
needed.
The data is licensed CC BY-NC 3.0 (per the Research Data link above) -- attribution required (see the citation above), non-commercial use only. This is separate from and stricter than the license on the code in this repository.