Skip to content

Latest commit

 

History

14 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 

Repository files navigation

Gene-expression-analysis, The codes are related to the following publication

Bioinformatics analysis of the genes involved in the extension of prostate cancer to adjacent lymph nodes by supervised and unsupervised machine learning methods: The role of SPAG1 and PLEKHF2 https://www.sciencedirect.com/science/article/pii/S0888754320300963?via%3Dihub

@article{shamsara2020genomics,
  author  = {Shamsara, Elham and Shamsara, Jamal},
  title   = {Bioinformatics analysis of the genes involved in the extension of prostate cancer to adjacent lymph nodes by supervised and unsupervised machine learning methods: The role of {SPAG1} and {PLEKHF2}},
  journal = {Genomics},
  year    = {2020},
  volume  = {112},
  number  = {6},
  pages   = {3871--3882},
  doi     = {10.1016/j.ygeno.2020.06.035}
}

Built on mlmolprop, pinned to the version this notebook was verified against: v1.1.0, on which Prad.ipynb runs end-to-end with no errors. Later releases have not been re-verified, so the pin is what reproduces the published results.

pip install "mlmolprop @ git+https://github.com/shamsaraj/mlmolprop.git@v1.1.0"
pip install seaborn

Requires Python >= 3.12. seaborn is used directly in the notebook for the composition histogram and the clustermap.

Data

data_Prad/ has the real data behind the analysis, from the paper's Research Data supplementary file on ScienceDirect (489 TCGA-PRAD patients, clinical fields plus 45,606 expression/mutation columns):

  • data_curated3-only_Mtrans.csv: clinical columns (PATIENT_ID, PATH_N_STAGE, PATH_T_STAGE, PRIOR_DX, RADIATION_THERAPY) plus the ~20,000 M_-prefixed expression columns -- the input the notebook's (commented-out) data_prep() call was built from.
  • M_1000_T_noscale.data / M_30_T_noscale.data: the pickled, already-processed train/test splits that Prad.ipynb actually loads via file2object (mlmolprop.data_prep, mean-imputed, PATH_N_STAGE == 3 rows dropped, top 1000/30 features by ANOVA F-test, TARGET=PATH_N_STAGE, rs=10). Regenerate them by uncommenting the data_prep() call in the second code cell.

The supplementary file's other two tables -- the full combined table (data_curated_trans.csv, 127 MB, over GitHub's per-file limit) and its non-M_-prefixed complement -- aren't used by this notebook and aren't included here; get them directly from the Research Data link above if needed.

The data is licensed CC BY-NC 3.0 (per the Research Data link above) -- attribution required (see the citation above), non-commercial use only. This is separate from and stricter than the license on the code in this repository.

About

Supervised/unsupervised ML on TCGA-PRAD expression data, identifying SPAG1/PLEKHF2 in prostate-cancer lymph-node extension

Topics

Resources

Stars

0 stars

Watchers

2 watching

Forks

Releases

Packages

Contributors

Languages