Using Causal Graph Model variable selection for BERT models Prediction of Patient Survival in a Clinical Text Discharge Dataset

Authors

DOI:

https://doi.org/10.62411/faith.3048-3719-61

Keywords:

BERT prediction, BERT prediction comparison, Causal DAG, Clinical text analysis, Predictor selection

Abstract

Feature selection in most black-box machine learning algorithms, such as BERT, is based on the cor-relations between features and the target variable rather than causal relationships in the dataset. This makes their predictive power and decisions questionable because of their potential bias. This paper presents novel BERT models that learn from causal variables in a clinical discharge dataset. The causal-directed acyclic Graphs (DAG) identify input variables for patients’ survival rate prediction and decisions. The core idea behind our model lies in the ability of the BERT-based model to learn from the causal DAG semi-synthetic dataset, enabling it to model the underlying causal structure accurately in-stead of the generic spurious correlations devoid of causation. The results from Causal DAG Conditional Independence Test (CIT) validation metrics showed that the conceptual assumptions of the causal DAG were supported, the Pearson correlation coefficient ranges between -1 and 1, the p-value was (>0.05), and the confidence interval of 95% and 25% were satisfied. We further mapped the semi-synthetic dataset that evolved from the Causal DAG to three BERT models. Two metrics, pre-diction accuracy, and AUC score, were used to compare the performance of the BERT models. The accuracy of the BERT models showed that the regular BERT has a performance of 96%, while Clinical-BERT performance was 90%, and Clinical-BERT-Discharge-summary was 92%. On the other hand, the AUC score for BERT was 79%, ClinicalBERT was 77%, while ClinicalBERT-discharge summary was 84%. Our experiments on the synthetic dataset for the patient’s survival rate from the causal DAG datasets demonstrate high predictive performance and explainable input variables for human under-standing to justify prediction.

Downloads

Download data is not yet available.

Author Biographies

Omachi Okolo, Modibbo Adama University

Department of Information Technology, Modibbo Adama University, Yola, Nigeria

B. Y. Baha, Modibbo Adama University

Department of Information Technology, Modibbo Adama University Yola, Nigeria

M.D. Philemon, Modibbo Adama University

Department of Information Technology, Modibbo Adama University Yola, Nigeria

References

G. T. Ayem, O. Asilkan, and A. Iorliam, “Design and Validation of Structural Causal Model: A Focus on EGRA Dataset,” J. Comput. Theor. Appl., vol. 1, no. 2, pp. 86–103, Nov. 2023, doi: 10.33633/jcta.v1i2.9304.

K. Yu et al., “Causality-based Feature Selection,” ACM Comput. Surv., vol. 53, no. 5, pp. 1–36, Sep. 2021, doi: 10.1145/3409382.

A. Feder, N. Oved, U. Shalit, and R. Reichart, “CausaLM: Causal Model Explanation Through Counterfactual Language Models,” Comput. Linguist., vol. 47, no. 2, pp. 333–386, 2021, doi: 10.1162/coli_a_00404.

K. A. Keith, D. Jensen, and B. O’Connor, “Text and Causal Inference: A Review of Using Text to Remove Confounding from Causal Estimates,” ArXiv. May 01, 2020. [Online]. Available: https://arxiv.org/abs/2005.00649

R. Guidotti, A. Monreale, S. Ruggieri, F. Turini, F. Giannotti, and D. Pedreschi, “A Survey of Methods for Explaining Black Box Models,” ACM Comput. Surv., vol. 51, no. 5, pp. 1–42, Sep. 2019, doi: 10.1145/3236009.

M. J. Vowels, “Trying to outrun causality with machine learning: Limitations of model explainability techniques for exploratory research.,” Psychol. Methods, Sep. 2024, doi: 10.1037/met0000699.

A. Molak and A. Jaokar, Causal Inference and Discovery in Python: Unlock the secrets of modern causal machine learning with DoWhy, EconML, PyTorch and more. Packt Publishing, 2023. [Online]. Available: http://ieeexplore.ieee.org/document/10251331

K. Lyu and others, “Causal knowledge graph construction and evaluation for clinical decision support of diabetic nephropathy,” J. Biomed. Inform., vol. 139, p. 104298, 2023, doi: 10.1016/j.jbi.2023.104298.

M. Piccininni, S. Konigorski, J. L. Rohmann, and T. Kurth, “Directed acyclic graphs and causal thinking in clinical risk prediction modeling,” BMC Med. Res. Methodol., vol. 20, no. 1, p. 179, Dec. 2020, doi: 10.1186/s12874-020-01058-z.

M. Liu, D. R. Bellamy, and A. L. Beam, “DAG-aware Transformer for Causal Effect Estimation,” ArXiv. Oct. 13, 2024. [Online]. Available: https://arxiv.org/abs/2410.10044

J. Zhang, J. Jennings, A. Hilmkil, N. Pawlowski, C. Zhang, and C. Ma, “Towards Causal Foundation Model: on Duality between Causal Inference and Attention,” ArXiv. Oct. 01, 2023. [Online]. Available: https://arxiv.org/abs/2310.00809

J. Medori and C. Fairon, “Machine learning and features selection for semi-automatic ICD-9-CM encoding,” in Proceedings of the NAACL HLT 2010 Second Louhi Workshop on Text and Data Mining of Health Documents, 2010, pp. 84–89. [Online]. Available: https://aclanthology.org/W10-1113/

B. I. Igoche, O. Matthew, P. Bednar, and A. Gegov, “Integrating Structural Causal Model Ontologies with LIME for Fair Machine Learning Explanations in Educational Admissions,” J. Comput. Theor. Appl., vol. 2, no. 1, pp. 65–85, Jun. 2024, doi: 10.62411/jcta.10501.

K. Huang, J. Altosaar, and R. Ranganath, “ClinicalBERT: Modeling Clinical Notes and Predicting Hospital Readmission,” ArXiv. Apr. 10, 2019. [Online]. Available: http://arxiv.org/abs/1904.05342

E. Alsentzer et al., “Publicly available clinical BERT embeddings,” in Proceedings of the 2nd Clinical Natural Language Processing Workshop, 2019, pp. 72–78. [Online]. Available: https://aclanthology.org/W19-1909/

S. Gopalakrishnan, V. Z. Chen, W. Dou, G. Hahn-Powell, S. Nedunuri, and W. Zadrozny, “Text to Causal Knowledge Graph: A Framework to Synthesize Knowledge from Unstructured Business Texts into Causal Graphs,” Information, vol. 14, no. 7, p. 367, Jun. 2023, doi: 10.3390/info14070367.

S. Khanna, “A Comprehensive Guide to Train-Test-Validation Split in 2024,” Analytics Vidhya, 2024. https://www.analyticsvidhya.com/back-channel/download-pdf.php?pid=134366&next=

R. Pryzant, D. Card, D. Jurafsky, V. Veitch, and D. Sridhar, “Causal Effects of Linguistic Properties,” ArXiv. Oct. 24, 2020. [Online]. Available: http://arxiv.org/abs/2010.12919

LLM, Large Language Models (LLMs) Interview Question. Medium, 2024. [Online]. Available: https://masteringllm.medium.com/recent-11-large-language-models-llms-interview-questions-

A. Turchin, S. Masharsky, and M. Zitnik, “Comparison of BERT implementations for natural language processing of narrative medical documents,” Informatics Med. Unlocked, vol. 36, p. 101139, 2023, doi: 10.1016/j.imu.2022.101139.

H. Alkattan, S. K. Towfek, and M. Y. Shams, “Tapping into Knowledge: Ontological Data Mining Approach for Detecting Cardiovascular Disease Risk Causes Among Diabetes Patients,” J. Artif. Intell. Metaheuristics, vol. 4, no. 1, pp. 08–15, 2023, doi: 10.54216/JAIM.040101.

A. Ankan, I. M. N. Wortel, and J. Textor, “Testing Graphical Causal Models Using the R Package ‘dagitty,’” Curr. Protoc., vol. 1, no. 2, Feb. 2021, doi: 10.1002/cpz1.45.

A. S. Maiya, “CausalNLP: A Practical Toolkit for Causal Inference with Text,” ArXiv. Computer Science - Computation and Language, Jun. 15, 2021. [Online]. Available: http://arxiv.org/abs/2106.08043

V. Veitch, D. Sridhar, and D. M. Blei, “Adapting Text Embeddings for Causal Inference,” in Conference on Uncertainty in Artificial Intelligence, 2020, May 2019, pp. 919–928. [Online]. Available: http://arxiv.org/abs/1905.12741

S. Sheikholeslami, “Ablation Programming for Machine Learning,” KTH-Royal Institute of Technology, 2019. [Online]. Available: https://www.diva-portal.org/smash/get/diva2:1349978/FULLTEXT01.pdf

X. Shen, S. Ma, P. Vemuri, M. R. Castro, P. J. Caraballo, and G. J. Simon, “A novel method for causal structure discovery from EHR data and its application to type-2 diabetes mellitus,” Sci. Rep., vol. 11, no. 1, p. 21025, Oct. 2021, doi: 10.1038/s41598-021-99990-7.

C. Si, W. Chen, W. Wang, L. Wang, and T. Tan, “An Attention Enhanced Graph Convolutional LSTM Network for Skeleton-Based Action Recognition,” in 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Jun. 2019, pp. 1227–1236. doi: 10.1109/CVPR.2019.00132.

Downloads

Published

2025-03-10

How to Cite

[1]
O. Okolo, B. Y. Baha, and M. Philemon, “Using Causal Graph Model variable selection for BERT models Prediction of Patient Survival in a Clinical Text Discharge Dataset”, J. Fut. Artif. Intell. Tech., vol. 1, no. 4, pp. 455–473, Mar. 2025.

Issue

Section

Articles

Similar Articles

1 2 3 4 5 6 7 8 > >> 

You may also start an advanced similarity search for this article.