Automated Feature Engineering for Large-Scale Healthcare Data

Authors

  • Tanja Schultz Author
    Competing Interests

    AI,ML

DOI:

https://doi.org/10.5281/zenodo.22172954

Keywords:

Clinical Feature Engineering, Automated Feature Engineering Systems, Healthcare Machine Learning, Predictive Model Pipelines, Electronic Health Records (EHR), Multimodal Clinical Data, Data Governance in Healthcare AI, Privacy and Regulatory Compliance, Prospective Model Validation, Bias-Aware Feature Design, Data Quality and Preprocessing Pipelines, Candidate Feature Generation and Selection, Feature Scoring Strategies, Reproducible Clinical ML, Biosensor and Time-Series Data, Imaging-Derived Biomarkers, Genomic and Sequencing Features, Point-of-Care Laboratory Data, Clinical Model Generalization, Production-Grade Healthcare AI.

Abstract

A rapidly proliferating corpus of clinical research harnessing the power of machine learning has substantial implications for healthcare feature engineering. As a broad umbrella encompassing data preprocessing, quality control, transformation, and generation, feature engineering addresses a major bottleneck in the production of predictive models. Automated feature engineering systems are increasingly deployed at scale to meet the challenges of generating the vast quantity of predictive features necessary for successful, generalizable, and clinically useful machine learning systems. Such clinical feature engineering systems produce features that are applied in a predictive setting after the fact and not explicitly linked to clinical care, but nevertheless involve substantial risk. A principled examination of a clinical feature engineering system can be framed in terms of six core components: data governance; compliance with privacy and regulatory constraints; appropriate validation and prospective evaluation; consideration of data biases; the use of effective data-quality and preprocessing pipelines; and sound candidate feature generation, scoring, and selection strategies. Health systems typically possess an assemblage of rich and diverse, yet underutilized, information with the potential to contribute meaningfully to clinical prediction problems. Electronic health record (EHR) data, comprising clinical notes, laboratory values, medication orders, and procedure codes; over a decade’s worth of length and width dataset and point-of-care laboratory test results; continuous biosensor measurements; DNA sequencing data; biomarkers derived from imaging; and drug compounds targeting genotypes provide raw material for hundreds of prediction problems in diverse specialties. However, machine learning in healthcare exhibits a stunning lack of reproducibility: many predictive models fail to retain their accuracy in different cohorts, and those that do are seldom incorporated into routine clinical care. A significant bottleneck underlying this failure lies with the feature engineering step.

References

1. Rashidi, H. H., Tran, N., Albahra, S., & Dang, L. T. (2021). Machine learning in health care and laboratory medicine: General overview of supervised learning and Auto-ML. International Journal of Laboratory Hematology, 43(S1), 15–22.

2. Brnabic, A., & Hess, L. M. (2021). Systematic literature review of machine learning methods used in the analysis of real-world data for patient-provider decision making. BMC Medical Informatics and Decision Making, 21, 54.

3. Pamisetty, A., Paleti, S., Adusupalli, B., Singireddy, J., Inala, R., & Nagabhyru, K. C. (2025). Explainable AI Systems for Credit Scoring and Loan Risk Assessment in Digital Banking Platforms. In 2025 IEEE 13th International Conference on Intelligent Data Acquisition and Advanced Computing Systems: Technology and Applications (IDAACS) (pp. 1478–1483). IEEE. 2025 IEEE 13th International Conference on Intelligent Data Acquisition and Advanced Computing Systems: Technology and Applications (IDAACS). https://doi.org/10.1109/idaacs68557.2025.11322144

4. Chen, Y., Liu, X., Liu, Y., & others. (2021). Benchmarking feature selection methods with different prediction models on large-scale healthcare event data. BenchCouncil Transactions on Benchmarks, Standards and Evaluations, 2(4), 100004.

5. Wang, Y., Li, Y., & others. (2021). Curvature-based feature selection with application in classifying electronic health records. Technological Forecasting and Social Change, 173, 121127.

6. Zöller, M.-A., & Huber, M. F. (2021). Benchmark and survey of automated machine learning frameworks. Journal of Artificial Intelligence Research, 70, 409–472.

7. Feurer, M., & Hutter, F. (2021). Hyperparameter optimization. In F. Hutter, L. Kotthoff, & J. Vanschoren (Eds.), Automated machine learning: Methods, systems, challenges (pp. 3–33). Springer.

8. Alshar, M. M., Shahdadpuri, N., Rajeshwari, M., Gupta, M., Joshi, N. R., & Singireddy, J. (2025). Enhanced Management & Performance of Remote Workforce with Cloud and AI-Driven HR Analytics. In 2025 3rd International Conference on Advances in Computation, Communication and Information Technology (ICAICCIT) (pp. 631–636). IEEE. 2025 3rd International Conference on Advances in Computation, Communication and Information Technology (ICAICCIT). https://doi.org/10.1109/icaiccit68829.2025.11434104

9. Liu, B., Wang, Y., & others. (2021). Evolving fully automated machine learning via life-long knowledge anchors. IEEE Transactions on Pattern Analysis and Machine Intelligence, 44(9), 5310–5324.

10. Karmaker, S. K., Hassan, M. M., Smith, M. J., Xu, L., Zhai, C., & Veeramachaneni, K. (2021). Automl-kg: A knowledge graph approach for automated machine learning. Proceedings of the AAAI Conference on Artificial Intelligence, 35(17), 15597–15604.

11. Chary, D. V., Meda, R., C, J. S. Mary., Narasimhachari, J. P., & A S, Y. (2025). TriFusionFormer: Tri-Modal Fusion Transformer Using Gated Modality Control and Multi-Scale Attention for Emotion Recognition. In 2025 International Conference on Communication, Computer, and Information Technology (IC3IT) (pp. 1–8). IEEE. 2025 International Conference on Communication, Computer, and Information Technology (IC3IT). https://doi.org/10.1109/ic3it66137.2025.11341646

12. de Sá, A. G. C., Pinto, W. J. G., de Oliveira, L. O. V. B., & Pappa, G. L. (2021). RECIPE: A grammar-based framework for automatically generating recommendations for feature engineering. Knowledge-Based Systems, 226, 107156.

13. Thirunavukarasu, A. J., Elangovan, K., Gutierrez, L., Hassan, R., Li, Y., Tan, T. F., Cheng, H., Teo, Z. L., Lim, G., & Ting, D. S. W. (2022). Clinical performance of automated machine learning: A systematic review. Annals of the Academy of Medicine, Singapore, 53(3), 187–207.

14. Naik, A. V., Sheelam, G. K., Panchakatla, N., Muthukumaran, K., & Saranya, K. (2025). Comprehensive Analysis on Depression Detection From Social Media Using Deep Learning and Transformer Architectures. In 2025 International Conference on Communication, Computer, and Information Technology (IC3IT) (pp. 1–8). IEEE. 2025 International Conference on Communication, Computer, and Information Technology (IC3IT). https://doi.org/10.1109/ic3it66137.2025.11341160

15. Romero, R. A. A., Deypalan, M. N. Y., Mehrotra, S., Jungao, J. T., Sheils, N. E., Manduchi, E., & Moore, J. H. (2022). Benchmarking AutoML frameworks for disease prediction using medical claims. BioData Mining, 15, 15.

16. Lagani, V., Athineou, G., Farcomeni, A., Tsagris, M., & Tsamardinos, I. (2022). Just Add Data: Automated predictive modeling for knowledge discovery and feature selection. npj Precision Oncology, 6, 38.

17. Krishnan, M., Nandan, B. P., Rongali, S. K., Meda, R., Kalisetty, S., & Singireddy, J. (2026, June). AI-Driven Data Engineering and Predictive Analytics Framework for Semiconductor Supply Chain Optimization and Digital Infrastructure Modernization. In 2026 6th International Conference on Intelligent Technologies (CONIT) (pp. 1-6). IEEE.

18. Musigmann, M., Akkurt, B. H., Krähling, H., Nacul, N. G., Remonda, L., Sartoretti, T., Henssen, D., Brokinkel, B., Stummer, W., Heindel, W., & Mannil, M. (2022). Testing the applicability and performance of AutoML for potential applications in diagnostic neuroradiology. Scientific Reports, 12, 13648.

19. Siriborvornratanakul, T. (2022). Human behavior in image-based road health inspection systems despite the emerging AutoML. Journal of Big Data, 9, 96.

20. Challa, K., Challa, S. R., Pamisetty, A., Kaulwar, P. K., & Koppolu, H. K. R. (2025, December). Transforming Payments: The Role of AI and Big Data in Fraud Alerts, Credit Monitoring, and Secure Transactions. In 2025 IEEE International Conference on Communication Networks and Computing (CNC) (pp. 1406-1414). IEEE.

21. Bertsimas, D., Dunn, J., Velmahos, G. C., & Kaafarani, H. M. A. (2022). Surgical risk is not linear: Predicting postoperative complications using machine learning. Annals of Surgery, 275(2), 270–278.

22. Johnson, A. E. W., Bulgarelli, L., Shen, L., Gayles, A., Shukla, K., Horng, S., Pollard, T. J., Hao, S., Moody, B., Gow, B., Lehman, L.-w. H., Celi, L. A., & Mark, R. G. (2023). MIMIC-IV, a freely accessible electronic health record dataset. Scientific Data, 10, 1.

23. Vadisetty, R., Nuka, S. T., Kalisetty, S., Pandugula, C., Burugulla, J. K. R., & Annapareddy, V. N. (2026). and Polypropylene Manufacturing. In Proceedings of Sixth International Ethical Hacking Conference: AI and Law (eHaCON 2025) (p. 269). Springer Nature.

24. Imrie, F., Cebere, B., McKinney, E. F., & van der Schaar, M. (2023). AutoPrognosis 2.0: Democratizing diagnostic and prognostic modeling in healthcare with automated machine learning. PLOS Digital Health, 2(6), e0000276.

25. Stevens, C. A. T., Lyons, A. R. M., Dharmayat, K. I., Mahani, A., Ray, K. K., Vallejo-Vaz, A. J., & Sharabiani, M. T. A. (2023). Ensemble machine learning methods in screening electronic health records: A scoping review. Digital Health, 9, 20552076231173225.

26. Rajamanickam, V., Singireddy, S., Davuluri, P. N., Sheelam, G. K., Aitha, A. R., & Vakkalagadda, T. (2026). AI-Enabled Visual Evidence Intelligence for Detecting Manipulated Digital Media. International Journal of Special Education, 41(14s), 115-124.

27. Gronsbell, J. L., Minnier, J., Yu, J., & others. (2023). Machine learning approaches for electronic health records phenotyping: A methodical review. Journal of the American Medical Informatics Association, 30(2), 334–347.

28. Kolasa, K., Admassu, B., Hołownia-Voloskova, M., Kędzior, K. J., Poirrier, J.-E., & Perni, S. (2024). Systematic reviews of machine learning in healthcare: A literature review. Expert Review of Pharmacoeconomics & Outcomes Research, 24(1), 63–115.

29. Yuan, H., Yu, K., Xie, F., Liu, M., & Sun, S. (2024). Automated machine learning with interpretation: A systematic review of methodologies and applications in healthcare. Medicine Advances, 2, e75.

30. Segireddy, A. R., Nagabhyru, K. C., Gadi, A. L., Pandiri, L., Paleti, S., Nandan, B. P., ... & Meda, R. (2026). U.S. Patent Application No. 19/389,116.

31. Yuan, H., Yu, K., Xie, F., Liu, M., & Sun, S. (2024). Human-in-the-loop machine learning for healthcare: Current progress and future opportunities in electronic health records. Medicine Advances, 2, e70.

32. Tayebi Arasteh, S., Han, T., Lotfinia, M., Kuhl, C., Kather, J. N., Truhn, D., & Nebelung, S. (2024). Large language models streamline automated machine learning for clinical studies. Nature Communications, 15, 1603.

33. Maguluri, K. K. (2026). Cloud-Integrated Machine Learning System for Ebola Virus Disease Prediction and Epidemic Intelligence in Smart Healthcare Systems. Journal of Advances in Management, Engineering and Science (JAMES), 1(03), 1-9.

34. Barras, M., & others. (2024). Evaluating automated machine learning platforms for use in healthcare. JAMIA Open, 7(2), ooae031.

35. An, J., Kim, I. S., Kim, K.-J., Park, J. H., Kang, H., Kim, H. J., Kim, Y. S., & Ahn, J. H. (2024). Efficacy of automated machine learning models and feature engineering for diagnosis of equivocal appendicitis using clinical and computed tomography findings. Scientific Reports, 14, 22658.

36. Davuluri, P. S. L. (2023). AI-Augmented Sanctions Screening: Enhancing Accuracy and Latency in Real Time Compliance Systems. AI-Augmented Sanctions Screening: Enhancing Accuracy and Latency in Real Time Compliance Systems (December 15, 2023).

37. Bandi, V. D. V. K. (2024). Automated feature engineering systems in large-scale healthcare data environments. Journal of Neonatal Surgery, 13(1), 2127–2141.

38. Liu, Z., Zhang, D., Liu, H., Dong, Z., Jia, W., & Tan, J. (2024). Visible-hidden hybrid automatic feature engineering via multi-agent reinforcement learning. Knowledge-Based Systems, 299, 111941.

39. Wang, H., Zhang, M., Mai, L., Li, X., Bellou, A., & Wu, L. (2025). An effective multi-step feature selection framework for clinical outcome prediction using electronic medical records. BMC Medical Informatics and Decision Making, 25, 84.

40. Bedi, B., Yandamuri, U. S., Kummari, D. N., Nagubandi, A. R., & Amistapuram, K. (2026, April). AI-Driven Sentiment and Behavior Analysis for Sustainable Business Growth. In 2026 International Conference on Emerging Research in Smart Electronics and Machine Informatics (ECMI) (pp. 1-11). IEEE.

41. Voskergian, D., Bakir-Gungor, B., & Yousef, M. (2025). Engineering novel features for diabetes complication prediction using synthetic electronic health records. Frontiers in Genetics, 16, 1451290.

42. Wyss, R., Yang, J., Schneeweiss, S., Plasek, J. M., Zhou, L., Deramus, T., Weberpals, J. G., Ngan, K., Tsacogianis, T. N., & Lin, K. J. (2025). Natural language processing for scalable feature engineering and ultra-high-dimensional confounding adjustment in healthcare database studies. Journal of Biomedical Informatics, 169, 104882.

43. de Winter, C., Frasincar, F., de Peuter, B., Matsiiako, V., Ido, E., & Klinkhamer, J. (2025). Automated feature engineering for automated machine learning. Knowledge-Based Systems, 321, 113671.

44. Vancalster, B., & others. (2025). Challenges and recommendations for electronic health records data extraction and preparation for dynamic prediction modeling in hospitalized patients: Practical guide and tutorial. Journal of Medical Internet Research, 27, e73987.

45. El Shawi, R., & Jamel, L. (2025). Leveraging ChatGPT and explainable AI for enhancing clinical decision support. Scientific Reports, 15, 38786.

46. Vankayalapati, R. K., Polineni, T. N. S., Ahammad, S. H., Pandugula, C., & Selvan, R. S. (2026). IoT-Enabled Augmented Reality for Real-Time Equipment Diagnosis. In Virtual Reality, Real Emergency (pp. 104-121). CRC Press.

47. Wang, H., Zhang, M., Mai, L., Li, X., Bellou, A., & Wu, L. (2025). Feature selection and machine learning for high-dimensional electronic medical records: Clinical prediction and interpretability considerations. BMC Medical Informatics and Decision Making, 25.

48. Francia, R., Leone, M., Leonardi, G., Montani, S., Pennisi, M., Striani, M., & D'Alfonso, S. (2025). AutoML-Med: A framework for automated machine learning in medical tabular data. arXiv.

49. Zhang, R., Yang, Q., Wang, X., Wang, T., Zhou, Q., Deng, Z., Li, K., Wang, Y., Fan, Y., Zhang, J., Huang, L., Liu, C., & Zhou, F. (2026). DeepSelective: Interpretable prognosis prediction via feature selection and compression in EHR data. Pattern Recognition, 174, 112970.

50. Zhang, Y., Liu, X., Chen, J., & others. (2023). Predicting disease onset from electronic health records for population health management: A scalable and explainable deep learning approach. Frontiers in Artificial Intelligence, 6.

51. Bellou, A., Li, X., Wu, L., & others. (2025). Explainable feature selection for high-dimensional clinical outcome prediction using electronic medical records. BMC Medical Informatics and Decision Making.

52. Inala, R., Garapati, R. S., Aitha, A. R., Komaragiri, V. B., Gottimukkala, V. R. R., Recharla, M., ... & Varri, D. B. S. (2026). U.S. Patent Application No. 19/389,108.

53. Sun, Y., Wang, Y., & others. (2023). Machine learning for clinical prediction using large-scale electronic health record data: Methods, challenges, and opportunities. Journal of Biomedical Informatics, 139.

54. Chen, R. J., Lu, M. Y., Williamson, D. F. K., & others. (2024). Towards a general-purpose foundation model for computational pathology. Nature Medicine, 30, 850–862.

55. Moor, M., Banerjee, O., Abad, Z. S. H., Krumholz, H. M., Leskovec, J., Topol, E. J., & Rajpurkar, P. (2023). Foundation models for generalist medical artificial intelligence. Nature, 616, 259–265.

56. Acosta, J. N., Falcone, G. J., Rajpurkar, P., & Topol, E. J. (2022). Multimodal biomedical AI. Nature Medicine, 28, 1773–1784.

57. Rajpurkar, P., Chen, E., Banerjee, O., & Topol, E. J. (2022). AI in health and medicine. Nature Medicine, 28, 31–38.

58. Kather, J. N., Calderaro, J., & others. (2021). Artificial intelligence in cancer diagnosis and prognosis: Current applications and future perspectives. Nature Reviews Cancer, 21, 747–762.

59. Topol, E. J. (2024). High-performance medicine: The convergence of human and artificial intelligence. Nature Medicine, 30, 1–7.

Additional Files

Published

2026-03-13

Data Availability Statement

None

How to Cite

Automated Feature Engineering for Large-Scale Healthcare Data. (2026). The American Journal of Analytics and Artificial Intelligence (AJAAI), 4(01). https://doi.org/10.5281/zenodo.22172954

Most read articles by the same author(s)

Similar Articles

1-10 of 38

You may also start an advanced similarity search for this article.