Research
On-Going Research
TIME-TO-EVENT PREDICTION OF CLINICAL EVENTS IN HEART FAILURE PATIENTS USING LONGITUDINAL EHR DATA
Faculty: Dr. Jose L. Zayas-Castro
Ph.D. student: Maryam Jafaripakzad
Heart failure affects over 6.7 million Americans and represents the leading cause of hospitalization with a 5-year mortality rate approaching 50%. Predicting adverse clinical events such as hospitalization, heart transplant, LVAD implantation, and mortality remains a critical challenge in cardiovascular care. This research develops a time-to-event prediction framework using longitudinal Electronic Health Record (EHR) data to forecast severe clinical outcomes in heart failure patients. The approach leverages Temporal Convolutional Networks (TCN) to model the complex, multivariate, and irregularly sampled nature of clinical data. Patient encounters, including clinic visits and hospitalizations, are aggregated into 3-month intervals (calendar quarters), capturing temporal patterns in demographics, medical history, family history, medications, laboratory results, vital signs, and diagnostic test outcomes. A novel Hybrid Fusion TCN (HF-TCN) architecture processes both time-varying sequences and time-invariant patient characteristics through parallel pathways, merging them via a fusion layer to enhance predictive accuracy. To address the highly imbalanced nature of adverse event data, a multi-branch training strategy is employed. This work aims to support proactive clinical decision-making and personalized risk stratification for heart failure management.

Figure: Longitudinal EHR Data Analysis Pipeline with Multi-Branch Hybrid Fusion TCN
Balancing Equity and Efficiency in Kidney Allocation
Faculty: Dr. Jose L. Zayas-Castro
PhD student: Daniela Cantarino
In 2022, 88,901 patients waited for a kidney transplant in the US while only 25,499 received one. The current allocation system faces a persistent trade-off: prioritizing post-transplant survival (efficiency) may overlook critically ill patients, while emphasizing medical urgency (equity) can reduce long-term graft outcomes. US recipients also face 25% higher graft-failure rates than comparable systems. Cold ischemia time during organ transport adds further logistical complexity.
KEY CONTRIBUTION
This paper introduces an intelligent decision-support system that integrates stochastic optimization and game-theoretic reasoning to balance equity and efficiency under uncertainty. It is the first multi-objective kidney allocation model that simultaneously balances urgency and long-term outcomes while controlling regional graft-failure risk. The framework operates at national scale using comprehensive OPTN/SRTR data covering 96,321 candidates and 1,837 deceased-donor kidneys.
METHODOLOGY

Graph of National Cumulative waitlist additions by week and year
1. Bi-Objective Model: Maximizes expected post-transplant survival (Obj. 1) and pre-transplant mortality risk capture (Obj. 2) subject to medical compatibility, transport logistics, and cold ischemia time constraints.
2. Chance Constraints: Regional graft-failure risk is bounded probabilistically using KDPI distributions, ensuring failure rates do not exceed current benchmarks with tunable confidence.
3. Nash Bargaining Solution: Selects a fair compromise on the Pareto frontier by maximizing the product of each objective鈥檚 improvement over the status quo, yielding balanced equity-efficiency weights.
4. Online Two-Stage Stochastic Program: Assigns each newly available kidney in real time using NBS-derived weights, historical allocations, and scenario-based forecasts of future organ supply (low/medium/high).

Interpretability of Data Structures with Clusterability
Faculty: Kaixun Hua
Current clustering algorithms face issues such as erroneous results with datasets lacking a clear clustering structure, poor performance with complex datasets, and scalability problems with large datasets. To overcome these, we proposed a clusterability index, linked to ultrametricity measurement, to gauge if a dataset exists meaningful clusters. We are developing a special minmax-based matrix multiplication approach for the dissimilarity matrix, enabling it to converge to an ultrametric dissimilarity matrix via power operations. The clusterability degree is determined by the minimum power operations needed for convergence. This method, applicable to any dissimilarity matrix, provides a value in the final matrix that signifies the minimum edge between two nodes in a weighted graph. The resulting ultrametric distance matrix effectively captures complex structures in challenging datasets, enhancing the performance of centroid-based clustering algorithms.

Explainable AI with Scalable Deterministic Global Optimal Training
Faculty: Kaixun Hua

Given the NP-hard nature of numerous machine learning (ML) challenges like clustering, decision trees, and neural networks, it's often considered that achieving global optimality in ML tasks is computationally prohibitive. Consequently, prevalent methods lean on either heuristic or local optimization algorithms, resulting in less than optimal solutions. Another prevalent assumption in the ML field suggests that in the big data era, simple, highly interpretable models (like decision trees) can't match the predictive prowess of black-box models. However, in our project, we tend to construct a set of scalable training algorithms via a reduced-space branch and bound framework offering guaranteed global optimality, and thus challenging these assumptions. Our preliminary findings are twofold: 1) Detecting problem structure makes solving large-scale ML tasks to global optimality computationally possible. 2) Achieving global optimality significantly enhances the performance of simple, interpretable ML models on large datasets.
A deep reinforcement learning approach for power management of battery-assisted fast-charging EV hubs participating in day-ahead and real-time electricity market
Faculty: Tapas Das
Ph.D. student: Diwas Paudel

Fast-charging EV hubs solely focus on swift EV charging, akin to traditional gas stations. The EVs entering these hubs receive their requested charge and leave promptly without parking. Present Tesla supercharging stations are fast-charging examples, though smaller than the study's envisaged scale. Fast-charging hubs are intricately linked with power markets and transportation networks. Profitable hub operations necessitate an encompassing model that factors in power market dynamics, pricing mechanisms, and the random EV charging demand. These charging hubs source power from the grid via committed hourly purchases from the day-ahead (DA) market (low price volatility) and supplement from the real-time (RT) market (higher price volatility). Many of these hubs leverage battery storage systems (BSS) for profit optimization via arbitrage. We consider the fast-charging hub power management via a two-step methodology: in the first step, a mixed-integer linear optimization model determines DA commitments, then in the second step, a deep reinforcement learning based real-time power management model optimizes the balance of DA and RT power, along with BSS supply to meet the charging demand in the hub.

DOMAIN-KNOWLEDGE INTEGRATION IN DOMAIN-INFORMED MACHINE LEARNING AND MULTISCALE MODELING OF NONLINEAR DYNAMICS IN COMPLEX SYSTEMS
Faculty: Trung Le
Ph.D. student: Phat Kim Huynh
Nonlinear dynamical systems are widely used to model various phenomena in science and engineering. Recent advancements in sensing technologies have enabled the collection of large amounts of multi-modal sensor data. This data allows us to gain insights into complex system dynamics and build data-driven models without needing the underlying equations. However, due to data complexity and multi-modality, there is a need for efficient data sampling strategies and sophisticated data-driven methods for analysis and modeling.
The combination of big data analytics and data-driven machine learning has transformed the traditional analysis of dynamical systems. However, it's important to consider the laws of physics and domain knowledge to avoid ill-posed problems or non-physical solutions.
Multiscale modeling integrates multi-scale and multi-physics knowledge into data-driven frameworks, providing insights into complex system behaviors. Machine learning and multiscale modeling can complement each other to build robust forecasting models that incorporate underlying physics and address ill-posed problems.
This dissertation focuses on advancing data-driven methods for complex system dynamics modeling, integrating domain knowledge into physics-informed machine learning, and developing mathematical frameworks and theories for multi-resolution systems modeling.
The proposed methods aim to achieve three goals: (1) efficiently represent and integrate domain knowledge in physics-informed models, (2) address multi-scale and multi-physics problems by advancing state-of-the-art models, and (3) establish a theoretical foundation for optimal sampling strategies, error and convergence analysis, and nonlinear dynamical models for complex systems.

Developing a Reliable and Trustworthy AI System through IoT Data Collection, Federated Learning, and Data-Centric AI
Faculty: Trung Le
Ph.D. Student: Quoc H. Nguyen
This research addresses the challenges of maintaining data privacy and security while upholding the performance and accuracy of AI systems. It explores the application of Federated Learning techniques, enabling model training on decentralized data sources without the need to share raw data. This approach ensures that sensitive information remains localized while harnessing the collective knowledge of distributed models. Moreover, the study delves into the concept of Data-Centric AI, emphasizing the significance of data quality, provenance, and reliability. By incorporating data governance principles, robust data preprocessing techniques, and accountability measures, the transparency and trustworthiness of AI systems can be enhanced. Additionally, an IoT system is being developed for seamless data collection. This IoT system leverages diverse IoT devices and sensors to gather real-time and contextual data from various sources. By collecting a comprehensive and diverse dataset, it facilitates the training of robust AI models. In summary, this research aims to develop a robust framework for constructing reliable and trustworthy AI systems by integrating Federated Learning and Data-Centric AI principles. By addressing concerns regarding data privacy and trust in AI applications, this work has the potential to provide significant contributions in the field.

OPTIMAL TEST DESIGN FOR RELIABILITY DEMONSTRATION UNDER MULTI-STAGE ACCEPTANCE UNCERTAINTIES
Faculty: Mingyang Li
Ph.D. student: Bingjie Wang
A reliability demonstration test (RDT) plays a critical role in ensuring product reliability and verifying its compliance with the designated requirements. Existing RDT designs primarily consider the cost of RDT itself or over a single demonstration stage, and hence fail to fully address the uncertainty of multi-stage RDTs and various pathways the product may go through. In practical, striking a delicate balance is crucial when it comes to deciding whether to enforce a stringent RDT to ensure superior reliability and potentially minimize future warranty service costs, or to avoid making the RDT excessively strict in order to expedite the product's market release and prevent additional reliability growth costs. Determining the optimal balance is challenging, not to mention that this decision-making process may be occurring multiple times in different stages. Our research aims to address these limitations by proposing a Bayesian multi-stage binomial RDT design framework. The research focuses on minimizing the total expected cost while incorporating prior knowledge and quantifying acceptance uncertainties across multiple stages (Figure 1). A recursive information propagation algorithm is developed to update product reliability knowledge, considering prior information and test results. A comprehensive sensitivity analysis explores the impact of cost structures and reliability growth rates, demonstrating the robustness of the proposed method. By emphasizing proactive planning and cost-benefit maximization, this research aims to provide a holistic approach to multi-stage RDT planning.

Descriptive diagram for illustrating the difference among different BRDT designs
Optimization of Energy-Efficient Neural Networks
Faculty: Susana Lai-Yuen
Ph.D. student: Esmat Ghasemi Saghand
Artificial neural networks (ANNs) have been very successful in many applications including image recognition, and autonomous vehicle control. However, ANNs require high computational power and resources to be deployed so they are mostly processed on large-scale computers or in the cloud. Consequently, there is a growing need for low-power neural systems that can be directly deployed on edge devices. We are investigating the next generation of neural networks called spiking neural networks (SNNs), which can provide a more energy-efficient and biologically plausible alternative to conventional ANNs. Given the inherent discrete and non-differentiable nature of SNNs, we are investigating and developing efficient training strategies and architectures for SNNs to improve their design and performance.
