Date of Defense

25-6-2026 2:00 PM

Location

F1-1117

Document Type

Thesis Defense

Degree Name

Master of Science in Civil Engineering (MSCE)

College

COE

Department

Civil and Environmental Engineering

Keywords

Pharmaceutically Active Compounds; Sorption Distribution Coefficient; Machine Learning; Gradient Boosting; Random Forest; Recursive Feature Elimination; Environmental Risk Assessment.

Abstract

The growing environmental presence of pharmaceutically active compounds (PACs) creates a critical need for rapid, cost-effective methods to predict their movement through soils and sediments. Traditional laboratory batch experiments are time-consuming, and previous predictive models often relied on limited data or ignored complex chemical behavior. The primary objective of this thesis is to develop robust machine learning models for predicting the sorption distribution coefficient of PACs. The study aims to determine whether separating data by ionic state improves predictive accuracy and to identify the key physical and chemical factors governing the sorption of PACs onto soils and sediments. A large dataset of 1,253 data points, encompassing linear and nonlinear sorption, was compiled and split into training and independent external validation sets. To assess the influence of ionic charge, these sets were further divided into three subsets: the complete aggregated data, the cationic subset, and the neutral subset. Two machine learning approaches, gradient boosting (GB) and random forest (RF), were trained and compared. Additionally, a recursive feature elimination (RFE) process was used to simplify the models by removing less important variables. Internal and external validation showed that the GB model consistently outperformed the RF model across the complete dataset, although the RF model demonstrated a slight advantage for the cationic subset during internal validation and for the neutral subset during external validation. Notably, both models consistently underestimated the sorption of cationic PACs. The feature selection process demonstrated that, regardless of the specific machine learning algorithm applied or the ionic state of the compounds, six core descriptors were consistently retained across every model: acid dissociation constant (pKa1), molar volume (Vm), sand content (%), cation exchange capacity (CEC), pH, and water solubility (log Sw). This study successfully develops optimized GB and RF models that serve as computationally efficient alternatives to traditional laboratory experiments. Specifically, the 10-feature GB model provides a faster tool for general environmental risk assessments, achieving the highest overall predictive accuracy on the complete dataset with R² of 0.9509. Meanwhile, the 10-feature RF model serves as a reliable alternative that demonstrated a slight predictive advantage specifically for neutral PACs. Segmenting the compounds by their ionic states clarifies the specific variables that drive sorption. The results also highlight the inherent complexities of cationic sorption, indicating that predicting the behavior of positively charged molecules requires further refinement in future research.

Share

COinS
 
Jun 25th, 2:00 PM

Developing Machine Learning Models for Sorption of Pharmaceuticals to Soils and Sediments

F1-1117

The growing environmental presence of pharmaceutically active compounds (PACs) creates a critical need for rapid, cost-effective methods to predict their movement through soils and sediments. Traditional laboratory batch experiments are time-consuming, and previous predictive models often relied on limited data or ignored complex chemical behavior. The primary objective of this thesis is to develop robust machine learning models for predicting the sorption distribution coefficient of PACs. The study aims to determine whether separating data by ionic state improves predictive accuracy and to identify the key physical and chemical factors governing the sorption of PACs onto soils and sediments. A large dataset of 1,253 data points, encompassing linear and nonlinear sorption, was compiled and split into training and independent external validation sets. To assess the influence of ionic charge, these sets were further divided into three subsets: the complete aggregated data, the cationic subset, and the neutral subset. Two machine learning approaches, gradient boosting (GB) and random forest (RF), were trained and compared. Additionally, a recursive feature elimination (RFE) process was used to simplify the models by removing less important variables. Internal and external validation showed that the GB model consistently outperformed the RF model across the complete dataset, although the RF model demonstrated a slight advantage for the cationic subset during internal validation and for the neutral subset during external validation. Notably, both models consistently underestimated the sorption of cationic PACs. The feature selection process demonstrated that, regardless of the specific machine learning algorithm applied or the ionic state of the compounds, six core descriptors were consistently retained across every model: acid dissociation constant (pKa1), molar volume (Vm), sand content (%), cation exchange capacity (CEC), pH, and water solubility (log Sw). This study successfully develops optimized GB and RF models that serve as computationally efficient alternatives to traditional laboratory experiments. Specifically, the 10-feature GB model provides a faster tool for general environmental risk assessments, achieving the highest overall predictive accuracy on the complete dataset with R² of 0.9509. Meanwhile, the 10-feature RF model serves as a reliable alternative that demonstrated a slight predictive advantage specifically for neutral PACs. Segmenting the compounds by their ionic states clarifies the specific variables that drive sorption. The results also highlight the inherent complexities of cationic sorption, indicating that predicting the behavior of positively charged molecules requires further refinement in future research.