Drinking water disinfection can result in the formation disinfection byproducts (DBPs, > 700 have been identified to date), many of them are reportedly cytotoxic, genotoxic, or developmentally toxic. Analyzing the toxicity levels of these contaminants experimentally is challenging, however, a predictive model could rapidly and effectively assess their toxicity. In this study, machine learning models were developed to predict DBP cytotoxicity based on their chemical information and exposure experiments. The Random Forest model achieved the best performance (coefficient of determination of 0.62 and root mean square error of 0.63) among all the algorithms screened. Also, the results of a probabilistic model demonstrated reliable model predictions. According to the model interpretation, halogen atoms are the most prominent features for DBP cytotoxicity compared to other chemical substructures. The presence of iodine and bromine is associated with increased cytotoxicity levels, while the presence of chlorine is linked to a reduction in cytotoxicity levels. Other factors including chemical substructures (CC, N, CN, and 6-member ring), cell line, and exposure duration can significantly affect the cytotoxicity of DBPs. The similarity calculation indicated that the model has a large applicability domain and can provide reliable predictions for DBPs with unknown cytotoxicity. Finally, this study showed the effectiveness of data augmentation in the scenario of data scarcity.
Download full-text PDF |
Source |
---|---|
http://dx.doi.org/10.1016/j.jhazmat.2024.133989 | DOI Listing |
Enter search terms and have AI summaries delivered each week - change queries or unsubscribe any time!