BOOK OF ABSTRACTS 103 IHSS&IWA26 / BRNO / CZECHIA / 23–28 August 2026 Thursday, 27 August 2026 / Hall C NOM and Aquatic Systems SL48 Molecular Characterization of Dissolved Organic Matter in Deep Groundwater of Marine Deposits by Data Fusion of EEM and FT-ICR-MS with Machine Learning Hayato Sato1, Huiyun Xue1, Kanako Toda2, Takumi Saito1,2* 1 Department of Nuclear Engineering and Management, School of Engineering, The University of Tokyo, 7-3-1 Hongo, Bunkyo-ku, Tokyo 113-8656, Japan 2 Nuclear Professional School, School of Engineering, The University of Tokyo, 2-22 Shirakata Shirane, Tokai-mura, Ibaraki, 319-1188, JAPAN * saito@n.t.u-tokyo.ac.jp Dissolved organic matter (DOM) in deep groundwater modulates the migration of radionuclides relevant to the geological disposal of high-level radioactive waste (HLW) through complexation, redox reactions, and competitive sorption. Understanding the molecular controls of these interactions, however, has been hindered by the chemical complexity of DOM and the limited access to deep aquifers. Here we combine fluorescence excitation– emission matrix (EEM) spectroscopy and Fourier-transform ion cyclotron resonance mass spectrometry (FTICR-MS) with data fusion and machine learning to resolve the endmembers, molecular fingerprints, and metal-binding behavior of DOM in saline groundwater from the Horonobe Underground Research Laboratory (URL), Hokkaido, Japan. [1,2] Groundwater was sampled at depths of 140–500 m below ground level from packer-isolated boreholes within the URL. DOM was concentrated by PPL solid-phase extraction (SPE) and analysed by EEM and negative-mode ESI FT-ICR-MS (Bruker ScimaX, 7 T). DOM samples were argumented with EEM data of all samples spiked with Eu³⁺ or UO2²⁺ (0–0.08 mM) were decomposed by parallel factor analysis (PARAFAC) into four components, and Stern–Volmer constants quantified their interactions with the two metal ions. Coupled matrix and tensor factorization (CMTF) was used to jointly factorize the EEM tensor and the FT-ICR-MS matrix under a shared sample-loading mode, directly linking fluorescence endmembers to molecular formulae. [3] To extend the data fusion analysis, we developed a supervised machine learning that fuses the two complementary data streams in an early-fusion scheme. PARAFAC component scores were combined with molecular-level descriptors derived from FT-ICR-MS—intensity-weighted H/C, O/C, double bond equivalent (DBE), modified aromaticity index (AI moₔ ), and class fractions (lipid-, protein-, lignin-, tannin-, and CRAM-like)—to form input features for Random Forest and XGBoost regressors trained to predict endmember
RkJQdWJsaXNoZXIy NDA4Mjc=