Predicting solution conformation and aggregation of conjugated polymers remains a bottleneck for translating solution processing into controlled film microstructures and for closing the loop in self-driving laboratories. We construct a cleaned, machine-readable dataset of 256 entries that links polymer chain length, polymer–solvent compatibility, formulation, and sample history to the radius of gyration, Rg. We evaluate molecular representations and common ML algorithms under both independent and identically distributed (IID) and out-of-distribution (OOD) regimes, including leave-one-polymer-out and polymer–solvent interaction splits. Under IID, models that combine size, chemistry, formulation, and history achieve strong accuracy (RMSE ≈ 0.18; R2 ≈ 0.90), while structure-only models add little beyond polymer chain length. Generalization is the central limit: OOD errors rise sharply for polar sidechain polymers and formulations with significant Hansen descriptor mismatches, and global distribution shift alone does not explain failure—alignment within key feature subspaces matters. Model uncertainty is reasonably calibrated overall, but degrades under strong distribution shift, limiting reliable experiment selection without targeted data. We identify concrete actions to improve robustness for future machine learning: standardized, machine-readable reporting of processing history (times, temperatures, and stimuli) and more deliberate sampling of chemical space. Together, these results highlight dataset design, robust generalization, and uncertainty-aware modeling as the foundations for accelerating discovery and control of conjugated polymer assemblies, capabilities that directly enable autonomous, data-driven experimentation in this field.
Dehghan-Toranposhti et al. (Wed,) studied this question.