This research aims to determine how to predict catalytic activity through learned protein fitness landscapes.
Utilized machine learning techniques
Applied extreme-value theory
Analyzed sparse sequence–activity data
Identified potential for higher-activity protein variants
Established a predictive framework for activity extrapolation
Abstract
Rugged protein fitness landscapes can be learned from sparse sequence–activity data. Machine learning plus extreme-value theory can guide directed evolution toward rarer, higher-activity variants.