The arbitrary-order hidden Markov model (α-HMM) is a nontrivial generalization of the standard HMM, designed to model stochastic processes with higher-order dependences among arbitrarily distant random events. The α-HMM admits an efficient Viterbi-style optimal decoding algorithm, making it feasible to discover higher-order dependences among data objects in observed sequential data. Because the α-HMM exceeds the expressive power of standard HMMs, fixed kth-order HMMs, and stochastic context-free grammars, effective probabilistic parameter estimation approaches are required to translate this theoretical expressiveness of the α-HMM into practical utility. This paper introduces a principled methodology for effective estimation of probabilistic parameters of the α-HMM from observed data. In large-scale sequential datasets, higher-order dependencies can vary widely across instances, so a single global parameter set may be inadequate. Instead, an amortized parameter inference approach is proposed for the α-HMM, in which an input-conditioned parameter estimator is learned from data and used to infer instance-specific parameters for each input instance to the decoding algorithm. Specifically, the neural parameter estimator is trained using a composite learning objective that is partially enabled by the optimal decoding algorithm. The effectiveness of the proposed parameter estimation method is demonstrated through empirical results of the application of the α-HMM in biomolecular structure modeling and prediction.
Zhang et al. (Tue,) studied this question.