Prior work showed that a three-step serial implementation of the two-point visual steering model in ACT-R can predict the upper bound of human path-keeping performance, but it is unknown whether the cognitive architecture or the number of processing steps drives this ability. This study compares four implementations: no cognitive architecture, ACT-R, and two variants of QN-MHP, each parameterised to minimise average path-keeping error without fitting to human data. Validation against humans reveals that implementations lacking a cognitive architecture or separate processing of the two visual points exceed human performance, failing to predict realistic upper bounds. Two-step serial processing within a cognitive architecture is necessary and sufficient to retain the predictive capabilities for an upper bound of human steering performance.