Long-Horizon State Tracking in LLMs: Executing MD5 through a Deep Sequence of Dependent Tool Calls
arXiv:2609.00012v1 Announce Type: new Abstract: Longhorizon tasks remain uncommon in large language model LLM evaluation, and for a reason: when each step depends on the last, perstep accuracy that looks excellent in isolation decays catastrophically, as errors cascade and the endtoend failure...
