[02.1]
All for One: LLMs Solve Mental Math at the Last Token With Information Transferred From Other Tokens
A mechanistic study of mental-math behavior in language models, arguing that the final token solution depends on information routed forward from earlier token positions.