Under Review

  • Which Characters Need Context? Decomposing Context Dependence into Trajectory and Magnitude

    fDT1sioSnn

    We introduce a per-character context-gain profiling framework that decomposes context dependence into trajectory and Mean Context Gain, revealing a structural-trajectory effect in natural-language prose, with a nominally significant cross-corpus result in Reuters; code corpora serve as contrast cases.
  • Shorter Chains via Macro Transitions: Constant-Factor Compression for Chain-of-Thought Counter Simulation

    HGmCGa0Q3O

    Transformers can simulate complex computations through chain of thought, but doing so one step per token makes generation slow and expensive. We introduce two ways to compress these reasoning traces by combining multiple transitions or counting operations into fewer tokens while preserving the original computation. Experiments show that moderate compression maintains strong length generalization while reducing trace length, latency, attention cost, and memory use. However, overly aggressive compression can hurt accuracy and require more training data.
  • Attenuated Friction Sensitivity in Imagined Rollouts: A Kinematic-Consistency Diagnosis of a DreamerV3-Class Locomotion World Model

    D5gNaP6QHN

    A kinematic-consistency diagnostic finds attenuated friction sensitivity in free imagined rollouts relative to real physics and observation-driven reconstruction, with weak resolved responses under action replay or long conditioning.
  • Diagnosing the Sparse Mechanism Shift Hypothesis: A Graph-Free Invariance Test

    6AGOgvdze1

    Many methods for distribution shift under causal structure assume the sparse mechanism shift (SMS) hypothesis: that across environments only a few causal conditionals change. This assumption drives mechanism-shift scoring, causal discovery in heterogeneous data, and transportable prediction, yet it is almost never tested on the data at hand. This paper makes SMS diagnosable. We first ask why it is hard to tell which mechanisms changed once the causal graph must be estimated rather than assumed known. A controlled ablation locates the cause: the false positives that limit precision arise at the truly invariant nodes, because their parent sets are mis-estimated; a better global skeleton, repairing the changed nodes, and conditioning-set voting do not remove them. We then give a graph-free, label-free invariance test that flags a node only when no conditioning subset makes its conditional invariant across environments (an inverse use of invariant causal prediction); this is robust to the same failure mode and matches an oracle that knows the true graph. Building on it, we define a data-referenced SMS diagnostic: a discovery statistic (the fraction of mechanisms flagged as changed) with a data-driven null floor obtained by splitting one environment in half, and a bootstrap verdict (no shift / sparse / dense / inconclusive). On controlled synthetic data the verdict is broadly consistent with the true sparsity; on real protein-signalling interventions it returns a dense verdict that clears the data-referenced no-shift floor, which we read as evidence against SMS rather than as within-environment detector noise, and a paired atomic-versus-fat-hand study offers a mechanistic reading and predicts when SMS should hold. We discuss the conditions under which the invariance test localises causal mechanism changes and the statistical interpretation of the verdict. We treat the verdict as a data-calibrated diagnostic rather than a hypothesis test with formal error control, and state the assumptions this requires.
  • The Feynman Trap: When Chain-of-Thought Reasoning Undermines Correct Intuitions in Language Models

    JNhGH8NVUL

    We systematically quantify a phenomenon where LLMs flip correct zero-shot answers to incorrect ones when prompted to reason step-by-step, and characterize when this cost outweighs the benefit across models and task domains.