Investigación y modelos abiertos

Artículos destacados por la comunidad de Hugging Face y los modelos abiertos que más atención reciben esta semana.

Artículos

  • ▲ 5527 sept 2026

    How Far Are We from Removing the Visual Encoder? Scaling Laws for Encoder-Free Multimodal Pretraining

    Most modern multimodal large language models (MLLMs) build on a pretrained visual encoder that provides a strong visual prior. Encoder-free MLLMs instead learn visual representations directly from raw pixels, offering a simple and unified architecture, but their scaling behavior has not been…

    Lin Chen, Bolin Ni, Qi Yang, Lan JiangarXiv ↗
  • ▲ 2527 sept 2026

    SentZero: An Enhanced Sentence-Centric Vision-Language Pretraining for Multi-Task Zero-Shot Chest X-Ray Analysis

    Vision-language (VL) pretraining using paired chest X-ray (CXR) images and radiology reports has shown strong potential for medical image understanding. However, existing methods often remain dependent on task-specific finetuning because radiology reports are lengthy, clinically dense, and…

    Hangyul Yoon, Hyungyung Lee, Edward Choi, Eunho YangarXiv ↗
  • ▲ 1527 sept 2026

    Hard Vision, Easy Vision: What GPT-6 Astra Reveals Across Computer Vision

    Frontier general-purpose systems are rapidly expanding beyond visual understanding into capabilities traditionally handled by dedicated computer-vision models. As these capabilities expand, a central question for the computer-vision community is how far this reach extends, and what remains hard. We…

    Hanoona Rasheed, Mohammed Irfan Kurpath, Bin Ren, Hisham CholakkalarXiv ↗
  • ▲ 1527 sept 2026

    FlowTool: Controlling Tool Parameter in Image Retouching via Flow Matching

    Tool-based image editing (image retouching) is commonly formulated with autoregressive multimodal large language models (MLLMs) that sequentially generate reasoning, tool selections, and parameter values. In this work, we present a novel approach to tool-based image editing by framing the task as a…

    Thanh-Long V. Le, Steven Walton, Seunghyun Yoon, Branislav KvetonarXiv ↗
  • ▲ 1127 sept 2026

    Imprint Reader: From Weight-Update Readout to Behavioral Intervention

    As language models take a growing role in AI development, a natural aspiration is for them to reflect on their own learning process, as humans do, and use that reflection to improve themselves. At the same time, these models have an advantage that human learners lack, since training leaves…

    Guanxu Chen, Qihao Lin, Jing ShaoarXiv ↗
  • ▲ 1027 sept 2026

    An RL View of OPD: Least Square Policy Distillation for Sample-Efficient LLM Reasoning

    We study on-policy distillation (OPD) through the lens of reinforcement learning, establishing a connection between the reverse-KL objective in OPD and KL-regularized policy optimization. Building on this connection, we introduce Least-Square Policy Distillation (LSPD), an RL-inspired framework…

    Shangzhe Li, Yuxiao Yang, Tianrun Yu, Kaixiang ZhaoarXiv ↗
  • ▲ 927 sept 2026

    Nereus: Adaptive Parallelism for LLM Post-Training

    Reinforcement learning (RL) post-training for large language models (LLMs) coordinates multiple models across generation, inference, and training on GPU clusters. Several factors may change during a run, including resource availability, sequence length, memory pressure, and stage bottlenecks. As a…

    Songlin Jiang, Tuo Shi, Sitong Zhang, Zeke WangarXiv ↗
  • ▲ 427 sept 2026

    ColNanoVDR: Document-Free Query Distillation for Multi-Vector Visual Document Retrieval via Optimal Transport

    Multi-vector retrievers built on vision-language models lead visual document retrieval (VDR), but they run a multi-billion-parameter query encoder on every search. Distilling this encoder into a small student that queries the teacher's existing index would remove the bottleneck. The standard…

    Zhuchenyang Liu, Ziyi Wang, Yao Zhang, Yu XiaoarXiv ↗
  • ▲ 327 sept 2026

    Can We Trust the Teacher? Decoupled Credit Direction-Magnitude for Self-Distillation

    RLVR provides reliable trajectory-level credit, while OPSD offers dense supervision for token-level credit. This exposes a fundamental coupling when updating step-level credit direction and magnitude with teacher supervision, preventing steps from receiving reliable credit directions and…

    Yugu Li, Zehong Cao, Peizhen Li, Yang ZhangarXiv ↗
  • ▲ 327 sept 2026

    Distillation Defenses Easily Break After Reinforcement Learning

    Distillation attacks copy the reasoning capabilities of closed-source large language models, allowing bad actors to replicate state-of-the-art performance at low cost. Attackers systematically collect a large volume of frontier model reasoning traces and then train (i.e., "distill") their own…

    Shidan Javaheri, Alexander Panfilov, Oliver Britton, Yarin GalarXiv ↗
  • ▲ 327 sept 2026

    FactorEngram: Factorized N-gram Memory with Basis-Level Gating for Language Models

    Lookup-based memory has been a promising way to scale the parameters of large language models (LLMs). It retrieves learned representations of local token patterns, such as n-grams, instead of reconstructing them through successive layers of computation. However, existing designs such as Engram…

    Bowen Yang, Jingbo Zhou, Qinghong Miao, Hua WuarXiv ↗
  • ▲ 227 sept 2026

    Who Gets a Token, and What Does It Carry? Unequal Name Support and Concept Access in Large Language Models

    Names are personal identifiers, but they also carry social meaning and are widely used to evaluate how language models treat different people. Such evaluations typically assume that matched names are comparable model inputs. We show that this assumption often fails at the lexical interface: matched…

    Mir Tafseer Nayeem, Davood RafieiarXiv ↗
  • ▲ 227 sept 2026

    Reinforcing Agentic Creativity in Scientific Ideation with Night Science

    Large language models (LLMs) excel at structured, verifiable tasks, but their low-entropy bias can produce homogeneous and predictable outputs, limiting their utility for open-ended scientific ideation. Effective discovery, however, spans a broader creative spectrum: from structured day science to…

    Priyanka Kargupta, Silviu Cucerzan, Shweti Mahajan, Allen HerringarXiv ↗
  • ▲ 227 sept 2026

    Geometry as Address: Routing Attention to Visual Memory for Long-Horizon Camera-Controlled Video Generation

    Long-horizon camera-controlled video generation requires recovering previously observed content from an ever-growing visual history. Existing approaches either search historical context implicitly or reconstruct it into persistent 3D memory, facing inefficient memory access or accumulated geometric…

    Zesong Yang, Weikai Chen, Liyuan Cui, Lutao JiangarXiv ↗
  • ▲ 227 sept 2026

    SCOPD: Sparse-Context On-Policy Self-Distillation for Efficient Vision-Language Models

    Reasoning vision-language models (VLMs) process images and videos as long sequences of visual tokens, making inference expensive. Training-free token pruning reduces this cost, but aggressive compression can sharply degrade performance, often attributed to irreversible loss of task-relevant visual…

    Ahmadreza Jeddi, Enming Zhang, Jasper Gerigk, Hakki KaraimerarXiv ↗
  • ▲ 227 sept 2026

    AnswerMap: Faithful Spatial Interpretability of VLMs from Answer Posteriors

    When a VLM answers a visual query, current interpretability tools rely on text rationales, which use a mismatched modality, or on internal read-outs, which originate too early to reflect the final output and require white-box access to the model. We introduce AnswerMap, a training-free,…

    Mohamed Eltahir, Fardows Adam, Duaa M. Tahir, Lama AlamoudiarXiv ↗
  • ▲ 10726 sept 2026

    VisionHOPE: Visual Backbones as Self-Modifying Learning Systems

    Visual backbones have evolved from Convolutional Neural Networks (CNNs) with local aggregation to Vision Transformers (ViTs) with global interactions, State-Space Models (SSMs) with input-dependent state transitions, and Test-Time Training (TTT) layers that adapt an inner learner while processing…

    Siran Peng, Tianshuo Zhang, Tianyu Fu, Weisong ZhaoarXiv ↗
  • ▲ 2926 sept 2026

    Recursive Harness Distillation across Agents for Robot Manipulation

    A central goal in robotics is to enable manipulation across changing tasks and environments. Vision-language-action (VLA) models provide broad manipulation capabilities but can struggle when execution requires diagnosing failures and adapting behavior. Strong agents can discover effective…

    Seungyeon Kim, Junhoo Lee, Minkyu Kim, Baekseung KimarXiv ↗
  • ▲ 2726 sept 2026

    QwenGyre: An Elastic Reinforcement Learning Framework for Training xLong-Horizon Agents

    Large language model (LLM) agents increasingly undertake extreme-long (xlong) horizon tasks, where a single execution can span hours, hundreds of model--environment interactions, and nearly 1M tokens per rollout. Applying online reinforcement learning (RL) to such executions poses two fundamental…

    Weiqi Wang, Yuxin Zhou, Mouxiang Chen, Siyuan ZhangarXiv ↗
  • ▲ 1526 sept 2026

    Structured Residual Connectivity Matters for Diffusion Transformers

    Diffusion Transformers (DiTs) have established themselves as a scalable backbone for high-fidelity image synthesis. However, unlike U-Net based diffusion models that rely on rigid, hand-crafted skip connections, DiTs predominantly use a uniform residual stream that integrates all preceding layers…

    Yuhe Liu, Xinyin Ma, Gongfan Fang, Songhua LiuarXiv ↗
  • ▲ 326 sept 2026

    Pruned CTC for Memory-Efficient Large-Vocabulary ASR Training

    Connectionist temporal classification (CTC) naturally supports offline and streaming speech recognition with utterance-level supervision, but conventional implementations materialize frame-by-vocabulary activations in memory, making CTC training with native LLM vocabularies prohibitively…

    Yifan Yang, Xiaoyu Yang, Zengrui Jin, Xian ShiarXiv ↗
  • ▲ 326 sept 2026

    KernelZero: Co-Evolving Proposer and Coder for Continuously Improved GPU Kernel Generation

    High-performance GPU kernels are essential to modern machine learning systems, yet automatically generating kernels that are both correct and efficient remains challenging. Existing LLM-based approaches face two major limitations: the scarcity of high-quality training data aligned with the model's…

    Changxin Ke, Rui Zhang, Zixiang Fang, Zhenghong LiarXiv ↗
  • ▲ 226 sept 2026

    WhiteMatter: All-to-All Cross-Layer Connections via KV Source Mixing

    When generating text, a Transformer produces representations of past tokens at every layer, but each layer can normally use only representations from the same depth. This restriction prevents the model from fully reusing information it has already computed. We introduce WhiteMatter, which allows…

    Wenbo Zhang, Xiang RenarXiv ↗
  • ▲ 226 sept 2026

    DISCO: Distributed Long Context Scaling with Grounding-Reasoning Disaggregation

    While Large Language Models (LLMs) advertise million-token context windows, reasoning quality often collapses as inputs grow -- a phenomenon termed context rot. This failure stems from a structural entanglement in monolithic architectures, where the massive search burden of contextual grounding…

    Guanzheng Chen, Viet Dac Lai, Subhojyoti Mukherjee, Branislav KvetonarXiv ↗
  • ▲ 226 sept 2026

    SMAT: Simple and Efficient Merge-Aware Training

    Model merging integrates the capabilities of multiple experts without joint retraining, but standard expert training optimizes task loss alone and does not guarantee good performance after merging. Merge-aware training (MAT) aims to improve merged performance, but existing methods do not fully…

    Yanggan Gu, Yuanyi Wang, Zhen Li, Shuo CaiarXiv ↗
  • ▲ 126 sept 2026

    What masking geometry works best for EEG foundation models?

    EEG foundation models hold promise for scalable brain-signal decoding across clinical and cognitive neuroscience applications, yet their pre-training pipelines remain poorly understood. Among design choices, the masking strategy is particularly critical: it determines what the network must predict…

    Pierre Guetschel, Bruno Aristimunha, Yassine El Ouahidi, Arnaud DelormearXiv ↗
  • ▲ 026 sept 2026

    Do We Really Need KL Divergence for On-Policy Distillation of Large Language Models?

    Since the advent of knowledge distillation, KL divergence has been the standard loss in distillation. Recently, on-policy distillation (OPD) has emerged as an efficient post-training paradigm for LLMs. As a distillation method, OPD naturally inherits KL divergence as its standard loss. However, in…

    Wenze Lin, Jiyuan Long, Jiale Zhao, Shenzhi WangarXiv ↗
  • ▲ 9625 sept 2026

    Groupwise Agentic Grading and Advantage Redistribution for Code Agent RL

    Reinforcement learning (RL) for code agents often uses executable tests to provide binary rewards. With these rewards, Group Relative Policy Optimization (GRPO) assigns identical advantages to test-passing trajectories within each rollout group, overlooking differences in implementation quality and…

    Jinhao Dong, Liang Zhao, Zihao Yue, Wenhan MaarXiv ↗
  • ▲ 2025 sept 2026

    In-Flight KV Cache with Clean Anchors for Faster Autoregressive Video Diffusion

    Few-step autoregressive video diffusion generates a long video by splitting the video into temporal chunks and generating chunk-by-chunk, each through a short sequence of denoising stages. To memorize chunks that are already generated, previous methods reconstruct a clean or less-noisy key--value…

    Yikai Wang, Xiao Han, Mengmeng Xu, Juan Camilo PerezarXiv ↗
  • ▲ 1725 sept 2026

    Change the Product, Keep the Parameters: Associative Algebra Layers for Transformers

    Fast matrix multiplication algorithms keep the product fixed and search for a cheaper way to evaluate it. We instead ask whether a Transformer's learned projections can use a different, cheaper product altogether. Building on an associative-algebra construction that replaces ordinary matrix…

    Ilya Koziev, Ivan OseledetsarXiv ↗
  • ▲ 825 sept 2026

    ExpVoyager: Direct Experience Navigation for Dynamic Agent Skill Synthesis

    Learning from experience in LLM agents has become a key paradigm for developing self-evolving agents that continuously learn and expand their capabilities. Within this paradigm, synthesizing the agent skill has emerged as a promising solution for transforming accumulated experience into reusable…

    Kwangwook Seo, Dongha LeearXiv ↗
  • ▲ 525 sept 2026

    Relic: From Multi-Agent Collaboration to Persistent Organizational Capability

    Multiple agents may often conflict in an organization: for example, one coding agent changes an interface in a repository, but another continues to develop on the old version where existing tests become stale. A conversation can resolve the episode, but when the participants change, what makes the…

    Hongyi Du, Tianyi Zhang, Weijia Zhang, Yi YangarXiv ↗
  • ▲ 225 sept 2026

    Allspark: Weak to Strong Transfer via Alternating Chain of Thought

    Recent progress in frontier models has renewed interest in large-scale reinforcement learning (RL), but the cost of generating large-model rollouts makes even testing RL recipes expensive. We ask whether reasoning improvements learned by a small, weak model can benefit a larger, stronger model…

    Kaizhao Liang, Junxiong Wang, Chen Liang, Zhendong WangarXiv ↗
  • ▲ 225 sept 2026

    Playing to Par: Reinforcement Learning for Provably Optimal Quadrilateral Block Decompositions

    A quadrilateral block decomposition of a planar domain is judged by whether it is complete, whether its elements are well shaped, and how many of its vertices are irregular. The last has a provable floor: the discrete Gauss-Bonnet identity enforces a lower bound on the total vertex irregularity of…

    Arjun Narayanan, Per-Olof PerssonarXiv ↗
  • ▲ 324 sept 2026

    NVAlign: Direct-Gradient Optimization for Non-Verbal Control in Continuous Autoregressive Flow Matching Text-to-Speech

    While modern text-to-speech (TTS) systems generate highly natural speech and support inline non-verbal vocalization (NVV) tags, accurate control over these events remains challenging. A key gap is the lack of established post-training methods for non-verbal control in continuous autoregressive…

    Qiaolin Wang, Pedro Sandoval-Segura, Anunaya Joshi, Edvardas JurkonisarXiv ↗
  • ▲ 224 sept 2026

    G^2PTQ: Improving LLM Post-Training Quantization with Generalized Gradient Compensation

    Post-training quantization (PTQ) is a practical approach to reducing the memory and computational footprint of large language models (LLMs) without retraining. GPTQ-based methods have become the de facto standard, yet they suffer from two complementary limitations. Methods with local, layer-wise…

    Ruikang Liu, Haoli Bai, Yuxuan Sun, Qian ZhangarXiv ↗
  • ▲ 3522 sept 2026

    The Past Frames the Future: Memory for Autoregressive Video Generation

    Advances in generative models have improved video fidelity, enabling long-horizon generation, interactive world modeling, and evolving visual environments. Autoregressive (AR) video generation extends visual sequences through causal rollouts. However, a fundamental bottleneck emerges: as the…

    Harold Haodong Chen, Rongjin Guo, Disen Lan, Wen-Jie ShuarXiv ↗
  • ▲ 3122 sept 2026

    Just-in-Time Memory: Learning to Curate Task-Adaptive Memory for LLM Agents

    Agentic memory systems reuse past experience to improve future performance, yet most existing designs curate memory at write time: once a task is completed, its trajectory is distilled into a fixed artifact, such as a reflection, workflow, skill, or reasoning strategy, that is later retrieved by…

    Yefan Zhou, Yang Li, Zeyu Leo Liu, Semih YavuzarXiv ↗
  • ▲ 1022 sept 2026

    MemBodied: Recurrent Associative Memory for Vision-Language-Action Models

    Vision-Language-Action models provide a strong foundation for general-purpose robot control, yet a vast majority of policies do not preserve and leverage episode-level information beyond the current observation. This limitation is consequential in history-dependent manipulation tasks that depend on…

    Tej Deep Pala, Navonil Majumder, Bryce Goh, Raphael YeearXiv ↗
  • ▲ 822 sept 2026

    WhatWorkedBench: Benchmarking Experimental Understanding in AI Agents

    AI research agents need reliable knowledge of how their experiments change outcomes. We introduce WhatWorkedBench to measure experimental understanding, the accuracy of predictions about component changes after budgeted experimentation. Agents inspect code, select measurements, and submit a…

    Jingjie Ning, Xueqi Li, Yibo Kong, Dongting LiarXiv ↗
  • ▲ 722 sept 2026

    Hunyuan-A13B Technical Report

    We present Hunyuan-A13B, an open-source large language model based on a Mixture-of-Experts architecture. It contains 80 billion total parameters but activates only 13 billion during inference, balancing model capability, computational efficiency, and deployment cost. The model is pretrained on a…

    Tencent Hunyuan Team, Ao Liu, Botong Zhou, Can XuarXiv ↗
  • ▲ 622 sept 2026

    Verifiable Hidden Dynamics Play: Generating Agentic RL Environments from Solved Mechanisms

    Language-model agents increasingly face long-horizon tasks with evolving state, interdependent decisions, and delayed outcomes. Scaling their training requires diverse agentic environments, dependable outcome signals, and low extension cost. Existing generation pipelines commonly construct an…

    Xinjie Shen, Wei Fan, Xudong Guo, Jianhong TuarXiv ↗
  • ▲ 522 sept 2026

    All modalities are equal, but video is more equal: Closing the Cross-Attention Gap in Joint Video Generation

    Video is a rich representation of a physical event, capturing appearance, geometry, motion, and temporal evolution. Other modalities, such as 3D body motion or audio, encode narrower aspects of the same event. We find that joint multimodal diffusion transformers exhibit a corresponding asymmetry in…

    Ohad Rahamim, Dvir Samuel, Idan Schwartz, Gal ChechikarXiv ↗
  • ▲ 422 sept 2026

    InternW0: A Foundational Physical World Model for Efficient Real-World Interactions

    Physical intelligence requires more than predicting how the world may evolve: predictions must remain actionable as the world continues to change. We introduce InternW0, the first instantiation of the InternW physical world model series from Shanghai AI Laboratory, built around omnimodal…

    Jisong Cai, Yao Mu, Ganlin Yang, Zhe CaoarXiv ↗
  • ▲ 322 sept 2026

    On the Diffusibility of High-Dimensional Latents

    Representation Autoencoders (RAEs) enable diffusion models to operate in the feature spaces of pretrained visual encoders. However, many off-the-shelf encoders are not optimized for faithful reconstruction, discarding fine-grained visual details. As expected, finetuning these encoders for image…

    Chao Feng, Zhiyang Xu, Bowei Chen, Yuanjun XiongarXiv ↗
  • ▲ 222 sept 2026

    FLEET: From Logits Entropy to Enhanced Trajectories in Text Generation

    Solutions based on large language models (LLMs) often rely on temperature sampling to improve accuracy and stability by aggregating multiple samples from the completion distribution. However, this memoryless approach is inherently suboptimal: because it lacks awareness of prior generations and…

    Oleksii Streltsov, Oleksandra VitkoarXiv ↗
  • ▲ 222 sept 2026

    Six Layers Less: Encoder Pruning for Whisper with Label-Free Recovery

    Pruning large pre-trained transformer-based ASR models such as OpenAI's Whisper has seen great adoption, as pruning the decoder led to significant end-to-end transcription speedups. For instance, the {\tt whisper-large-v3-turbo} variant reduced the decoder from 32 to 4 layers, while Distill-Whisper…

    Rasmus Aagaard, Nicki Skafte DetlefsenarXiv ↗
  • ▲ 222 sept 2026

    EmbodiedSWE: Coding Agents for Long Horizon Dexterous Robotics

    We study coding agents for long-horizon, dexterous robotics and ask whether their solutions can provide scalable supervision for learning general robot policies. To test this, we develop EMBODIEDSWE-BENCH, a simulation benchmark for coding agents spanning contact-rich manipulation, deformable…

    Haoxiang You, Zeyu Shen, Yilang Liu, Zhicheng ZhengarXiv ↗
  • ▲ 122 sept 2026

    Uranus: Building the Next-Generation Simulation Infrastructure for Embodied AI

    Scalable simulation is essential for robot data generation, policy training, evaluation, and safe iteration, yet real-world interaction is costly and conventional simulators require labor-intensive construction. We present Uranus, a data-driven robot simulator built around a…

    Wenkang Qin, Yukun Zhou, Noah Shen, Jisong CaiarXiv ↗
  • ▲ 122 sept 2026

    StudentBench: AI and human tutoring yield equivalent GRE learning gains

    Artificial intelligence offers an unprecedented opportunity to augment human capabilities, yet progress at the frontier has focused primarily on advancing model capabilities. We introduce StudentBench, a suite of AI teaching evaluations and a public platform that enables large-scale data collection…

    Curtis Northcutt, Inaara Hasmani, Kevin Feng, Trevor KhangiarXiv ↗
  • ▲ 7621 sept 2026

    SpeakerMem-R1: Speaker-Centered Dual-Track Memory for Multi-Party Dialogue

    Long-term conversational memory in multi-party settings requires more than retrieving relevant content from long-term conversations: it must distinguish who said what, whom each statement concerns, how individuals perceive one another, what information is shared by the group, and how states change…

    Haobo Zheng, Tan Tang, Yan Chen, Weijie WangarXiv ↗
  • ▲ 2521 sept 2026

    JEV-as-a-Judge: Accept When Confident, Escalate When Unsure

    LLM-as-a-judge enables evaluation across diverse tasks, but inference cost and confidence reliability become critical at scale. We study whether a decision-only judge can provide an economical first pass and identify when stronger evaluation is needed. Comparing jev-as-a-judge with sixteen…

    Yubo Li, Yidi Miao, Ramayya Krishnan, Rema PadmanarXiv ↗
  • ▲ 1521 sept 2026

    PACT: From Credit Assignment to Critic Alignment

    Reinforcement learning has become a central component of large language model (LLM) post-training, yet token-level credit lacks a generally accepted mathematical definition, leaving its relationship to commonly used training signals unclear. We formulate three regularity conditions, namely…

    Jiayan Fu, Hang Xu, Yong Zhang, Zhaokai LuoarXiv ↗
  • ▲ 1421 sept 2026

    Agensh: Scaling Organizational Intelligence to 1,024 Agents

    A multi-agent system can reduce latency on complex tasks by executing work concurrently. Several pioneering harness frameworks support multi-agent systems. However, the scalability of current multi-agent harnesses is often constrained by a central orchestrator's capacity to allocate tasks and…

    Zhihao Zhan, Ting Song, Li Dong, Shaohan HuangarXiv ↗
  • ▲ 1021 sept 2026

    GeoPair: Geometry-Preserving Cross-Layer Factorization for Training-Free Transformer Compression

    Transformer architectures exhibit cross-layer redundancies, yet post-training compression pipelines typically optimize layers in isolation or rely on heuristic grouping strategies that disregard layer-specific activation geometries. We introduce a principled, training-free framework that…

    Baher Mohammad, Ammar Ali, Stamatios LefkimmiatisarXiv ↗
  • ▲ 421 sept 2026

    RoboFollow: Unveiling the Instruction Following Mirage in Embodied Agents

    Modern embodied agents achieve impressive success rates, yet their actual instruction-following ability is far weaker than these numbers suggest. We trace this illusion to a structural property we term low scene entropy: when a visual scene admits only one valid task, language becomes redundant and…

    Chang Guo, Yukun Xie, Bohan Tan, Zheng ChangarXiv ↗
  • ▲ 421 sept 2026

    Blaming Across the Aisle: Political Contrasting and Blame Attribution in the Danish Parliament

    Political discourse is widely perceived to be growing more hostile, yet robust evidence remains scarce. This study examines blame attribution in the Danish Parliament from 1997 to 2026, combining a purpose-built classifier, BlameBERT (F1: 0.80), with multilevel statistical modeling. The classifier…

    Markus Lundsfryd Jensen, Rune Egeskov Trust, Kenneth Christian Enevoldsen, Sara KoldingarXiv ↗
  • ▲ 221 sept 2026

    Capable yet Parsimonious: Extracting and Characterizing Hidden Chain-of-Thought in Frontier Models

    The rapid capability gains of frontier language models are widely attributed to improved reasoning abilities, yet this cannot be verified as raw CoT traces in closed-source systems are hidden. By registering a simple custom tool through a standard API feature, we induce frontier models to…

    Xiaoyu Luo, Tao Ren, Wenrui Yu, Xiao LiarXiv ↗
  • ▲ 221 sept 2026

    Calibration as a First-Class Criterion in LLM Evaluation

    Calibration of language models -- the alignment between expressed or implicit confidence and empirical correctness -- is a well-studied subfield within NLP. Methods to measure it already exist. The problem is adoption: outside this subfield, NLP research regularly introduces new models, datasets,…

    Mario Sanz-Guerrero, Katharina von der WensearXiv ↗
  • ▲ 221 sept 2026

    MemoryAthena: Adaptive Routing over Latent and Generated Memories

    Learned-memory methods store information in an explicit table and consume it through a separate reader, allowing addressing, storage, and reading to be modified independently. We study whether useful memory can also be generated rather than only retrieved. MemoryAthena uses three pathways: direct…

    Mingyuan Li, Guangsheng Yu, Juyuan Zhang, Xu WangarXiv ↗