Research Papers
arXiv · cs.AI, cs.LG, cs.CL · 50 papers
Skill-Space Shooting for Autonomous Robot Policy Improvement
Zihang Rui, Renhao Wang, Haoxu Huang +1
Robots deployed in the physical world must be able to improve beyond their initial training as they encounter new situations and failures. For this improvement to scale across tasks, it must make effective use of experience without requiring human de
Imagine3D-LLM: Teaching MLLMs to Imagine 3D Scenes Before Answering
Jaewoo Jung, Hyeonseo Yu, Honggyu An +10
Reasoning about the 3D world from multi-view images remains a fundamental challenge for Multimodal Large Language Models (MLLMs). While modern MLLMs handle single-image inputs effectively, they struggle to integrate evidence across viewpoints into a
Breakdown of Local Denoising as Semantic Speciation
Guangkuo Liu, Mert Okyay, Yifan F. Zhang +3
The dynamics of generative models exhibit two apparently distinct temporal windows: a speciation window, in which a sample commits to a semantic class, and a nonlocality window, in which local context windows become insufficient for generation. Motiv
STEPQuant: When and Where Errors Matter in Delta-Rule Recurrent State Quantization
Bingchen Yao, Haobo Xu, Haokun Lin +6
Linear attention replaces growing KV caches with fixed-size recurrent states, yet these persistent states can become a substantial memory bottleneck under concurrent serving. Directly quantizing recurrent states to low precision often leads to severe
LeapQuant: Efficient Linear Attention with Accurate Recurrent State Quantization
Yi Pan, Haocheng Xi, Kan Zhu +10
Recent LLMs increasingly adopt hybrid designs that replace standard attention with linear attention, such as Gated DeltaNet (GDN) and Kimi Delta Attention (KDA). Although they compress the context into a fixed-size recurrent state and substantially r
Cropland PAtteRNS: Parallel Dimensional Attention Networks and Attention to Dataset Disparity for Crop Segmentation in Satellite Imagery Time Series Data
Joseph Metcalfe, Sara Sharifzadeh, Fabio Caraffini
The landscape of satellite imagery time series datasets and boundary-pushing architectures for cropland segmentation has never been richer. However, in this gold rush, important truths are being missed on both fronts, as a drive for the most novel co
A Spectral Theory of Distortion in LLM Graph Reconstruction: Sharp Bounds and Empirical Characterization
Jianru Shen
Evaluations of graph reconstruction by language models typically report a single aggregate distance between the original and the reconstructed graph. We prove that for the Wasserstein distance between Laplacian spectra such a summary is bracketed by
EmoRES-TTS: Residual-Enhanced Vector Steering for Emotional Speech Generation
Kuan-Po Huang, Haohe Liu, Puyuan Peng +5
Emotion-conditioned text-to-speech (TTS) models may fail to express the requested emotion reliably, and improving controllability by additional training is costly in both computation and emotion-labeled speech training data. We therefore study vector
Beyond the Timeline: Augmenting Long-Video Memory with Grounded Entity Biographies
Hui Ren, Lei Fan, Henry Pao +5
Answering questions about long videos often requires connecting events involving the same objects across hours or days. Chronological descriptions and text-derived entities can leave physical identity unresolved: different objects may share a descrip
Pretraining Latent Information Feedback Transformers with Teacher Supervision
Dor Tirosh, Ido Amos, Mor Geva
Transformer language models (LMs) are feed-forward: deep-layer representations are never fed back to shallower layers, and the only pathway for information to flow downward across generation steps is the decoded token. This narrow channel forces mode
Thinking Before Thinking: Scaling Agentic Inference Through Meta-Reasoning
Paras Dahal, Anton Bakhtin, Taco Cohen +9
As agents take on longer and more complex problems, controlling the execution becomes a task in its own right. Each step in the run brings new control choices, like which partial work to build on, whether to start fresh, or when to stop. We introduce
Learning Meta-Skills for Agent Harness Design in Test-Time AI4AI
Cheng Qian, Kunlun Zhu, Beibin Li +2
Agent performance depends on both reasoning ability and the environment in which it acts. We study test-time AI-for-AI, asking how a Builder can learn to construct better execution environments for a Target while both models' weights remain fixed. To
AdviSD: Learning to Advise Frontier LLMs via Targeted Multi-Turn Self-Distillation
Rishabh Agrawal, Hejie Cui, Shasha Li +2
A small trainable advisor can steer a frozen language-model executor using natural-language advice. In addition to learning from task rewards, the advisor can use feedback from completed interactions to improve its advice. However, a plausible correc
Breaking the Uniformity Trap: Scaling Video Diffusion Model via SplitMoE
Yu Xu, Yuxin Zhang, Xiao Yang +7
Mixture-of-Experts (MoE), popularized by large language models, is a promising paradigm for scaling visual generative models. However, conventional token-wise MoE routes tokens independently within a homogeneous expert pool and regularizes expert usa
LongHarness Bench: Stress-Testing Language Model Harnesses for Long-Context Reasoning
Quang Hieu Pham, Thuy Duong Nguyen, Jocelyn Qiaochu Chen +1
Language-model (LM) harnesses enable LMs to operate effectively over long contexts using additional compute. However, existing long-context evaluations are insufficient for distinguishing modern harnesses, reflected by saturated accuracy across harne
Multi-Agent Flow Matching with Decoupled Generative Guidance
Ruoyu Lin, Magnus Egerstedt, Fabio Pasqualetti
Generative modeling is widely used for producing diverse objects from complex, multimodal distributions. However, its expressivity does not, in general, come with formal guarantees that the generated objects satisfy hard constraints or requirements.
Achieving an $O(1/N)$ Optimality Gap in Average-Reward Weakly-Coupled MDPs
Yige Hong, Xiangcheng Zhang, Qiaomin Xie +2
We study average-reward weakly-coupled Markov decision processes (WCMDPs), where a WCMDP consists of $N$ smaller MDPs, called arms, that share multiple per-step budget constraints. We consider the setting where the arms have identical model parameter
WUSH-KV: KV Cache Quantization with Data-Adaptive Transforms
Jiale Chen, Vage Egiazarian, Eldar Kurtić +2
KV cache memory and bandwidth costs grow with context length and batch size, which limits efficient long-context inference. To address this bottleneck, we introduce WUSH-KV for low-bit KV-cache quantization. It adapts WUSH, which constructs a data-aw
Stochastic World Models for Verifying Vision-Based Neural Feedback Systems
I. Samuel Akinwande, Mykel J. Kochenderfer, Clark Barrett
Verifying a vision-based neural feedback system requires a model of the observations its controller acts upon. Such a model must capture the variation the sensor produces, while remaining tractable for closed-loop analysis. Generative adversarial net
ReCIRC: Rectified Conformal Risk Control
Bruno Marcondes e Resende, Helton Graziadei, Thiago Rodrigo Ramos +1
Many applications of black-box predictive models require controlling task-relevant error rates, such as missed lesion pixels in segmentation or missed labels in multilabel classification. Conformal risk control (CRC; Angelopoulos et al., arXiv:2208.0
From Routing Signals to Selective Review: Visual regrounding in MoE VLMs
Hongzhu Guo, Mohsen Fayyaz, Nanyun Peng
Vision-language models (VLMs) may accept false visual premises, answering questions about a target object's color, count, location, or state even when it is absent. We call this reliability-critical behavior a target-absence grounding failure. Existi
How Local Mixing Encodes Relative Position in Global NoPE Attention
Cutter Dawes, Nick Alonso, Tom Figliolia +1
The attention operation is naively position invariant. However, positional information is fundamental to natural language, and therefore a variety of explicit position encodings have been developed in transformer-based models, such as rotary position
Do LLM Agents Execute the Plans They Declare? From Planning-Mode Declaration to Pattern-Specific Execution
Subba Reddy Oota, Francisco Herrera, Jordi Cabot Sagrera +2
Large language models (LLMs) enable agents to solve long-horizon tasks by generating a plan and then executing it in an environment. However, successful planning requires two distinct capabilities: selecting an appropriate plan for the task and execu
Correct Answers, Invalid Traces: What Verifiable Grade-School Math Reveals About Chain-of-Thought Traces
Ratish Puduppully, Pranabendu Misra, Paarth Iyer +3
Chain-of-thought traces are widely read as records of how models reach their answers, informing debugging, agent auditing, and claims about reasoning. Testing this interpretation is difficult because natural-language thinking traces are rarely mechan
Pruning for Efficiency, Paying in Fairness: Demographic Disparities in Pruned Speech-LLMs
Ganesh Pavan Kartikeya Bharadwaj Kolluri, Michael Kampouridis, Ravi Shekhar
Speech-LLMs are expensive to run, making compression important for real-world deployment. However, compressed models are usually selected using aggregate word error rate (WER), which can hide how pruning affects different demographic groups. In this
Explore Broadly, Reason Sharply: Push Small Models toward the Frontier via Sampling
Panagiotis Theodoropoulos, Nan Jiang, Xintong Duan +4
Power-sharpened sampling is an inference-time alternative to reinforcement-learning (RL) post-training for enhancing reasoning in large language models (LLMs). High-probability sequences are amplified under the base model without parameter updates or
Effective Dense Retrieval using Only In-Context Examples
Nour Jedidi, Abdul Basit Ali, Hang Li +1
Turning decoder-only large language models (LLMs) into strong dense retrievers typically requires some form of retriever training. In this paper, we ask whether LLMs can instead be prompted to produce effective representations for dense retrieval giv
NeuronEye: Query-Guided Visual Concept Activation for Vision-Language Reasoning
Ruiyu Yan, Bowen Chen, Shaowen Wan +1
Current vision-language models (VLMs) encode visual information in dense hidden states where object identity, spatial layout, and local attributes are implicitly entangled rather than explicitly disentangled, limiting their ability to isolate and mod
Tail-Influence Sampling for CVaR Policy Evaluation
Pauline Bourigault, Xiaotong Ji, Matthieu Zimmer +2
Policies with similar mean returns can differ sharply in rare failures, yet estimating lower-tail conditional value-at-risk (CVaR) accurately can require many costly rollouts. When different conditional components of a stochastic workflow can be quer
Probe-Space Preconditioning for Fast and Stable Zero-Order Training
Francois Chaubard, Mykel J. Kochenderfer, Chris Ré
Backpropagation (BP) dominates deep learning but imposes a massive memory tax. For example, training OPT-30B with Adam requires $\approx$ 600GB of GPU memory (assuming batch size 8 and sequence length 2048). Alternatively, zero-order optimization (ZO
Dimensionally consistent surrogate modelling through dimensional analysis and harmonic expansions
Ernest Tarrus, Hector Gisbert
Dimensional homogeneity is a fundamental constraint on physically meaningful models, requiring invariance under changes of units. We present a data-driven method for constructing surrogate models that satisfy this constraint at the level of the hypot
Character Training for Risk-Averse Agents
Arav Dhoot, Punya Syon Pandey, Jamie Johnson +3
Risk aversion in resources could prevent misaligned AI agents from causing catastrophic harm. Misaligned but risk-averse agents would tend to favor safer strategies like making deals with humans over riskier strategies like rebelling. We train agents
Mira: Memory-Efficient MoE Inference Using Adaptive Caching and Predictive Expert Staging
Sanjali Yadav, Bahar Asgari
Mixture-of-Experts (MoE) models are a compelling architecture for scaling model capacity, making them especially attractive for deployment on resource-constrained, single-GPU systems. However, this benefit is difficult to realize because expert param
Neural topology optimization of ship structures under propulsion machinery vibrations
Shengyu Yan, Muhammad Muztahidul Hakim Zareer, Jasmin Jelovica
Ship structural vibrations contribute to noise, fatigue, and equipment damage, while dynamic-compliance topology optimization can produce pathological designs near resonance. This study extends neural-reparameterized topology optimization using a con
Traversing the solution space of neural networks with Hessian Null Space Continuation
Ann Huang, Mitchell Ostrow, Zhouyang Lu +3
On a single task, deep networks can learn many solutions, depending on their optimizer, training data, architecture, and hyperparameters. Many of these solutions are mode-connected: rather than isolated points in weight space, they are connected by l
Optimal Quantum-Classical Separations for Exact Learning
Srinivasan Arunachalam, Amin Shiraz Gilani, Nikhil S. Mande
We study exact learning with membership queries for concept classes $\mathcal C\subseteq\{0,1\}^N$, focusing on the relationships among their deterministic, randomized, and quantum query complexities, denoted $\mathsf{D}(\mathcal C)$, $\mathsf{R}(\ma
Probability is Not Enough: Exploring and Counting Divergent Tokens for Reasoning Uncertainty Quantification in LLMs
Feiyang Li, Shengjing Liu, Qi Zhan +7
As the chain-of-thought reasoning capabilities of large language models improve, evaluating and calibrating their reasoning confidence is becoming increasingly important for quantifying the uncertainty of their answers. Current methods for estimating
A foundation model for energy and radiation systems built on heterogeneous scientific interfaces
Samrendra Roy, Tapas Tripura, Yoon Pyo Lee +2
Scientific foundation models are commonly evaluated after heterogeneous physical problems have already been translated into a compatible gridded, tokenized or symbolic representation. This leaves the scientific interface outside both the pretrained m
Alpha Diffusion Language Models: Factorization Alone Is Not the Problem
Nikita Gushchin, Dmitry Baranchuk, Alexander Korotin
Discrete diffusion language models can generate multiple tokens in parallel, but reducing the number of denoising steps can lead to inconsistent predictions. Standard cross-entropy training fits conditional token marginals, whereas parallel generatio
Jaxolotl: A Unified High-Performance Benchmark Suite for LTL-Based Multi-Task RL
Mathias Jackermeier, Jacques Cloete, Alessandro Abate
Training agents to follow arbitrary instructions is an important goal of multi-task reinforcement learning (RL). Linear temporal logic (LTL) provides a precise and structured formalism for specifying instructions to agents, and has been successfully
Latent Inference-Time Guidance of Time Series Foundation Models
Chloé Hashimoto-Cullen, Amaury Durand, Laurent Bozzi +3
Time Series Foundation Models (TSFMs) currently provide state-of-the-art results in forecasting tasks. They are available out-of-the-box and rely on in-context learning to make their predictions, which makes the quality of their performance highly se
Improving Function Space Flow Matching with Kernel Optimal Transport
Fred Xu, Thomas Markovich, Barbora Barancikova +1
Generative models for function-valued data, such as time series and solutions of partial differential equations, must learn distributions over infinite-dimensional spaces. Functional Flow Matching (FFM) extends Flow Matching to this setting, learning
UserProxyBench: Evaluating LLM User Simulators for Agent Benchmarks and Training
Ashish Jain, Armaan Sandhu
Interactive agent benchmarks and multi-turn reinforcement learning increasingly place a second language model in the role of the user. This simulated user controls what information the agent receives and when, yet current benchmarks score only the ag
Gender bias across LLMs is common and highly heterogenous
Edoardo Bolzoni, Valerio Capraro
Understanding gender biases in large language models (LLMs) is increasingly important as these systems become embedded in decision-support tools with real consequences. Prior research has focused only on a small set of models, leaving open the extent
The finite-horizon five-expert prediction problem
Jeff Calder, Nadejda Drenska
We give an explicit solution to the five expert prediction with expert advice partial differential equation (PDE) in the finite-time horizon setting. The solution formula establishes that the adversary's rank strategy $(1,0,1,0,0)$ is globally optima
doPlan: A Variable-Horizon Dataset for Multi-Stage Language-Conditioned Planning in Autonomous Driving
Parthib Roy, Yash Tandon, Marcus Blennemann +4
Autonomous vehicles interacting with passengers through natural language must reason beyond immediate commands. Passenger intent may span multiple stages of behavior, depend on future events, refer to surrounding agents or landmarks, and remain relev
Dr. OPD: Learning What to Follow for Optimal On-Policy Distillation of Large Language Models
Zhenyu Wang, Tianze Wang, Linjun Zhang +1
On-policy distillation (OPD) trains a student on its own generated responses using dense, token-level supervision from a stronger teacher. Vanilla OPD treats all teacher signals equally, assuming that the teacher's supervision is equally important fo
Retrieval-Augmented Skill Optimization via Cross-Harness Adaptation
Jaewon Chu, Ji Soo Lee, Jihwan Park +8
An agent skill is a reusable, actionable natural-language artifact that guides an agent to perform a task effectively under a given harness. Recent studies have explored the optimization of agent skills, contributing to a growing collection of public
PE-EK-PINN: Physics Embedding with Evolving Kernel for Scalable Physics-Informed Neural Networks
Huiwen Zhang, Feng Ye, Chu Ma
Physics-Informed Neural Networks (PINNs) embed governing equations into deep learning, but enforce them only through loss residuals, leaving highly oscillatory wave behavior to be discovered by optimization. As a result, methods that achieve relative
Auditable Long-Term Memory: A Deterministic Retrieval Chain Measured at 479/475 of 500 on LongMemEval-S
Christopher J. Chanhnourack
We evaluate an auditable long-term memory system on LongMemEval-S. Its retrieval chain uses hybrid candidate retrieval, cross-encoder reranking, coverage-first packet compilation, and deterministic reasoning scaffolds; an LLM is used only as a replac