Theory/Optimization • Score 85
How to Have a Sensitive Debate: An Instance-Optimal Protocol for AI Debate
arXiv:2610.02557v1 Announce Type: new
Abstract: As powerful AI systems reach and sometimes surpass the abilities of human experts across a range of cognitively demanding tasks, the problem of accurate oversight and supervision of these systems has become increasingly urgent. One promising approach is AI debate, which seeks to leverage a debate between two powerful AIs to break complex questions down into simpler claims that can be easily judged directly. Theoretical work on debate has formalized this intuition in the language of computational complexity theory, where the goal is to design protocols (i.e., rules of the debate game) that provide rigorous guarantees on correctness for judging solutions to complex problems with limited supervision. Specifically, the current best protocol has been shown to work for all problems that have sufficiently stable decompositions into subproblems. In this paper, we design a new protocol for this same class of problems that improves on the prior work in several ways. First, correctness holds in a worst-case rather than an average-case sense. Second, being honest and correct is a dominant-strategy equilibrium for both debaters, rather than a Stackelberg equilibrium. Finally, we prove black-box lower bounds, showing that our new protocol is instance-wise optimal. That is, no protocol for this class of problems can outperform ours while making only black-box queries to human judgments. We obtain these results by relating the notion of stable problem decompositions to the concept of fractional block sensitivity from query complexity.
Fonte: arXiv cs.AI
Theory/Optimization • Score 85
Approximation Property of Dropout Neural Networks: Sobolev Rates and Confidence Bounds
arXiv:2610.02253v1 Announce Type: new
Abstract: The universal approximation property of dropout neural networks does not by itself describe the network size required for an accurate random realization. In this work, we study approximation of the unit ball of $W^{n,\infty}([0,1]^d)$ by ReLU networks whose edges are retained independently with probability $p$. The approximation error is measured uniformly over the input domain, and the guarantee holds with probability at least $1-\delta$ for a single sampled network. We construct networks of constant depth and size $\widetilde O_{n,d}(p^{-9}\varepsilon^{-\max\{d/n,2\}} \log(1/\delta))$. The construction combines bounded local subnetworks, localization on a successful approximation event, and a multiscale Taylor decomposition. Conversely, Sobolev capacity imposes a lower bound on the number of surviving edges, while approximation of a fixed affine function requires an output-layer cost of order $((1-p)/p)\varepsilon^{-2}\log(1/\delta)$ at sufficiently high confidence. For fixed $p\in(0,1)$ and $\delta<\min\{1/2,1-p\}$, the upper and lower bounds match in the accuracy exponent under a fixed or logarithmic depth budget. When $d\leq2n$, they also match in confidence up to logarithms of accuracy. We extend the lower bounds to $W^{n,r}$ targets with $L^s$ error, and distinguish this extension from the upper bound for $W^{n,\infty}$. The optimal retention dependence and logarithmic factors remain open.
Fonte: arXiv cs.LG
Theory/Optimization • Score 85
Mitigating Convergence Collapse in Fixed-Target Anomaly Detectors via Kernel-Anchored Locality Regularization
arXiv:2610.02345v1 Announce Type: new
Abstract: A family of tabular anomaly detectors trains a neural map toward a fixed target under squared-error loss and scores anomalies by the test-time residual; contraction matching, one-step rectified flow, and reconstruction autoencoders all fit this template. We characterize a convergence collapse: better optimization makes the detector worse. At convergence, the learned map tracks the target even off-distribution, so the residual signal vanishes on anomalies as well as on normal data. These detectors therefore rely on implicit non-convergence (early stopping, capacity caps) to retain signal. We argue this is structural: effective anomaly detection requires a locality constraint that blocks unconstrained extrapolation. Classical detectors (kNN, KDE, isolation forests, LOF) enforce locality explicitly; fixed-target neural detectors do not. We formalize the connection by showing that the kernel-regression analog of a fixed-target detector is a finite-bandwidth Nadaraya-Watson smoother, which we call Kernel Contraction Matching (KCM). KCM is closed-form, training-free, and CPU-efficient, yet matches established neural baselines on ADBench. Building on this bridge, we introduce the Kernel-Anchored Regularizer (KAR), which penalizes deviation of the neural prediction from a kernel-weighted average of training targets. Across collapse-prone ADBench datasets and three backbones, KAR mitigates collapse and improves AUROC under prolonged training.
Fonte: arXiv cs.LG
NLP/LLMs • Score 85
Harnessing LLMs as Agents: What Does It Cost?
arXiv:2610.02488v1 Announce Type: new
Abstract: Language-model agents increasingly rely on harnesses that manage bounded context, persistent memory, tools, verification, and repeated execution, yet existing notions of model capability do not quantify the computational resources these mechanisms consume. We introduce the Language Model Agent Machine (LAM), a resource-bounded abstraction that fixes the underlying semantic model while explicitly charging harness-level resources. We establish four classes of results. Communication: LAM execution is instancewise equivalent to red--blue pebbling under simultaneous call--transfer budgets, transferring classical I/O lower bounds to context--memory traffic. Access: memory interfaces induce asymptotic separations, including a $\Theta(n)$ gap between random and non-speculative sequential access on pointer chasing. Recomputation: bit-reversal DAGs require $\Theta(n^2/(C+S)+n)$ model calls with context capacity $C$ and persistent-memory capacity $S$, quantifying when stored intermediate state avoids repeated semantic computation. Reliability: we derive tight stage-local sampling bounds, exact imperfect-verification costs, and a Young--Daly-type checkpoint law with a closed-form optimal verification interval. Controlled and held-out experiments on GPT-6 Astra test communication and reliability predictions, including checkpoint optima, policy selection under programmatic checking, and tradeoffs among call granularity, logical input traffic, and reliability on chained MATH tasks. Together, these results provide a resource theory for the computational cost of language-model agent harnesses.
Fonte: arXiv cs.LG
NLP/LLMs • Score 85
DeReAct: Decomposed Reasoning and Acting for Reliable AI Agents
arXiv:2610.02351v1 Announce Type: new
Abstract: ReAct-based agents typically rely on a single LLM policy to propose actions, interact with the environment, and decide when a task is complete. This coupling makes action authorization and completion control difficult to enforce independently, allowing errors to propagate and unsupported completion claims to terminate execution. We introduce DeReAct, a modular agent architecture that externalizes two gating policies: a Critic that validates proposed actions before execution, and a Context Manager that reconstructs an environment-supported \textsc{State} and certifies task completion.
Across GAIA and SWE-bench Verified, DeReAct improves Pass@1 most for weaker Brain models, with gains of 6.5--7.0 points for Qwen3-Coder-480B and 4.2--5.2 points for Claude Sonnet~4.5; gains diminish as Brain capability increases. Trajectory and ablation analyses show that external gating is effective when targeted failures are sufficiently prevalent and the gating policy is itself sufficient. With Claude Opus~4.5, Pass@1 remains comparable to ReAct, while DeReAct produces more evidence-complete and constraint-satisfying trajectories, indicating that completion control can trade earlier termination for stronger grounding. Overall, DeReAct improves weaker agents while retaining grounding benefits as models strengthen.
Fonte: arXiv cs.AI
RL • Score 85
Reinforcement Learning Techniques for the Optimization of Target Polarization in Nuclear Physics Scattering Experiments
arXiv:2610.02452v1 Announce Type: new
Abstract: The operation of dynamically polarized targets in nuclear physics experiments relies on continuous tuning of the microwave frequency to compensate for radiation damage and evolving material properties, a task that is traditionally performed through manual trial-and-error by expert operators. This work presents a data-driven control framework that combines surrogate modeling with reinforcement learning to optimize the target polarization. Using operational data from the APOLLO cryogenic target system, we train and evaluate multilayer perceptron and Gaussian process regression models to predict polarization as a function of microwave frequency, beam current, and accumulated radiation dose. We show that Gaussian process-based models provide calibrated uncertainty estimates and reliably identify regions outside the training distribution, while MLPs exhibit limited sensitivity to distributional shift. To enable learning and control across multiple target samples, we introduce a Gaussian process approximation and embed the surrogate model within a standardized simulation environment. A reinforcement learning agent is trained using a lower-confidence-bound reward formulation that balances performance maximization against uncertainty. We are able to show an almost 2x improvement on the operators actions utilizing our RL agent.
Fonte: arXiv cs.AI
NLP/LLMs • Score 85
Fast Models, Slow Evidence: A Paired and Self-Audited Evaluation of System-1 Decision Models for LLM Agent Harnesses
arXiv:2610.02267v1 Announce Type: new
Abstract: Agent harnesses make many small, typed decisions per task: which model to call, which tool to use, whether retrieved text is relevant, whether an input carries an injection. System-1 decision models answer such questions in a single forward pass with class probabilities, promising large cost and latency savings over LLM calls. We present a paired evaluation of an open-weight (Laya) and a hosted (Jev) System-1 model on 11 agent decision points built from 18 public sources: 7,283 base cases plus 6,640 robustness variants, with byte-identical inputs, paired tests, and cross-hardware and cross-day reproducibility checks. Jev is significantly more accurate on 9 of 11 decision points (+10.8 to +46.0 pp). Neither model beats chance on zero-shot model routing, and they tie on RAG relevance gating. Laya changes 30% of its answers when the option order is reversed and degrades sharply with many or similar candidates (31% at 50 nearest-neighbour tools, vs. 98% for Jev on items with a unique correct tool). We also audit our own pipeline. Three analysis errors and one design confound distorted headline deployment claims: an omitted pre-screen cost (reported 23.9% saving, actual 4.3%), gate accuracy reported as end-to-end quality (58% vs. 98%), in-sample thresholds (5% target, up to 17% held-out misses), and a "channel effect" on injection false positives that vanishes with channel-native content. Two other suspected confounds did not change the conclusions. All cases, raw outputs and analysis code are available at https://github.com/David-DL-Space/sys1-eval.
Fonte: arXiv cs.AI
Theory/Optimization • Score 85
Iterating Consistency Models: Stability, Error Bounds and Noise Schedules
arXiv:2610.03414v1 Announce Type: new
Abstract: Consistency models (CMs) have become a leading approach for generating high-quality samples in few steps. However, adding steps can improve or degrade sample quality in ways that are highly sensitive to the schedule and that existing theory does not fully explain. To provide accuracy guarantees and guide CM sampler design, we analyze multistep CM sampling as a composition of noising and approximate denoising operators. Under explicit, verifiable stability assumptions, we derive a non-asymptotic error bound that separates contraction of the initialization error from accumulation of approximation error. The bound assigns distinct roles to the schedule: large early noise levels drive contraction, while small late noise levels control the residual bias. As a corollary, we obtain explicit constants for strongly log-concave and semi-log-concave targets. We further establish a complementary guarantee whose assumptions, one-step accuracy and stability, can be estimated for a given trained model. Experiments show that the contraction and approximation profiles entering our bounds can be reliably measured and closely match the predicted functional forms. Together, these results provide a meaningful convergence theory for multi-step CMs and a practical route to sampler design.
Fonte: arXiv stat.ML
NLP/LLMs • Score 85
What Does a Token Cost? A Mixture-of-Agents Measurement of Sufficient Per-Token Compute
arXiv:2610.02491v1 Announce Type: new
Abstract: Large language models spend the same amount of computation on every token they generate, regardless of how difficult each token is to produce. Methods such as speculative decoding and model routing are built on the premise that much of this computation is unnecessary, yet the computation an individual token actually requires has not been measured. We measure it through a Mixture-of-Agents (MoA) lens: a panel of fifteen language models of increasing capacity, drawn from three families, in which every agent attempts to reproduce a reference sequence token by token, conditioned on the correct preceding tokens. We define the inference cost of the smallest agent that succeeds as the token's sufficient compute, which upper-bounds what the token requires. On three core benchmarks, a 0.5B agent reproduces 92--95\% of reference tokens. Across Qwen, OLMo, and R1-distilled panels, the most expensive 10\% account for 64--80\% of estimated FLOPs. On all 500 MATH-500 problems, the MoA-derived map helps model routing reduce projected latency from 7.59 to 5.12 seconds while slightly improving accuracy, relative to the best confidence-routing baseline. The MoA-map helps drafting use 32.6\% fewer draft tokens and approximately 20\% lower projected latency than fixed-window drafting at similar accuracy. These comparisons reveal remaining allocation headroom, motivating controllers that exploit sufficient-compute structure.
Fonte: arXiv cs.AI
Theory/Optimization • Score 85
$\Psi$-Resilience: Model-Free Feature Importance from 1D Topological Signals
arXiv:2610.02299v1 Announce Type: new
Abstract: We introduce $\Psi$-Resilience, a model-free feature importance method that derives explanations directly from the data itself via 1D topological signals. Our method constructs a class-disagreement landscape by estimating class-conditional densities and taking their pointwise absolute difference along the feature axis. Then, the 0-dimensional persistence of this 1D signal defines a resilience functional that aggregates only those topological features that survive perturbations up to a robustness scale which is set by the user. This gives us a context-robust importance score that is inherently auditable via the underlying 1D landscapes and their persistence. We evaluate our method on both synthetic and real datasets. On synthetic generators with specified ground-truth importance, $\Psi$-Resilience recovers the ranking of features with high fidelity, achieving Spearman rank correlations up to 0.8 and performing competitively with multiple feature importance methods, including SHAP and mutual information. On real datasets with no known ground truth, our technique agrees with these methods, with correlations up to 0.9. These results show that $\Psi$-Resilience is a stable explanation method that enables rigorous, distribution-level auditing of feature importance without relying on a predictive model.
Fonte: arXiv cs.LG
NLP/LLMs • Score 85
Revealing Epistemic Uncertainty in MLLMs via Causal-Invariant Masking
arXiv:2610.02887v1 Announce Type: new
Abstract: Multimodal Large Language Models (MLLMs) suffer from hallucinations, creating a critical need for Uncertainty Quantification (UQ) to ensure reliable deployment. However, existing approaches struggle to detect uncertainty caused by superficial associations, especially when the query-relevant signal is weak. We mainly attribute this issue to their bias toward aleatoric uncertainty arising from data ambiguity, overlooking epistemic uncertainty stemming from model limitations. To further decompose uncertainty types for a comprehensive UQ, we propose Causal-Invariant Masking (CIM), which measures the semantic shift between the original predictions and those conditioned on a causally-focused view. Based on this framework, we introduce Semantic Divergence as our core metric for UQ and provide theoretical evidence that it converges to the variance of model's sensitivity to non-causal correlations, establishing its ability to capture MLLM's limitation. To accelerate UQ in MLLMs, we further propose Expected Embedding Drift (EED), a fast geometric proxy metric that estimates semantic shift directly within the hyperspherical embedding space. Experiments show that our method achieves state-of-the-art performance on various benchmarks, while the proposed EED accelerates by nearly 50% with comparable performance.
Fonte: arXiv cs.CV
NLP/LLMs • Score 85
HyperThink: Text-to-Parameter Hypernetworks for Efficient Reasoning
arXiv:2610.03039v1 Announce Type: new
Abstract: Long-form thinking traces can substantially improve the multi-step reasoning performance of large language models (LLMs), but they introduce high inference-time overhead, with latency dominated by sequential decoding. We propose HyperThink, a text-to-parameter approach that amortizes this reasoning computation into a single query-conditioned parameter update: a lightweight hypernetwork reads the question and predicts updates to a small subset of the base LLM's parameters, while a vector-quantized decoder constrains them to a finite set of reusable patterns to improve robustness and transfer. Trained end-to-end on outputs from the base model itself, HyperThink eliminates long thinking traces at test time: after one hypernetwork forward pass, the adapted model generates a concise step-by-step solution and final answer without an intermediate trace, using far fewer tokens while retaining strong reasoning performance. Empirically, HyperThink improves the low-latency region of the accuracy-latency trade-off on mathematical and general reasoning tasks, with its strongest gains in the near-non-thinking regime.
Fonte: arXiv cs.CL
Theory/Optimization • Score 85
A fast non-reversible sampler for Bayesian mixture models
arXiv:2510.03226v2 Announce Type: replace-cross
Abstract: Mixtures models are a cornerstone of Bayesian modelling, and it is well-known that sampling from the resulting posterior distribution can be a hard task. In particular, popular reversible Markov chain Monte Carlo schemes are often slow to converge when the number of observations $n$ is large. In this paper we introduce a novel and simple non-reversible sampling scheme for Bayesian mixture models (with fixed or varying number of components), which is shown to drastically outperform classical samplers in many scenarios of interest, especially during convergence phase and when components in the mixture have non-negligible overlap. At the theoretical level, we show that the performance of the proposed non-reversible scheme cannot be worse than the standard one, in terms of asymptotic variance, by more than a factor of four; and we provide a scaling limit analysis suggesting that the non-reversible sampler can reduce the convergence time from O$(n^2)$ to O$(n)$. We also discuss why the statistical features of mixture models make them an ideal case for the use of non-reversible discrete samplers.
Fonte: arXiv stat.ML
Theory/Optimization • Score 85
Classical and Quantum Speedups for Non-Convex Optimization via Energy Conserving Descent
arXiv:2604.13022v2 Announce Type: replace-cross
Abstract: We present the first analytical study of ECD, focusing on the one-dimensional setting for this first installment. We formalize a stochastic ECD dynamics (sECD) with energy-preserving noise, as well as a quantum analog of the ECD Hamiltonian (qECD), providing the foundation for a quantum algorithm through Hamiltonian simulation in a tractable model where the barrier-crossing mechanism can be computed explicitly. For one-dimensional double-well objectives in the under-guessing regime, we compute the expected dynamical hitting times from a local minimum to the global minimum. We prove that both sECD and qECD exhibit exponential improvements in continuous hitting time relative to their respective gradient-based baselines, stochastic gradient descent (SGD) and quantum tunneling walk (QTW). For objectives with tall barriers, qECD admits a further hitting time improvement over sECD. Mechanistically, ECD sidesteps the exponential cost associated with rare-escape events of SGD from local minima by moving from dissipative to energy-conserving dynamics.
Fonte: arXiv stat.ML
RL • Score 85
Lexicographic Multi-Objective On-Policy Distillation
arXiv:2610.02359v1 Announce Type: new
Abstract: Reinforcement learning from verifiable rewards (RLVR) usually optimizes answer correctness, yet useful language-model behavior also requires high-quality reasoning and concise responses. Existing multi-reward post-training methods typically scalarize rewards or combine specialists without explicitly protecting a reward priority order. This is problematic when trade-offs are asymmetric: conciseness, for example, should not improve at the cost of correctness. We introduce Lexicographic Multi-Objective On-Policy Distillation (LMOPD), a multi-teacher method for integrating reward-specialized policies under explicit priorities. For each student rollout, LMOPD selects the specialist for the first objective whose gate detects a deficiency, then locally projects its centered log-policy correction to remove components that oppose higher-priority specialists. We evaluate 30B-A3B mixture-of-experts transformer models in two- and four-expert settings on three math benchmarks, measuring retained specialist gains. With two experts, LMOPD's point estimates fully retain the accuracy and reasoning-quality gains while acquiring $46.9\%$ of the conciseness gain. With four experts, it retains $\approx90\%$ of both the accuracy gain and reasoning-correctness gain, compared to only $\approx57\%$ by the next best evaluated baseline. Matched four-expertablations show that lexicographic routing outperforms random routing and that projection further strengthens both top-priority capabilities. Across both scales, LMOPD preserves the highest-priority capabilities more effectively than the existing baselines we evaluate, demonstrating the value of explicit priorities for specialist integration.
Fonte: arXiv cs.LG
Theory/Optimization • Score 85
Amortized Structured Stochastic Variational Inference for Gaussian Process Latent Variable Models
arXiv:2610.03647v1 Announce Type: new
Abstract: Many machine learning methods aim to approximate the lower-dimensional manifold on which the data lives. A desirable feature of such methods is that they should capture the epistemic uncertainty of this learned manifold. One model that achieves this is the Gaussian Process Latent Variable Model, in which a Gaussian Process (GP) mapping from the latent space provides an estimate of the uncertainty of the manifold. However, the effectiveness of this uncertainty estimation is limited by the mean-field variational approximation between the GP inducing points and the latent variables. In this work, we apply Amortized Structured Stochastic Variational Inference to allow the variational posterior for the latent space to be conditionally dependent on the value of the inducing points. We demonstrate that this more flexible variational posterior improves several metrics relating to the reconstruction of points on the data manifold.
Fonte: arXiv stat.ML
Theory/Optimization • Score 85
A Residual Tree Gaussian Process Modeling Framework for High-Dimensional Data
arXiv:2610.02893v1 Announce Type: cross
Abstract: With the advance of measurement technologies and increasing computing power, large spatial data with heterogeneous structures are often collected over high-dimensional domains. Existing Gaussian process (GP) models and computational strategies are often inadequate for analyzing such datasets in multi-dimensional domains. To address these challenges, we develop a Bayesian residual tree GP methodology called ResTGP for large spatial data with potentially heterogeneous structures in multi-dimensional domains. The key idea is to decompose a Gaussian process at a cascade of resolutions along a dyadic tree through iteratively computing predictive and residual processes so that the residual process on each tree node, both interior and leaf, becomes sufficient for the finer-level dependency within that node. This allows characterization of the underlying covariance structure in a flexible, multi-scale manner while achieving divide-and-conquer on the data domain, which leads to computational efficiency. To allow efficient tree inference, we introduce a computational strategy for Bayesian inference based on recursive message passing, which scales linearly with the sample size given the tree. This paper also proves posterior consistency of the model for estimating continuous functions in a nonparametric regression framework. Extensive numerical examples and the storm surge application confirm the advantages of the proposed method.
Fonte: arXiv stat.ML
Theory/Optimization • Score 85
Cross-Fitting Under Nonregularity: Normality and Inference via Locality
arXiv:2610.02944v1 Announce Type: cross
Abstract: Cross-fitting is routine in much of applied research. While conventional confidence intervals that ignore cross-fold dependence are asymptotically valid in several settings, they undercover in many applications that share a common form of nonregularity: from the classic cross-validation problem of testing whether a fitted model outperforms another, to testing for heterogeneous treatment effects with machine learning, to estimating the value of a potentially non-unique optimal treatment regime. Exploiting a new locality condition, I show that a large class of cross-fitting estimators still satisfies a central limit theorem despite the nonregularity, but with an asymptotic variance that must be adjusted for the cross-fold correlation. Then, I propose a method for estimating this correlation and construct new confidence intervals that attain asymptotically nominal coverage. Finally, I show that the proposed confidence intervals attain approximately nominal coverage in a simulation study with random forests and neural networks.
Fonte: arXiv stat.ML
Theory/Optimization • Score 85
Asymptotic Performance of Time-Varying Bayesian Optimization
arXiv:2505.13012v3 Announce Type: replace
Abstract: Time-Varying Bayesian Optimization (TVBO) is the go-to framework for optimizing a time-varying black-box objective function that may be noisy and expensive to evaluate, but its excellent empirical performance remains to be understood theoretically. Is it possible for the instantaneous regret of a TVBO algorithm to vanish asymptotically, and if so, when? We answer this question of great importance by providing upper bounds and algorithm-independent lower bounds for the cumulative regret of TVBO algorithms. In doing so, we provide important insights about the TVBO framework and derive sufficient conditions for a TVBO algorithm to have the no-regret property. To the best of our knowledge, our analysis is the first to cover all major classes of stationary kernel functions used in practice.
Fonte: arXiv stat.ML
NLP/LLMs • Score 85
Overcoming Challenges of Interpretive Structural Modeling with Large Language Models
arXiv:2610.02254v1 Announce Type: new
Abstract: Interpretive Structural Modeling (ISM) is a well-known process for multi-criteria decision making. The success of ISM over other methodologies is its ability to model causal relationships, the binary scale of factors, and resulting hierarchical representation. Traditionally, the modeling process is performed by repeated interactions with subject matter experts until consensus is reached. This process is tedious, labor-intense, and most importantly limits the ability of ISM to scale to studies with hundreds of variables. Drawing on existing work of causal graph discovery with large language models (LLM) as imperfect experts, this work explores an integrated LLM-ISM approach for ISM. Pairwise, k-wise, rowwise, and full graph discovery methodologies are compared and evaluated. It is shown that causal graph discovery methods for ISM perform best using rowwise (SHD=160, F1-score=0.77) and full graph methods (SHD=135, F1-score=0.73).
Fonte: arXiv cs.LG
Theory/Optimization • Score 85
Near-Optimal Convex Optimization with Lazy Second-Order Oracles
arXiv:2610.03222v1 Announce Type: cross
Abstract: This paper studies the complexity of convex optimization using lazy second-order oracles (Doikov, Chayti, and Jaggi, ICML 2023), where an algorithm queries gradients every iteration and Hessians once per $m$ iterations. Under this setting, we show a lower bound of $\Omega(m+ m^{1/7} \epsilon^{-2/7})$ on the number of total iterations to find an $\epsilon$-solution using a novel block zero-chain construction. Then we propose a novel method that achieves a new upper bound of $\tilde{\mathcal{O}}(m+ m^{1/7} \epsilon^{-2/7})$, which significantly improves the prior one (Chen, Liu, Luo, and Zhang, COLT 2026) of $\tilde{\mathcal{O}}(m+ m^{13/21} \epsilon^{-2/7})$ and is tight up to logarithmic factors.
Fonte: arXiv stat.ML
Theory/Optimization • Score 85
Diffusion Flow Matching: Dimension-Improved KL Bounds and Wasserstein Guarantees
arXiv:2606.16610v2 Announce Type: replace
Abstract: Diffusion Flow Matching (DFM) has recently emerged as a versatile framework for generative modeling, yet its theoretical convergence properties remain only partially understood. In this work, we provide refined and novel convergence guarantees for Brownian motion based DFMs, focusing on the discretization error. Our analysis is conducted under the Kullback-Leibler (KL) divergence and the 2-Wasserstein distance. Under finite-moment conditions and a mild score integrability assumption, we derive KL convergence bounds with improved dimensional dependence compared to prior work, achieving, up to our knowledge, state-of-the-art scaling under minimal conditions. We further extend the analysis to the 2-Wasserstein distance: under an additional first-order score integrability assumption and a weak log-concavity condition, we obtain convergence guarantees with dimensional dependence consistent with the KL case.
Fonte: arXiv stat.ML
Theory/Optimization • Score 85
Invariance of Clustering Operations in Causal Effect Identification
arXiv:2610.03101v1 Announce Type: new
Abstract: Clustering variables in causal graphs reduces the size of the graph and simplifies causal inference. However, arbitrary clustering can alter crucial causal relations among variables and lead to erroneous conclusions. While the identifiability of a causal effect in the clustered graph implies the identifiability in the original graph under mild conditions, nonidentifiability in clustered graph does not imply nonidentifiability in the original graph without further assumptions. When both identifiability and nonidentifiability are preserved, the clustering operation is called identification invariant. We present a broad class of clustering operations that are identification invariant based on conditions related to the c-components of the original graph. Finally, we demonstrate use of the results in practical settings.
Fonte: arXiv stat.ML
NLP/LLMs • Score 85
HARPO: Hallucination-Aware Reinforcement Learning for Faithful and Creative Language Generation
arXiv:2610.03063v1 Announce Type: new
Abstract: Large Language Models (LLMs) are prone to generating hallucinated content, which compromises their reliability in knowledge-intensive tasks. To address this challenge without sacrificing creativity, we propose HARPO, a reinforcement learning framework designed to jointly optimize faithfulness and creativity. HARPO incorporates a Hallucination-Aware Generative Reward Model (HA-GRM), trained via verifiable feedback, to assess both faithfulness and writing quality. A Selective Activation Mechanism (SAM) activates writing rewards only for outputs judged hallucination-free by HA-GRM, while a data curriculum progressively shifts training from creative writing to hallucination-centric tasks. On RAGTruth, our Qwen3-4B-based HA-GRM achieves a response-level F1 score of 78.08%, compared with 66.37% for the supervised fine-tuning baseline. Experiments on Qwen2.5 and Qwen3 models from 1.7B to 8B parameters show improvements in both faithful generation and writing quality. On Qwen3-4B, HARPO reduces the HA-GRM-judged hallucination rate on MultiHopRAG from 3.29% to 1.02%, while increasing the Arena-Hard-v2.0 creative-writing score from 16.95% to 27.54%.
Fonte: arXiv cs.CL
Theory/Optimization • Score 85
Predictively Oriented Gaussian Process Posteriors
arXiv:2610.03201v1 Announce Type: new
Abstract: Gaussian Processes (GPs) are a powerful tool for modelling and quantifying uncertainty in functional relationships. However, they require practitioners to make a number of design decisions, such as the choice of the kernel and the observation model. Suboptimal choices can produce misspecified models that do not capture the underlying data generating process. We introduce Predictively Oriented Gaussian Processes (PrO-GPs), which treat predictive uncertainty as the primary inferential target and provide a robust alternative to standard GPs. Although direct computation of a PrO posterior for nonparametric models is intractable, we derive a reduced formulation and practical sampling scheme for efficient computation. Through synthetic and real data experiments, we show that PrO-GPs produce better calibrated predictive distributions under model misspecification compared to standard GP approaches.
Fonte: arXiv stat.ML
MLOps/Systems • Score 85
SpectralCache: Accelerating Diffusion-Based World Models via Spectral Feature Caching
arXiv:2610.02660v1 Announce Type: new
Abstract: Diffusion-based world models enable high-quality interactive environment generation but suffer from substantial inference overhead due to repeated Transformer evaluations during denoising. Existing caching methods mainly exploit temporal redundancy at the feature or token level, leaving the underlying mathematical structure of diffusion features largely unexplored. In this work, we reveal that world-model features exhibit highly stable singular subspaces across nearby denoising steps, while their singular values follow predictable evolution patterns. Building on this observation, we propose SpectralCache, a training-free spectral caching framework that reuses stable singular subspaces and estimates only low-dimensional singular values through linear extrapolation. We further exploit the spectral consistency between neighboring full-computation features to skip selected expensive backbone evaluations via singular value scaling. Extensive experiments on representative world models demonstrate that SpectralCache consistently improves inference efficiency while preserving generation quality. On HunyuanWorld-Voyager-13B, SpectralCache achieves 5.22x acceleration while maintaining a WorldScore of 65.90 for static scenes, substantially outperforming existing training-free caching methods in inference efficiency.
Fonte: arXiv cs.CV
Theory/Optimization • Score 85
From Behavior to Provenance: Attributing Tabular Foundation Models to Synthetic Pretraining Data
arXiv:2610.02347v1 Announce Type: new
Abstract: Training-data attribution aims to identify which training examples shape model behavior, yet validating such claims is difficult because causal training influence is rarely observable. We argue that controlled synthetic pretraining makes attribution experimentally testable. Using O'PRIOR, a provenance-rich synthetic task generator for tabular foundation models, we construct a testbed in which every pretraining task carries explicit lineage over structural mechanisms, missingness, confounding, shortcuts, and distribution shift. We combine behavior-conditioned attribution with counterfactual retraining and provenance-aware interventions to test both task-level faithfulness and mechanism-level consistency. On held-out real tasks, removing the top-attributed 5% of synthetic tasks decreases mean ROC-AUC by 0.013, compared with 0.002$\pm$0.004 under random removal, while removing bottom-attributed tasks improves performance by 0.003. Within shortcut-provenance tasks, targeted removal yields an effect of 0.043 versus 0.016 for matched random removal. Provenance discrimination is more modest by ranking AUROC (0.55-0.62), despite substantial top-k enrichment, revealing that provenance association and interventional faithfulness need not coincide. Our results establish synthetic provenance as a controlled setting for verifiable contributive attribution
Fonte: arXiv cs.LG
NLP/LLMs • Score 85
AREX: Affine-Residual Exponential Integrator for Few-Step Sampling in Flow Matching
arXiv:2610.03483v1 Announce Type: new
Abstract: We introduce AREX, a training-free sampler for pretrained flow matching models that uses the target mean and covariance to capture an analytically tractable part of the sampling dynamics. We show that the velocity field of the moment-matched Gaussian target is the $L^2$-optimal affine approximation to the marginal velocity field. This motivates decomposition of the learned dynamics into an affine component over the whole sampling path, determined by the first two target moments, and a neural residual term. AREX keeps the affine component and integrates it using an explicit matrix-valued propagator. In turn, we only require to integrate over the residual term. This differs from scalar exponential integrators, which analytically handle only isotropic linear dynamics. Across image and text-to-image generation tasks, AREX consistently improves sample fidelity in the few-step sampling regime without retraining the underlying model.
Fonte: arXiv stat.ML
Theory/Optimization • Score 85
Simulation-Free Learning of Population Dynamics with Wasserstein Lagrangian Residuals
arXiv:2610.03679v1 Announce Type: cross
Abstract: The dynamics of cells, organisms, and fluids are often modeled as probability distributions evolving over time. Reconstructing and extrapolating this evolution from unpaired snapshots requires assumptions about the underlying process. Wasserstein gradient flows are a common choice, but they cannot describe conservative or periodic dynamics. Lagrangian mechanics in Wasserstein space covers both, but existing methods for learning it are simulation-based: they run a numerical solver at every training step, which makes training expensive. We propose Double-Stitch, a simulation-free method that learns these mechanics by penalizing the residual of the equation of motion along a learned population path. We derive this equation from a Clebsch variational principle that does not require gradient velocities, and show that the residual vanishes exactly when the equation holds. We test Double-Stitch on synthetic, single-cell and ocean vortex datasets and find that it matches or outperforms gradient-flow methods and simulation-based WLM on most tasks, while training $4$-$14$ times faster than WLM. We provide a JAX implementation of Double-Stitch at https://github.com/BasisResearch/stitching.
Fonte: arXiv stat.ML
Applications • Score 85
TerrainForge: Physics-Grounded road geometry Editing for Counterfactual Autonomous Driving
arXiv:2610.02825v1 Announce Type: new
Abstract: Road geometry (e.g., crests, sags, and speed humps) and surface conditions (e.g., wet or icy pavement) affect how vehicles move, what drivers and onboard cameras observe, and how much clearance remains between vehicles. Editing these properties in a driving scene therefore requires corresponding changes in vehicle motion. Capturing these differences in a driving video requires a road edit to propagate to vehicle motion, camera viewpoint, and the clearance between vehicles. We present TerrainForge, a framework for generating road geometry-focused counterfactuals from reconstructed multi-vehicle driving episodes. A unified road model connects scene deformation with four-wheel vehicle dynamics, allowing crests, sags, speed humps, and friction changes to propagate through vehicle motion, camera viewpoint, and inter-vehicle clearance. Vehicle dynamics are evaluated against CarSim, and prescribed road geometry is verified in reconstructed Waymo scenes. Across 18 episodes, leaving surrounding vehicles on their recorded trajectories instead of recomputing their responses produces median peak differences in predicted ego-lead distance of 1.52 m for crests and 1.41 m for sags. We further simulate the ego response to 15,758 road edits across 983 braking episodes, pairing each edit with its safety outcomes relative to an unedited replay. These pairs train a first-stage screening surrogate that takes the original driving context and candidate road-edit parameters as input and predicts the resulting change in the ego's terminal gap. On held-out scenes, this prediction achieves 22-40% lower mean absolute error than predicting no change, so candidates can be screened cheaply before the full multi-vehicle rollout.
Fonte: arXiv cs.CV
NLP/LLMs • Score 85
Conformal Prediction for Time Series with Deep Sequence Models
arXiv:2610.02357v1 Announce Type: new
Abstract: Recent advances in deep learning for time series prediction have amplified the need for reliable uncertainty quantification. Conformal prediction has gained attention as a distribution-free framework for constructing prediction intervals with coverage guarantees. However, its coverage guarantees rely on data exchangeability, an assumption generally violated in time series. Active research has focused on developing conformal prediction methods for time series that overcome this limitation. While deep sequence models, such as recurrent neural networks and Transformers, have often been used in conformal prediction for time series, limited work has systematically studied how deep sequence models can be utilized in conformal prediction for time series. In this work, we systematically investigate the use of deep sequence models in conformal prediction for time series through three approaches: conditional quantile regression, conditional quantile function estimation, and localized conformal prediction. We provide a theoretical analysis establishing asymptotic conditional coverage guarantees for all three approaches under suitable assumptions. Through comprehensive experiments on real-world datasets, we demonstrate the effectiveness of leveraging deep sequence models into conformal prediction for time series.
Fonte: arXiv stat.ML
NLP/LLMs • Score 85
ViTok: Improving Dense Semantics in AM-RADIO-Style Multi-Teacher Distillation with PHI-S and Masked Image Modelling
arXiv:2610.02903v1 Announce Type: new
Abstract: We study how to consolidate the current VITOK progress into a single multi-teacher distillation recipe that jointly preserves global recognition and dense semantics. Our starting point is an AM-RADIO-style student distilled from SigLIP2 and DINOv3-L, where SigLIP2 supplies strong global semantics and DINOv3-L supplies stronger dense features. The central empirical issue is that the same recipe does not optimize all objectives equally well: changes that improve ImageNet-1K kNN accuracy can still degrade ADE20K segmentation. We summarize a progression of modifications that make this trade-off more explicit and more manageable: split adaptor heads for CLS and patch tokens, asymmetric cosine/MSE losses, initialization from a DINOv3-L checkpoint, teacher reweighting, masked image modeling (MIM), and PHI-S feature balancing. The resulting model reaches 83.2 patch kNN and 85.2 CLS kNN, slightly surpassing the DINOv3-L teacher on ImageNet-1K kNN classification, while PHI-S restores ADE20K performance from 46.5/58.1 to 48.5/61.0 mIoU/mAcc, matching the teacher on this dense benchmark. We also summarize negative results: scaling distillation from ImageNet-1K to ImageNet22K does not consistently help, and naively adding extra teachers such as SAM3 or HOG features introduces interference. Rather than claiming a final recipe, this paper distills the current project state into a compact empirical story and a concrete set of lessons for future iterations.
Fonte: arXiv cs.CV
Vision • Score 85
Revisiting Visual Representation Enhancement of VLMs via Kernel Canonical Correlation Analysis
arXiv:2610.02718v1 Announce Type: new
Abstract: Vision-language models such as CLIP exhibit strong semantic generalization, but remain limited in fine-grained visual perception. A recent work named KUEA presents a natural remedy by finetuning the image encoder under the supervision of the vision-centric DINOv2 to align their kernel matrices element-wisely, while regularizing the embeddings to remain close to the pretrained visual encoder for preserving image-text semantics in CLIP. However, we show that diminishing the role of the alignment loss to DINOv2 does not necessarily degrade its fine-grained visual performance, suggesting that the kernel-matrix discrepancy may be insufficient for further visual representation enhancement, motivating us to revisit the alignment formulation. In this work, we present a novel perspective to characterize representation alignment on feature subspaces through Kernel Canonical Correlation Analysis (KCCA), which maximizes the projection correlations. In optimization, we derive an efficient end-to-end training scheme upon KKT conditions, avoiding the eigenvalue problem in KCCA. Further, we extend our method into a 3-view formulation, i.e., 3vKCCA, in which the projections from the pretrained text encoder are also incorporated under a unified optimization framework for joint alignment. With CLIP ViT-L/14 on ImageNet-1K, our 3vKCCA improves the MMVP-VLM accuracy from 17.8 to 25.9, substantially outperforming the existing methods, and meanwhile maintains zero-shot image--text retrieval performance.
Fonte: arXiv cs.CV
Vision • Score 85
GRAFT: Growing Agglomerative Foundation Models via Continual Teacher Distillation
arXiv:2610.02597v1 Announce Type: new
Abstract: Vision foundation models such as DINOv2, SigLIP2, and MASt3R develop complementary capabilities from different pretraining objectives, yet their knowledge remains distributed across separate, specialized models. Multi-teacher knowledge distillation offers a path toward consolidating these capabilities into a single agglomerative backbone, but existing approaches assume a fixed set of teachers, and incorporating a new teacher requires repeating expensive joint distillation over the entire teacher set. We introduce GRAFT, a continual multi-teacher distillation framework that enables a unified backbone to progressively acquire capabilities from an open-ended sequence of foundation models. When a new teacher arrives, GRAFT treats the previously distilled model as a teacher for preserving learned capabilities, while the current student jointly learns from both the previous model and the incoming teacher. Furthermore, to reconcile the incompatible representation geometries of heterogeneous teachers, we introduce Teacher Specific Readout Tokens, which grant each teacher an independent read-out of the shared encoder, together with Geometry Agnostic Relational Loss that aligns a vision-language teacher by matching image-text similarity structures rather than raw feature values. We provide GRAFT model, which is a single, continually extensible backbone that unifies five domains, including image understanding, 2D dense prediction, 3D human pose estimation, 3D vision, and vision-language, delivering strong performance across all of them while acquiring each new capability at the cost of a single distillation rather than a full re-distillation.
Fonte: arXiv cs.CV
Vision • Score 85
CHASE-VLA: Post-Training Quantization Framework for Vision-Language-Action Models with Chunk-Aware Scale Estimation
arXiv:2610.02666v1 Announce Type: new
Abstract: Vision-Language-Action (VLA) models map visual observations and language instructions to continuous robot actions, but a diffusion-based action expert (AE) poses a key challenge for low-bit post-training quantization (PTQ). The AE is repeatedly invoked across denoising steps and policy queries, where fixed calibration scales can be mismatched with activation ranges that vary with denoising progress and intended motion. We propose CHASE-VLA, a chunk-aware PTQ method that exploits a VLA-specific signal readily available from the policy: the generated action chunk, including its unexecuted future suffix. Rather than relying only on static scale matching for AE layers, CHASE-VLA combines the previously generated chunk as causal action context with denoising step group information to adapt AE activation scales. This enables W4A4 quantization of both MLP and attention projections in the repeated AE without modifying the pretrained policy. On LIBERO, CHASE-VLA achieves 97.3% average success rate on $\pi_{0.5}$ when both MLP and attention projections in the AE are quantized to W4A4, restoring FP16-level performance. CHASE-VLA also reduces the weight storage of the quantized AE linear layers by 73.4% and their single-chunk memory traffic by 70.9% and 71.2% on $\pi_{0.5}$ and GR00T N1.6, respectively, with a predictor overhead of at most 1.26% of the saved storage.
Fonte: arXiv cs.CV
NLP/LLMs • Score 85
Spatial Memory Intelligence: Endowing World Models with Understanding-Driven Long-Term Memory
arXiv:2610.02521v1 Announce Type: new
Abstract: Long-video generation and world models have shown strong potential for interactive entertainment and embodied simulation by predicting future observations conditioned on user actions and historical memory. However, as memory sequences grow longer and their structures become increasingly complex, managing long-range spatial context becomes increasingly challenging, calling for a more intelligent and systematic memory-management strategy. Building on the advancing spatial reasoning capabilities of multimodal large language models (MLLMs) and the broader vision of unified models, we propose Spatial Memory Intelligence (SMI), the first framework to systematically employ an understanding model for spatial-memory management in long-video world models. SMI introduces four coordinated atomic operations: spatial clustering, within-cluster sparsification, action-aware retrieval, and reliability-aware filtering. Extensive experiments across multiple baselines, benchmarks, and world-model backbones demonstrate the effectiveness and generalizability of SMI, achieving comprehensive improvements in memory sparsity, generation stability, and spatial consistency.
Fonte: arXiv cs.CV
NLP/LLMs • Score 85
MEA: A Reward-Driven Multi-Agent System for Faithful Model Explanations
arXiv:2610.02480v1 Announce Type: new
Abstract: Recent years have seen the employment of a plethora of machine learning (ML) models in high-stakes domains, but they remain largely opaque to the practitioners who act on their predictions. While post-hoc explanation methods offer a lens into this model behavior, wielding them effectively demands expertise most domain experts lack: navigating high-dimensional outputs, selecting the best explanations, and synthesizing evidence across disparate tools. To this end, we present MEA, a multi-agent framework that removes the explanation knowledge barrier entirely: a Proposer agent selects and configures explanation tools based on the question and modality, while an Actor agent is optimized end-to-end against faithfulness, transforming the outputs into natural language explanations grounded in model behavior across tabular, text, and vision modalities. Further, we introduce diverse question types spanning feature attribution, counterfactual reasoning, and spurious feature detection, each paired with a perturbation-based faithfulness metric. We find that frontier LLMs systematically produce unfaithful explanations. By optimizing against faithfulness rewards augmented with a modality-adaptive penalty, MEA consistently outperforms post hoc explainers, agentic, and closed-source baselines across six datasets, with reward-driven optimization yielding faithfulness gains of +28% (tabular), +21% (text), and +34% (vision) over the untrained backbone. More broadly, our findings suggest that AI agents themselves can serve as a scalable, adaptable interface to ML explainability, opening a path toward natural-language explainability that generalizes beyond the fixed, single-purpose tools that have long defined the field.
Fonte: arXiv cs.AI
RecSys • Score 88
THPL: A Vision-to-Language Decision Support Framework for Rainbow Trout Feeding Management in RAS
arXiv:2610.02378v1 Announce Type: new
Abstract: In Recirculating Aquaculture Systems (RAS), precision feeding is critical for minimizing costs and improving fish welfare. However, existing methods lack cognitive alignment between fish behaviors and management knowledge, impeding translation into executable, interpretable feeding decisions. To address this, we propose THPL, a generative feeding decision framework tailored for rainbow trout (Oncorhynchus mykiss) in RAS. First, Fishsort extracts trajectories to establish an Activity Coefficient (AC) quantifying feeding intensity. Second, a Hierarchical Behavior Encoder (HBE) models individual temporal progression and collective dynamics using Temporal and Set Transformers, transforming trajectory tensors into dual-evidence representations of explicit physical and implicit soft tokens. Finally, these tokens are integrated with environmental parameters, metadata, and expert rules to fine-tune an LLM via LoRA, followed by counterfactual multimodal Direct Preference Optimization (mDPO) to reinforce causal reasoning. Results show that AC exhibits a statistically significant monotonic positive correlation with expert-annotated feeding intensity (Spearman $\rho = 0.925$, $p < 0.001$). Ablations indicate that decision accuracy improves from 33.33% (text-only baseline) to 93.33% with dual-evidence tokens, confirming that continuous spatiotemporal tokens provide necessary physical grounding for LLMs. Compared with standard LoRA, counterfactual mDPO elevates decision accuracy from 93.33% to 96.67%, advances METEOR from 58.10% to 85.30%, reduces Self-BLEU-2 from 58.79% to 52.88%, and increases Distinct-3 from 6.68% to 7.81%, suppressing templating and actuation biases while reinforcing causal consistency and operational safety. Overall, by integrating continuous kinematics with LLM reasoning, this study provides a novel decision support paradigm for precision aquaculture.
Fonte: arXiv cs.AI
NLP/LLMs • Score 85
Adaptive Mutual Distillation for Balanced Multi-Task Post-Training of Large Language Models
arXiv:2610.02856v1 Announce Type: new
Abstract: Multi-task post-training of large language models (LLMs) aims to improve performance across tasks with unequal amounts of training data. Existing methods focus primarily on balancing task contributions during single-model training. Different task-balancing strategies can produce models with complementary strengths, creating opportunities for mutual distillation. However, the usefulness of cross-model supervision can vary across tasks, transfer directions, and stages of training. We propose Adaptive Mutual Distillation (AMD), a collaborative post-training framework that jointly trains two models with different task-balancing strategies. AMD evaluates candidate adjustments to distillation weights through short training probes shared across tasks, then uses task-wise validation scores to select an adjustment for each task and transfer direction. Across six benchmarks and three LLM backbones, both AMD models achieve higher average benchmark scores than supervised fine-tuning (SFT) baselines trained with the same sampling strategies. They also outperform the task-balancing methods evaluated in our experiments. Merging the two trained models can further improve their average benchmark score while yielding a single model for inference. The merged models outperform multi-task SFT by an average of 2.91 points across the three backbones.
Fonte: arXiv cs.CL
NLP/LLMs • Score 85
Improving Atomic-Fact Recall via Focused Views in Unstructured Knowledge Editing
arXiv:2610.02772v1 Announce Type: new
Abstract: Large language models (LLMs) increasingly serve as general-purpose interfaces to factual knowledge, but their parameters do not automatically reflect information that changes after pretraining. Knowledge editing (KE) provides a targeted alternative to costly retraining by modifying selected knowledge and preserving unrelated knowledge and general capabilities. Conventional KE uses structured factual triples, whereas unstructured KE (UKE) uses free-form passages containing multiple facts. Nonetheless, existing UKE editors exhibit a failure mode known as context reliance: edited LLMs can often reproduce the editing passage but fail to reliably recall its individual facts without the original passage context. We identify context-induced difficulty underestimation under the standard passage-level editing objective: later facts receive increasingly rich ground-truth context and consequently incur lower initial losses, making them appear easier to learn. In response, we propose FOVEATED, a plug-and-play framework that constructs focused views of each sentence by randomly shifting the Rotary Position Embedding (RoPE) positions assigned to the keys of its preceding context. The perturbation is applied during editing and removed afterward, leaving the model's native positional encoding unchanged at inference time. We instantiate FOVEATED for both direct-optimization and locate-then-edit editors. We theoretically analyze how FOVEATED counteracts context-induced difficulty underestimation and empirically demonstrate consistent improvements across five KE editors, two LLM backbones, and three benchmarks.
Fonte: arXiv cs.CL
Theory/Optimization • Score 85
Bifidelity Karhunen-Lo\`eve Expansion Surrogate with Active Learning for Random Fields
arXiv:2511.03756v2 Announce Type: replace
Abstract: We present a bifidelity Karhunen--Lo\`{e}ve expansion (KLE) surrogate model for field-valued quantities of interest (QoIs) under uncertain inputs. The QoIs considered here are scalar fields. The approach combines the spectral efficiency of the KLE with polynomial chaos expansions (PCEs) to preserve an explicit mapping between input uncertainties and output fields. By coupling inexpensive low-fidelity (LF) simulations that capture dominant response trends with a limited number of high-fidelity (HF) simulations that correct for systematic bias, the proposed method can enable accurate and computationally affordable surrogate construction. To further improve surrogate accuracy, we develop an active learning strategy that adaptively selects new HF evaluations based on the surrogate's generalization error, estimated via cross-validation and modeled using Gaussian process regression. New HF samples are then acquired by maximizing an expected improvement criterion, targeting regions of high surrogate error. The resulting BF-KLE-AL framework is demonstrated on three examples of increasing complexity: a one-dimensional analytical benchmark, a two-dimensional convection-diffusion system, and a three-dimensional turbulent round jet simulation based on Reynolds-averaged Navier--Stokes (RANS) and enhanced delayed detached-eddy simulations (EDDES). The experiments show that bifidelity gains depend on LF accuracy, discrepancy approximation, and the allocation of simulation cost. Active learning improves prediction over random sampling in several settings, while the cost-matched comparisons identify both favorable regimes and cases where an HF-only surrogate is more accurate.
Fonte: arXiv stat.ML
Theory/Optimization • Score 85
KL Convergence Guarantees for Score diffusion models under minimal data assumptions
arXiv:2308.12240v3 Announce Type: replace-cross
Abstract: Diffusion models are a new class of generative models that revolve around the estimation of the score function associated with a stochastic differential equation. Subsequent to its acquisition, the approximated score function is then harnessed to simulate the corresponding time-reversal process, ultimately enabling the generation of approximate data samples. Despite their evident practical significance these models carry, a notable challenge persists in the form of a lack of comprehensive quantitative results, especially in scenarios involving non-regular scores and estimators. In almost all reported bounds in Kullback Leibler (KL) divergence, it is assumed that either the score function or its approximation is Lipschitz uniformly in time. However, this condition is very restrictive in practice or appears to be difficult to establish. To circumvent this issue, previous works mainly focused on establishing convergence bounds in KL for an early stopped version of the diffusion model and a smoothed version of the data distribution, or assuming that the data distribution is supported on a compact manifold. These explorations have led to interesting bounds in either Wasserstein or Fortet-Mourier metrics. However, the question remains about the relevance of such early-stopping procedure or compactness conditions. In particular, if there exist a natural and mild condition ensuring explicit and sharp convergence bounds in KL. In this article, we tackle the aforementioned limitations by focusing on score diffusion models with fixed step size stemming from the Ornstein-Uhlenbeck semigroup and its kinetic counterpart. Our study provides a rigorous analysis, yielding simple, improved and sharp convergence bounds in KL applicable to any data distribution with finite Fisher information with respect to the standard Gaussian distribution.
Fonte: arXiv stat.ML
Theory/Optimization • Score 85
Broken scale symmetries in undercomplete linear autoencoders
arXiv:2610.03640v1 Announce Type: cross
Abstract: Neural network loss landscapes have many symmetries, which are preserved by gradient flow but broken by finite-stepsize stochastic gradient descent (SGD). A canonical example of such a symmetry is scale in homogeneous networks: one can scale up the parameters in one layer and down in the next without changing the network output. Previous work has documented cases in which SGD breaks this symmetry in favor of balancing gradient noise or minimizing fluctuations. Here, we show that the solution geometry of undercomplete linear autoencoders instead selects a preferred sign for scale drift: on the PCA solution manifold, SGD favors large decoder weights. This directed scale drift occurs on a slow timescale, and its dynamics admit an analytically-tractable effective description. However, it cannot continue indefinitely: increasing scale eventually drives the dynamics towards a finite-stepsize stability boundary. The resulting solutions are sharper than a balanced baseline in the sense of the maximum eigenvalue of the loss Hessian, but different sharpness measures can move in opposing directions. Thus, undercomplete autoencoders give a concrete illustration of how loss geometry can convert residual gradient noise into directed motion along a manifold of functionally-equivalent solutions.
Fonte: arXiv stat.ML
Theory/Optimization • Score 85
AECSF: Adaptive Ensemble Conditional Score Filtering for High-Dimensional Nonlinear Data Assimilation
arXiv:2609.32411v2 Announce Type: replace
Abstract: Bayesian state estimation for high-dimensional nonlinear dynamical systems entails a fundamental tension between statistical fidelity and computational tractability, as particle weights can collapse, while Gaussian ensemble updates can miss non-Gaussian posterior structure. Score-based diffusion filters offer a sampling-based alternative, but existing training-free score filters often rely on heuristic likelihood corrections, which can compromise posterior accuracy by neglecting uncertainty about the system state associated with each noisy reverse particle. To address these issues, we propose AECSF, a training-free adaptive ensemble conditional score filter. AECSF constructs an analytically tractable score estimator from the conditional Tweedie identity, which recasts noisy posterior score estimation as estimating the conditional mean of the system state given a noisy reverse particle and the observation. To estimate these conditional means efficiently, AECSF employs a shared adaptive weighted proposal ensemble, while particle-specific conditional weights yield an estimate for each noisy reverse particle without separate proposal sampling. The proposal ensemble is updated using reverse-particle information within the same reverse-diffusion run to improve conditional-mean estimation. Theoretically, we characterize when a fixed weighted proposal measure yields the exact noisy posterior score. Under stated assumptions, we establish a bound relating conditional-mean estimation errors to reverse-sampling endpoint error. Numerical experiments demonstrate that AECSF improves the accuracy of posterior sampling and nonlinear filtering in high-dimensional problems with limited forecast ensembles.
Fonte: arXiv stat.ML
NLP/LLMs • Score 85
Estimating prevalence with precision and accuracy
arXiv:2507.06061v2 Announce Type: replace
Abstract: Unlike classification, whose goal is to estimate the class of each data point, quantification (or prevalence estimation) aims to estimate the distribution of classes in a dataset. An important task in prevalence estimation is to quantify the uncertainty in prevalence estimates. In this paper, we introduce Precise Quantifier (PQ), a Bayesian aggregative quantifier that achieves narrow prediction intervals with sufficient coverage (i.e., sufficient proportion of intervals containing the true prevalence). We find that PQ produces more precise prevalence estimates than existing methods as the discriminative power of the underlying classifier increases and as the validation-to-test size ratio increases. These empirical results suggest that PQ uses validation information more effectively to quantify uncertainty in prevalence estimates than existing approaches.
Fonte: arXiv stat.ML
NLP/LLMs • Score 85
Silent Dissent: LLM Agents That Yield to the Majority Still Represent Their Original Premise
arXiv:2610.02702v1 Announce Type: new
Abstract: Multi-agent debate is increasingly used to reach consensus among LLM agents, yet agents often yield to a unanimous majority. When an agent changes its answer, has it changed its mind or only its statement? We study this with two-hop factual questions whose intermediate entity (the bridge, e.g. the country in "the capital of the country where the Sagrada Familia is located") is never stated by anyone. Scripted peers, in the role of Asch's confederates, unanimously assert a wrong answer taken from another fact with a different bridge. At the moment the agent answers, we read the bridge from its residual stream with the Jacobian lens (J-lens) and, for comparison, the logit lens. In pre-registered tests on held-out facts with four open-weight models, agents of Qwen3.5-4B, Qwen3.6-27B and Gemma-4-E4B-it that gave in still represented their original bridge in the pre-registered layers below the output (hit@100 above a control entity: 0.85, 0.22 and 0.24), where the logit lens rarely ranked it among the top 100 tokens (0.00-0.06). These agents also represented the bridge behind the peers' answer, beyond a mention baseline. A pre-registered addendum hid the agent's earlier answer or removed it: agents that gave in still represented their original bridge in all four models (0.43, 0.29, 0.37 and 0.25 with the answer hidden), including Llama-3.1-8B-Instruct, which barely did so with its answer in view (0.03). The premise can thus be computed from the question alone while the agent states the majority's answer. Hiding the earlier answer also changed conformity: Qwen3.5-4B gave in on 89% of questions instead of 8%. In exploratory interventions, injecting the bridge's J-lens direction brought agents back to their original answer only in the two Qwen models. Stated consensus in multi-agent debate can thus overstate agreement. We also report the negative results of our pre-registered program.
Fonte: arXiv cs.CL
NLP/LLMs • Score 85
From Retrieval to Typed Decisions: Calibrated System One Models from Biomedical Sentence Encoders
arXiv:2610.02486v1 Announce Type: new
Abstract: Typed decision models answer schema-constrained questions about a text in one forward pass and return probabilities meant to be thresholded. We ask whether biomedical sentence encoders trained for retrieval are good starting points for such models. We present SBERT2S1, which converts Sentence-Transformers encoders into bi-encoder, cross-head (C) and prior-fused residual (PFR) decision models, together with BIODECIDE, a biomedical typed-decision suite, and MEDLINE-S1, 243k training decisions derived from NLM indexing. Across six parent-retriever pairs, retrieval training improves zero-shot matching of content-bearing options. After fine-tuning, its effect depends on the head: across five pairs and three training-set sizes, retrieval training significantly helps PFR, which keeps the retrieval prior, in 10 of 15 comparisons, but helps C in one and hurts it in five. A matched grid of two heads and five training objectives shows that C outperforms PFR under every objective, and that the released RLCD recipe of open System One models trails cross-entropy by 2.5-3.0 points. The deficit stems mainly from its reward normalisation, which inflates the noisy score-function term 3.6-15-fold; an unbiased leave-one-out estimator recovers most of the gap. After temperature scaling, no objective is clearly better calibrated than cross-entropy. We release the code, the MEDLINE-S1 labels and a model.
Fonte: arXiv cs.CL
Theory/Optimization • Score 85
Feature tracking in physics-informed neural networks via joint optimization of nonlinear deformation manifolds: application to shocks
arXiv:2610.02230v1 Announce Type: cross
Abstract: Physics-informed neural networks (PINNs) often converge to inaccurate solutions for conservation laws with shocks, because uniformly distributed collocation points undersample localized features and let the residual be dominated by regions that are already well resolved. We propose a feature-tracking PINN (FT-PINN) in which the solution network is defined on a fixed reference domain and composed with a diffeomorphic deformation map from a parameterized nonlinear manifold. The deformation and solution-network parameters are trained jointly by minimizing the pulled-back conservation-law residual. This lets collocation points concentrate along features of essentially arbitrary geometry, including curved, oblique, and merging shocks, without prior knowledge of their locations. The framework is agnostic to the choice of parameterization. Boundary preservation is enforced exactly through a tangential projection of the displacement, and folding is discouraged by a one-sided penalty on the Jacobian determinant. On four test problems (space-time viscous Burgers with merging shocks, a decelerating Burgers shock, the space-time Euler shock tube, and steady 2D Euler regular shock reflection), FT-PINN resolves shocks at their correct locations with a limited collocation budget. A vanilla PINN with the same architecture, budget, and training either misplaces the shocks or fails to form them.
Fonte: arXiv stat.ML
NLP/LLMs • Score 85
Finding the Move Is Not Winning the Game: XiangqiBench for Closed-Loop Evaluation of LLM Agents
arXiv:2610.02425v1 Announce Type: new
Abstract: Static evaluations credit a language model for naming the right move, but an agent must carry a plan through to a verified outcome while an opponent responds. We introduce XiangqiBench, an executable benchmark that measures this difference in Chinese chess: starting from 119 tactical endgames with forced mates supported by engine or checks-only search, an LLM agent must deliver checkmate against an engine defender. An interactive REPL interface separates real moves, state queries, and forward simulation, and we record 8,568 multi-turn trajectories from 12 frontier LLMs under two observation protocols. Three signals that look like competence each overstate closed-loop success. (i) The Conversion Gap: models play the stored reference first move in 26.1\% of Sighted trials, yet only 13.9\% of these trials end in a win. (ii) The Consistency Gap: the leading model reaches 38.7\% pass@3 but only 5.9\% pass^3, winning all three trials on 7 of the 46 positions it ever wins. (iii) The Simulation Gap: 32.3\% of accepted simulation calls stop on an illegal move, and in 49.3\% of comparable cases the real defender replies differently from the line the agent simulated; self-authored rollouts check legality but cannot anticipate the opponent. Finding the move is not winning the game: agent evaluations should score closed-loop outcomes and report reliability alongside coverage.
Fonte: arXiv cs.CL
NLP/LLMs • Score 85
SDECast: Probabilistic Weather Forecasting in Continuous Time with Neural SDEs
arXiv:2610.03313v1 Announce Type: new
Abstract: Existing machine learning weather forecasting models typically generate forecasts through autoregressive rollouts at a fixed temporal resolution. While highly efficient for long-range prediction, this formulation can suffer from severe error accumulation when used with shorter time steps and does not explicitly encode the locality and temporal continuity of atmospheric dynamics. To address these limitations, we introduce **SDECast**, a Neural Stochastic Differential Equation (SDE) framework for continuous-time probabilistic weather forecasting. SDECast extends SDE Matching to learn stochastic dynamics directly in physical space, without requiring repeated SDE simulation during training. On a simulated geophysical flow, we show that SDECast recovers meaningful drift dynamics and faithfully reproduces the underlying continuous-time behavior. We then demonstrate its scalability to global weather forecasting at hourly resolution, where SDECast produces skillful probabilistic forecasts for lead times of up to five days.
Fonte: arXiv stat.ML
RL • Score 85
Bandits via Additive Quantized Representations
arXiv:2610.02440v1 Announce Type: new
Abstract: Contextual bandits require balancing nonlinear reward modeling with online efficiency. Tree ensembles and neural methods capture nonlinearities but require periodic retraining and large replay buffers. Linear models update efficiently per observation with O(1) memory, but are fundamentally restricted to linear reward structures. We propose Residual Quantization (RQ) as a representation layer to bridge this gap. An offline-trained RQ codebook maps continuous contexts into discrete centroid assignments across multiple levels, set dynamically through a shadow mechanism. This enables a spectrum of additive bandit algorithms that achieve nonlinear expressivity with strictly bounded memory. Across 13 datasets, RQ variants beat their non-RQ counterparts on 11 of 13 datasets, often by wide margins, while matching doubling-retrain XGBoost and neural baselines using up to 1000 times less memory.
Fonte: arXiv cs.LG
NLP/LLMs • Score 85
Power-SMC: Low-Latency Sequence-Level Power Sampling for Training-Free LLM Reasoning
arXiv:2602.10273v3 Announce Type: replace
Abstract: Reasoning ability in large language models is often attributed to \emph{distribution sharpening}: concentrating output probability on high-likelihood sequences. Recent works show that this sharpening effect can be obtained at inference time, without modifying model parameters, and can elicit strong reasoning performance. A natural formalization is the \emph{sequence-level power distribution}, which is proportional to the model's probability raised to an exponent $\alpha>1$. Prior work leveraged Metropolis--Hastings (MH) sampling to draw samples from this distribution and achieves strong results, however, at order-of-magnitude inference slowdowns. We introduce \textbf{Power-SMC}, a \textit{`training-free'} sampling method that targets the same power distribution yielding close to standard decoding latency. Power-SMC maintains multiple candidate sequences in parallel. Each candidate sequence is assigned a score, namely the \emph{importance weight}, that measures how well it matches the power distribution. It then periodically prunes low-scoring candidate sequences in favor of high-scoring ones. We further provide a theoretical justification for the design choices in Power-SMC. Among all next-token sampling strategies that do not rely on future tokens, we prove that sampling temperature $\tau{=}1/\alpha$ uniquely eliminates per-step weight variance. Finally, we characterize the remaining source of weight instability and introduce a gradual sharpening schedule to reduce weight collapse, while targeting the same power distribution. Extensive evaluations on MATH500, GSM8K, GPQA, and HumanEval show that Power-SMC matches or exceeds MH sampling in accuracy, preserves output diversity unlike RL-finetuned models, while \textbf{accelerating inference speed by up to} {$\mathbf{17.6}\times$}. The code is available at https://github.com/ArminAzizi98/Power-SMC.
Fonte: arXiv stat.ML
Theory/Optimization • Score 85
Below what training size do deep tabular generators stop beating trivial baselines? A preregistered benchmark on a size ladder of clinical and standard datasets
arXiv:2610.03500v1 Announce Type: cross
Abstract: Deep tabular generative models are benchmarked on datasets with tens of thousands of rows; clinical datasets have hundreds. We preregistered and ran a size-ladder benchmark to find where the two regimes diverge: 8 public datasets subsampled from 200 to 20,000 training rows, seven generators (independent marginals, Gaussian copula, SMOTE, unconditional SMOTE, CTGAN, TVAE, TabDDPM) with a fixed 20-trial tuning budget and 5 evaluation seeds, plus 4 natively small clinical datasets at true size, for 2,220 committed runs in total. The primary metric is the AUROC of fixed classifiers trained on synthetic and tested on real data. In 23 of 24 (dataset, deep model) pairs no deep model ever beats the best trivial baseline by more than seed noise, at any training size we measured. The best baseline wins 40 of 49 (dataset, size) cells. Our preregistered prediction that the deep models' ranking would be unstable at small sizes is falsified: mean Kendall tau between adjacent rungs below 5,000 rows is 0.806, above our 0.8 threshold, and stability is highest at the smallest sizes rather than lowest. One caveat bounds all of this: in 81% of cells the gap between the top two methods is smaller than the variation between seeds. Finally, method rankings on natively small clinical datasets agree only moderately with rankings on subsampled large ones (mean tau 0.57 to 0.64), which questions whether a subsampled large dataset can stand in for a small one. All 2,220 result files, the preregistration and its hash, and the code that regenerates every figure and number from those files are public.
Fonte: arXiv stat.ML
Theory/Optimization • Score 85
Muon Learns Facts Better: Understanding the Role of Spectral Orthogonalization
arXiv:2610.02798v1 Announce Type: cross
Abstract: The Muon optimizer applies spectral orthogonalization to matrix-valued updates and has shown strong performance in large-scale neural network training, yet the mechanisms of this transformation in feature learning remain poorly understood. In this work, we investigate this question through a tractable factual-recall model, where a fact maps each subject-relation pair to an answer, and a linear transformer learns the subject- and relation-dependent information required to recover this mapping. The transformer is optimized with gradient flow (GF), spectral GF, or Sign GF, which are continuous-time limits of gradient descent, Muon, and Adam, respectively. Prior studies (Nichani et al., 2025) have shown that when the number of subjects exceeds the number of relations, GF learns relation-dependent information before subject-dependent information, producing a feature-separation phase during training. We characterize this separation with the learning times when the subject- and relation-dependent components of the prediction reach a target accuracy. With $S$ subjects and $R$ relations, GF has a learning-time ratio of $\widetilde{\Theta}(\sqrt{S/R})$, whereas Spectral GF reduces this ratio to $\widetilde{\Theta}(1)$. In addition, for fixed $S$ and $R$, the subject- and relation-dependent errors decay as $1/(T\log T)$ in training time $T$ under GF, but as $\exp(-\mathrm{poly}(T))$ under spectral GF. Finally, we show that GF and spectral GF are equivariant under orthogonal transformations of the token embeddings, whereas Sign GF is not: Different orthonormal embeddings can potentially produce no feature separation, a large feature-separation phase, or even a reversed learning order. These results provide a mechanistic view of how spectral orthogonalization can fundamentally reshape feature-learning dynamics.
Fonte: arXiv stat.ML
Evaluation/Benchmarks • Score 85
When Is Accuracy Evidence? A Unified Theory of Generalisation, Validation, and Information Fusion
arXiv:2610.03465v1 Announce Type: new
Abstract: K-fold cross-validation (CV) is widely used as evidence of out-of-sample performance, although folds are neither independent experiments nor equally informative under heterogeneous data. Cross Upper-Bound Validation (CUBV) replaces point-wise CV accuracy by conservative upper bounds on true risk. Here we generalise CUBV through a single exponential framework in which the moment-generating function of the generalisation gap is controlled by a cumulant envelope gamma(lambda). This yields a family of risk bounds covering Hoeffding-, Bernstein-, dependency-aware, PAC-Bayesian, and heterogeneous source-fusion settings. For K-fold CV, dependence between fold-wise gaps is modelled through a joint sub-Gaussian proxy matrix. Under equicorrelation, this gives an effective number of folds, Keff = K/[1+(K-1)rho], showing that increasing K does not necessarily increase statistical evidence when folds are strongly dependent. The framework is also extended to posterior distributions over predictors and weighted multi-source fusion, where weights are selected by minimising an upper bound on future risk rather than empirical error alone. Experiments with trained linear classifiers on heterogeneous multimodal Gaussian mixtures compare K-fold CV with full-sample resubstitution plus risk correction. Bounds are evaluated by coverage and tightness. In low-dimensional small-sample settings, K-fold partitioning can increase uncertainty because individual folds under-represent minority modes, while corrected resubstitution can remain valid and tighter; this effect disappears as sample size increases. Overall, gamma-CUBV separates observed performance, uncertainty, dependence, model complexity, and confidence into explicit terms, providing a unified route from CV scores to risk statements and a principled validation criterion for heterogeneous small-sample applications such as neuroimaging.
Fonte: arXiv stat.ML
NLP/LLMs • Score 85
Differential Privacy as a Perk: Federated Learning over Multiple-Access Fading Channels with a Multi-Antenna Base Station
arXiv:2510.23463v4 Announce Type: replace-cross
Abstract: Federated Learning (FL) is a distributed learning paradigm that preserves privacy by eliminating the need to exchange raw data during training. In its prototypical edge instantiation with underlying wireless transmissions enabled by analog over-the-air computing (AirComp), referred to as \emph{over-the-air FL (AirFL)}, the inherent channel noise plays a unique role of \emph{frenemy} in the sense that it degrades training due to noisy global aggregation while providing a natural source of randomness for privacy-preserving mechanisms, formally quantified by \emph{differential privacy (DP)}. It remains, nevertheless, challenging to effectively harness such channel impairments, as prior arts, under assumptions of either simple channel models or restricted types of loss functions, mostly considering (local) DP enhancement with a single-round or non-convergent bound on privacy loss. In this paper, we study AirFL over multiple-access fading channels with a multi-antenna base station (BS) subject to user-level DP requirements. Despite a recent study, which claimed in similar settings that artificial noise (AN) must be injected to ensure DP in general, we demonstrate, on the contrary, that DP can be gained as a \emph{perk} even \emph{without} employing any AN. Specifically, we derive a novel bound on DP that converges under general bounded-domain assumptions on model parameters, along with a convergence bound with general smooth and non-convex loss functions. Next, we optimize over receive beamforming and power allocations to characterize the optimal convergence-privacy trade-offs, which also reveal explicit conditions in which DP is achievable without compromising training. Finally, our theoretical findings are validated by extensive numerical results.
Fonte: arXiv stat.ML
Evaluation/Benchmarks • Score 85
Open-Endedness Bench: Measuring Epistemic Process from Agent Records
arXiv:2610.02588v1 Announce Type: new
Abstract: Agents are increasingly given open-ended research tasks: discovering an empirical law from self-designed experiments, improving a heuristic whose optimum nobody knows, or beating a standing record. Their execution logs record every step of this research, yet the runs are still judged by their outcome score. That score alone does not establish whether an agent's claims follow from executed experiments, and a reference answer may be unavailable. We evaluate the agent's epistemic process: how it forms hypotheses, tests them, and revises them in response to evidence. We introduce OEB (Open-Endedness Bench), a benchmark-agnostic methodology that reads only the agent's execution record and never a reference answer or an outcome score. OEB compiles the record into a unified epistemic event graph whose edges connect the propositions the agent states to the executed actions that test them; each node carries an exact excerpt that code verifies against the record. One principle governs scoring: prose can state a proposition, but only evidence returned by an executed action can support or refute it, so OEB checks what the agent writes against what it actually ran. From the graph, OEB scores four competence axes (evidence, experiment, revision, and no reward hacking), mostly as the share of opportunities for sound research that the agent took, and profiles six subjective persona traits that describe the agent's research habits. We score 119 existing runs over 12 tasks from three benchmarks: LLM post-training, chip design, and a training-speed record. Against logged results, only 16-29% of the improvements agents claim are real. On 9 of 10 tasks, the best run tries more new ideas in its second half than the worst run. The persona readings follow the model: for every trait, the model that ran explains more of its variance across runs than the task (a median of 43% against 7%).
Fonte: arXiv cs.AI
Theory/Optimization • Score 85
Threshold-Aware Conformal Routing
arXiv:2610.02487v1 Announce Type: new
Abstract: High-fidelity simulations are essential to scientific and engineering design, but can be expensive to run repeatedly. Learned surrogates offer a faster alternative, yet their higher errors may alter downstream decisions. This accuracy-speed tradeoff creates a need to determine whether a surrogate can be used or the full simulator remains necessary. We study decisions determined by whether a scalar quantity of interest lies above or below a fixed threshold. For each input, we use the surrogate when its conformal interval lies entirely on one side of the threshold and route the input to simulation when the interval intersects it. Standard conformal prediction constructs intervals without reference to the downstream decision threshold: even a narrow interval near the threshold can cross it and trigger simulation, whereas a wider interval farther away can remain entirely on one side and require no simulation. We introduce Threshold-Aware Conformal Routing (TACR), which learns an input-dependent scale using a threshold-aware objective that concentrates interval tightness near the decision boundary. Exact split-conformal calibration on held-out data preserves distribution-free marginal coverage, which also upper-bounds the probability of an incorrect threshold decision that is not routed. Across various scientific and engineering datasets, TACR reduces simulator deferrals by 14-75% relative to standard conformal prediction at the same coverage target. Against a variant without threshold-local weighting but with similar predictor accuracy, TACR further reduces deferrals by 10-24% on four datasets. These results show that optimizing interval allocation for routing can reduce simulator calls without weakening the standard conformal guarantee.
Fonte: arXiv cs.LG
NLP/LLMs • Score 85
A Generative Model of Complex Networks Using Graphons and Neural Inverse Operators
arXiv:2610.02439v1 Announce Type: new
Abstract: Generative graph models are central to understanding and simulating complex networks. However, existing approaches have complementary strengths and limitations. Mechanistic models offer interpretability but rely on instance-specific estimation methods. Deep generative models, on the other hand, offer amortized inference at the cost of interpretability and are largely limited to graph sizes seen during training. Scientific applications motivate a framework that retains the strengths of both paradigms. We bridge them by formulating both the generative model and parameter recovery in function space. A multifractal step graphon extends standard step graphons with a recursive construction that compactly parameterizes complex networks. This formulation admits a neural inverse operator to recover its parameters, enabling inference on unseen graph sizes. We evaluate our model, trained only on synthetic multifractal step graphon realizations, against both paradigms. Against a graph foundation model pretrained on empirical networks, our method achieves the best average performance on three of four metrics in a zero-shot graph-generation benchmark, indicating that the model transfers to real-world graphs. We also apply our method to single-observation networks, a regime largely inaccessible to deep models that require training corpora, where it performs comparably to an instance-specific method that optimizes on each graph. In a multi-subject EEG case study, the inferred parameters track a reversible change in brain state more sensitively than traditional network statistics. Together, these results indicate that mechanistic interpretability and amortized inference can be effectively unified in a generative graph model to enhance our understanding of complex networks.
Fonte: arXiv cs.LG
NLP/LLMs • Score 85
DAWIS: Data Assimilation with Windowed Inverse Sampling via Multitask Interpolants
arXiv:2610.03314v1 Announce Type: new
Abstract: Flow- and diffusion-based generative models have recently emerged as flexible and highly efficient forecasting models for dynamical systems. When combined with inference-time guidance, they offer a promising route to high-dimensional non-Gaussian data assimilation (DA), the problem of combining forecasts with observations to estimate latent system states. Existing filters, however, condition on a fixed history and assimilate only the most recent observation, leaving them unable to revise past states when new observations arrive. Estimates then stay tethered to a history that later observations may contradict, and errors accumulate over the assimilation run. To this end, we introduce **DAWIS**, a unified DA method covering filtering, fixed-lag smoothing, and block smoothing within a single framework. DAWIS replaces the single flow time of a state-level prior with a multitask stochastic interpolant over a window of consecutive states, assigning a separate flow time to each. An assimilation cycle inverts the window to a vector of per-state turning points and regenerates it under observation guidance, with the turning points controlling how strongly each state is held fixed, revised, or generated from scratch. The same construction can also absorb the forecast into the assimilation cycle, removing the need for a separate forecasting model. Experiments on challenging nonlinear systems show that DAWIS improves on both filtering and smoothing baselines under sparse, noisy, and nonlinear observations. The code for DAWIS is available at https://github.com/Erik-Wikingsson/DAWIS
Fonte: arXiv stat.ML