
category.学术arXiv cs.LG (机器学习)
arXiv:2610.04038v1 Announce Type: new Abstract: Synthetic data are increasingly used as an alternative to sharing sensitive records. However, synthetic data generation does not guarantee privacy, as diffusion models trained or adapted on sensitive data remain susceptible to reconstruction attacks. Moreover, while approaches that use differential privacy (DP), such as DP-SGD, achieve provably priva
10月6日 04:00

category.学术arXiv cs.LG (机器学习)
arXiv:2610.03830v1 Announce Type: new Abstract: In partially observable Markov decision processes, the optimal policy generally depends on the history of observations and past actions. Asymmetric actor-critic methods have become popular to learn such policies when additional information, such as the true state of the environment, is available during training. The critic, which is not needed at exe
10月6日 04:00

category.学术arXiv cs.LG (机器学习)
arXiv:2610.03831v1 Announce Type: new Abstract: Capable systems often show jagged competence: low average error alongside failures on particular inputs. We ask where it comes from in a task built from two known layers. A lower layer A computes five per-slot sums from records; an upper layer B uses the slot-1 sum and a mode carried over from earlier boards to predict the next board's category, so B
10月6日 04:00

category.学术arXiv cs.CV (计算机视觉)
arXiv:2610.03980v1 Announce Type: new Abstract: Text-to-video (T2V) diffusion models can reproduce copyrighted, violent, or explicit content, which motivates concept erasure: removing designated concepts from a pretrained model while preserving its behavior on everything else. Existing T2V erasure methods leave two problems open. Their frame-agnostic suppression can leave isolated frames in which
10月6日 04:00

category.学术arXiv cs.LG (机器学习)
arXiv:2610.04034v1 Announce Type: new Abstract: Data assimilation combines model forecasts with noisy, incomplete observations to estimate the evolving state of a dynamical system. Existing methods face two compounding challenges: high-dimensional nonlinear dynamics make repeated forward simulation computationally expensive, while sparse observations provide limited direct information about the fu
10月6日 04:00

category.学术arXiv cs.CL (自然语言处理)
arXiv:2610.04400v1 Announce Type: new Abstract: General turn-taking behavior in real-time dialogue systems requires deciding whether to keep listening or start responding while listening, and whether to continue or stop while speaking. Existing turn detectors use heterogeneous, task-specific label spaces and are often trained on limited annotations or evaluated on isolated utterances, making them
10月6日 04:00

category.学术arXiv cs.CV (计算机视觉)
arXiv:2610.03799v1 Announce Type: new Abstract: Recovering executable CAD programs from 3D meshes is challenging due to the compositional nature of CAD construction and the interaction between discrete modeling choices and continuous parameters. Many learning-based methods predict complete programs in a single pass and rely predominantly on sketch-extrude representations, limiting operation divers
10月6日 04:00

category.学术arXiv cs.CL (自然语言处理)
arXiv:2610.04401v1 Announce Type: new Abstract: Public vs. private universities is a debatable issue, and it creates polarization on social media in Bangladesh. Debate on quality, jobs, and prestige is passionate among the students, parents, and graduates, the majority of whom speak Bengali, a low-resource language. To measure this polarization, this paper introduces a manually annotated dataset o
10月6日 04:00

category.学术arXiv cs.CL (自然语言处理)
arXiv:2610.04347v1 Announce Type: new Abstract: Identical utterance choices can arise from different communicative causes, and identical interpretations can leave different traces in what a listener learns. Rational Speech Act (RSA) models treat interpretation as inference over speaker meaning, but standard one-shot RSA does not intrinsically distinguish these causal update targets. We develop a c
10月6日 04:00

category.学术arXiv cs.CV (计算机视觉)
arXiv:2610.03991v1 Announce Type: new Abstract: Multimodal skin lesion classification combines clinical images with patient metadata to improve diagnostic accuracy. However, complete metadata available during training may be only partially accessible at deployment, and resource-constrained settings additionally require computational efficiency. We address these challenges with a privileged-informa
10月6日 04:00

category.学术arXiv cs.CV (计算机视觉)
arXiv:2610.03797v1 Announce Type: new Abstract: World Action Models (WAMs) leverage pretrained video foundation models for robot manipulation, but their large backbones and video-action co-prediction are expensive. Although existing acceleration techniques offer many ways to reduce this cost, selecting and composing them requires substantial engineering for each model and hardware platform. To tac
10月6日 04:00

category.学术arXiv cs.CL (自然语言处理)
arXiv:2610.04403v1 Announce Type: new Abstract: Post-training quantization to the GGUF format's mixed-precision K-quants is commonly how open-weight language models reach consumer hardware, yet its effect on fine-grained lexical competence is uncharacterized. We audit 27 quantized artifacts across 13 families and four architecture backbones, 0.35B-14B parameters, evaluated down their published lad
10月6日 04:00

category.学术arXiv cs.LG (机器学习)
arXiv:2610.03951v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used as classifiers, yet they operate as opaque systems whose decisions are difficult to interpret, which complicates their use in regulated domains such as credit scoring or medical diagnosis. We propose an evolutionary framework that iteratively discovers natural language feature definitions (rubrics) f
10月6日 04:00

category.学术arXiv cs.LG (机器学习)
arXiv:2610.03983v1 Announce Type: new Abstract: We study constrained online convex optimization with adversarial convex losses and constraints ($\mathsf{COCO}$). At each round \(t\in[T]\), a learner selects \(x_t\) from a \(d\)-dimensional convex decision set \(\mathcal X\), after which an adaptive adversary reveals a convex cost function \(f_t\) and constraint function \(g_t\). Consequently, the
10月6日 04:00

category.学术arXiv cs.LG (机器学习)
arXiv:2610.03896v1 Announce Type: new Abstract: Large Language Models (LLMs) are increasingly used for automated algorithm design. However the computational cost of evaluating the generated algorithms can be excessive. We consider the common LLM-driven automated algorithm design (LLM4AD) setting in which a candidate algorithm is evaluated by aggregating its performance over a shared set of trainin
10月6日 04:00

category.学术arXiv cs.LG (机器学习)
arXiv:2610.03880v1 Announce Type: new Abstract: Inverse materials design is a long-standing goal of computational materials discovery. Generative models for crystalline materials are typically trained to match the distribution of a structure database, while nothing in their training objective points them at specific design goals such as targeted properties. We use group-relative policy optimizatio
10月6日 04:00

category.学术arXiv cs.CL (自然语言处理)
arXiv:2610.04370v1 Announce Type: new Abstract: Data quality now matters as much as compute for training language models. Much training data comes from human annotation of text, and interpretive annotation has no ground truth that could settle what is "accurate". Two lines of work respond to this. One combines annotators into a "ground truth" and measures how well they agree with each other; the o
10月6日 04:00

category.学术arXiv cs.LG (机器学习)
arXiv:2610.03952v1 Announce Type: new Abstract: A one-step curvature probe evaluates an initial direction, whereas continual learners are judged after reaching comparable new-task fit. In an overparameterized linearization, projected gradient descent converges to $\Delta_P=PJ^\top(JPJ^\top)^{-1}r$, and its squared-displacement inflation is exactly the reciprocal of the retained fitting capacity $c
10月6日 04:00

category.学术arXiv cs.CV (计算机视觉)
arXiv:2610.03792v1 Announce Type: new Abstract: Reinforcement learning from verifiable rewards (RLVR) has produced large reasoning gains in language models, and verifiable video benchmarks make it applicable to causal-temporal video question answering. We study what RLVR teaches video-language models about time. We fine-tune four open models (Qwen3-VL-8B/4B, Qwen2.5-VL-7B, Gemma-3-12B) with group
10月6日 04:00

category.学术arXiv cs.CV (计算机视觉)
arXiv:2610.03928v1 Announce Type: new Abstract: Screen localization and pointing estimation are key to low-cost interactive devices. Yet developing and evaluating these algorithms requires realistic data: synthetic captures cannot fully reproduce the optical distortion, rolling shutter, motion blur, and display processing of a physical acquisition, and most existing datasets provide static images
10月6日 04:00

category.学术arXiv cs.CL (自然语言处理)
arXiv:2610.04239v1 Announce Type: new Abstract: Uncertainty quantification (UQ) for large language models (LLMs) aims to provide reliable measures of predictive confidence, yet current methods are often unstable under meaning-preserving perturbations. Semantically equivalent paraphrases can induce substantial variability in predictive confidence, even for methods with formal guarantees, such as co
10月6日 04:00

category.学术arXiv cs.CL (自然语言处理)
arXiv:2610.04156v1 Announce Type: new Abstract: LLM agents for clinical text-to-SQL applications reason autonomously over multiple steps but cannot assess whether their own reasoning or outputs can be trusted. In high leverage applications such as healthcare, this presents a critical risk where system mistakes can be costly. These reliability failures are also resource failures: an incorrect reaso
10月6日 04:00

category.学术arXiv cs.LG (机器学习)
arXiv:2610.03939v1 Announce Type: new Abstract: Many inference workloads evaluate a trained tree ensemble on row groups that share feature values: discrete-time survival models expand each patient into $G$ time steps, click-through-rate models score every item in a search session, and scenario analyses vary a few inputs while holding the rest fixed. Standard inference treats each row independently
10月6日 04:00

category.学术arXiv cs.CL (自然语言处理)
arXiv:2610.03827v1 Announce Type: new Abstract: Researchers often support the claim that a model shares structure with the brain, or across languages, by reporting a similarity score. We ask what such a score reads when the shared structure is absent, or when the tool that measures it does not work. We check two settings, and in both the score is not what it appears. First, a probe trained to tell
10月6日 04:00

category.学术arXiv cs.AI
arXiv:2610.03966v1 Announce Type: new Abstract: Each run of an AI-driven research system (ADRS) is an expensive search over a vast solution space, and dependable evaluation requires many runs, making run data both costly to produce and valuable to retain for large-scale analysis. Yet this data remains fragmented: teams operate in isolation, ADRS frameworks emit results in different formats, and no
10月6日 04:00

category.学术arXiv cs.CL (自然语言处理)
arXiv:2610.03825v1 Announce Type: new Abstract: A benchmark score compares a system output against a reference, and methodological attention falls almost entirely on the first term. We measure the second. The retained annotation record of a six-language benchmark for personally identifiable information contains two independent annotator labellings, the aggregate shipped as gold, and a reviewer gol
10月6日 04:00

category.学术arXiv cs.CL (自然语言处理)
arXiv:2610.04125v1 Announce Type: new Abstract: Self-report is an appealing low-cost probe of an LLM's dispositions, but recent work finds only selective agreement between what models report and how they behave. Prior accounts establish these patterns by prompting black-box LLMs, leaving open whether the gap is a prompting artefact or a fact about how the underlying constructs are represented inte
10月6日 04:00

category.学术arXiv cs.AI
arXiv:2610.03998v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used as synthetic personas representing survey respondents. Their validity as substitutes for particular respondents depends on whether they reproduce individuals' decisions. We examine what information helps synthetic respondents predict each individual's later choices, using five conditions that add pro
10月6日 04:00

category.学术arXiv cs.CL (自然语言处理)
arXiv:2610.04132v1 Announce Type: new Abstract: Building LLMs that behave well socially, not merely correctly, requires Building LLMs that behave well socially, not merely correctly, requires more than producing locally helpful responses. A socially competent agent must infer users' unstated goals, respect their preferences, and adapt as the conversation unfolds. These behaviors are inherently mul
10月6日 04:00

category.学术arXiv cs.CL (自然语言处理)
arXiv:2610.03829v1 Announce Type: new Abstract: Real-world outpatient oncology notes contain specialised terminology, tumour staging expressions, treatment names, toxicity descriptions, and institution-specific de-identification markers that may not be represented efficiently by general biomedical or adjacent clinical language models. We developed and evaluated oncology-specific BERT-style encoder
10月6日 04:00

category.学术arXiv cs.CL (自然语言处理)
arXiv:2610.03802v1 Announce Type: new Abstract: Context: Fine-grained emotion classification of mobile app reviews enables requirements engineering activities that go beyond polarity-based opinion mining, including emotionally informed issue prioritisation and feature-oriented feedback analysis. However, automatic fine-grained emotion extraction from app reviews remains understudied. Objectives: B
10月6日 04:00

category.学术arXiv cs.AI
arXiv:2610.03959v1 Announce Type: new Abstract: Recent world action models (WAMs) reuse pretrained video VAEs whose encoder latents directly condition downstream action policies. Quantization must therefore preserve not only reconstruction fidelity but also the policy-facing latent contract expected by the frozen policy. Direct NVFP4 leaves W4A4 quantization error uncompensated, whereas joint quan
10月6日 04:00

category.学术arXiv cs.AI
arXiv:2610.03984v1 Announce Type: new Abstract: Autonomous coding agents solve repository issues by reading code, running commands, editing files, and submitting patches. Extra inference-time compute yields gains only when it produces a useful repair and supplies reliable evidence for choosing one. Three behaviors decide both, and we argue they are teachable rather than byproducts of scale, so a p
10月6日 04:00

category.学术arXiv cs.CL (自然语言处理)
arXiv:2610.04074v1 Announce Type: new Abstract: Despite rapid progress in automating scientific research, generating promising and well grounded research solutions remains a central challenge. We isolate research ideation as a standalone task and build our solution on the intuition that a challenge in one field can often be addressed by a mechanism that solved an analogous challenge in another. Ac
10月6日 04:00

category.学术arXiv cs.AI
arXiv:2610.04012v1 Announce Type: new Abstract: Language-model systems can separate contextual computation, persistent storage, and exact execution instead of updating all capabilities through one shared parameter system. We investigate FEM-ASM, a finite-element-method-inspired organization in which independently constructed document states and deterministic executable skills contribute typed prop
10月6日 04:00

category.学术arXiv cs.AI
arXiv:2610.04008v1 Announce Type: new Abstract: Executable Agent Skills combine natural-language instructions and scripts into reusable packages for LLM agents, and revising them requires fixing errors without breaking correct behavior. Existing benchmarks do not systematically distinguish documentation repair, script repair, and preservation when evaluating skill self-evolution. We introduce Skil
10月6日 04:00

category.学术arXiv cs.AI
arXiv:2610.04088v1 Announce Type: new Abstract: Before autonomous driving systems can be deployed on public roads, it is vital that these systems comply with safety standards, traffic rules, and social norms. Although neural networks trained on large amounts of driving data perform well in routine driving tasks, these models often struggle in novel situations that are not well-represented in the d
10月6日 04:00

category.学术arXiv cs.AI
arXiv:2610.03872v1 Announce Type: new Abstract: AI agents are becoming increasingly capable of generating scientific code, but generating code is not the same as improving the algorithms behind it. For numerical solvers, execution feedback can expose poor performance, but rarely reveals its underlying cause and how to address it. We introduce Auto-Diagnosis and Skill Discovery (ADSD), a framework
10月6日 04:00

category.学术arXiv cs.AI
arXiv:2610.04011v1 Announce Type: new Abstract: Reinforcement learning with verifiable rewards improves reasoning, while the allocation of learning signal shapes which solutions remain accessible under repeated sampling. Group-relative objectives assign equal advantages to equally rewarded responses, making aggregate credit proportional to sampled mode frequency. We introduce Exploration-Preservin
10月6日 04:00

category.学术arXiv cs.CL (自然语言处理)
arXiv:2610.03935v1 Announce Type: new Abstract: General decision models, such as Jev, have recently emerged as efficient alternatives to LLMs for structured judgment and selection. But what kinds of decisions can these models reliably make, and how does their behavior change when individual decisions are composed into larger systems? To study this, we introduce JEVal, a bilingual benchmark compris
10月6日 04:00

category.学术arXiv cs.CL (自然语言处理)
arXiv:2610.03940v1 Announce Type: new Abstract: During fine-tuning, a language model can assign less probability to previously learned answers even when the current gradient acts to preserve that probability. With momentum, each update also carries gradients computed at earlier model states, and these stored contributions can push the model in the opposite direction. We investigate how this optimi
10月6日 04:00

category.学术arXiv cs.AI
arXiv:2610.04019v1 Announce Type: new Abstract: Graph based cyber attack detection studies employ various graph construction and representation strategies across different cybersecurity application domains. This diversity motivates a quantitative examination of how representation strategies are distributed across these application domains. This study presents a quantitative analysis of 37 original
10月6日 04:00

category.学术arXiv cs.AI
arXiv:2610.04072v1 Announce Type: new Abstract: Developing socially intelligent AI remains heavily dependent on human-annotated data, limiting the scale and breadth of social understanding models can acquire. Methods that derive training signals from unlabeled data offer a path beyond this dependence, but social predictions lack the verification oracles available in mathematics and coding. Moreove
10月6日 04:00

category.学术arXiv cs.CV (计算机视觉)
arXiv:2610.04073v1 Announce Type: new Abstract: Recent advances in attention-based deep learning have motivated their adoption for remote sensing image classification; however, their benefits for cryospheric imagery, where surface states are dominated by fine-grained textures and class imbalance, remain unclear. In this work, we revisit a benchmark Greenland Ice Sheet image dataset, previously sho
10月6日 04:00

category.学术arXiv cs.AI
arXiv:2610.04133v1 Announce Type: new Abstract: Multi-agent systems built on large language models (LLMs) are increasingly applied to scientific discovery and hypothesis generation. Both the effect of refinement and the diversity of the delivered set are hard to interpret before experimental ground truth exists, and both are typically reported by deciding whether pairs of generated hypotheses desc
10月6日 04:00

category.学术arXiv cs.CV (计算机视觉)
arXiv:2610.04066v1 Announce Type: new Abstract: Automatic calving-front delineation from synthetic aperture radar imagery is challenging because the front is a thin and often ambiguous boundary between glacier ice, ocean, and surrounding rock or terrain. The CAlving Fronts and where to Find thEm (CaFFe) dataset provides both binary calving-front masks and broader semantic zone masks, making it pos
10月6日 04:00

category.学术arXiv cs.CV (计算机视觉)
arXiv:2610.04014v1 Announce Type: new Abstract: We introduce the first learned articulated body simulator based on a selective state space model (SSM), called PhysMamba. PhysMamba predicts next-frame full-body state from position, rotation, and joint-action history, without velocity inputs. We compare four architectures under partial- and full-observation inputs and three training protocols. The f
10月6日 04:00

category.学术arXiv cs.CV (计算机视觉)
arXiv:2610.04104v1 Announce Type: new Abstract: Semantic Knowledge-driven Contrastive Learning (SKCL) uses a language model to decide which classes are related, and pulls each image towards the prototypes of its semantic neighbours. On CIFAR-100-LT (beta = 100) it reports 54.02% top-1 accuracy, 2.01 points above Balanced Contrastive Learning (BCL), the method it builds on. The code and the class d
10月6日 04:00