Read Less. Know More.

全部资讯

754 条资讯

category.学术arXiv cs.LG (机器学习)

Protecting Sensitive Data in Image Synthesis via PAC-Private Adaptation for Diffusion Models

arXiv:2610.04038v1 Announce Type: new Abstract: Synthetic data are increasingly used as an alternative to sharing sensitive records. However, synthetic data generation does not guarantee privacy, as diffusion models trained or adapted on sensitive data remain susceptible to reconstruction attacks. Moreover, while approaches that use differential privacy (DP), such as DP-SGD, achieve provably priva

10月6日 04:00
category.学术arXiv cs.LG (机器学习)

Memory-State Critic for Asymmetric Actor-Critic with Application to Vision-Based Pursuit-Evasion

arXiv:2610.03830v1 Announce Type: new Abstract: In partially observable Markov decision processes, the optimal policy generally depends on the history of observations and past actions. Asymmetric actor-critic methods have become popular to learn such policies when additional information, such as the true state of the environment, is available during training. The critic, which is not needed at exe

10月6日 04:00
category.学术arXiv cs.LG (机器学习)

Where Does Jagged Competence Come From?

arXiv:2610.03831v1 Announce Type: new Abstract: Capable systems often show jagged competence: low average error alongside failures on particular inputs. We ask where it comes from in a task built from two known layers. A lower layer A computes five per-slot sums from records; an upper layer B uses the slot-1 sum and a mode carried over from earlier boards to predict the next board's category, so B

10月6日 04:00
category.学术arXiv cs.CV (计算机视觉)

FADE: Frame-Aware Diffusion-Transformer-based Multi-Concept Erasure for Video Unlearning

arXiv:2610.03980v1 Announce Type: new Abstract: Text-to-video (T2V) diffusion models can reproduce copyrighted, violent, or explicit content, which motivates concept erasure: removing designated concepts from a pretrained model while preserving its behavior on everything else. Existing T2V erasure methods leave two problems open. Their frame-agnostic suppression can leave isolated frames in which

10月6日 04:00
category.学术arXiv cs.LG (机器学习)

LD-EnFF: Latent-Dynamics Ensemble Flow Filtering for Data Assimilation with Sparse Observations

arXiv:2610.04034v1 Announce Type: new Abstract: Data assimilation combines model forecasts with noisy, incomplete observations to estimate the evolving state of a dynamical system. Existing methods face two compounding challenges: high-dimensional nonlinear dynamics make repeated forward simulation computationally expensive, while sparse observations provide limited direct information about the fu

10月6日 04:00
category.学术arXiv cs.CL (自然语言处理)

XTurnix: Large-Scale Self-Supervised Turn Control through Two-State Binary Decisions

arXiv:2610.04400v1 Announce Type: new Abstract: General turn-taking behavior in real-time dialogue systems requires deciding whether to keep listening or start responding while listening, and whether to continue or stop while speaking. Existing turn detectors use heterogeneous, task-specific label spaces and are often trained on limited annotations or evaluated on isolated utterances, making them

10月6日 04:00
category.学术arXiv cs.CV (计算机视觉)

StepCAD: Mesh-to-CAD Code Generation via LLM Policy and Geometry-Guided Search

arXiv:2610.03799v1 Announce Type: new Abstract: Recovering executable CAD programs from 3D meshes is challenging due to the compositional nature of CAD construction and the interaction between discrete modeling choices and continuous parameters. Many learning-based methods predict complete programs in a single pass and rely predominantly on sketch-extrude representations, limiting operation divers

10月6日 04:00
category.学术arXiv cs.CL (自然语言处理)

Ideological Stance Detection in a Low-Resource Language: Polarization in Bangladeshi Public vs Private University Discourse on Social Media

arXiv:2610.04401v1 Announce Type: new Abstract: Public vs. private universities is a debatable issue, and it creates polarization on social media in Bangladesh. Debate on quality, jobs, and prestige is passionate among the students, parents, and graduates, the majority of whom speak Bengali, a low-resource language. To measure this polarization, this paper introduces a manually annotated dataset o

10月6日 04:00
category.学术arXiv cs.CL (自然语言处理)

Evidence and Intervention: A Coupled Active-Inference Extension of Rational Speech Act Models

arXiv:2610.04347v1 Announce Type: new Abstract: Identical utterance choices can arise from different communicative causes, and identical interpretations can leave different traces in what a listener learns. Rational Speech Act (RSA) models treat interpretation as inference over speaker meaning, but standard one-shot RSA does not intrinsically distinguish these causal update targets. We develop a c

10月6日 04:00
category.学术arXiv cs.CV (计算机视觉)

Masked Privileged-Information Distillation for Multimodal Skin Lesion Classification Under Missing Clinical Metadata

arXiv:2610.03991v1 Announce Type: new Abstract: Multimodal skin lesion classification combines clinical images with patient metadata to improve diagnostic accuracy. However, complete metadata available during training may be only partially accessible at deployment, and resource-constrained settings additionally require computational efficiency. We address these challenges with a privileged-informa

10月6日 04:00
category.学术arXiv cs.CV (计算机视觉)

WAMJET: A Harness for World Action Model Acceleration

arXiv:2610.03797v1 Announce Type: new Abstract: World Action Models (WAMs) leverage pretrained video foundation models for robot manipulation, but their large backbones and video-action co-prediction are expensive. Although existing acceleration techniques offer many ways to reduce this cost, selecting and composing them requires substantial engineering for each model and hardware platform. To tac

10月6日 04:00
category.学术arXiv cs.CL (自然语言处理)

Saying, Not Knowing: Aggressively GGUF-Quantized Small Language Models Still Write Rare Words They Can No Longer Define

arXiv:2610.04403v1 Announce Type: new Abstract: Post-training quantization to the GGUF format's mixed-precision K-quants is commonly how open-weight language models reach consumer hardware, yet its effect on fine-grained lexical competence is uncharacterized. We audit 27 quantized artifacts across 13 families and four architecture backbones, 0.35B-14B parameters, evaluated down their published lad

10月6日 04:00
category.学术arXiv cs.LG (机器学习)

Evolving LLM-Generated Features for Interpretable Classification

arXiv:2610.03951v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used as classifiers, yet they operate as opaque systems whose decisions are difficult to interpret, which complicates their use in regulated domains such as credit scoring or medical diagnosis. We propose an evolutionary framework that iteratively discovers natural language feature definitions (rubrics) f

10月6日 04:00
category.学术arXiv cs.LG (机器学习)

$\tilde{O}(\sqrt{T})$ Regret and Polylogarithmic Constraint Violation for COCO

arXiv:2610.03983v1 Announce Type: new Abstract: We study constrained online convex optimization with adversarial convex losses and constraints ($\mathsf{COCO}$). At each round \(t\in[T]\), a learner selects \(x_t\) from a \(d\)-dimensional convex decision set \(\mathcal X\), after which an adaptive adversary reveals a convex cost function \(f_t\) and constraint function \(g_t\). Consequently, the

10月6日 04:00
category.学术arXiv cs.LG (机器学习)

AdaEva: Accelerating LLM-Driven Algorithm Design with Adaptive Partial Evaluation

arXiv:2610.03896v1 Announce Type: new Abstract: Large Language Models (LLMs) are increasingly used for automated algorithm design. However the computational cost of evaluating the generated algorithms can be excessive. We consider the common LLM-driven automated algorithm design (LLM4AD) setting in which a candidate algorithm is evaluated by aggregating its performance over a shared set of trainin

10月6日 04:00
category.学术arXiv cs.LG (机器学习)

Reinforcement Learning on the Discrete Composition Channel of a Crystal Generator: Validated Gains and Reward Hacking

arXiv:2610.03880v1 Announce Type: new Abstract: Inverse materials design is a long-standing goal of computational materials discovery. Generative models for crystalline materials are typically trained to match the distribution of a structure database, while nothing in their training objective points them at specific design goals such as targeted properties. We use group-relative policy optimizatio

10月6日 04:00
category.学术arXiv cs.CL (自然语言处理)

Boundaries Agree, Labels Do Not: Intra-Annotator Dynamics as a Kind of Training Data

arXiv:2610.04370v1 Announce Type: new Abstract: Data quality now matters as much as compute for training language models. Much training data comes from human annotation of text, and interpretive annotation has no ground truth that could settle what is "accurate". Two lines of work respond to this. One combines annotators into a "ground truth" and measures how well they agree with each other; the o

10月6日 04:00
category.学术arXiv cs.LG (机器学习)

One-Step Curvature Probes Miss the Fitting Operator: Retained Capacity and Terminal Null-Space Correction for Continual Learning

arXiv:2610.03952v1 Announce Type: new Abstract: A one-step curvature probe evaluates an initial direction, whereas continual learners are judged after reaching comparable new-task fit. In an overparameterized linearization, projected gradient descent converges to $\Delta_P=PJ^\top(JPJ^\top)^{-1}r$, and its squared-displacement inflation is exactly the reciprocal of the retained fitting capacity $c

10月6日 04:00
category.学术arXiv cs.CV (计算机视觉)

What Do Verifiable Rewards Teach Video-Language Models About Time? A Controlled Multi-Model Study

arXiv:2610.03792v1 Announce Type: new Abstract: Reinforcement learning from verifiable rewards (RLVR) has produced large reasoning gains in language models, and verifiable video benchmarks make it applicable to causal-temporal video question answering. We study what RLVR teaches video-language models about time. We fine-tune four open models (Qwen3-VL-8B/4B, Qwen2.5-VL-7B, Gemma-3-12B) with group

10月6日 04:00
category.学术arXiv cs.CV (计算机视觉)

DABACO: A Multi-Camera Dataset and Benchmark for Screen Localization and Pointing Estimation

arXiv:2610.03928v1 Announce Type: new Abstract: Screen localization and pointing estimation are key to low-cost interactive devices. Yet developing and evaluating these algorithms requires realistic data: synthetic captures cannot fully reproduce the optical distortion, rolling shutter, motion blur, and display processing of a physical acquisition, and most existing datasets provide static images

10月6日 04:00
category.学术arXiv cs.CL (自然语言处理)

Conformal Prediction with Paraphrase-Aware Scoring for LLM Uncertainty Quantification

arXiv:2610.04239v1 Announce Type: new Abstract: Uncertainty quantification (UQ) for large language models (LLMs) aims to provide reliable measures of predictive confidence, yet current methods are often unstable under meaning-preserving perturbations. Semantically equivalent paraphrases can induce substantial variability in predictive confidence, even for methods with formal guarantees, such as co

10月6日 04:00
category.学术arXiv cs.CL (自然语言处理)

Trajectory-Derived Confidence for Reliable, Resource-Aware Clinical Text-to-SQL Agents

arXiv:2610.04156v1 Announce Type: new Abstract: LLM agents for clinical text-to-SQL applications reason autonomously over multiple steps but cannot assess whether their own reasoning or outputs can be trusted. In high leverage applications such as healthcare, this presents a critical risk where system mistakes can be costly. These reliability failures are also resource failures: an incorrect reaso

10月6日 04:00
category.学术arXiv cs.LG (机器学习)

TreeWalker: Partial Evaluation for Grouped Tree-Ensemble Inference

arXiv:2610.03939v1 Announce Type: new Abstract: Many inference workloads evaluate a trained tree ensemble on row groups that share feature values: discrete-time survival models expand each patient into $G$ time steps, click-through-rate models score every item in a search session, and scenario analyses vary a few inputs while holding the rest fixed. Standard inference treats each row independently

10月6日 04:00
category.学术arXiv cs.CL (自然语言处理)

The Score Is Not the Structure: Brain Alignment and Cross-Lingual Transfer

arXiv:2610.03827v1 Announce Type: new Abstract: Researchers often support the claim that a model shares structure with the brain, or across languages, by reporting a similarity score. We ask what such a score reads when the shared structure is absent, or when the tool that measures it does not work. We check two settings, and in both the score is not what it appears. First, a probe trained to tell

10月6日 04:00
category.学术arXiv cs.AI

ROAR: Unifying Runs across Heterogeneous AI-Driven Research Systems

arXiv:2610.03966v1 Announce Type: new Abstract: Each run of an AI-driven research system (ADRS) is an expensive search over a vast solution space, and dependable evaluation requires many runs, making run data both costly to produce and valuable to retain for large-scale analysis. Yet this data remains fragmented: teams operate in isolation, ADRS frameworks emit results in different formats, and no

10月6日 04:00
category.学术arXiv cs.CL (自然语言处理)

Same Output, Different Gold: Measuring How Reference Choice Moves a Multilingual Benchmark Score

arXiv:2610.03825v1 Announce Type: new Abstract: A benchmark score compares a system output against a reference, and methodological attention falls almost entirely on the first term. We measure the second. The retained annotation record of a six-language benchmark for personally identifiable information contains two independent annotator labellings, the aggregate shipped as gold, and a reviewer gol

10月6日 04:00
category.学术arXiv cs.CL (自然语言处理)

Representational Control over Self-Report & Behavior Coherence in LLM Risk-Taking

arXiv:2610.04125v1 Announce Type: new Abstract: Self-report is an appealing low-cost probe of an LLM's dispositions, but recent work finds only selective agreement between what models report and how they behave. Prior accounts establish these patterns by prompting black-box LLMs, leaving open whether the gap is a prompting artefact or a fact about how the underlying constructs are represented inte

10月6日 04:00
category.学术arXiv cs.AI

Behavioral History Outperforms Descriptions of the Person for LLM Synthetic Personas

arXiv:2610.03998v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used as synthetic personas representing survey respondents. Their validity as substitutes for particular respondents depends on whether they reproduce individuals' decisions. We examine what information helps synthetic respondents predict each individual's later choices, using five conditions that add pro

10月6日 04:00
category.学术arXiv cs.CL (自然语言处理)

PB-GRPO: Learning Socially Adaptive LLM Agents from Persona-Driven Simulation with Preference-Batched GRPO

arXiv:2610.04132v1 Announce Type: new Abstract: Building LLMs that behave well socially, not merely correctly, requires Building LLMs that behave well socially, not merely correctly, requires more than producing locally helpful responses. A socially competent agent must infer users' unstated goals, respect their preferences, and adapt as the conversation unfolds. These behaviors are inherently mul

10月6日 04:00
category.学术arXiv cs.CL (自然语言处理)

OncoNoteBERT: A Foundation Representation Model for Natural Language Processing of Real-World Outpatient Oncology Notes

arXiv:2610.03829v1 Announce Type: new Abstract: Real-world outpatient oncology notes contain specialised terminology, tumour staging expressions, treatment names, toxicity descriptions, and institution-specific de-identification markers that may not be represented efficiently by general biomedical or adjacent clinical language models. We developed and evaluated oncology-specific BERT-style encoder

10月6日 04:00
category.学术arXiv cs.CL (自然语言处理)

Fine-Grained Emotion Classification from Mobile App Reviews: An Empirical Study with Large Language Models

arXiv:2610.03802v1 Announce Type: new Abstract: Context: Fine-grained emotion classification of mobile app reviews enables requirements engineering activities that go beyond polarity-based opinion mining, including emotionally informed issue prioritisation and feature-oriented feedback analysis. However, automatic fine-grained emotion extraction from app reviews remains understudied. Objectives: B

10月6日 04:00
category.学术arXiv cs.AI

LatentQuant: Preserving the Policy-Facing Latent Contract under NVFP4 VAE Quantization

arXiv:2610.03959v1 Announce Type: new Abstract: Recent world action models (WAMs) reuse pretrained video VAEs whose encoder latents directly condition downstream action policies. Quantization must therefore preserve not only reconstruction fidelity but also the policy-facing latent contract expected by the frozen policy. Direct NVFP4 leaves W4A4 quantization error uncompensated, whereas joint quan

10月6日 04:00
category.学术arXiv cs.AI

Teaching Agents to Code Reliably

arXiv:2610.03984v1 Announce Type: new Abstract: Autonomous coding agents solve repository issues by reading code, running commands, editing files, and submitting patches. Extra inference-time compute yields gains only when it produces a useful repair and supplies reliable evidence for choosing one. Three behaviors decide both, and we argue they are teachable rather than byproducts of scale, so a p

10月6日 04:00
category.学术arXiv cs.CL (自然语言处理)

IdeaScientist: Orchestrating Agents for Grounded Scientific Ideation

arXiv:2610.04074v1 Announce Type: new Abstract: Despite rapid progress in automating scientific research, generating promising and well grounded research solutions remains a central challenge. We isolate research ideation as a standalone task and build our solution on the intuition that a challenge in one field can often be addressed by a mechanism that solved an analogous challenge in another. Ac

10月6日 04:00
category.学术arXiv cs.AI

Beyond the Parameter Monolith: Reconstructive Memories, Executable Skills, and Residual Assembly for Language Models

arXiv:2610.04012v1 Announce Type: new Abstract: Language-model systems can separate contextual computation, persistent storage, and exact execution instead of updating all capabilities through one shared parameter system. We investigate FEM-ASM, a finite-element-method-inspired organization in which independently constructed document states and deterministic executable skills contribute typed prop

10月6日 04:00
category.学术arXiv cs.AI

SkillScriptBench: Benchmarking Self-Evolution of Executable Agent Skill Packages Beyond Markdown

arXiv:2610.04008v1 Announce Type: new Abstract: Executable Agent Skills combine natural-language instructions and scripts into reusable packages for LLM agents, and revising them requires fixing errors without breaking correct behavior. Existing benchmarks do not systematically distinguish documentation repair, script repair, and preservation when evaluating skill self-evolution. We introduce Skil

10月6日 04:00
category.学术arXiv cs.AI

Towards Safer Autonomous Driving in an Open World: A Dual-Process Approach

arXiv:2610.04088v1 Announce Type: new Abstract: Before autonomous driving systems can be deployed on public roads, it is vital that these systems comply with safety standards, traffic rules, and social norms. Although neural networks trained on large amounts of driving data perform well in routine driving tasks, these models often struggle in novel situations that are not well-represented in the d

10月6日 04:00
category.学术arXiv cs.AI

Training Numerical Intelligence via Auto-Diagnosis and Skill Discovery

arXiv:2610.03872v1 Announce Type: new Abstract: AI agents are becoming increasingly capable of generating scientific code, but generating code is not the same as improving the algorithms behind it. For numerical solvers, execution feedback can expose poor performance, but rarely reveals its underlying cause and how to address it. We introduce Auto-Diagnosis and Skill Discovery (ADSD), a framework

10月6日 04:00
category.学术arXiv cs.AI

Exploration-Preserving Policy Optimization

arXiv:2610.04011v1 Announce Type: new Abstract: Reinforcement learning with verifiable rewards improves reasoning, while the allocation of learning signal shapes which solutions remain accessible under repeated sampling. Group-relative objectives assign equal advantages to equally rewarded responses, making aggregate credit proportional to sampled mode frequency. We introduce Exploration-Preservin

10月6日 04:00
category.学术arXiv cs.CL (自然语言处理)

General Decision Models: Benchmarking and Insights Beyond Jev

arXiv:2610.03935v1 Announce Type: new Abstract: General decision models, such as Jev, have recently emerged as efficient alternatives to LLMs for structured judgment and selection. But what kinds of decisions can these models reliably make, and how does their behavior change when individual decisions are composed into larger systems? To study this, we introduce JEVal, a bilingual benchmark compris

10月6日 04:00
category.学术arXiv cs.CL (自然语言处理)

A Step Towards Forgetting: Optimiser History and the Loss of Answer Mass

arXiv:2610.03940v1 Announce Type: new Abstract: During fine-tuning, a language model can assign less probability to previously learned answers even when the current gradient acts to preserve that probability. With momentum, each update also carries gradients computed at earlier model states, and these stored contributions can push the model in the opposite direction. We investigate how this optimi

10月6日 04:00
category.学术arXiv cs.AI

A Quantitative Analysis of Graph Representation Strategies for Cyber Attack Detection

arXiv:2610.04019v1 Announce Type: new Abstract: Graph based cyber attack detection studies employ various graph construction and representation strategies across different cybersecurity application domains. This diversity motivates a quantitative examination of how representation strategies are distributed across these application domains. This study presents a quantitative analysis of 37 original

10月6日 04:00
category.学术arXiv cs.AI

Reinforcement Learning with Comparative Evidence for Social Intelligence

arXiv:2610.04072v1 Announce Type: new Abstract: Developing socially intelligent AI remains heavily dependent on human-annotated data, limiting the scale and breadth of social understanding models can acquire. Methods that derive training signals from unlabeled data offer a path beyond this dependence, but social predictions lack the verification oracles available in mathematics and coding. Moreove

10月6日 04:00
category.学术arXiv cs.CV (计算机视觉)

Why Convolution Still Matters: Evaluating Inductive Biases in Cryospheric Image Classification

arXiv:2610.04073v1 Announce Type: new Abstract: Recent advances in attention-based deep learning have motivated their adoption for remote sensing image classification; however, their benefits for cryospheric imagery, where surface states are dominated by fine-grained textures and class imbalance, remain unclear. In this work, we revisit a benchmark Greenland Ice Sheet image dataset, previously sho

10月6日 04:00
category.学术arXiv cs.AI

Auditing Pairwise Equivalence Judgments: Self-Critique Effects and Diversity Measurement in Multi-Agent Hypothesis Generation

arXiv:2610.04133v1 Announce Type: new Abstract: Multi-agent systems built on large language models (LLMs) are increasingly applied to scientific discovery and hypothesis generation. Both the effect of refinement and the diversity of the delivered set are hard to interpret before experimental ground truth exists, and both are typically reported by deciding whether pairs of generated hypotheses desc

10月6日 04:00
category.学术arXiv cs.CV (计算机视觉)

Evaluating Zone-Guided Front Extraction for Glacier Calving-Front Delineation in SAR Imagery

arXiv:2610.04066v1 Announce Type: new Abstract: Automatic calving-front delineation from synthetic aperture radar imagery is challenging because the front is a thin and often ambiguous boundary between glacier ice, ocean, and surrounding rock or terrain. The CAlving Fronts and where to Find thEm (CaFFe) dataset provides both binary calving-front masks and broader semantic zone masks, making it pos

10月6日 04:00
category.学术arXiv cs.CV (计算机视觉)

PhysMamba: Selective State Space Models as Learned Articulated Body Simulators

arXiv:2610.04014v1 Announce Type: new Abstract: We introduce the first learned articulated body simulator based on a selective state space model (SSM), called PhysMamba. PhysMamba predicts next-frame full-body state from position, rotation, and joint-action history, without velocity inputs. We compare four architectures under partial- and full-observation inputs and three training protocols. The f

10月6日 04:00
category.学术arXiv cs.CV (计算机视觉)

Where Does the Semantic Gain Come From? A Reproduction and Extension of Semantic Knowledge-driven Contrastive Learning for Long-Tailed Recognition

arXiv:2610.04104v1 Announce Type: new Abstract: Semantic Knowledge-driven Contrastive Learning (SKCL) uses a language model to decide which classes are related, and pulls each image towards the prototypes of its semantic neighbours. On CIFAR-100-LT (beta = 100) it reports 54.02% top-1 accuracy, 2.01 points above Balanced Contrastive Learning (BCL), the method it builds on. The code and the class d

10月6日 04:00