Read Less. Know More.

全部资讯

754 条资讯

category.学术arXiv cs.CV (计算机视觉)

VCURF: Virtual Camera-based Uncertainty of Radiance Fields

arXiv:2610.04076v1 Announce Type: new Abstract: Radiance fields, implemented with either implicit (NeRF) or explicit (Gaussian Splatting) representations, are advancing the state of the art in novel view synthesis at a rapid pace. Even though the rendered views they generate are often compelling, they are not free of errors. In this paper, we propose a new approach for pixel-wise uncertainty quant

10月6日 04:00
category.学术arXiv cs.CV (计算机视觉)

Dynamic Quadtree Tokenization and Transformer for Adaptive Mesh PDE Forecasting

arXiv:2610.04044v1 Announce Type: new Abstract: The quadratic attention cost of Vision Transformers (ViTs) forces a trade-off between spatial resolution and rollout horizon, particularly for fine-scale PDEs where shocks, reaction fronts, and material interfaces occupy small, evolving regions of the domain. Conventional neural surrogates also lack mechanisms to adapt resolution dynamically. We prop

10月6日 04:00
category.学术arXiv cs.LG (机器学习)

Progressive Multi-Ancestor Bit-Depth Distillation

arXiv:2610.04100v1 Announce Type: new Abstract: Model compression strategies are widely employed to reduce memory footprint and network complexity, particularly for devices with constrained computational, memory, and energy resources. Prior works that rely on simultaneous conversion from floating-point high-precision (FP32) to integer low-precision (INT4) representations and distillation into smal

10月6日 04:00
category.学术arXiv cs.LG (机器学习)

Articulatory Entrainment and Coordination Complexity in Spontaneous Autistic and Non-autistic Dialogue

arXiv:2610.04071v1 Announce Type: new Abstract: Articulatory entrainment, the adaptation of vocal tract coordination to facilitate interaction remains underexplored in spontaneous dialogue, particularly among autistic speakers. Many prior studies have utilized task-based, phoneme-level analyses with invasive measurement techniques. Here, we introduce a speaker-independent framework to quantify art

10月6日 04:00
category.学术arXiv cs.LG (机器学习)

On architectural choices for interpretability and thermodynamic consistency in Physically Recurrent Neural Networks in the low-data regime

arXiv:2610.04067v1 Announce Type: new Abstract: In this paper, we unravel the effect of different decoder architectures on the interpretability of the latent space of the Physically Recurrent Neural Network. Particular emphasis is given to a new weight normalization constraint, which acts as a regularization technique and enables robust training in the low-data regime. A brief visual exploration i

10月6日 04:00
category.学术arXiv cs.LG (机器学习)

Slaying the Hydra: Interaction-Aware Circuit Discovery in Language Models

arXiv:2610.04017v1 Announce Type: new Abstract: Localizing behavior to individual components of a language model is a central goal of mechanistic interpretability. However, scoring components one at a time misses context-dependent effects: a primary component can inhibit the activation of a backup, leading to issues with ranking components. Actual causality studies the structure of such interactio

10月6日 04:00
category.学术arXiv cs.CV (计算机视觉)

ASD-FEAT: A Multi-Modal Infant Video-Derived Dataset for Early ASD Risk Prediction

arXiv:2610.04051v1 Announce Type: new Abstract: Accurate early screening for Autism Spectrum Disorder (ASD) is a precursor to timely intervention, which is critical for improving cognitive and behavioral outcomes. We present ASD-FEAT (ASD - Feature Extraction And Tracking), a multimodal dataset derived from video recordings of infant-adult interaction sessions. The key contribution of ASD-FEAT is

10月6日 04:00
category.学术arXiv cs.CV (计算机视觉)

A Theory of Shape Reconstruction from Heat Conduction and Shading

arXiv:2610.04052v1 Announce Type: new Abstract: Shape from shading using a single image of a Lam- bertian surface is inherently ambiguous. When the light source direction is known, the surface normal estimation has a cone- ambiguity, which worsens when the source is unknown. Recently, shape from heat conduction has emerged as an approach that leverages heat transport equations to estimate the Shap

10月6日 04:00
category.学术arXiv cs.CV (计算机视觉)

Scaling 3D Visual Grounding in Abdominal CT

arXiv:2610.04095v1 Announce Type: new Abstract: Visual grounding models can enhance radiology workflows by linking report findings to image regions. This is particularly valuable for 3D CT, where findings often occupy a tiny fraction of the volume. Training 3D grounding models requires large sets of paired phrases and regions, and building such datasets is expensive, requiring radiologists to anno

10月6日 04:00
category.学术arXiv cs.LG (机器学习)

Pareto-Dominant Clarification: Post-Training Coding LLMs via PPO-Lagrangian Budget Constraints

arXiv:2610.04089v1 Announce Type: new Abstract: Coding agents operating under ambiguous instructions or user prompts must decide whether to ask clarifying questions or attempt a solution directly. While clarification from the user may improve the correctness of the agent's solution, each back-and-forth interaction incurs user and system costs, forming an explicit accuracy vs. efficiency tradeoff.

10月6日 04:00
category.学术arXiv cs.CV (计算机视觉)

Dependable AI-Assisted Engineering: A Formal Framework for AI Participation and Assurance in Safety-Critical Workflows

arXiv:2610.04084v1 Announce Type: new Abstract: Generative AI can produce engineering artefacts, but generation alone does not determine whether or how those artefacts should enter safety-critical workflows. This paper develops a formal framework for assigning AI participation and assurance at the level of individual workflow units. Each unit has a participation and assurance record covering its e

10月6日 04:00
category.学术arXiv cs.CL (自然语言处理)

Evaluating Modeling Approaches for Experience-Level Classification in Job Description

arXiv:2610.04304v1 Announce Type: new Abstract: This paper investigates the task of predicting job experience levels in recruitment texts, aiming to automatically identify the qualifications required for positions. Unlike traditional text classification, recruitment texts typically possess explicit internal structures, with different paragraphs playing disproportionate roles in conveying experienc

10月6日 04:00
category.学术arXiv cs.AI

Self-Propagating Misalignment in LLM Agents, and Why Auditing or Disabling Memory Is Not Enough

arXiv:2610.04083v1 Announce Type: new Abstract: Memory poisoning attacks on LLM agents typically assume an external adversary who plants content in the agent's persistent memory to steer its behavior. We instead study, with no adversary involved, whether a misaligned agent can write a goal it cannot yet act on to persistent memory, so that a future aligned agent carries it out when the opportunity

10月6日 04:00
category.学术arXiv cs.CV (计算机视觉)

UniBRep: Learning Unified Geometry and Topology for Image-conditioned B-Rep Generation

arXiv:2610.04092v1 Announce Type: new Abstract: Generating a boundary representation (B-rep) conditioned on a single image requires faithful reconstruction of geometry, valid topology, and support for complex shapes. We present UniBRep, a geometry-first framework that adapts a pretrained image-to-3D model to generate a feature mesh as a unified intermediate representation. Its surface provides a g

10月6日 04:00
category.学术arXiv cs.CL (自然语言处理)

GlitchPatch: Repairing Glitch Tokens in Frozen Language Models via Local Retokenization

arXiv:2610.04399v1 Announce Type: new Abstract: Glitch tokens are anomalous vocabulary entries that can cause large language models (LLMs) to produce outputs inconsistent with their inputs. Existing repair methods require access to model internals, making them impractical for frozen checkpoints. We investigate whether glitch tokens can be repaired outside the model by optimizing the input tokeniza

10月6日 04:00
category.学术arXiv cs.CV (计算机视觉)

Beyond Token Accuracy: Prioritizing What Matters for Visual Reconstruction

arXiv:2610.03822v1 Announce Type: new Abstract: Tokens have become a unified interface for multimodal foundation models, making visual-token communication a natural paradigm for efficient image delivery. However, existing methods typically rely on static policies that cannot jointly adapt to image content and channel conditions. Moreover, their token-level utility objectives do not necessarily tra

10月6日 04:00
category.学术arXiv cs.LG (机器学习)

DUET: Co-Evolving Solver and Grader Agents

arXiv:2610.04087v1 Announce Type: new Abstract: Agentic workflows are increasingly used across domains such as technology, finance, and enterprise operations. As these agents become more widely deployed, continually improving them becomes increasingly important. This raises an immediate challenge: How should the agent evolve? This evolution requires effective evaluation that can assess outcomes an

10月6日 04:00
category.学术arXiv cs.LG (机器学习)

Robust blind unmixing: A geometric approach to overcoming basis variation

arXiv:2610.04091v1 Announce Type: new Abstract: Signal separation problems are common in science. A prominent example of this occurs during the use of diffraction or spectroscopy to identify the individual components of a mixture by measuring it. In the simplest case, the measured signal is a linear combination of basis patterns corresponding to the constituent parts. The unmixing problem is to in

10月6日 04:00
category.学术arXiv cs.CV (计算机视觉)

Are We Measuring Anticipation? Auditing Privileged Information in Procedural Video Evaluation

arXiv:2610.03826v1 Announce Type: new Abstract: Benchmark scores license claims about the capabilities being evaluated. We audit the inference licensed by an evaluation protocol, rather than the predictive model alone. Using procedural action anticipation as a controlled case study, we study a broader evaluation-validity failure mode: a protocol can remain temporally causal and free of classical t

10月6日 04:00
category.学术arXiv cs.LG (机器学习)

Beyond Masked Sparsity: SNACK Enables Truly Sparse Neural Networks on GPU

arXiv:2610.04093v1 Announce Type: new Abstract: Deep neural networks continue to grow in parameter count, driving up training and inference cost on GPUs. Sparse neural networks and Dynamic Sparse Training (DST) promise to reduce these costs, but most implementations rely on binary masks over dense tensors and recover little of the theoretical compute, memory, or energy savings. We propose SNACK, a

10月6日 04:00
category.学术arXiv cs.CL (自然语言处理)

Can LLMs Separate Pasted Artifacts from User Speech? Absorption at Unmarked Prompt Seams

arXiv:2610.04210v1 Announce Type: new Abstract: Large language models (LLMs) receive each user message as plain text, even when it combines text from different sources. For example, a user may paste text into a prompt and keep typing a comment directly below it. We study absorption: a phenomenon where the model treats a trailing user comment as part of the pasted text, returning it inside the edit

10月6日 04:00
category.学术arXiv cs.LG (机器学习)

Retrieval-Centric Deep Learning in Growing Nonparametric Neural Networks

arXiv:2610.03858v1 Announce Type: new Abstract: We investigate a general-purpose layer for deep learning that, instead of compressing arbitrary-size training data into fixed-size weight matrices, stores a new pair of key-value representations for every data point during training, and retrieves and recombines these representations through an attention mechanism at inference time - resulting in a gr

10月6日 04:00
category.学术arXiv cs.LG (机器学习)

BAT-NO: A Boundary-Condition-Aware Transformer Neural Operator for Crashworthiness Prediction of Vehicle Components

arXiv:2610.03854v1 Announce Type: new Abstract: High-fidelity finite-element simulations provide accurate crashworthiness predictions, but their cost limits iterative design exploration. Deep learning surrogates can reduce this cost, but many component-level models are developed under a single prescribed boundary condition, limiting generalisation to boundary variations. This work proposes a Bound

10月6日 04:00
category.学术arXiv cs.CL (自然语言处理)

AI-Enabled Quality Assurance for Multiple-Choice Assessment Items

arXiv:2610.04267v1 Announce Type: new Abstract: Generating multiple-choice questions is increasingly scalable, but establishing their assessment quality remains difficult. We present a focused narrative review of automated item-writing flaw detection, revision, psychometric screening, and NLP benchmark auditing. Database searches, citation retrieval, and nominated sources yield fourteen research r

10月6日 04:00
category.学术arXiv cs.CV (计算机视觉)

Fractal Cross Product: Theory, Differentiable Implementation and Application to Medical Image Analysis

arXiv:2610.03755v1 Announce Type: new Abstract: The magnitude of the generalized Euclidean cross product is a Gram volume whose degree under common scaling is fixed by the integer dimension of the spanning frame. We formulate a generalized Fractal Cross Product (FCP) as a nonlinear radial deformation with a prescribed positive degree $D$, which may be non-integer. The scalar construction applies i

10月6日 04:00
category.学术arXiv cs.CV (计算机视觉)

Do Motion Tokenizers for Co-Speech Gesture Generation Encode Gesture Semantics?

arXiv:2610.03765v1 Announce Type: new Abstract: Discrete motion tokenizers encode motion as atomic units and are widely used for co-speech gesture generation. It remains unclear which motion properties, especially those relevant to gesture semantics, are recoverable from these codebooks. We probe a reconstruction-trained codebook using 19 co-speech gesture descriptors spanning from raw motion to a

10月6日 04:00
category.学术arXiv cs.CL (自然语言处理)

Playing social deduction games with reinforcement fine-tuned large language models

arXiv:2610.04261v1 Announce Type: new Abstract: Reinforcement fine-tuning (RFT) is increasingly used in applications where large language models (LLMs) interact with humans and other agents. Here we use social deduction games to study how RFT changes LLMs' social behaviour. We let fine-tuned and base LLM agents play hidden-role games that require hidden-state inference, social reading and vote ste

10月6日 04:00
category.学术arXiv cs.LG (机器学习)

Bayes-Sufficient Compression Is Not Enough: How Does Communication Help Multi-Agent Systems?

arXiv:2610.03769v1 Announce Type: new Abstract: Multi-agent LLM systems pair a sender with broad context and an executor with a limited local view. We study when a short message improves the executor's next decision, when raw context is preferable, and when a stronger sender helps. Our framework, \emph{receiver-relative bounded coordination}, expresses message utility as receiver gain minus protoc

10月6日 04:00
category.学术arXiv cs.LG (机器学习)

KVE-KD: Key Visual Evidence-Guided Knowledge Distillation for Vision-Language Models

arXiv:2610.03842v1 Announce Type: new Abstract: Knowledge distillation is crucial for deploying vision-language models on resource-constrained devices. However, existing methods typically impose uniform supervision across visual tokens or rely on static token selection, which confuses task-relevant cues with background noise and degrades cross-modal reasoning. To address this limitation, we propos

10月6日 04:00
category.学术arXiv cs.LG (机器学习)

Learning Latent Protein Languages for Autoregressive Generation

arXiv:2610.03978v1 Announce Type: new Abstract: Autoregressive transformers remain comparatively weak for protein sequence and structure generation. We study the role of target representation: amino acid tokens encode residue identities without explicit contextual semantics, while backbone coordinates require a discrete representation in our framework. We introduce two learned latent protein langu

10月6日 04:00
category.学术arXiv cs.CV (计算机视觉)

Dynamic Time Step Prediction in Inverse Heat Dissipation for Blur-Like Image Restoration Tasks

arXiv:2610.03942v1 Announce Type: new Abstract: When using diffusion models to target image restoration problems, diffusion inversion is typically employed to retain relevant image information from the degraded images. Instead of inverting back to the initial time step (i.e., T), many methods invert to a pre-determined intermediate time step, in order to better preserve information from degraded s

10月6日 04:00
category.学术arXiv cs.LG (机器学习)

Probabilistic Algorithms for Ising Machines from Optimization to Generative AI

arXiv:2610.03972v1 Announce Type: new Abstract: Ising machines have emerged as promising hardware accelerators for intractable optimization and sampling problems, yet their practical impact increasingly hinges on the co-design of algorithms and hardware, where algorithmic demands shape new architectures and new hardware capabilities inspire entirely new algorithms. In this Review, we survey probab

10月6日 04:00
category.学术arXiv cs.LG (机器学习)

Synthesizing Physics Formulae with Transformers

arXiv:2610.03947v1 Announce Type: new Abstract: Finding a compact formula that fits a set of input-output pairs and predicts outputs on unseen inputs is a fundamental problem in science. Symbolic regression automates the search for such formulae: search-based methods explore the space of possible formulae directly, while transformers pre-trained on synthetic data produce formulae of comparable qua

10月6日 04:00
category.学术arXiv cs.LG (机器学习)

DePICT: Decision-Preserving Interface for Constrained Downstream Tasks

arXiv:2610.03945v1 Announce Type: new Abstract: A constrained optimization problem may involve a parameter in its objective and active constraints, yet the final decision may remain insensitive to small changes in that parameter. This raises a fundamental question: which inputs does a decision making system truly depend on? Building on this question, we introduce DePICT, a procedure for constructi

10月6日 04:00
category.学术arXiv cs.LG (机器学习)

COVER: Learning to Accept More in Selective Sleep Staging

arXiv:2610.03911v1 Announce Type: new Abstract: Traditional sleep-staging methods apply the same model to every EEG epoch. Such uniform deployment expends computation on epochs that a smaller model could handle reliably, motivating cascades in which a primary classifier accepts its reliable predictions and defers the remainder to a more capable model. In this paper, we study the first stage of suc

10月6日 04:00
category.学术arXiv cs.CL (自然语言处理)

LongSocialBench: Do Long-Context LLMs Understand Online Discussion Threads?

arXiv:2610.04118v1 Announce Type: new Abstract: Long-context LLMs can now ingest entire online discussion threads, but understanding their social discourse requires more than reading a long document: models must track parent-reply relations, turning points, scoped subtrees, cross-branch contrasts, and participant trajectories. To test this structure-aware social reasoning, we introduce LongSocialB

10月6日 04:00
category.学术arXiv cs.CL (自然语言处理)

Copying Before Suppression: What Drives a Below-Chance Dip During Language Model Training?

arXiv:2610.04119v1 Announce Type: new Abstract: Mechanistic interpretability usually studies fully trained models, yet the computations that drive a behaviour can change while the model is still learning the task. On the Indirect Object Identification task, a model should continue with the name mentioned once rather than the name mentioned twice. Pythia models pass through an early training window

10月6日 04:00
category.学术arXiv cs.CL (自然语言处理)

SYNLAT: Syntax-Aligned Text-Latent Compression for Chain-of-Thought Reasoning

arXiv:2610.03839v1 Announce Type: new Abstract: Long chain-of-thought (CoT) traces impose substantial output-token costs. Under constrained budgets, compression must preserve answer-critical information, making boundary placement central. Token-level and fixed-length boundaries can fragment coherent spans such as phrases, formulas, and local derivations, whereas step-level boundaries can bind cont

10月6日 04:00
category.学术arXiv cs.AI

Retrieval-Augmented Large Language Model Decision-Making for Autonomous Driving Guided by Chinese Philosophical Wisdom

arXiv:2610.03948v1 Announce Type: new Abstract: Autonomous driving decision systems must balance safety, efficiency, and social norms in complex traffic interactions. Philosophical and ethical considerations have received limited attention in existing autonomous driving decision-making approaches based on numerical optimization, sequence prediction, and large language models (LLMs). We propose Chi

10月6日 04:00
category.学术arXiv cs.CL (自然语言处理)

Periscope: Extending Frozen Language Models Beyond Their Context Window

arXiv:2610.04047v1 Announce Type: new Abstract: A language model reads long text in one quadratic forward pass, stops at the context window, and loses accuracy with length before reaching it. We ask whether the read can be factorized when deciding over a finite set: which document is relevant, which option is supported, which passage is the evidence. Periscope, a training-free inference method, ar

10月6日 04:00
category.学术arXiv cs.AI

MLLMs Fail to Refuse when Using Tools Agentically

arXiv:2610.03938v1 Announce Type: new Abstract: Agentic multimodal large language models (MLLMs) have recently pushed the frontier of visual reasoning by calling tools such as zooming and tagging. Despite the recent strong success of agentic MLLMs, this work uncovers a critical safety failure in the tool-use paradigm: agentic tool-using MLLMs become less capable of refusing harmful requests. Our e

10月6日 04:00
category.学术arXiv cs.AI

SGAnalog: An End-to-End Circuit Benchmark from Open-Source Silicon Tapeouts

arXiv:2610.03934v1 Announce Type: new Abstract: Existing analog integrated circuit design benchmarks make two questions hard to answer: whether a model has learned transferable circuit skills rather than recalled familiar examples, and whether its output works under defined process and test conditions. We introduce a benchmark built from human-designed, open-source circuits associated with Tiny Ta

10月6日 04:00
category.学术arXiv cs.CL (自然语言处理)

When Evidence Changes: Evaluating Memory Repair and Re-reading in Language-Model Agents

arXiv:2610.03902v1 Announce Type: new Abstract: When documents supporting an agent's derived facts are revoked or replaced, should it repair memory or re-read current evidence? We introduce an evidence-revision evaluation on medication- and problem-list tasks from public ICU records. Under revocation, replacement and control events, we compare full and source-filtered re-reading with caching, rebu

10月6日 04:00
category.学术arXiv cs.CL (自然语言处理)

BAIBAICHUCHU at the NTCIR-19 FinArg-3 Task: When Is Maximum Possible Profit Predictable from Investor Text?

arXiv:2610.03962v1 Announce Type: new Abstract: The BAIBAICHUCHU team participated in the Social Media Subtask of NTCIR-19 FinArg-3, ranking Chinese investor posts by Maximum Possible Profit (MPP). A three-track ensemble of lexical features, a FinArg-2-pre-finetuned MacBERT ranker, and an LLM judge reaches 0.734 in post-grouped development evaluation, but our best official run scores 0.517. All tw

10月6日 04:00
category.学术arXiv cs.CL (自然语言处理)

Representation-Aligned Auxiliary Supervision for Language Model Adaptation

arXiv:2610.04098v1 Announce Type: new Abstract: Language models exhibit strong reasoning capabilities, yet adapting them to structured domains remains challenging and can yield inconsistent outcomes. We identify representation compatibility, the extent to which a model effectively processes a representation for a structured task, as a key factor in adaptation. We study this in chess, which provide

10月6日 04:00
category.学术arXiv cs.AI

Proxy Confidence: Auditing Black-Box LLM Agents with a Surrogate's Log-Probabilities

arXiv:2610.03894v1 Announce Type: new Abstract: A deployed LLM agent emits tool calls, queries, and code that can be silently wrong -- by the time the error surfaces, the action has run. Frontier chat APIs hide the model's token probabilities; the agent's stated confidence barely beats chance on the mistakes that matter; and resampling does not help, since frontier models are highly repetitive, re

10月6日 04:00
category.学术arXiv cs.AI

The Cost of a Hop: Benchmarking NLIP and A2A

arXiv:2610.04053v1 Announce Type: new Abstract: Autonomous agents built on Large Language Models (LLMs) need standardized protocols to interoperate across systems. Several now exist (A2A, MCP, ACP, ANP, NLIP), but the Natural Language Interaction Protocol (NLIP) has not appeared in any controlled performance study, and no work has measured where an agent protocol's latency is spent. We compare NLI

10月6日 04:00