Read Less. Know More.

全部资讯

754 条资讯

category.学术arXiv cs.CV (计算机视觉)

DTFormer: Text-Guided Semantic Alignment for RGB-D Segmentation

arXiv:2610.07014v1 Announce Type: new Abstract: RGB-D semantic segmentation has made notable progress by fusing RGB and Depth, yet mainstream models still learn features almost exclusively from pixel-level supervision, lacking direct high-level semantic constraints. This raises a central question-can external knowledge such as language priors inject stronger semantic discriminability into mainstre

10月7日 04:00
category.学术arXiv cs.CV (计算机视觉)

Anchor and Adapt: Asymmetric Prompt Adaptation for Few-Shot Industrial Anomaly Detection

arXiv:2610.07016v1 Announce Type: new Abstract: In few-shot industrial anomaly detection, the few normal target images provide no direct defect supervision, making anomaly prompts difficult to learn from these samples alone. Some vision-language methods therefore use manually specified descriptions to supply explicit anomaly semantics. However, constructing these descriptions requires product-spec

10月7日 04:00
category.学术arXiv cs.CV (计算机视觉)

Decomposition-Guided Curvelet Thresholding for Sharp-to-Soft CT Kernel Conversion

arXiv:2610.07067v1 Announce Type: new Abstract: Image denoising is a crucial task in image processing, focused on improving image quality by minimizing noise while maintaining essential structural elements. This study presents a hybrid denoising framework that combines several decomposition techniques, including empirical mode decomposition (EMD), variational mode decomposition (VMD), multichannel

10月7日 04:00
category.学术arXiv cs.AI

LatentQuant: Preserving the Policy-Facing Latent Contract under NVFP4 VAE Quantization

arXiv:2610.03959v1 Announce Type: new Abstract: Recent world action models (WAMs) reuse pretrained video VAEs whose encoder latents directly condition downstream action policies. Quantization must therefore preserve not only reconstruction fidelity but also the policy-facing latent contract expected by the frozen policy. Direct NVFP4 leaves W4A4 quantization error uncompensated, whereas joint quan

10月7日 04:00
category.学术arXiv cs.AI

Teaching Agents to Code Reliably

arXiv:2610.03984v1 Announce Type: new Abstract: Autonomous coding agents solve repository issues by reading code, running commands, editing files, and submitting patches. Extra inference-time compute yields gains only when it produces a useful repair and supplies reliable evidence for choosing one. Three behaviors decide both, and we argue they are teachable rather than byproducts of scale, so a p

10月7日 04:00
category.学术arXiv cs.CL (自然语言处理)

Identifying Introspection From the Inside

arXiv:2610.07186v1 Announce Type: new Abstract: Large language models make claims about themselves that are both consequential and increasingly difficult to verify from behavior alone. How can we distinguish plausible confabulations from genuine introspection? In this paper, we identify mechanistic signatures of faithful self-report in a controlled setting. Using low-rank adapters, we train models

10月7日 04:00
category.学术arXiv cs.AI

Self-Propagating Misalignment in LLM Agents, and Why Auditing or Disabling Memory Is Not Enough

arXiv:2610.04083v1 Announce Type: new Abstract: Memory poisoning attacks on LLM agents typically assume an external adversary who plants content in the agent's persistent memory to steer its behavior. We instead study, with no adversary involved, whether a misaligned agent can write a goal it cannot yet act on to persistent memory, so that a future aligned agent carries it out when the opportunity

10月7日 04:00
category.学术arXiv cs.LG (机器学习)

DART-ES: Difficulty-Aware Reweighting and Targeted Replay for Fine-Tuning LLMs with Evolution Strategies

arXiv:2610.06993v1 Announce Type: new Abstract: Evolution Strategies (ES) enable memory efficient full parameter fine-tuning of large language models (LLMs) using only forward computation. However, standard ES uniformly averages rewards across problems and compresses problem level population feedback into a single scalar, making it difficult to capture how the learning value of each problem change

10月7日 04:00
category.学术arXiv cs.LG (机器学习)

Learning to Remember: Distilling Memory Retention for Compact Recurrent Neural Networks

arXiv:2610.06942v1 Announce Type: new Abstract: Deep learning models, particularly recurrent neural networks and their variants, such as long short-term memory, have significantly advanced time series analysis. These models capture complex, sequential patterns in time series, enabling real-time assessments. However, their high computational complexity and large model sizes pose challenges for depl

10月7日 04:00
category.学术arXiv cs.CL (自然语言处理)

CroissantMiner: Automated Extraction and Validation of Croissant Metadata for ML Datasets

arXiv:2610.07132v1 Announce Type: new Abstract: Croissant has emerged as a standard for machine-readable dataset metadata, yet populating its fields remains labor-intensive and requires careful reading of accompanying dataset documentation. We present the first benchmark enabling end-to-end evaluation of metadata extraction aligned with a community-standard schema. The benchmark comprises 602 pape

10月7日 04:00
category.学术arXiv cs.CV (计算机视觉)

MoonGS: High-quality Representation of the Lunar Surface via Gaussian Splatting Using Robust Depth Features from Image Pairs

arXiv:2610.07110v1 Announce Type: new Abstract: High-quality 3D reconstruction of lunar terrain from sparse rover images is indispensable for autonomous lunar exploration, but remains challenging because viewpoint overlap is insufficient, surface textures are weak, and data volume is limited. We propose MoonGS, the first feed-forward 3D Gaussian Splatting framework tailored to lunar scenes. Given

10月7日 04:00
category.学术arXiv cs.LG (机器学习)

Mask-Guided KV Cache Eviction in Block Diffusion Language Models

arXiv:2610.06996v1 Announce Type: new Abstract: Block diffusion language models keep a large key-value (KV) cache throughout generation and attend to it at every denoising step, limiting both memory capacity and generation speed. Reducing these costs requires deciding which past tokens to use for denoising the current block (selection) and which to keep in memory for future blocks (eviction). We p

10月7日 04:00
category.学术arXiv cs.LG (机器学习)

QiYao-I: A Manifold Based Foundation Model for Irregular Multivariate Time Series Forecasting

arXiv:2610.06936v1 Announce Type: new Abstract: Irregular multivariate time series forecasting is a challenging yet important problem in real-world applications, where observations are often irregularly sampled and asynchronously recorded across variables. Existing time series foundation models are mostly built on regularly sampled sequences, making them difficult to generalize to irregular time i

10月7日 04:00
category.学术arXiv cs.LG (机器学习)

AttSVD:Prompt-Adaptive Low-Rank KV Cache Compression via Attention-Guided SVD

arXiv:2610.06927v1 Announce Type: new Abstract: The key-value (KV) cache of autoregressive transformers grows linearly with context length and dominates memory at long context. Most training-free remedies evict low-importance tokens, an irreversible choice along the sequence axis. We instead keep every token and store it more cheaply along the "feature" axis. We therefore propose AttSVD, a new "in

10月7日 04:00
category.学术arXiv cs.CV (计算机视觉)

A Data-Centric Review of Plant Disease Datasets: Taxonomy, Critical Analysis, Environmental Variability, and Implications for Precision Agriculture

arXiv:2610.07087v1 Announce Type: new Abstract: Despite rapid advances in artificial intelligence, reliable real-world plant disease detection remains a persistent challenge. Visual and deep learning approaches have shown promising results, but their deployment under field conditions remains limited. A key bottleneck is the reliance on laboratory-generated datasets that lack environmental diversit

10月7日 04:00
category.学术arXiv cs.AI

Beyond the Parameter Monolith: Reconstructive Memories, Executable Skills, and Residual Assembly for Language Models

arXiv:2610.04012v1 Announce Type: new Abstract: Language-model systems can separate contextual computation, persistent storage, and exact execution instead of updating all capabilities through one shared parameter system. We investigate FEM-ASM, a finite-element-method-inspired organization in which independently constructed document states and deterministic executable skills contribute typed prop

10月7日 04:00
category.学术arXiv cs.AI

SkillScriptBench: Benchmarking Self-Evolution of Executable Agent Skill Packages Beyond Markdown

arXiv:2610.04008v1 Announce Type: new Abstract: Executable Agent Skills combine natural-language instructions and scripts into reusable packages for LLM agents, and revising them requires fixing errors without breaking correct behavior. Existing benchmarks do not systematically distinguish documentation repair, script repair, and preservation when evaluating skill self-evolution. We introduce Skil

10月7日 04:00
category.学术arXiv cs.LG (机器学习)

Near-Optimal Sample Complexity for Recursive Entropic Risk Reinforcement Learning with a Generative Model

arXiv:2610.06931v1 Announce Type: new Abstract: In this paper, we study the sample complexities of value and policy learning in finite discounted Markov decision processes (MDPs) under recursive entropic risk preferences with risk parameter \(\beta\neq 0\), assuming access to a generative model of the MDP. We provide a refined analysis of model-based risk-sensitive Q-value iteration (MB-RS-QVI), a

10月7日 04:00
category.学术arXiv cs.AI

Towards Safer Autonomous Driving in an Open World: A Dual-Process Approach

arXiv:2610.04088v1 Announce Type: new Abstract: Before autonomous driving systems can be deployed on public roads, it is vital that these systems comply with safety standards, traffic rules, and social norms. Although neural networks trained on large amounts of driving data perform well in routine driving tasks, these models often struggle in novel situations that are not well-represented in the d

10月7日 04:00
category.学术arXiv cs.AI

Training Numerical Intelligence via Auto-Diagnosis and Skill Discovery

arXiv:2610.03872v1 Announce Type: new Abstract: AI agents are becoming increasingly capable of generating scientific code, but generating code is not the same as improving the algorithms behind it. For numerical solvers, execution feedback can expose poor performance, but rarely reveals its underlying cause and how to address it. We introduce Auto-Diagnosis and Skill Discovery (ADSD), a framework

10月7日 04:00
category.学术arXiv cs.CV (计算机视觉)

R2RI: A Multi-View Event and RGB Dataset for Robot-to-Robot Interaction

arXiv:2610.07117v1 Announce Type: new Abstract: Understanding and modeling interactions between autonomous agents is a fundamental challenge in robotics, with broad implications for collaborative systems, social robotics, and human-robot coexistence. Although the study of robot interactions has emerged as a compelling research direction, progress has been severely hampered by the absence of large-

10月7日 04:00
category.学术arXiv cs.CV (计算机视觉)

On Color Alignment in VAE Latent Spaces and Its Applications

arXiv:2610.07072v1 Announce Type: new Abstract: Variational autoencoders (VAEs) are a key part of modern text-to-image models, which generate images within their latent space. VAEs are known to disentangle the main factors of variation in the data, and color is known to be one of the most structured of these in natural images: decorrelating it yields one luminance axis and two opponent-color axes.

10月7日 04:00
category.学术arXiv cs.AI

Exploration-Preserving Policy Optimization

arXiv:2610.04011v1 Announce Type: new Abstract: Reinforcement learning with verifiable rewards improves reasoning, while the allocation of learning signal shapes which solutions remain accessible under repeated sampling. Group-relative objectives assign equal advantages to equally rewarded responses, making aggregate credit proportional to sampled mode frequency. We introduce Exploration-Preservin

10月7日 04:00
category.学术arXiv cs.CV (计算机视觉)

Graph-Based Recognition of Simulated Train-Driver States From Facial and Upper-Body Keypoints

arXiv:2610.07083v1 Announce Type: new Abstract: Driver fatigue poses a significant challenge to railway safety, with traditional systems like the dead-man switch offering limited and basic alertness checks. This study presents a vision-based monitoring system that relies solely on a single front-facing RGB camera and a graph neural network to classify simulated train-driver states into alert, not-

10月7日 04:00
category.学术arXiv cs.CV (计算机视觉)

CALR: Continuous Anchored Latent Reasoning via Render-of-Thought Compression

arXiv:2610.07175v1 Announce Type: new Abstract: Visual latent reasoning compresses rendered derivations into compact intermediate states, reducing textual reasoning overhead. Existing approaches differ in how they represent these states: continuous methods avoid vocabulary constraints, whereas discrete methods improve accuracy through quantization into a finite codebook. Our analysis of representa

10月7日 04:00
category.学术arXiv cs.CV (计算机视觉)

A BEMD-Based Quaternion Filtering Approach Sharp-to-Soft Kernel CT Image Conversion

arXiv:2610.07071v1 Announce Type: new Abstract: The quality of computed tomography (CT) images is significantly affected by the selection of reconstruction kernels: sharp kernels improve spatial resolution but increase noise, whereas soft kernels diminish noise at the expense of edge clarity. This study presents an innovative enhancement framework utilising Bidimensional Empirical Mode Decompositi

10月7日 04:00
category.学术arXiv cs.AI

A Quantitative Analysis of Graph Representation Strategies for Cyber Attack Detection

arXiv:2610.04019v1 Announce Type: new Abstract: Graph based cyber attack detection studies employ various graph construction and representation strategies across different cybersecurity application domains. This diversity motivates a quantitative examination of how representation strategies are distributed across these application domains. This study presents a quantitative analysis of 37 original

10月7日 04:00
category.学术arXiv cs.CL (自然语言处理)

WavePrune: One period is often enough for RoPE

arXiv:2610.06963v1 Announce Type: new Abstract: Rotary Position Embedding (RoPE) encodes token positions by rotating each two-dimensional channel of the query and key vectors at a channel-specific frequency, making the attention logits invariant to a common shift of positions. However, this rotation is periodic, and it leads to position aliasing where relative positions separated by a full rotatio

10月7日 04:00
category.学术arXiv cs.CL (自然语言处理)

Verdicts Without Annotated Evidence: Rejection Sampling or Label-Only Post-Training for Evidence Recovery?

arXiv:2610.06962v1 Announce Type: new Abstract: In many review workflows the verdict is the only thing retained. The passages behind it are not marked, because that annotation costs far more than recording the decision. We measure how much of that evidence a small language model can recover when it is post-trained on the verdicts alone, with no human evidence labels at any stage. On ContractNLI th

10月7日 04:00
category.学术arXiv cs.AI

Reinforcement Learning with Comparative Evidence for Social Intelligence

arXiv:2610.04072v1 Announce Type: new Abstract: Developing socially intelligent AI remains heavily dependent on human-annotated data, limiting the scale and breadth of social understanding models can acquire. Methods that derive training signals from unlabeled data offer a path beyond this dependence, but social predictions lack the verification oracles available in mathematics and coding. Moreove

10月7日 04:00
category.学术arXiv cs.LG (机器学习)

STOCK-JEPA: Prior-Anchored Latent Revision Representation Learning in Equity Markets

arXiv:2610.07006v1 Announce Type: new Abstract: Learning effective representations helps characterize the structure and dynamics of equity markets from financial data with a low signal-to-noise ratio. Black-box deep models can capture complex patterns but may overfit sample noise and lack explicit economic structure. Meanwhile, classic linear financial models provide interpretable references, but

10月7日 04:00
category.学术arXiv cs.LG (机器学习)

Should We Skip Diffusion?

arXiv:2610.07002v1 Announce Type: new Abstract: Diffusion models learn semantic representations while generating images. In the Decoupled Diffusion Transformer (DDT), a condition encoder provides features that guide a velocity decoder in denoising. To enable effective denoising at all noise levels, these features must capture both high-level abstract structures and low-level details. However, skip

10月7日 04:00
category.学术arXiv cs.LG (机器学习)

The Premise Is the Problem: Exchangeability Failure in Self-Monitored Test-Time Adaptation

arXiv:2610.07038v1 Announce Type: new Abstract: Modern forecasting models are often updated after deployment so they can respond to changing data. These updates can also make predictions worse, so practical systems need a reliable monitor that can detect harmful changes and trigger protection. A natural design is to monitor the same prediction errors that guide the updates. This paper asks whether

10月7日 04:00
category.学术arXiv cs.AI

Auditing Pairwise Equivalence Judgments: Self-Critique Effects and Diversity Measurement in Multi-Agent Hypothesis Generation

arXiv:2610.04133v1 Announce Type: new Abstract: Multi-agent systems built on large language models (LLMs) are increasingly applied to scientific discovery and hypothesis generation. Both the effect of refinement and the diversity of the delivered set are hard to interpret before experimental ground truth exists, and both are typically reported by deciding whether pairs of generated hypotheses desc

10月7日 04:00
category.学术arXiv cs.CL (自然语言处理)

Does Steering Break Your Model? A Multi-Dimensional Evaluation Suite for LLM Steering Methods

arXiv:2610.07722v1 Announce Type: new Abstract: Activation steering provides a lightweight and flexible way to control large language model (LLM) behavior. However, effective steering requires more than inducing the intended behavior: it should also limit unintended changes and remain robust across inputs and training data. Existing evaluations cover these dimensions only in fragments. As a result

10月7日 04:00
category.学术arXiv cs.CL (自然语言处理)

Stabilizing language models under continual learning via condition-anchored distillation

arXiv:2610.06940v1 Announce Type: new Abstract: Continual adaptation of language models can change their output distribution on prompts learned earlier, while retaining every old prompt-answer pair may be undesirable or impossible. We study condition-anchored generative distillation (CAGD): retain a small set of old prompts, use a frozen previous model to reconstruct completions and generation sta

10月7日 04:00
category.学术arXiv cs.LG (机器学习)

Identifiable World Models from Pretrained Diffusion Representations

arXiv:2610.07028v1 Announce Type: new Abstract: Diffusion-based world models can generate and predict trajectories in high-dimensional dynamical systems, but predictive accuracy does not imply that their latent coordinates recover the underlying state variables or causal interactions. We ask whether a frozen pretrained diffusion model can be equipped with identifiable coordinates without retrainin

10月7日 04:00
category.学术arXiv cs.LG (机器学习)

TRIAGE: Direction-Aware Mismatch Stabilization of Native NVFP4 Reinforcement Learning

arXiv:2610.07043v1 Announce Type: new Abstract: Low-precision execution can substantially accelerate reinforcement learning (RL) for large language models, but discrepancies between learner and sampler execution can destabilize policy optimization. In this paper, we characterize the interaction between mismatch and the policy-gradient direction, distinguishing locally amplifying from contracting u

10月7日 04:00
category.学术arXiv cs.LG (机器学习)

Shaping the Wind: Nested Potentials for Kinematically Admissible Urban Wind Prediction

arXiv:2610.07033v1 Announce Type: new Abstract: Predicting transient urban winds is fundamental to understanding urban microclimates and designing climate-resilient cities. Building-resolving large-eddy simulation produces detailed incompressible urban wind fields at substantial computational cost for each layout. Neural surrogates offer a faster alternative by learning to predict the evolution of

10月7日 04:00
category.学术arXiv cs.CV (计算机视觉)

Localize Any Object in X-Ray Security Scans without Human Annotation

arXiv:2610.07326v1 Announce Type: new Abstract: Universal object localization in X-ray security inspection is critical for automated threat detection in safety-critical venues. However, unlike everyday RGB images that dominate web-scale visual data, X-ray scans exhibit distinct color patterns, ambiguous boundaries, and compositional structures caused by volumetric superposition. These gaps hinder

10月7日 04:00
category.学术arXiv cs.LG (机器学习)

Event-Driven ML Pipeline Orchestration for Manufacturing: An AWS Industry Experience

arXiv:2610.06890v1 Announce Type: new Abstract: We present an industry experience report on three years of operating an event-driven cloud infrastructure for continuous machine learning training in automotive manufacturing. Our system orchestrates GPU-accelerated training of product-specialized model pairs, a physics prediction model and a reinforcement-learning control policy, across multiple pla

10月7日 04:00
category.学术arXiv cs.CV (计算机视觉)

Compositional Concept Erasure in Text-to-Image Diffusion Models via Hierarchically Grounded Semantic Surgery

arXiv:2610.07337v1 Announce Type: new Abstract: Removing copyrighted, unsafe, or user-specified concepts from a deployed text-to-image diffusion model is now a practical requirement. Weight-editing methods can suppress fixed targets, but they require per-target retraining and modify the model checkpoint. Training-free methods, on the other hand, are deployment-friendly, but they suffer from text-s

10月7日 04:00
category.学术arXiv cs.CV (计算机视觉)

Deep Learning Based Illegal Bowling Action Detection

arXiv:2610.07223v1 Announce Type: new Abstract: Cricket, often referred to as the "gentleman's game," adheres to a strict rule set for both batsmen and bowlers, where each delivery can significantly impact the match outcome. Detecting illegal bowling actions is crucial for maintaining fair play, yet it remains challenging for umpires to monitor in real time. Existing sensor-based solutions have li

10月7日 04:00
category.学术arXiv cs.CL (自然语言处理)

TIDE 2.0: an open, model-agnostic engine for keyed de-identification of clinical notes

arXiv:2610.07224v1 Announce Type: new Abstract: Clinical notes capture most of what is documented about a patient's care, but they cannot be used for research until protected health information (PHI) is removed. De-identification is often treated as a detection problem. Detection alone is not sufficient: redaction strips clinical content along with identifiers, date blanking destroys the temporal

10月7日 04:00
category.学术arXiv cs.CL (自然语言处理)

Closing Ambient Clinical Documentation Gaps with Automated Provider Queries

arXiv:2610.07502v1 Announce Type: new Abstract: Provider queries are clarifying requests sent by clinical documentation specialists to physicians to close gaps in the clinical note and ensure accurate billing. Prior work automates note drafting, ICD-10 coding, and order extraction assuming a complete transcript, leaving these gaps unaddressed. We study whether an LLM can automate the query loop, t

10月7日 04:00
category.学术arXiv cs.CL (自然语言处理)

HouseholdBench: Evaluating Large Language Models as Predictors of Household Economic Behavior

arXiv:2610.07563v1 Announce Type: new Abstract: Large language models (LLMs) have the potential to meet a key goal in economics: a quantitative model of household decision making, across a variety of settings. Yet existing evaluations cover few surveys and outcomes, and do not study how households adjust to changing economic conditions. We introduce a new evaluation, HouseholdBench, which unites 6

10月7日 04:00
category.学术arXiv cs.CL (自然语言处理)

Two Vectors Replace In-Context Demos: Structured Task Adaptation via Embeddings

arXiv:2610.07572v1 Announce Type: new Abstract: In-context learning (ICL) adapts frozen large multimodal models (LMMs) to new tasks from a few demonstrations (demos), but re-encodes them at every query, where each demo image adds up to hundreds of visual tokens. Demo-free methods remove this cost with a compact task state. However, they add it at locations searched per task or at every decoder lay

10月7日 04:00
category.学术arXiv cs.CV (计算机视觉)

State-Aware Interaction MIL for Rare Joint Molecular Phenotype Prediction in Colorectal Cancer and Lung Adenocarcinoma

arXiv:2610.06991v1 Announce Type: new Abstract: Joint molecular phenotype prediction is complicated by small joint-positive populations and overlapping histological features across alternative molecular states. Existing computational pathology approaches typically predict biomarkers independently or formulate the joint-positive phenotype as a binary endpoint. Independent prediction does not model

10月7日 04:00