AiAnyTool - Best AI Tools Directory and Artificial Intelligence Software Hub Logo
Loading theme toggle
Real-Time Coverage

AI News Today

Live

32030 stories from 30+ sources, refreshed continuously.

arXiv cs.AIResearch

Diversifying to Verify: When Task-Equivalent Programs Differ in Verifiability

arXiv:2607.09366v1 Announce Type: cross Abstract: Program verification is crucial for software correctness, but producing fully verified programs remains difficult in practice. This paper studies whether implementation structure affects automated verifiability when multiple generated programs are intended to satisfy the same task-level semantics. We present Diversify2Verify, a staged LLM-based pipeline for Why3 that infers representation-specific contracts, generates and tests diverse recursive

Read source article
arXiv cs.AIResearch

When Routes Run Out: Adversarial Co-Learning and Explainable Robustness in Quantum Repeater Networks

arXiv:2607.09378v1 Announce Type: cross Abstract: We study an adversarial bandit problem for entanglement-based quantum-network routing over a modest graph corpus. Alice selects an end-to-end repeater route for an Ekert-91 protocol (E91) representing her move, while Eve selects an attack surface, either edge intercept--resend or repeater memory degradation. Payoffs are drawn from cached SeQUeNCe-simulated E91 transcripts, and Alice accepts a turn when the finite-sample statistic violates the Cla

Read source article
arXiv cs.AIResearch

STEEL: Sparsity-Aware Fused Attention for Energy-Efficient Long-Sequence Inference on AMD's XDNA NPU

arXiv:2607.09385v1 Announce Type: cross Abstract: The growing adoption of large language model-based agents within operating system workflows has increased the importance of energy-efficient inference on laptop-class systems-on-chip (SoCs). While cloud offloading remains common, it introduces reliability and privacy concerns that are particularly problematic for agentic workloads. Recent laptop SoCs, therefore, incorporate neural processing engines (NPUs) optimized for energy efficiency; however

Read source article
arXiv cs.CVResearch

ProSGNeRF: Progressive Dynamic Neural Scene Graph with Frequency Modulated Foundation Model in Urban Scenes

arXiv:2312.09076v4 Announce Type: replace Abstract: Implicit neural representation has demonstrated promising results in 3D reconstruction on various scenes. However, existing approaches either struggle to model fast-moving objects or are incapable of handling large-scale camera ego-motions in urban environments. This leads to low-quality synthesized views of the large-scale urban scenes. In this paper, we aim to jointly solve the problems caused by large-scale scenes and fast-moving vehicles, w

Read source article
arXiv cs.CVResearch

Zero-shot 3D General Obstacle Detection via Multimodal Foundation Models and Geometry

arXiv:2408.12322v2 Announce Type: replace Abstract: Detecting general obstacles is critical for autonomous driving, especially in long-tail scenarios with rare or unseen objects. Existing methods rely on supervision or predefined categories, limiting generalization. We propose a training-free approach that combines multimodal foundation models with geometric reasoning for 3D obstacle detection. Our key idea is to detect obstacles as deviations from the road surface, segmented in 2D and localized

Read source article
arXiv cs.CVResearch

Prototypical Few-Shot Medical Image Semantic Segmentation with Background Fusion

arXiv:2412.02983v2 Announce Type: replace Abstract: Few-shot Semantic Segmentation (FSS) aims to adapt a pre-trained model to new classes with as few as a single labeled training sample per class. The existing prototypical work used in natural image scenarios biasedly focus on capturing foreground's discrimination while employing a simplistic representation for background, grounded on the inherent observation separation between foreground and background. However, a frequency spectrum entropy ana

Read source article
arXiv cs.CVResearch

On Motion Blur and Deblurring in Visual Place Recognition

arXiv:2412.07751v2 Announce Type: replace Abstract: Visual Place Recognition (VPR) in mobile robotics enables robots to localize themselves by recognizing previously visited locations using visual data. While the reliability of VPR methods has been extensively studied under conditions such as changes in illumination, season, weather and viewpoint, the impact of motion blur is relatively unexplored despite its relevance not only in rapid motion scenarios but also in low-light conditions where lon

Read source article
arXiv cs.CVResearch

VerteNet -- A Multi-Context Hybrid CNN Transformer for Accurate Vertebral Landmark Localization in Lateral Spine DXA Images

arXiv:2502.02097v4 Announce Type: replace Abstract: Vertebral Landmarks Localization in Dual-Energy X-ray Absorptiometry based Lateral Spine Imaging plays a critical role in evaluating spinal alignment, Vertebral Fracture Assessment, and facilitating intervertebral guide placement for Abdominal Aortic Calcification quantification. While lateral spine DXA scans offer advantages such as reduced cost and lower radiation exposure, its analysis remains challenging due to a low signal-to-noise ratio a

Read source article
arXiv cs.CVResearch

Looking Locally: Object-Centric Vision Transformers as Foundation Models for Efficient Segmentation

arXiv:2502.02763v3 Announce Type: replace Abstract: Current state-of-the-art segmentation models encode entire images before focusing on specific objects. This wastes computational resources. We introduce FLIP (Fovea-Like Input Patching), a parameter-efficient vision model that realizes object segmentation through biologically-inspired top-down attention. FLIP selectively samples multi-resolution patches centered on objects of interest from the input. As a result, it allocates high-resolution pr

Read source article
arXiv cs.CVResearch

AffordanceSAM: Segment Anything Once More in Affordance Grounding

arXiv:2504.15650v3 Announce Type: replace Abstract: Building a generalized affordance grounding model to identify actionable regions on objects is vital for real-world applications. Existing methods to train the model can be divided into weakly and fully supervised ways. However, the former method requires a complex training framework design and can not infer new actions without an auxiliary prior. While the latter often struggle with limited annotated data and components trained from scratch de

Read source article
arXiv cs.CVResearch

Memory-SAM: Human-Prompt-Free Tongue Segmentation via Retrieval-to-Prompt

arXiv:2510.15849v3 Announce Type: replace Abstract: Accurate tongue segmentation is crucial for reliable TCM analysis. Supervised models require large annotated datasets, while SAM-family models remain prompt-driven. We present Memory-SAM, a training-free, human-prompt-free pipeline that automatically generates effective prompts from a small memory of prior cases via dense DINOv3 features and FAISS retrieval. Given a query image, mask-constrained correspondences to the retrieved exemplar are dis

Read source article
arXiv cs.CVResearch

Self-transcendence: Is External Feature Guidance Indispensable for Accelerating Diffusion Transformer Training?

arXiv:2601.07773v3 Announce Type: replace Abstract: Recent works such as REPA have shown that guiding diffusion models with external semantic features (e.g., DINO) can significantly accelerate the training of diffusion transformers (DiTs). However, the use of pretrained external features as guidance signals introduces additional dependencies. We argue that DiTs actually have the power to guide the training of themselves, and propose SelfTranscendence, an effective method that achieves fast conve

Read source article
arXiv cs.CVResearch

Robust Self-Supervised Cross-Modal Super-Resolution against Real-World Misaligned Observations

arXiv:2602.18822v3 Announce Type: replace Abstract: Cross-modal super-resolution (SR) on real-world misaligned data is challenging, as only unlabeled low-resolution (LR) source and high-resolution (HR) guide images with complex spatial misalignment are available. Previous methods either rely on simulated training data or adopt suboptimal alignment strategies that overlook cross-modal dependencies, limiting their practical performance. To address these issues, we propose RobSelf, a self-supervise

Read source article
arXiv cs.CVResearch

LUMOS: Latent Universal Medical Priors for Segmentation

arXiv:2603.01115v2 Announce Type: replace Abstract: General vision foundation models (VFMs) have been primarily developed on natural images, and their utility for medical image segmentation is therefore often considered to depend on costly adaptation or domain-specific fine-tuning. In this paper, we revisit this assumption from a different perspective: rather than requiring VFM segmentors to relearn visual regularities, we investigate whether the low-level visual priors necessary for anatomical

Read source article
arXiv cs.CVResearch

Any to Full: Prompting Depth Anything for Depth Completion in One Stage

arXiv:2603.05711v2 Announce Type: replace Abstract: Accurate, dense depth estimation is crucial for robotic perception, but commodity sensors often yield sparse or incomplete measurements due to hardware limitations. Existing RGBD-fused depth completion methods learn priors jointly conditioned on training RGB distribution and specific depth patterns, limiting domain generalization and robustness to various depth patterns. Recent efforts leverage monocular depth estimation (MDE) models to introdu

Read source article
arXiv cs.CVResearch

Predictive Photometric Uncertainty in Gaussian Splatting for Novel View Synthesis

arXiv:2603.22786v2 Announce Type: replace Abstract: Recent advances in 3D Gaussian Splatting have enabled impressive photorealistic novel view synthesis. However, to transition from a pure rendering engine to a reliable spatial map for autonomous agents and safety-critical applications, knowing where the representation is uncertain is as important as the rendering fidelity itself. We bridge this critical gap by introducing a lightweight, plug-and-play framework for pixel-wise, view-dependent pre

Read source article
arXiv cs.CVResearch

RehearsalNeRF: Decoupling Intrinsic Neural Fields of Dynamic Illuminations for Scene Editing

arXiv:2603.27948v3 Announce Type: replace Abstract: Although there has been significant progress in neural radiance fields, an issue on dynamic illumination changes still remains unsolved. Different from relevant works that parameterize time-variant/-invariant components in scenes, subjects' radiance is highly entangled with their own emitted radiance and lighting colors in spatio-temporal domain. In this paper, we present a new effective method to learn disentangled neural fields under the seve

Read source article
arXiv cs.CVResearch

Diagnosing Long-Video Quantitative Reasoning in Multimodal LLMs via Enumeration and Counting

arXiv:2603.29943v2 Announce Type: replace Abstract: Final-answer video QA can show whether a model predicts the right number, but not which instances it counted, when the supporting evidence occurs, or why it failed. We diagnose long-video quantitative reasoning in multimodal large language models (MLLMs) through three coupled abilities: enumerating query-relevant instances, temporally grounding supporting evidence, and aggregating the evidence into counts. To support this analysis, we build EC-

Read source article
arXiv cs.CVResearch

Render-in-the-Loop: Vector Graphics Generation via Visual Self-Feedback

arXiv:2604.20730v3 Announce Type: replace Abstract: Multimodal Large Language Models (MLLMs) have shown promising capabilities in generating Scalable Vector Graphics (SVG) via direct code synthesis. However, existing paradigms typically adopt an open-loop "blind drawing" approach, where models generate symbolic code sequences without perceiving intermediate visual outcomes. This methodology severely underutilizes the powerful visual priors embedded in MLLMs vision encoders, treating SVG generati

Read source article
arXiv cs.CVResearch

AS-Bridge: A Bidirectional Generative Framework Bridging Next-Generation Astronomical Surveys

arXiv:2603.11928v2 Announce Type: replace-cross Abstract: The upcoming decade of observational cosmology will be shaped by large sky surveys, such as the ground-based LSST at the Vera C. Rubin Observatory and the space-based Euclid mission. While they promise an unprecedented view of the Universe across depth, resolution, and wavelength, their differences in observational modality, sky coverage, point-spread function, and scanning cadence make joint analysis beneficial, but also challenging. To

Read source article
arXiv cs.CVResearch

Policy-based Tuning of Autoregressive Image Models with Instance- and Distribution-Level Rewards

arXiv:2603.23086v2 Announce Type: replace-cross Abstract: Autoregressive (AR) models are highly effective for image generation, yet their standard maximum-likelihood estimation training lacks direct optimization for sample quality and diversity. While reinforcement learning (RL) has been used to align diffusion models, these methods typically suffer from output diversity collapse. Similarly, concurrent RL methods for AR models rely strictly on instance-level rewards, often trading off distributi

Read source article
arXiv cs.LGResearch

Learning More from Less: Reinforcement Learning from Hindsight

arXiv:2607.09042v1 Announce Type: new Abstract: Reinforcement learning (RL) is increasingly used to post-train vision-language-action (VLA) models, but every update consumes robot rollouts that are slow and costly to collect, making sample efficiency a central concern. Manipulation tasks typically provide only sparse rewards, so a weak policy fails almost every rollout early in training and has little to learn from, even when those failures execute coherent behavior. Such a failure, however, is

Read source article
arXiv cs.LGResearch

CoCoT-EEG: Contrastive-Pretrained Multiscale Convolutional Transformer for EEG Decoding

arXiv:2607.09543v1 Announce Type: new Abstract: Self-supervised pretrained foundation models (FM) have shown early promise for non-invasive electroencephalogram (EEG) decoding applications. Many recent large-scale models converged on the approach of tokenizing raw EEG followed by masked reconstruction pretraining. However, this recipe has been shown to be suboptimal for data, like EEG, with high noise amplitude and information confined to limited dimensions such as narrow frequency bands. Buildi

Read source article
arXiv cs.LGResearch

Clean2FX: Label-conditioned modeling for clean-to-effect guitar audio transformations

arXiv:2607.08863v1 Announce Type: cross Abstract: We present Clean2FX, a study and demo of label-conditioned clean-to-effect transformation for electric guitar audio. Given a clean guitar input and a target effect label, the task is to synthesize the corresponding effected signal while preserving the musical content. Training and evaluation pairs are constructed from EGFxSet real, single tone recordings by assembling matched clean/effected chords, melodies, and mixed timelines. This allows for c

Read source article
arXiv cs.LGResearch

FlowDAgger: Human-in-the-Loop Adaptation of Generative Robot Policies in Latent Space

arXiv:2607.08877v1 Announce Type: cross Abstract: Pretrained generative robot policies based on flow matching and diffusion have achieved impressive results across a wide range of manipulation tasks. Yet real-world deployments routinely expose failure modes outside the pretraining distribution. Closing these gaps typically requires large-scale data collection or online reinforcement learning on physical hardware, which is impractical for rapid and safe adaptation. We present FlowDAgger, a sample

Read source article
arXiv cs.LGResearch

RaMark: Radioactive Watermarking for Generated Tabular Data

arXiv:2607.09000v1 Announce Type: cross Abstract: Recent advances in generative modeling have made generated tabular data a practical solution for privacy-sensitive data sharing, where watermarking enables ownership verification. However, existing watermarking methods fundamentally fail under retraining attacks, in which an adversary retrains a generative model on a watermarked dataset and regenerates high-utility data that no longer carries the watermark. We address this challenge by introducin

Read source article
arXiv cs.LGResearch

Quantum-Enhanced Synthetic Data Generation Using Quantum Circuit Born Machines for Imbalanced Tabular Learning

arXiv:2607.09113v1 Announce Type: cross Abstract: Data scarcity and class imbalance are persistent challenges in machine learning that degrade model generalization and introduce predictive bias. We present a hybrid quantum-classical framework for synthetic data generation using a Quantum Circuit Born Machine (QCBM) to address these limitations. The proposed approach exploits quantum mechanical properties -- superposition and entanglement -- within a parameterized variational quantum circuit to m

Read source article
arXiv cs.LGResearch

Control Laguerre Tessellation: Semi-discrete Optimal Transport Over Control Systems

arXiv:2607.09139v1 Announce Type: cross Abstract: We study the optimal transport of optimally controlled agents from a compactly supported absolutely continuous source to a discrete target measure. The ground cost for the transport is induced by the optimal cost of the agents' motion. When this ground cost satisfies the twist condition, the optimal transport map is given almost everywhere in terms of a Laguerre tessellation of the state space. We refer to this control-theoretic generalization of

Read source article
arXiv cs.LGResearch

GenVid2Robot: From Video Generation to Robot Manipulation via Rigid-Geometric Consistency

arXiv:2607.09191v1 Announce Type: cross Abstract: Generated videos provide useful visual motion priors for robot manipulation, but their visual plausibility does not imply physical executability. A generated video usually lacks metric geometry, grasp grounding, robot kinematic feasibility, and execution-time feedback, which makes direct trajectory replay unreliable in real-world manipulation. This paper presents GenVid2Robot, a rigid-geometric consistency framework that converts generated video

Read source article
arXiv cs.LGResearch

Influence Diagnostics in High-dimensional M-estimation: Precise Asymptotics

arXiv:2607.09250v1 Announce Type: cross Abstract: The impact of a given training point on a statistical model is classically measured through its leave-one-out influence, which quantifies the effect of its removal from the training set on the model accuracy. While the statistics of leave-one-out influences are well understood in the low-dimensional, large sample limit $n\to \infty, d=O(1)$, they become more intricate in high dimensions, as the influence of a given sample develops non-trivial dep

Read source article
arXiv cs.LGResearch

Entropy-Constrained Machine Learning with Residual Data Augmentation for Modeling Chemical Kinetics

arXiv:2607.09582v1 Announce Type: cross Abstract: We present a physics-constrained machine learning framework for accelerating the direct numerical simulation (DNS) of turbulent reacting flows. The model replaces the direct evaluation of detailed chemical source terms with a surrogate that predicts reaction rates from a reduced thermochemical state. To improve physical consistency, the second law of thermodynamics is incorporated as a training constraint by enforcing non-negative entropy generat

Read source article
arXiv cs.LGResearch

LLM for EDA in Front-End Design: Challenges and Opportunities

arXiv:2607.09616v1 Announce Type: cross Abstract: As chip complexity increases and time-to-market pressures grow, front-end design has become a critical bottleneck in chip development. Recently, Large Language Models (LLMs) have shown great potential in Electronic Design Automation (EDA). Beyond specification understanding, LLMs show the potential to serve as a unified intelligent interface for hardware description language (HDL) generation, testbench construction, and design space exploration.

Read source article
arXiv cs.LGResearch

LDPKiT: Superimposing Remote Queries for Privacy-Preserving Distillation

arXiv:2405.16361v4 Announce Type: replace Abstract: To protect privacy in regulated domains such as healthcare and finance, model owners may allow only remote API access while keeping both the training data and model parameters private. However, model users performing inference on such remotely hosted models may be required to transmit potentially sensitive inputs, raising privacy concerns. In this work, we present LDPKiT, a framework for non-adversarial, privacy-preserving model distillation th

Read source article
arXiv cs.LGResearch

Regression-aware Continual Learning for Android Malware Detection

arXiv:2507.18313v2 Announce Type: replace Abstract: Malware evolves rapidly, forcing machine learning-based detectors to be continuously updated. With antivirus vendors processing hundreds of thousands of new samples daily, datasets can grow to billions of examples, making full retraining impractical. Continual learning (CL) has emerged as a scalable alternative, enabling incremental updates without full data access while mitigating catastrophic forgetting. In this work, we analyze a critical ye

Read source article
arXiv cs.LGResearch

Scalable Varied-Density Clustering via Graph Propagation

arXiv:2508.02989v2 Announce Type: replace Abstract: We propose a novel perspective on varied-density clustering for high-dimensional data by framing it as a label propagation process in neighborhood graphs that adapt to local density variations. Our method formally connects density-based clustering with graph connectivity, enabling the use of efficient graph propagation techniques developed in network science. To ensure scalability, we introduce a density-aware neighborhood propagation algorithm

Read source article
arXiv cs.LGResearch

Hair-Trigger Alignment: Black-Box Evaluation Cannot Guarantee Post-Update Alignment

arXiv:2601.22313v2 Announce Type: replace Abstract: Large Language Models (LLMs) are rarely static and are frequently updated in practice. A growing body of alignment research has shown that models initially deemed ``aligned'' can exhibit misaligned behavior after fine-tuning. These works typically assume that the initial model is aligned based on static black-box evaluation, i.e., the absence of undesired responses to a fixed set of queries. However, the limits of black-box evaluation for post-

Read source article
arXiv cs.LGResearch

Untied Ulysses: Memory-Efficient Context Parallelism via Headwise Chunking

arXiv:2602.21196v2 Announce Type: replace Abstract: Efficiently processing long sequences with Transformer models usually requires splitting the computations across accelerators via context parallelism. The dominant approaches in this family of methods, such as Ring Attention or DeepSpeed Ulysses, enable scaling over the context dimension but do not focus on memory efficiency, which limits the sequence lengths they can support. More advanced techniques, such as Fully Pipelined Distributed Transf

Read source article
arXiv cs.LGResearch

Learning Lineage-guided Geodesics with Finsler Geometry

arXiv:2603.16708v2 Announce Type: replace Abstract: Trajectory inference investigates how to interpolate paths between observed timepoints of dynamical systems, such as temporally resolved population distributions, with the goal of inferring trajectories at unseen times and better understanding system dynamics. Previous work has focused on continuous geometric priors, utilizing data-dependent spatial features to define a Riemannian metric. In many applications, there exists discrete, directed pr

Read source article
arXiv cs.LGResearch

Bridging the Gap Between Climate Science and Machine Learning in Climate Model Emulation

arXiv:2603.22320v2 Announce Type: replace Abstract: For decades, physics-based climate models have been used to provide insights for climate decision-making. Their application is, however, constrained by significant computational and technical demands. Machine learning (ML) emulators offer a way to reduce these high computational costs; yet, it remains challenging to use ML emulators effectively in climate research. In practice, climate scientists often bypass emulators altogether, and machine l

Read source article
arXiv cs.LGResearch

Kronecker-Structured Nonparametric Spatiotemporal Point Processes

arXiv:2603.23746v2 Announce Type: replace Abstract: Events in spatiotemporal domains arise in numerous real-world applications, where uncovering event relationships and enabling accurate prediction are central challenges. Classical Poisson and Hawkes processes rely on restrictive parametric assumptions that limit their ability to capture complex interaction patterns, while recent neural point process models increase representational capacity but integrate event information in a black-box manner,

Read source article