AiAnyTool - Best AI Tools Directory and Artificial Intelligence Software Hub Logo
Loading theme toggle
Real-Time Coverage

AI News Today

Live

31912 stories from 30+ sources, refreshed continuously.

arXiv cs.CVResearch

MicroCharNet: Less is More for License Plate Character Detection

arXiv:2607.11830v1 Announce Type: new Abstract: License plate character detection is a crucial component of intelligent transportation systems, where high accuracy and computational efficiency are required for real-time deployment. Although recent deep learning-based methods have substantially improved detection performance, many high-accuracy models rely on large-scale architectures that incur substantial computational overhead, limiting their applicability to resource-constrained devices. In t

Read source article
arXiv cs.CVResearch

Cycle-World: Mitigating Error Accumulation in Long-term Video World Models via Reverse-Prediction Cycle Consistency

arXiv:2607.11836v1 Announce Type: new Abstract: Autoregressive diffusion models have enabled high-quality video generation, yet their sequential nature inherently suffers from error accumulation. In long-horizon video synthesis, minor prediction deviations compound over time, inevitably leading to unconstrained generative drift, structural collapse, and severe visual degradation. To address this, we propose Cycle-World, a novel framework designed for stable and temporally consistent long-video g

Read source article
arXiv cs.CVResearch

HASTE: A Platform for Rapid Post-Disaster Building Damage Assessment

arXiv:2607.11838v1 Announce Type: new Abstract: When a large disaster strikes, responders need a map of which buildings are damaged within hours. The models that do well on public benchmarks assume matched before-and-after imagery and a training set drawn from similar past events, and neither is usually available for a new disaster in its first day. We present HASTE (High-speed Assessment and Satellite Tracking for Emergencies), a no-code web platform that lets analysts who are not machine learn

Read source article
arXiv cs.CVResearch

Beyond the Single Camera: Agentic Multi-View Reasoning in Sports Video Understanding

arXiv:2607.11844v1 Announce Type: new Abstract: Recent Multimodal Large Language Models (MLLMs) achieve strong performance on single-view video understanding benchmarks. However, sports videos involve dense occlusion, rapid motion, and complex interactions that are difficult to resolve from a single viewpoint. In practice, sports events are recorded from multiple camera angles, providing complementary evidence used by referees. Yet, no existing benchmark evaluates MLLMs on multi-view sports vide

Read source article
arXiv cs.CVResearch

Read It Back: Pretrained MLLMs Are Zero-Shot Reward Models for Text-to-Image Generation

arXiv:2607.11886v1 Announce Type: new Abstract: In this paper, we propose SpectraReward, a training-free reward function that turns pretrained MLLMs into off-the-shelf reward models for image-generation reinforcement learning. Instead of asking the MLLM to judge a generated image or answer decomposed verification questions, SpectraReward measures how well the original prompt can be recovered from the generated image through a single image-conditioned, teacher-forced forward pass. We use the aver

Read source article
arXiv cs.CVResearch

Calibrated Hybrid CNN-Transformer for Retinal OCT Classification

arXiv:2607.09809v1 Announce Type: cross Abstract: Deep models for retinal optical coherence tomography (OCT) classification report high accuracy but rarely report whether their confidence can be trusted -- a gap that matters when a wrong-but-confident reading delays sight-saving treatment. We pair a hybrid convolutional-Transformer encoder with a gradient-boosting (XGBoost) classification head and a three-part clinical safety layer: confidence calibration, out-of-distribution (OOD) rejection, an

Read source article
arXiv cs.CVResearch

CHM-Net: Center Heatmap-driven Macro-Micro Modeling Network for MRI-based Microbial Density Stratification

arXiv:2607.09812v1 Announce Type: cross Abstract: Microbial density is clinically important for tumor assessment and treatment decision-making, and recent advances in deep learning suggest that it can be non-invasively inferred from multimodal MRI. In this work, MRI-based Microbial Density Stratification (MRI-MDS) is first investigated as a patient-level representation learning task, and Center Heatmap-driven Macro-micro modeling Network (CHM-Net) is introduced for this task. CHM-Net first estab

Read source article
arXiv cs.CVResearch

RASR: Range-Aware Scale Recovery for Metric UAV Navigation

arXiv:2607.09815v1 Announce Type: cross Abstract: Under Global Navigation Satellite System (GNSS) denial, a UAV controller still needs a distance and heading command it can execute, making accurate metric last-meter navigation essential. Dense pair-geometry foundation models transfer relative structure well, yet the distance scale of their raw metric outputs remains poorly calibrated. Under the relative error metric of PairUAV, correcting only the average scale can still leave costly, distance-d

Read source article
arXiv cs.CVResearch

Performance Benchmarking and Optimisation of Clustering Algorithms for Local and Non-Local Similarity Measure in Medical Image Analysis

arXiv:2607.09821v1 Announce Type: cross Abstract: Medical imaging generates high-resolution images posing significant storage, transmission, and computational challenges. While low-rank matrix approximation (LoRMA) techniques offer efficient compression by exploiting structural redundancy, global approaches often fail to preserve local details critical for diagnosis. This paper focuses on clustering techniques that exploit non-local self-similarity to identify structurally similar regions in med

Read source article
arXiv cs.CVResearch

Tracking Intermittent Particles with Self-Learned Visual Features

arXiv:2607.09829v1 Announce Type: cross Abstract: In time-lapse fluorescence imaging, single-particle-tracking is a powerful tool to monitor the dynamics of objects of interest, and extract information about biological processes. However, tracked particles can be subject to occlusion and intermittent detectability. When these phenomena persist over a few frames, tracking algorithms tend to produce multiple tracklets for the same particle. In this work, we introduce self-supervised learning of vi

Read source article
arXiv cs.CVResearch

Slide-Level Active Learning Reduces Annotation Burden in H&E images

arXiv:2607.09831v1 Announce Type: cross Abstract: Deep learning-based segmentation of histopathology whole-slide images (WSIs) requires large amounts of pixel-level annotations, which are costly and time-consuming to obtain. Active learning (AL) has been proposed to reduce this effort, but existing methods exhibit three key limitations. Uncertainty estimation is unreliable on partially annotated WSIs, patch-level acquisition is inconsistent with slide-level annotation workflows, and class imbala

Read source article
arXiv cs.CVResearch

Neural Posterior Estimation for Inferring Weak Lensing Shear

arXiv:2607.09867v1 Announce Type: cross Abstract: The prevailing approach to inferring weak gravitational lensing shear from images involves detecting galaxies, estimating their ellipticities, and calibrating these estimates to correct for image noise, selection bias, and model misspecification. Characterizing the statistical model and assumptions underlying this pipeline is challenging, which makes it difficult to propagate uncertainty through its various stages. As an alternative, we propose t

Read source article
arXiv cs.CVResearch

PrismAD: Decoupled Planning via Semantic Mixture-of-Planners for End-to-End Autonomous Driving

arXiv:2607.10336v1 Announce Type: cross Abstract: This letter presents PrismAD, a decoupled end-to-end autonomous driving framework based on a Semantic Mixture-of-Planners. Existing planners usually aggregate heterogeneous scene tokens into a coupled representation space, forcing a single planning branch to jointly model agent interaction, road geometry, and driving intention. Such coupling may weaken factor-specific reasoning and obscure the contribution of different planning cues. To address t

Read source article
arXiv cs.CVResearch

Robotic Contextual Awareness for Human-Robot Collaboration and Environmental Understanding

arXiv:2607.10372v1 Announce Type: cross Abstract: The transition of autonomous mobile robots from controlled industrial settings to dynamic, human-centric environments, such as manufacturing, logistics, and healthcare, has made their safe and autonomous operation a critical area of research. These sophisticated machines must be capable of perceiving, understanding, and interacting with their surroundings to navigate freely and perform complex tasks. A significant obstacle to achieving this is th

Read source article
arXiv cs.LGResearch

NeuroMem-FHP: A Likelihood-Free Deep Learning Framework for Parameter Estimation of Fractional Hawkes Process

arXiv:2607.11177v1 Announce Type: new Abstract: In this paper, we propose deep learning based NeuroMem-FHP framework for estimating the parameters of the fractional Hawkes process (FHP), a self-exciting point process that captures long-range dependence through a fractional Mittag-Leffler excitation kernel. Two neural architectures, namely a Long Short-Term Memory (LSTM) network and a Transformer, are developed to estimate the model parameters $(\mu,\gamma,\alpha,\beta)$ directly from sequences o

Read source article
arXiv cs.LGResearch

FastTPS: An Optimized Method for LLM Token Phase for AI accelerators

arXiv:2607.11211v1 Announce Type: new Abstract: The popularity of large language models (LLMs) escalates an ongoing demand for effective inference. However, due to the sequential processing of tokens during the token phase in decoder-only LLMs inference, the inherent low parallelism leads to reduced throughput and suboptimal utilization of the computing units on artificial intelligence (AI) accelerators, particularly when handling long-sequence inputs that impose significant memory overhead. Rec

Read source article
arXiv cs.LGResearch

SPARC-Net: A Spectral, Causality-Aware, and Hard-Constrained Physics-Informed Architecture for Stiff and Shock-Dominated Partial Differential Equations

arXiv:2607.11310v1 Announce Type: new Abstract: Physics-Informed Neural Networks (PINNs) provide a meshless approach for solving partial differential equations (PDEs), but suffer severe degradation in stiff and shock-dominated problems, where small PDE residuals can correspond to globally inaccurate solutions. We show these failures are multi-causal, arising from the concurrent interplay of (i) spectral bias against sharp features, (ii) imbalanced multi-term optimization and loss-weight collapse

Read source article
arXiv cs.LGResearch

Surprisingly Simple and Effective Multi-Domain Graph Foundation Model through Graph-to-Table Alignment

arXiv:2607.11374v1 Announce Type: new Abstract: Graph Foundation Models (GFMs) have emerged as a promising paradigm for learning transferable representations across diverse graph domains. Recent advancements in GFMs have been largely dominated by two paradigms: Graph Neural Network and Large Language Model (LLM) based methods. However, these methods often face a fundamental dilemma between training with limited data and a heavy reliance on textual attributes. Tabular foundation models (TFMs) off

Read source article
arXiv cs.LGResearch

Physics-Aware Conditional SetGAN for Spatially Consistent Multi-User TR 38.901 Channel Generation

arXiv:2607.11429v1 Announce Type: new Abstract: TR 38.901-based channel models such as Sionna are reliable, but generating many multi-user channel realizations remains expensive. This paper asks a practical question: can a trained generative model produce multi-user TR 38.901 channels faster than Sionna without losing the spatial correlations imposed by user geometry? To answer this question, we propose a physics-aware, geometry-conditioned SetGAN trained on Sionna reference data. The method sep

Read source article
arXiv cs.LGResearch

Velocity Scheduled Flow Matching

arXiv:2607.11442v1 Announce Type: new Abstract: Flow matching trains a neural network to regress the conditional velocity along a linear interpolant between noise and data, and the number of network evaluations~(NFE) sets the cost of sampling. The straight-line interpolant carries an implicit choice: the sample moves at constant speed throughout the trajectory. We relax this choice and introduce Velocity Scheduled Flow Matching~(VSFM), which replaces the conditional target $x_1 - x_0$ with $v(t)

Read source article
arXiv cs.LGResearch

Event-based Neural Decoding for Neuroprosthetic Motor Control

arXiv:2607.11445v1 Announce Type: new Abstract: A substantial number of patients experience diminished mobility due to disabilities, diseases, or accidents. Although modern prostheses, powered by deep neural networks, hold the promise of significantly enhancing the quality of life for these individuals, their widespread adoption is hindered by significant latency, energy consumption, and spatial requirements. Wired connections to external high-performance processors restrict patient mobility, wh

Read source article
arXiv cs.LGResearch

HyperSafe: Inference-Time Safety Recovery for Fine-Tuned Language Models

arXiv:2607.11475v1 Announce Type: new Abstract: Safety alignment in large language models can be fragile under fine-tuning, as even benign task adaptation may increase harmful compliance. Existing defenses mainly follow two directions: they either intervene during or after fine-tuning through retraining or weight modification, which can be costly and may hurt task performance, or they use model-agnostic safety classifiers, which may miss failures specific to a given fine-tuned checkpoint. These

Read source article
arXiv cs.LGResearch

DAG-FM: A Foundation Model for Causal Discovery under Heterogeneous Causal Mechanisms

arXiv:2607.11510v1 Announce Type: new Abstract: Causal discovery from observational tabular data remains fundamentally challenging, primarily due to the heterogeneity of underlying causal mechanisms and the high-dimensional combinatorial search space of Directed Acyclic Graphs (DAGs). In this paper, we propose \textbf{DAG-FM}, a novel foundation model architecture that amortizes causal discovery. Unlike direct matrix prediction, DAG-FM decomposes the causal discovery process into two auto-regres

Read source article
arXiv cs.LGResearch

Random Label Prediction Heads for Studying Memorization in Deep Neural Networks

arXiv:2607.11541v1 Announce Type: new Abstract: We introduce a straightforward yet effective method to empirically study memorization in deep neural networks for classification tasks. Our approach augments each training sample with auxiliary random labels, which are then predicted by a random label prediction head (RLP-head). RLP-heads can be attached at arbitrary depths of a network, predicting random labels from the corresponding intermediate representation and thereby enabling analysis of how

Read source article
arXiv cs.LGResearch

Advancing Optimal Subset Oracle via Learning Relaxation of Neural Set Functions

arXiv:2607.11555v1 Announce Type: new Abstract: Learning neural set functions is pivotal to a wide range of important applications, including compound selection in AI-driven drug discovery and product recommendation. Recent work has introduced optimal subset oracles to implicitly learn set functions under practical weakly supervised settings, where model parameters are optimized through mean-field variational inference. However, these frameworks rely on Monte Carlo sampling to estimate gradients

Read source article
arXiv cs.LGResearch

Fundamental Limitations of Fixed-Budget Best-Arm Identification

arXiv:2607.11635v1 Announce Type: new Abstract: In fixed-budget best-arm identification, also known as ranking and selection, an algorithm has a sampling budget to distribute across $K$ arms. Each sample provides noisy feedback about that arm's mean, and the goal is to identify the arm with the largest mean. A common performance benchmark is the static oracle: a non-adaptive strategy that knows the means in advance and chooses fixed sampling proportions to maximize the exponential decay rate of

Read source article
arXiv cs.LGResearch

How to Tame Grokking: Representation Geometry as a Control Signal

arXiv:2607.11666v1 Announce Type: new Abstract: Grokking is a phenomenon in which neural networks initially memorize training data and only later exhibit strong generalization after prolonged optimization. Despite extensive recent study, the factors influencing the emergence and timing of grokking remain incompletely understood. We investigate the relationship between representation geometry and delayed generalization. We find that dimensionality collapse consistently precedes the onset of grokk

Read source article
arXiv cs.LGResearch

A multi-scale feature enhanced graph neural network for fluid dynamics prediction in complex geometries

arXiv:2607.11672v1 Announce Type: new Abstract: Industrial design in fields such as vehicle and aerospace engineering often relies on large-scale numerical simulations to evaluate fluid dynamics performance, which can incur substantial computational costs. Deep neural networks have shown promise in improving simulation efficiency, especially graph neural networks (GNNs), which demonstrate great potential due to their flexibility with unstructured data. However, GNNs face challenges when dealing

Read source article
arXiv cs.LGResearch

CatRetriever: Contrastive Representation Learning for Slab-to-Bulk Retrieval in Generative Catalyst Discovery

arXiv:2607.11712v1 Announce Type: new Abstract: Inverse design is an emerging data-driven paradigm for efficiently navigating vast chemical spaces to discover new materials with targeted properties, and in the context of heterogeneous catalysis, surface generative models have recently advanced this goal by directly generating catalyst surface-adsorbate structures. However, these models typically operate at the slab level and do not provide the corresponding parent bulk structure, making it diffi

Read source article
arXiv cs.LGResearch

HiFi-LLP: High-Fidelity, Low-Cost Latency Predictors with Confidence for Robust HW-NAS

arXiv:2607.11746v1 Announce Type: new Abstract: With deep neural networks (DNNs) increasingly deployed on edge devices, hardware (HW)-aware optimization techniques--such as HW-aware compression and HW-aware neural architecture search (HW-NAS)--have become essential. These methods rely on real feedback from the target hardware to tailor DNN architectures for efficient deployment. While the search can be parallelized, latency measurements via hardware-in-the-loop (HIL) remain a bottleneck due to t

Read source article
arXiv cs.LGResearch

From Global to Factor-Wise Expert Composition in Discrete Diffusion Models

arXiv:2607.11758v1 Announce Type: new Abstract: Discrete diffusion models offer a powerful framework for solving complex reasoning tasks, particularly through compositional generation, which combines multiple pre-trained experts to generalize beyond their individual training data. Recent theoretical corrections introduce time-dependent mixing weights to better align composed diffusion dynamics with the intended target. However, these methods are fundamentally limited by working on a per-sample b

Read source article
arXiv cs.LGResearch

From Expressivity to Sample Complexity: Narrow Teachers for Transformers via C-RASP

arXiv:2607.11760v1 Announce Type: new Abstract: A theoretical understanding of Transformers is crucial to better understand the capacities and limitations of large language models (LLMs). There is much work analyzing the expressivity of attention-based models. By proposing handcrafted weights or using computational complexity arguments, a large amount of past theoretical works have sought to characterize which tasks are and which are not in the hypothesis class of Transformer models. However, li

Read source article
arXiv cs.LGResearch

Relaxing Faithfulness with Intervention-Only Causal Discovery

arXiv:2607.11816v1 Announce Type: new Abstract: Causal discovery algorithms learn a network that describes the causal dependencies among random variables. A common workflow involves first utilizing conditional independence properties on observational data to determine partially directed causal relationships, then applying interventions to orient the unknown causal directions. A critical assumption for the first step is faithfulness: a requirement that causally linked variables exhibit statistica

Read source article
arXiv cs.LGResearch

Requential Coding: Pushing the Limits of Model Compression with Self-Generated Training Data

arXiv:2607.11883v1 Announce Type: new Abstract: Compression is fundamental to intelligence. A model that can represent its training data as a short code has discovered regularities that enable generalization. Large neural networks may learn functions far simpler than their parameter counts suggest, but it is challenging to construct codes that realize this simplicity. Parameter-based methods such as quantization produce code lengths that scale with model size, insensitive to how much information

Read source article
arXiv cs.CL (NLP)Research

The Paternalistic Filter: Epistemic Injustice and Differential Refusal in LLM-Mediated History Education for Marginalized Romanian Students

arXiv:2607.11292v1 Announce Type: cross Abstract: As Large Language Models (LLMs) are increasingly deployed as conversational tutors, they risk institutionalizing systemic inequalities. This study presents a systematic API audit of four LLMs acting as history tutors, evaluating 1,800 responses regarding the 1989 Romanian Revolution across five student personas varying by ethnicity and socio-economic tier. We uncover four interconnected patterns of \emph{epistemic paternalism}: (1)~\textbf{Differ

Read source article
arXiv cs.CL (NLP)Research

Forgetting Our Way to Shared Meaning: Effects of Forgetting on Conceptual Alignment in a Non-Partnership Coordination Game

arXiv:2607.11787v1 Announce Type: cross Abstract: Shared meaning in language requires people to learn and agree on categories. We ask how characteristics of agents' memories change the emergence and evolution of shared meaning. Without a coordination game, models of conceptual semantics cannot explain how shared meaning emerges and changes in groups of people; however, existing games assume that players share payoffs in a partnership setting. We model conceptual alignment as a non-partnership ga

Read source article
arXiv cs.CL (NLP)Research

Disentangling Feature Structure: A Mathematically Provable Two-Stage Training Dynamics in Transformers

arXiv:2502.20681v3 Announce Type: replace Abstract: Transformers may exhibit two-stage training dynamics during the real-world training process. For instance, when training GPT-2 on the Counterfact dataset, the answers progress from syntactically incorrect to syntactically correct to semantically correct. However, existing theoretical analyses hardly account for this feature-level two-stage phenomenon, which could be conceptually attributed to disentangled two-type features like syntax and seman

Read source article
arXiv cs.CL (NLP)Research

Sense and Sensitivity: Examining the Influence of Semantic Recall on Long Context Code Understanding

arXiv:2505.13353v5 Announce Type: replace Abstract: Large language models (LLMs) are increasingly deployed for understanding large codebases, but whether they understand operational semantics of long code context or rely on pattern matching shortcuts remains unclear. We distinguish between lexical recall (retrieving code verbatim) and semantic recall (understanding operational semantics). Evaluating 10 state-of-the-art LLMs, we find that while frontier models achieve near-perfect, position-indep

Read source article
arXiv cs.CL (NLP)Research

PM-KVQ: Progressive Mixed-precision KV Cache Quantization for Long-CoT LLMs

arXiv:2505.18610v2 Announce Type: replace Abstract: Recently, significant progress has been made in developing reasoning-capable Large Language Models (LLMs) through long Chain-of-Thought (CoT) techniques. However, this long-CoT reasoning process imposes substantial memory overhead due to the large Key-Value (KV) Cache memory overhead. Post-training KV Cache quantization has emerged as a promising compression technique and has been extensively studied in short-context scenarios. However, directl

Read source article
arXiv cs.CL (NLP)Research

REAL: Reading Out Transformer Activations for Precise Localization in Language Model Steering

arXiv:2506.08359v3 Announce Type: replace Abstract: Inference-time steering aims to alter a large language model's (LLM's) responses without changing its parameters, but a central challenge is identifying the internal modules that most strongly govern the target behavior. Existing approaches often rely on simplistic cues or ad hoc heuristics, leading to suboptimal or unintended effects. We introduce REAL, a framework for identifying behavior-relevant modules (attention heads or layers) in Transf

Read source article