AiAnyTool - Best AI Tools Directory and Artificial Intelligence Software Hub Logo
Loading theme toggle
Real-Time Coverage

AI News Today

Live

34185 stories from 30+ sources, refreshed continuously.

arXiv cs.CVResearch

InteractiveAvatar: Real-Time Streaming Video Generation for Consistent and Intent-Aware Avatars

arXiv:2606.22905v1 Announce Type: new Abstract: Recent diffusion-based models have enabled realistic audio-driven avatar generation in real-time streaming. However, existing approaches struggle to maintain visual temporal consistency and fail to explicitly perceive user intent in complex interactive streaming scenarios. To address these challenges, we propose InteractiveAvatar, a real-time infinite-streaming video generation framework that supports visually consistent avatar video generation and

Read source article
arXiv cs.CVResearch

Each Judge Its Own Yardstick: Discovering Per-VLM Taxonomies for Physical Video Evaluation

arXiv:2606.22918v1 Announce Type: new Abstract: Maintaining physical consistency in video generators and world models increasingly relies on vision-language models (VLMs) as automated judges that provide reward signals, ranking decisions, and data-filtering criteria. Yet VLMs differ substantially in training data and architecture, encoding physical phenomena through distinct internal representations. A single global evaluation schema therefore gives every VLM the same axes of competence, regardl

Read source article
arXiv cs.CVResearch

MythraGen: Two-Stage Retrieval Augmented Art Generation Framework

arXiv:2606.22924v1 Announce Type: new Abstract: Text-to-image generation has seen rapid advancements, especially with the development of generative models. However, challenges remain in achieving high-quality, contextually accurate image outputs that faithfully match the provided textual descriptions, especially in artistic generation. In this paper, we present a simple yet efficient retrieval augmented generation framework, namely MythraGen, for text-to-artistic image generation by integrating

Read source article
arXiv cs.CVResearch

Hybrid Compression: Integrating Pruning and Quantization for Optimized Neural Networks

arXiv:2606.22935v1 Announce Type: new Abstract: Deep neural networks have witnessed remarkable advancements in recent years and have become integral to various applications. However, alongside these developments, training and deployment of neural network models on embedding and edge devices face significant challenges due to limited memory and computational resources. These problems can be addressed with deep neural network compression, which involves a trade-off between model size and performan

Read source article
arXiv cs.CVResearch

Evo-RAD: Navigating Rare Retinal Disease Diagnosis via Self-Evolving Agentic Retrieval

arXiv:2606.22955v1 Announce Type: new Abstract: Large-scale pretrained foundation models have revolutionized general medical screening, but often falter on rare diseases because such conditions are underrepresented in real-world clinical datasets. While retrieval-augmented diagnosis attempts to mitigate this, conventional static methods frequently succumb to the hubness problem, retrieving visually similar but semantically incorrect common diseases. To address this, we propose Evo-RAD, a self-ev

Read source article
arXiv cs.CVResearch

Concept Alignment Contrast and Long-Short Prompt Memory for Test-Time Adaptation of SAM3 in Medical Image Segmentation

arXiv:2606.22963v1 Announce Type: new Abstract: Concept segmentation models like Segment Anything Model 3 (SAM3) show strong generalization on natural images, yet their performance degrades in medical imaging due to the domain gap caused by different imaging principles and styles. Test-Time Adaptation (TTA) is essential for improving the testing performance by updating the model on the fly without annotations. However, existing vision-language TTA methods are mainly driven by image-level uncerta

Read source article
arXiv cs.CVResearch

Can Single-View Mesh Reconstruction Generalize to Robot Camera Rotation?

arXiv:2606.22987v1 Announce Type: new Abstract: Single-view mesh reconstruction predicts object meshes and spatial layouts from a single observation, making it attractive for fast robot spatial reasoning and real-to-sim digital twins. However, robot-mounted cameras naturally rotate during manipulation and navigation, while learned single-view reconstruction models often rely on view-dependent priors and may generalize poorly to out-of-distribution camera rotations. Such rotations can introduce 3

Read source article
arXiv cs.CVResearch

MotionMAR: Multi-scale Auto-Regressive Human Motion Reconstruction from Sparse Observations

arXiv:2606.23000v1 Announce Type: new Abstract: Human motion follows a temporal hierarchical structure, transitioning from low-frequency global trajectories to high-frequency details. Inspired by the success of multi-level autoregressive models in computer vision, we propose MotionMAR, a coarse-to-fine framework for motion reconstruction from sparse observations. It first estimates the global trajectory of human motion and then gradually refines the temporal details. This architecture consists o

Read source article
arXiv cs.CVResearch

Boosting Neural Video Codec via Scale-Driven Online Flow Refinement

arXiv:2606.23023v1 Announce Type: new Abstract: Although state-of-the-art neural video codecs (NVCs) have achieved remarkable performance, they suffer from limited generalization when encountering complex motion patterns unseen during training. To bridge this domain gap without the expensive cost of online fine-tuning, we propose a Training-Free Scale-Driven Online Flow Refinement (SOFR) method. Serving as a plug-and-play module, SOFR integrates motion information from coarse and fine scales and

Read source article
arXiv cs.CVResearch

SPAR: Semantic-Pixel Self-Alignment and Adaptive Routing for Unified Multimodal Models

arXiv:2606.23041v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) have achieved remarkable success in visual understanding but remain constrained in visual generation due to the fundamental feature discrepancy between semantic perception and pixel-level reconstruction. Bridging this gap requires overcoming two core challenges: endowing semantic encoders with high-fidelity reconstruction capabilities, and effectively aligning generative models with semantic spaces without r

Read source article
arXiv cs.CVResearch

UECP: Uncertainty-Enhanced Collaborative Perception

arXiv:2606.23046v1 Announce Type: new Abstract: Collaborative perception serves as a pivotal solution to enhance the perception capability of individual agents in autonomous driving, where a core challenge lies in seeking reliable evidence to quantify and weight the contribution of each participating agent. Existing methods typically rely on a confidence map, which is co-trained with the detection head, but it is inherently correlated with the detection results and thus fails to provide unbiased

Read source article
arXiv cs.LGResearch

Encoder-Decoder Manifold Alignment for Idempotent Generation

arXiv:2606.22304v1 Announce Type: new Abstract: Recently, several learning paradigms have been introduced to enforce idempotency in generative models. The goal is to ensure that repeated application of a model leaves samples unchanged once they lie on the target data manifold. In practice, however, many of these approaches fail to achieve exact fixed points, leading to instability and drift under repeated application. In this work, we argue that a key reason for this failure is a geometric misma

Read source article
arXiv cs.LGResearch

Multigrid Training for Molecular Generation using Graph Neural Networks

arXiv:2606.22377v1 Announce Type: new Abstract: Deep learning has demonstrated significant success for modeling biochemical molecular systems, where inputs are commonly represented as graphs or 3D grids. A major challenge is that computational cost scales with resolution, making full graph/grid computation of molecular densities expensive and often unstable. We introduce a multigrid training strategy that leverages low-resolution optimization to accelerate learning at higher resolution through p

Read source article
arXiv cs.LGResearch

Bypassing Minimization Bias: A Shift-Invariant Variance Estimator for Off-Equilibrium Local Learning Coefficients

arXiv:2606.22389v1 Announce Type: new Abstract: Singular Learning Theory leverages the Local Learning Coefficient (LLC) to quantify the geometry of neural network loss landscapes. However, mean-energy LLC estimators depend explicitly on an additive loss baseline, typically an estimate of the local minimum. During transient, off-equilibrium training phases, this minimum is unknown; substituting it with the lowest noisy mini-batch loss induces a systematic minimization bias that distorts the geome

Read source article
arXiv cs.LGResearch

QeHDC: Hyperdimensional Computing based on Quantum-enhanced binding and SuperClass Construction

arXiv:2606.22421v1 Announce Type: new Abstract: Hyperdimensional Computing (HDC) is a robust computational framework inspired by human cognition characterized by simple and efficient operations within high-dimensional vector spaces. Quantum-enhanced Hyperdimensional Computing (QeHDC) extends classical HDC by leveraging quantum mechanical properties to enhance computational efficiency. In this paper, we propose a novel Quantum HDC framework featuring a one-pass training method, leveraging sinusoi

Read source article
arXiv cs.LGResearch

Enhancing LLMs for Graph Tasks via Graph-aware LoRA Generation

arXiv:2606.22429v1 Announce Type: new Abstract: Graph neural networks (GNNs) tightly couple their input-output parameters to dataset-specific feature spaces and target sets, exhibiting limited transferability across different datasets. In contrast, language models (LMs) generalize flexibly via a unified input-output interface, motivating recent attempts to adapt LMs to graph tasks. However, existing methods struggle to encode whole-graph information, leading to potential information loss and sub

Read source article
arXiv cs.LGResearch

Escaping the Variance Trap: Jacobian-Free Dynamics for Root-Finding Bilevel Optimization

arXiv:2606.22433v1 Announce Type: new Abstract: Many central machine learning tasks, from entropy tuning in reinforcement learning to equilibrating generative adversarial networks, are fundamentally stochastic root-finding problems rather than loss minimization. Yet, they are frequently forced into a minimization framework via squared residuals, introducing a critical flaw we identify as the Variance Trap. Standard bilevel minimization algorithms require estimating hypergradients involving impli

Read source article
arXiv cs.LGResearch

Adaptive Recurrent Message Passing for Test Time Computing on Graphs

arXiv:2606.22462v1 Announce Type: new Abstract: Pre-trained foundation models have demonstrated remarkable success in many domains, enabling a unified backbone to generalize across diverse downstream tasks. However, extending this paradigm to graph learning remains challenging due to the intrinsic mismatch between graph data and fixed architectural designs. In this work, we show that this limitation can be overcome via recurrent graph models. To achieve this, we conduct a systematic theoretical

Read source article
arXiv cs.LGResearch

Federated learning with heavy-tailed gradient noise and communication noise: a variance-reduction based algorithm

arXiv:2606.22466v1 Announce Type: new Abstract: Federated learning (FL) is an emerging distributed machine learning paradigm that enables local devices to jointly train a global model while keeping data decentralized and private. We propose a variance-reduction based algorithm, VRA-FedSGD, for FL in the presence of heavy-tailed gradient noise and communication noise, where these noises are prevalent in large-scale machine learning over wireless networks and Internet of Things deployments. VRA-Fe

Read source article
arXiv cs.LGResearch

Stationary Robust Mean-Field Games under Model Mismatches

arXiv:2606.22579v1 Announce Type: new Abstract: Deploying multi-agent reinforcement learning (MARL) in the real world is often limited by model mismatches between the training simulators and the true environment, which could be further amplified through strategic interactions and result in severe performance degradation upon deployment. Distributional robustness offers a principled response by optimizing policies against worst-case transition models drawn from an uncertainty set, but standard ro

Read source article
arXiv cs.LGResearch

Training-free Task Classification for Multi-Task Model Merging

arXiv:2606.22589v1 Announce Type: new Abstract: Ever since the advent of foundation models and the pre-training-finetuning paradigm, there have been numerous efforts to merge multiple task-specific experts into a single multi-task model. Prior work largely focuses on finding a single merged model, but it often underperforms individual experts due to parameter interference. To resolve this, dynamic model merging employs routing to activate task-relevant parameters per input. However, existing rou

Read source article
arXiv cs.LGResearch

Scalable Maximum Entropy Reinforcement Learning for Diffusion Policies via Adjoint Matching

arXiv:2606.22630v1 Announce Type: new Abstract: Diffusion policies have recently emerged as a powerful paradigm for representing complex action distributions in reinforcement learning (RL). However, their application to online RL remains limited by the challenge of scalable training in the absence of ground-truth data, where standard optimization techniques such as score matching are not directly applicable. In this work, we introduce a highly efficient algorithm for optimizing diffusion policie

Read source article
arXiv cs.LGResearch

A Markov Chain Approach to Preference Alignment

arXiv:2606.22652v1 Announce Type: new Abstract: We propose Markov Chain from Human Feedback (MCHF), an elementary approach for aligning generative models from pairwise human preferences. Unlike Reinforcement Learning from Human Feedback (RLHF), which reduces comparisons to a scalar reward, and Nash Learning from Human Feedback (NLHF), which preserves pairwise utilities through a KL-regularized minimax optimization, MCHF uses pairwise preferences directly to define a transition mechanism over mod

Read source article
arXiv cs.LGResearch

LSTM Variants for Chaotic Dynamical Systems: An Empirical Study on the Lorenz Attractor

arXiv:2606.22662v1 Announce Type: new Abstract: Forecasting chaotic dynamical systems such as the Lorenz attractor is notoriously difficult: small numerical errors are amplified exponentially over long autoregressive rollouts. We study seven recurrent and convolutional architectures for the AI-DEEDS 2026 Chaotic Systems Challenge: a vanilla LSTM, an LSTM with additive attention, a Bidirectional LSTM (BiLSTM), a BiLSTM trained with the Huber loss, a Temporal Convolutional Network (TCN), a CNN fro

Read source article
arXiv cs.LGResearch

GRADE: Graph Representation of LLM Agent Dependency and Execution

arXiv:2606.22741v1 Announce Type: new Abstract: Can one graph represent every kind of LLM agent's run? A trace records what each step did, never what it relied on, the state it read, and the results it reused. GRADE recovers that missing layer: it models any run as one graph over its step nodes with two edge layers, execution edges (what ran in what order) read from the trace for free, and dependency edges (what each step relied on) rarely logged, so each is graded by how it is known, observed,

Read source article
arXiv cs.LGResearch

One-Step Flow Matching for Generative Modeling of Path-Dependent Physical Fields

arXiv:2606.22752v1 Announce Type: new Abstract: Physical simulations for intricate geometries with path-dependent constitutive models face difficulties due to the enormous computational cost they require. Recently, the emergence of generative AI models, which succeed in image and video synthesis tasks, has provided a promise to further improve simulations. Although U-Net-based denoising diffusion probabilistic models (DDPMs) have been adopted for elastic stress field generation, they typically r

Read source article
arXiv cs.LGResearch

Factored Gossip DiLoCo: Reducing Blocking Communication in DiLoCo

arXiv:2606.22768v1 Announce Type: new Abstract: To make large-scale distributed training practical outside high-bandwidth datacenters, we must reduce blocking, high-volume synchronization. While DiLoCo communicates infrequently, its outer synchronization remains bandwidth-heavy and brittle to stragglers and transient failures. We relax exact synchronization to approximate synchronization via mixing/gossip, which degrades gracefully under delays and communication failures. This allows us to facto

Read source article
arXiv cs.LGResearch

Towards Robust Personalized Federated Learning: Vulnerability Assessment and Defense Co-Design

arXiv:2606.22782v1 Announce Type: new Abstract: The proliferation of IoT devices has fueled distributed edge systems to collect vast amounts of sensitive data, creating fertile ground for on-device machine learning applications. While federated learning (FL) mitigates privacy concerns by exchanging model parameters instead of raw data, we identify a critical blind spot in current research. We examine the most commonly used personalized federated learning (PFL) methods, which allow clients to mai

Read source article
arXiv cs.LGResearch

Retrieval-Augmented Multimodal Learning for Enzyme-Substrate Interaction Prediction Under Low-Homology Shift

arXiv:2606.22823v1 Announce Type: new Abstract: Enzyme substrate interaction (ESI) prediction is a fundamental computational task for biocatalyst discovery and reaction screening in large biochemical spaces. In practical settings, ESI prediction is challenged by sparse positive supervision and low-homology distribution shift, where test enzymes share limited sequence identity with those observed during training. To address these challenges, we propose RAMMESI, a retrieval-augmented multimodal fr

Read source article
arXiv cs.LGResearch

RLM-Cascade: Response-Level Speculative Decoding for Cost-Efficient LLM API Serving

arXiv:2606.22840v1 Announce Type: new Abstract: We present RLM-Cascade, a proxy-layer system that applies speculative decoding at the response level to reduce LLM API costs without requiring model architecture access or a shared vocabulary. A fast, inexpensive draft model generates a candidate response; a capable verify model accepts, enhances, or is bypassed entirely depending on a lightweight complexity router. On a real-world agentic coding workload (Claude Code), RLM-Cascade achieves a draft

Read source article
arXiv cs.LGResearch

When AUC 0.998 Is Not Enough: A Candidate Evaluation Protocol for Hidden-State Probes of Indirect Prompt Injection in Multimodal Computer-Use Agents

arXiv:2606.22864v1 Announce Type: new Abstract: Hidden-state probing -- a linear classifier on a frozen vision-language model's internal activations -- has emerged as an attractive evaluation tool for flagging indirect prompt injection (IPI) in multimodal computer-use agents before the agent emits a corrupted action. We argue, on a single-backbone cautionary case study (Qwen2.5-VL-7B on Mind2Web, teacher-forced replay), that a high probing AUC on a clean-vs-attack split is not, on its own, evide

Read source article
arXiv cs.CL (NLP)Research

Not All Claims Are Equally Risky: FACTOR for Adaptive Verification in Factual Long-Form Generation

arXiv:2606.22474v1 Announce Type: new Abstract: Large Language Models (LLMs) generate fluent long-form text, however, often add unsupported factual claims. Existing verification techniques improve factuality by grounding generation in external evidence. However, the same verification policy usually applies to all claims despite being differences in hallucination risks. We propose \textit{FACTOR} (\textit{FACTuality-Oriented Risk-aware Verification}), an inference-time model that adapts verificat

Read source article
arXiv cs.CL (NLP)Research

ROMEVA: Geometry-Preserving Vocabulary Expansion for Roman Urdu Language Models

arXiv:2606.22478v1 Announce Type: new Abstract: Multilingual Language Models like mBERT are widely used for low-resource NLP, yet their adaptation to morphologically inconsistent languages such as Roman Urdu remains underexplored. Roman Urdu spelling variation causes severe sub-word fragmentation, averaging 1.50 sub-words per token. We propose \textit{ROMEVA} (Roman Urdu Embedding-preserving Vocabulary Adaptation), which combines sub-word-average initialization and a PCA-guided anchor loss to st

Read source article
arXiv cs.CL (NLP)Research

Breaking the Likelihood Trap: Variance-Calibrated Modulation for Large Language Model Decoding

arXiv:2606.22511v1 Announce Type: new Abstract: In open-ended generation, LLMs frequently fall into the "likelihood trap", marked by repetitive degeneration and vocabulary dullness, creating a discrepancy between machine-generated and human-written text. While post-hoc tail truncation (e.g., Top-$p$, Min-$p$) avoids sampling from the unreliable tail, it can over-sample from the uncalibrated head and misalign generation with human lexical preferences; fixed scalar repetition penalties likewise ig

Read source article
arXiv cs.CL (NLP)Research

Look Light, Think Heavy: What Multimodal Chain-of-Thought Reasoning Can and Cannot Do

arXiv:2606.22565v1 Announce Type: new Abstract: Chain-of-Thought (CoT) has become a standard method for improving reasoning capabilities in large language models (LLMs) by eliciting step-by-step thinking, but its effectiveness in multimodal tasks remains unclear. In this paper, we aim to systematically investigate the key question: What can multimodal Chain-of-Thought reasoning do, and where and why does it fall short? To this end, we evaluate 12 multimodal tasks across perception and reasoning

Read source article
arXiv cs.CL (NLP)Research

What are Key Factors for Updates in RL for LLM Reasoning?

arXiv:2606.22570v1 Announce Type: new Abstract: Reinforcement Learning from Verifiable Rewards (RLVR) has emerged as a promising framework for enhancing the reasoning ability of large language models. However, much of the existing work is guided by heuristic intuition, leading to divergent algorithmic choices, even contradictory ones that nevertheless report empirical gains. To better understand this phenomenon, we conduct a theoretical analysis of RLVR updates. Our study reveals that difference

Read source article
arXiv cs.CL (NLP)Research

Context-Aware Distillation and Ablation for Text2DSL

arXiv:2606.22578v1 Announce Type: new Abstract: We extend our prior work on Text2DSL automatic generation of domain-specific language (DSL) code from natural language descriptions along two complementary axes. First, we replace prompt-only synthetic generation with context-aware distillation, in which a teacher large language model (DeepSeek-V4-Flash) operates under an explicitly defined structured context comprising a BNF grammar, an API specification, and a closed identifier vocabulary; the re

Read source article
arXiv cs.CL (NLP)Research

Sub-Billion, Super-Frontier: Small Language Models Rival Zero-Shot Frontier LLMs on General and Literary Relation Extraction

arXiv:2606.22606v1 Announce Type: new Abstract: Large language models (LLMs) achieve strong relation extraction (RE), but their computational demands and reliance on proprietary APIs limit deployment in resource-constrained or privacy-sensitive settings. We investigate how far small language models (SLMs) can close this gap across general-domain and literary text. We evaluate five models from 360M to 3B parameters under three domain-composition regimes and two prompt-conditioned tuning styles (3

Read source article
arXiv cs.CL (NLP)Research

Orthogonal Representation Editing: Decoupling Semantic Entanglement in Batch Knowledge Editing of LLMs

arXiv:2606.22627v1 Announce Type: new Abstract: Knowledge editing aims to efficiently update factual information in Large Language Models (LLMs) without full retraining. However, existing methods still suffer from performance degradation in batch knowledge editing. We identify that semantic representation entanglement, such as overlapping concepts and shared syntactic patterns, accumulates interference in the representation space and reduces editing precision. To bridge this gap, in this paper,

Read source article
arXiv cs.CL (NLP)Research

moBERTo: A Modern Encoder for Portuguese via Continued Pretraining of ModernBERT

arXiv:2606.22722v1 Announce Type: new Abstract: Encoder-only transformer models remain essential for production NLP pipelines. We introduce moBERTo, a Portuguese adaptation of ModernBERT obtained through continued pretraining of the ModernBERT-base checkpoint on 60 billion tokens (5 epochs over a 12-billion-token corpus curated from FineWeb2 and filtered with educational and STEM classifiers). We preserve the original architecture, including rotary positional embeddings, alternating local-global

Read source article