AiAnyTool - Best AI Tools Directory and Artificial Intelligence Software Hub Logo
Loading theme toggle
Real-Time Coverage

AI News Today

Live

34182 stories from 30+ sources, refreshed continuously.

arXiv cs.CVResearch

SeFi-Image: A Text-to-Image Foundation Model with Semantic-First Diffusion

arXiv:2606.22568v1 Announce Type: new Abstract: Training image generation foundation models consumes substantial resources. Previous methods have attempted to leverage semantic guidance to accelerate the training process, yet their experiments were only conducted on simple datasets such as ImageNet, at low resolutions, and with small-scale models. In this paper, we propose SeFi-Image, a text-to-image foundation model built upon semantic-first diffusion, a novel latent diffusion modeling paradigm

Read source article
arXiv cs.CVResearch

Automated sign detection across the Electronic Babylonian Library: A large-scale dataset and end-to-end cuneiform OCR pipeline

arXiv:2606.22608v1 Announce Type: new Abstract: Learning to read cuneiform tablets is an extremely demanding task; consequently, of the roughly half million excavated tablets, only a small fraction has been analysed by Assyriologists. Computer vision offers a promising avenue for decipherment but requires large, densely annotated datasets. To address this limitation, the largest annotated cuneiform sign dataset to date is used, and a Deformable Detection Transformer (DETR)-based object detection

Read source article
arXiv cs.CVResearch

OmniSpace: Efficient Geometry Awareness for Autonomous Vehicles MLLMs

arXiv:2606.22617v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) have achieved remarkable performance on 2D visual tasks, yet enhancing their spatial intelligence for real-world applications such as Autonomous Vehicles (AV) remains an open challenge. Existing geometry-aware MLLMs typically rely on auxiliary 3D models at inference time, introducing pipeline complexity and the risk of cascading failures. In this paper, we present OmniSpace, a simple yet effective plug-and-p

Read source article
arXiv cs.CVResearch

DR-Mamba: Automatic Inference-Time Domain Adaptation for Document Image Binarization via Sample-Conditioned Detail-Background Suppression

arXiv:2606.22625v1 Announce Type: new Abstract: Degraded document image binarization is sensitive to domain shifts caused by paper aging, bleed-through, stains, shadows, and uneven illumination, and the foreground-background separation of recent learning-based methods can become unstable on unseen degradation domains. We propose DR-Mamba, a sample-conditioned detail-background suppression framework that performs automatic inference-time domain adaptation for document image binarization. Unlike t

Read source article
arXiv cs.CVResearch

MaRS: Robust Out-of-Distribution Detection via Mahalanobis Residual Scoring

arXiv:2606.22649v1 Announce Type: new Abstract: Foundation models provide highly descriptive representations for medical images, yet their reliability degrades under distribution shifts arising from changes in patients, devices, or acquisition conditions. Reliable out-of-distribution (OOD) detection is therefore essential for safe deployment. Recent post-hoc detectors efficiently exploit frozen embeddings (\emph{e.g.}, kNN), whereas reconstruction-based OOD detection in latent feature space has

Read source article
arXiv cs.CVResearch

Prompting Diffusion Models for Zero-Shot Instance Segmentation

arXiv:2606.22660v1 Announce Type: new Abstract: Several disruptive research directions have recently emerged in computer vision, including foundation models achieving previously unseen zero-shot performance in scene understanding, even interactively, and generative models that synthesize extremely realistic images. The latter have also been shown to be highly effective in scene understanding tasks thanks to their rich priors. However, for promptable segmentation, foundation models struggle with

Read source article
arXiv cs.CVResearch

NullFlow: One-Step Generative Reconstruction

arXiv:2606.22696v1 Announce Type: new Abstract: We propose NullFlow, a principled framework for one-step generative image reconstruction. Our key idea is to confine the generative flow to a measurement-consistent subspace. Because the flow never leaves this subspace, NullFlow needs no separate data-fidelity corrections, unlike existing solvers. NullFlow samples in a single network evaluation by learning the flow's average velocity, avoiding the step-by-step integration of traditional flow matchi

Read source article
arXiv cs.CVResearch

Modular Diffusion Models for Structured Visual Recognition

arXiv:2606.22702v1 Announce Type: new Abstract: Traditional supervised methods for structured visual recognition tasks -- such as object detection, segmentation, and scene graph generation -- often produce deterministic, fixed outputs, limiting their ability to capture the inherent uncertainty in complex visual scenes. As a consequence, such point estimates are unable to capture the prediction uncertainty (or multi modality) intrinsic to these problems, often arising from natural ambiguities (e.

Read source article
arXiv cs.LGResearch

Towards Understanding the Power and Limits of the Muon Optimizer: A River-Valley Perspective

arXiv:2606.21514v1 Announce Type: new Abstract: Recently, Muon has gained substantial attention as an appealing alternative to Adam-like optimizers, with many works highlighting its advantages through spectral normalization and improved conditioning. Yet this positive theoretical narrative contrasts with its empirical performance in large language model (LLM) training, where Muon's gains over Adam/AdamW are often mixed, schedule-sensitive, and not uniformly superior. To address this gap, we deve

Read source article
arXiv cs.LGResearch

Embedding Linear Equality Constraints in Probabilistic Neural Networks for Dynamic Modelling

arXiv:2606.21728v1 Announce Type: new Abstract: Machine learning models are increasingly used to model chemical process systems, yet they often lack principled uncertainty quantification and mechanisms to enforce physical constraints. We propose a probabilistic neural network framework that guarantees satisfaction of linear equality constraints within a given tolerance, while capturing aleatoric uncertainty. Compared to state-of-the-art methods, our formulation demonstrates improved predictive a

Read source article
arXiv cs.LGResearch

AdaPrivate-TS: Private Thompson Sampling for Contextual Bandits with Privacy Amplification

arXiv:2606.21757v1 Announce Type: new Abstract: We present AdaPrivate-TS, a differentially private contextual bandit algorithm that combines Thompson Sampling with batched zCDP composition. Our key insight is that differential privacy noise inflates the posterior covariance in a structured way: adding Gaussian noise $N(0,\sigma^2 I)$ to $b$ yields sampling covariance $v^2 A^{-1} + \sigma^2 A^{-2}$, which Thompson Sampling interprets as increased uncertainty rather than pure corruption. Under eve

Read source article
arXiv cs.LGResearch

A Causal DAG Prior for Synthetic Time-Series Classification Datasets

arXiv:2606.21776v1 Announce Type: new Abstract: A Prior-data fitted Network learns the posterior predictive induced by its training prior; bringing this paradigm to multivariate time-series classification therefore calls for a synthetic generator that produces complete labelled datasets with temporal structure. We introduce a causal prior that synthesizes each dataset from a randomly sampled DAG over typed nodes across two modalities (tabular attributes and time series), natively producing multi

Read source article
arXiv cs.LGResearch

RocketPFN: Accurate Time Series Classification via In-Context Learning

arXiv:2606.21786v1 Announce Type: new Abstract: We introduce RocketPFN, a training-free pipeline for time series classification that combines random convolutional feature extraction (Rocket) with in-context classification via a pretrained tabular foundation model (TabPFN v2.5). On 92 UCR datasets (30-resample protocol), RocketPFN matches HC2, the strongest published method on the archive, in mean accuracy (both 0.900, Wilcoxon p=0.50), with no training on the target data and a median inference t

Read source article
arXiv cs.LGResearch

Causal Variational Deep Embedding: A Family of Interventional Generators for Confounded Images

arXiv:2606.21806v1 Announce Type: new Abstract: Deep generative models reproduce the observational distribution of their training data, inheriting any spurious associations it contains. A common source is an unobserved confounder that shapes both an attribute the user wants to control at sampling time and an attribute expected to vary in response. Existing causal generative approaches resolve the resulting ambiguity by imposing structural assumptions strong enough to single out one interventiona

Read source article
arXiv cs.LGResearch

Mat-Pref: Verifiable-Reward Training Improves Compositional Reasoning in Inorganic Materials

arXiv:2606.21830v1 Announce Type: new Abstract: Reinforcement learning from verifiable rewards (RLVR) has driven rapid progress in mathematical and code reasoning, but when extended to science, existing benchmarks do not decompose what generalizes: do gains reflect structural transfer, property transfer, or memorization? We introduce Mat-Pref, a benchmark of 10,837 ionic-substitution questions across 11 inorganic structure families, grounded in density functional theory calculations from the Mat

Read source article
arXiv cs.LGResearch

WiSP: A Working-Set View of Mixture-of-Experts Serving on Extremely Low-Resource Hardware

arXiv:2606.21868v1 Announce Type: new Abstract: Modern Mixture-of-Experts (MoE) models place most of their parameters in expert layers, yet only a small fraction of those experts are used for any token. The unused weights must still be stored where the GPU can reach them. On commodity GPUs the common fix is layer-level CPU offloading, which keeps memory low but streams all of a layer's experts across PCIe on every forward pass, losing much of MoE's sparsity benefit. We cast low-resource MoE serv

Read source article
arXiv cs.LGResearch

Data Pruning: Redundant, Problematic, and Interdependent Samples

arXiv:2606.21916v1 Announce Type: new Abstract: The performance of deep learning models is affected by not only data quantity but also data quality. Data pruning is a process by which practitioners can reduce the size of a dataset by only keeping the most important training data points, thereby achieving similar test set performance. We empirically investigate two popular data pruning methods under noisy and noiseless conditions and show that these methods fail in the presence of significant lab

Read source article
arXiv cs.LGResearch

Selective Ensemble Based on Preference-Directed Multi-Objective Bandits

arXiv:2606.21929v1 Announce Type: new Abstract: Selective ensemble for modern machine learning systems requires choosing promising model candidates under limited evaluation budgets, while downstream tasks often specify only partial preferences over capabilities such as accuracy, robustness, and reasoning. This setting naturally gives rise to a sequential decision problem under partially specified linear preferences. We formalize it as preference-directed multi-objective bandits (PDMOB), where ad

Read source article
arXiv cs.LGResearch

DevoTG: Temporal Graph Neural Networks for Modeling C. elegans Developmental Connectomics

arXiv:2606.21940v1 Announce Type: new Abstract: Understanding how a nervous system wires itself from birth to adulthood is a fundamental challenge in developmental neuroscience. We present DevoTG, a temporal graph framework that applies Temporal Graph Neural Networks (TGNs) to two complementary representations of C. elegans neural development: a Continuous-Time Dynamic Graph (CTDG) of cell division events derived from cell lineage data, and a Discrete-Time Dynamic Graph (DTDG) of the developing

Read source article
arXiv cs.LGResearch

A Standard Processing Pipeline for High-accuracy Measurement of Few-shot Regression on Laser Induced Breakdown Spectroscopy

arXiv:2606.21960v1 Announce Type: new Abstract: Laser-induced breakdown spectroscopy (LIBS) faces challenges in high-accuracy quantitative measurement under few-shot scenarios due to spectral noise and data scarcity. Traditional preprocessing methods often fail to preserve subtle spectral features or capture nonlinear correlations. This work proposes a standardized processing pipeline integrating diffusion-based denoising, attention-based autoencoder for dimensionality reduction, group shuffling

Read source article
arXiv cs.LGResearch

Load Testing for Machine Learning Model Serving Systems at Scale

arXiv:2606.22013v1 Announce Type: new Abstract: Machine learning (ML) model serving has become a dominant consumer of GPU infrastructure, yet capacity planning in these systems remains largely ad hoc. Under-provisioning leads to service-level objective (SLO) violations and production incidents, while over-provisioning results in substantial resource waste. This paper presents \sys, an industrial load testing framework for ML serving systems that systematically estimates serving capacity through

Read source article
arXiv cs.LGResearch

Provably Efficient Policy-Reward Co-Pretraining for Adversarial Imitation Learning

arXiv:2606.22056v1 Announce Type: new Abstract: Adversarial imitation learning (AIL) achieves high-quality imitation compared to behavioral cloning (BC), but demands substantial online environment interaction. Recent empirical work has explored initializing AIL algorithms with BC pretrained policies to address this limitation, yet a rigorous theoretical understanding of pretraining's role in AIL remains elusive. This paper provides a systematic theoretical analysis and introduces principled pret

Read source article
arXiv cs.LGResearch

Frequency-Domain Neural ODEs for Modeling Non-Linear Dynamical Systems

arXiv:2606.22075v1 Announce Type: new Abstract: Standard continuous-depth models, such as Neural Ordinary Differential Equations (NODEs), offer significant advantages in modeling physical systems by learning continuous vector fields rather than discrete temporal steps. However, when applied to complex dynamical systems, standard NODEs frequently struggle with highly nonlinear dynamics. This paper investigates the Frequency-domain Neural ODE (FNODE), an architecture that projects continuous tempo

Read source article
arXiv cs.LGResearch

Meta-Reinforcement Learning via Evolution for Multi-Objective Combinatorial Supply Chain Optimisation

arXiv:2606.22146v1 Announce Type: new Abstract: Meta-reinforcement learning is a promising approach to multi-objective optimisation because it enables rapid policy adaptation across changing environments and preference settings. However, conventional few-shot methods usually fine-tune from a single shared meta-policy, which can reduce solution diversity and limit exploration of the Pareto front, especially in high-dimensional combinatorial problems such as supply chain optimisation. We propose a

Read source article
arXiv cs.LGResearch

Parameterized Representations via Implicit Stochastic Modulation for High-Dimensional and High-Order Neural PDE Solvers

arXiv:2606.22150v1 Announce Type: new Abstract: Solving high-dimensional and high-order PDEs is challenged by the coupled growth of spatial dimensionality and derivative order. Recent stochastic derivative estimators reduce this cost by replacing full derivative tensors with randomized dimension or Taylor estimators, but they are mostly designed for fixed physical parameters and require retraining for each new parameter. We show that direct conditional parameterization of such solvers entangles

Read source article
arXiv cs.LGResearch

Drowning in Routine: Signal Dilution in Multi-Turn Agent Training

arXiv:2606.22164v1 Announce Type: new Abstract: Multi-turn agents interleave consequential decisions with routine execution: some actions change the downstream return distribution, while others are necessary but reward-equivalent. The cost of trajectory-level credit assignment, often attributed to long horizons, is in fact governed by decision density $\rho$: the fraction of turns whose actions affect the return. When decision density is low, routine turns create signal dilution: they add gradie

Read source article
arXiv cs.LGResearch

Early-Exit Graph Neural Networks for Link Prediction

arXiv:2606.22167v1 Announce Type: new Abstract: Graph Neural Networks are great for link prediction in various network-like structures; however, the question of their speed/quality tradeoff has been barely studied. While in practice the time it takes to do inference matters little for small benchmarks, the latency does limit applicability in large-scale domains. In this work, we explore early-exiting strategies that can be applied to Graph Neural Networks to solve the problem of link-prediction

Read source article
arXiv cs.LGResearch

SamatNext v0.2-B: An Exploratory Study of RMS-Normalized Hybrid Decoders for Curriculum Retention in Small Code Models

arXiv:2606.22248v1 Announce Type: new Abstract: Standard autoregressive Transformer decoders can often exhibit substantial forgetting under sequential fine-tuning on shifting curriculum distributions. This technical report evaluates SamatNext v0.2-B, an experimental 356M-parameter hybrid sequence decoder that alternates Differential-Attention-style layers with DeltaNet-inspired simplified linear-state mixer layers using RMS normalization and output scale calibration. We study the model under a c

Read source article
arXiv cs.CL (NLP)Research

Beyond Hooking Onto the World: Referential Profiles and the Numerical Structure of LLM Grounding

arXiv:2606.21195v1 Announce Type: new Abstract: This paper revisits the grounding problem for large language models in light of recent vector-grounding accounts. I accept the shift from classical symbol grounding to vector grounding, but argue that the current debate remains incomplete in two respects. First, reference is often treated too thinly, as if it were a fixed link between an isolated expression and an object. I argue instead that reference is profile-based, context-sensitive, discourse

Read source article
arXiv cs.CL (NLP)Research

Where Does the Signal Live? A Web Data Recipe for Medical Encoder Pretraining

arXiv:2606.22079v1 Announce Type: new Abstract: Web data curation has been widely studied for decoder Large Language Model (LLM) pretraining. Encoders for dense-terminology domains such as medicine, by contrast, are pretrained on small, manually-curated corpora that limit scalability and writing style diversity, a bottleneck even more severe in non-English clinical settings. Whether web-scale data curation also benefits encoder Masked Language Modeling (MLM) in a dense-terminology domain remains

Read source article
arXiv cs.CL (NLP)Research

Plurification in/of language technology -- The integration of culture in next-generation AI

arXiv:2606.22097v1 Announce Type: new Abstract: The paper explores how "culture" can be operationalised in Natural Language Processing (NLP) and what this reveals about the possibilities and limits of considering a plurality of cultural backgrounds in technological design. It proposes that cultural alignment cannot be achieved only by adding more examples of "other cultures", rather it requires plural epistemologies: allowing multiple, locally grounded ways of knowing. To analyze how this plural

Read source article
arXiv cs.CL (NLP)Research

From Recognition to Understanding: Unlocking Cognitive Time Series Reasoning with LLMs

arXiv:2606.22126v1 Announce Type: new Abstract: Time series analysis has recently been coupled with Large Language Models (LLMs) to leverage their reasoning and world knowledge capabilities, yet gains remain limited. We attribute this to a fundamental mismatch between existing task formulations and LLM strengths: most settings reduce time series understanding to curve-fitting systems, focusing on low-level prediction while ignoring the semantic, contextual, and reasoning-intensive nature of real

Read source article
arXiv cs.CL (NLP)Research

BioMatrix: Towards a Comprehensive Biological Foundation Model Spanning the Modality Matrix of Sequences, Structures, and Language

arXiv:2606.22138v1 Announce Type: new Abstract: We present BioMatrix, the first multimodal foundation model that natively integrates sequences, structures, and natural language for both molecules and proteins within a single decoder-only architecture. Existing biological foundation models pursue native multimodality and broad entity coverage separately: those that fuse multiple modalities under a shared objective remain confined to a single entity type, while those spanning multiple entity types

Read source article
arXiv cs.CL (NLP)Research

The Score Granularity Gap in Black-Box LLM Classification: A Comparative Study of Confidence Constructions

arXiv:2606.22179v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly deployed as black-box classifiers in pipelines that automate confident decisions and route uncertain ones to human review. Such selective prediction needs a confidence score that an operator can threshold at a chosen risk level. Prior work asks whether LLM confidence is well calibrated or well ranked; we ask a complementary, deployment-oriented question that has been largely overlooked: at what resoluti

Read source article
arXiv cs.CL (NLP)Research

When Is Emergent Consensus Real? A Measured Coupling Gain and a Validity Diagnostic for LLM Agent Societies

arXiv:2606.22203v1 Announce Type: new Abstract: LLM "agent societies" are studied via demonstrations of emergent consensus or polarization -- with no measurable control parameter, no theory of when each regime appears, and no test of whether an outcome is a genuine social dynamic or a model artifact. We introduce the coupling gain gamma, measured per-agent by counterfactually perturbing a neighbour's stated opinion. (i) gamma is stable and model-distinguishing -- across five frontier models it s

Read source article
arXiv cs.CL (NLP)Research

Lexical Consensus: Grounded Word Learning and Shared Meaning in Artificial Agents

arXiv:2606.22207v1 Announce Type: new Abstract: Artificial intelligence systems are commonly evaluated through task performance and behavioral imitation, but such evaluations leave open whether an artificial agent can acquire, stabilize, and use new lexical meanings from grounded experience. This paper introduces Lexical Consensus, an experimental framework for studying grounded word learning over a structured perceptual substrate. Using frozen DINOv2 visual embeddings, Carroll-style nonce words

Read source article
arXiv cs.CL (NLP)Research

Evaluating Large Language Models for Hausa and Fongbe Machine Translation: Benchmarks, Failures, and Metric Reliability

arXiv:2606.22269v1 Announce Type: new Abstract: We investigate the translation quality of current large language models (LLMs) for English-to-Hausa and English-to-Fongbe - two typologically distinct West African languages from the Afroasiatic and Niger-Congo families respectively - and evaluate whether standard automatic metrics reliably reflect human judgment for these low-resource languages. We evaluate four models (GPT-4o Mini, Claude Sonnet 4, Gemini 2.5 Flash, and Qwen2.5-7B) at progressive

Read source article
arXiv cs.CL (NLP)Research

MixedPEFT: Combining Multiple PEFT Methods with Mixed Objectives for Unsupervised Domain Adaptation

arXiv:2606.22272v1 Announce Type: new Abstract: Pre-trained language models struggle when applied to new domains, as full fine-tuning is computationally expensive and prone to catastrophic forgetting. This study addresses this challenge by presenting a novel parameter-efficient strategy for unsupervised domain adaptation that combines custom PEFT architectures with mixed-objective training. Our approach simultaneously optimizes classification performance on labeled source domain data and masked

Read source article
arXiv cs.CL (NLP)Research

From Speech to Text Corpora: Evaluating ASR-Based Data Acquisition for Low-Resource Fongbe and Hausa

arXiv:2606.22274v1 Announce Type: new Abstract: Low-resource African languages lack text corpora needed for language model training. We investigate whether ASR pipelines can extend text resources for two typologically distinct West African languages: Fongbe (tonal, diacritic-rich) and Hausa (non-tonal). We fine-tune MMS-300M on a curated 12.3-hour Fongbe dataset, achieving 9.48% WER on the ALFFA benchmark - a 78% relative reduction from the prior 44.04% baseline - while preserving tonal diacriti

Read source article
arXiv cs.CL (NLP)Research

Learning at the Right Pace: Adaptive Data Scheduling Improves LLM Reinforcement Learning

arXiv:2606.22305v1 Announce Type: new Abstract: Large Language Models (LLMs) achieve remarkable reasoning capabilities through reinforcement learning (RL) post-training. However, existing RL post-training commonly relies on uniform data sampling, which ignores the semantic structure of the training data and the changing capability of the training policy. To address these limitations, we propose Adaptive Data Scheduling (ADS), a dual-level data scheduling framework for pacing RL post-training tha

Read source article