AiAnyTool - Best AI Tools Directory and Artificial Intelligence Software Hub Logo
Loading theme toggle
Real-Time Coverage

AI News Today

Live

31834 stories from 30+ sources, refreshed continuously.

arXiv cs.CVResearch

LEGO-SLAM: Language-Embedded Gaussian Optimization SLAM

arXiv:2511.16144v2 Announce Type: replace Abstract: Recent advances in 3D Gaussian Splatting (3DGS) have enabled Simultaneous Localization and Mapping (SLAM) systems to build photorealistic maps. However, these maps lack the open-vocabulary semantic understanding required for robotic interaction. Integrating language features into SLAM remains a significant challenge, as storing high-dimensional features incurs excessive memory and rendering overhead, while existing methods with static models la

Read source article
arXiv cs.CVResearch

CASHEW: Stabilizing Multimodal Reasoning via Iterative Trajectory Aggregation

arXiv:2601.08010v3 Announce Type: replace Abstract: Vision-language models achieve strong performance across a wide range of multimodal understanding and reasoning tasks, yet their multi-step reasoning remains unstable. Repeated sampling over the same input often produces divergent reasoning trajectories and inconsistent final predictions. To address this, we introduce two complementary approaches inspired by test-time scaling: (1) CASHEW, an inference-time framework that stabilizes reasoning by

Read source article
arXiv cs.CVResearch

Where Not to Learn: Prior-Aligned Training with Subset-based Attribution Constraints for Reliable Decision-Making

arXiv:2602.07008v3 Announce Type: replace Abstract: Reliable models should not only predict correctly, but also justify decisions with acceptable evidence. Yet conventional supervised learning typically provides only class-level labels, allowing models to achieve high accuracy through shortcut correlations rather than the intended evidence. Human priors can help constrain such behavior, but aligning models to these priors remains challenging because learned representations often diverge from hum

Read source article
arXiv cs.CVResearch

Spherical-GOF: Geometry-Aware Panoramic Gaussian Opacity Fields for 3D Scene Reconstruction

arXiv:2603.08503v2 Announce Type: replace Abstract: Omnidirectional images are increasingly used in robotics and vision due to their wide field of view. However, extending 3D Gaussian Splatting (3DGS) to panoramic camera models remains challenging, as existing formulations are designed for perspective projections and naive adaptations often introduce distortion and geometric inconsistencies. We present Spherical-GOF, an omnidirectional Gaussian rendering framework built upon Gaussian Opacity Fie

Read source article
arXiv cs.CVResearch

Authoring for Living Worlds: Tool-Constrained LLM Agents for Executable Multi-Actor Scenarios

arXiv:2604.10383v2 Announce Type: replace Abstract: We use LLM agents to author executable specifications for a living world: formal Graphs of Events in Space and Time (GESTs) that a 3D game engine executes deterministically into multi-actor narrative videos, with per-frame spatial, temporal, and semantic ground truth as a byproduct of execution. This inverts the dominant paradigm of LLM agents driving neural video generators, which emit pixels with no semantic guarantees and no annotations. Aut

Read source article
arXiv cs.CVResearch

GTASA: Ground Truth Annotations for Spatiotemporal Analysis, Evaluation and Training of Video Models

arXiv:2604.10385v3 Announce Type: replace Abstract: Game engines hold what video models struggle to learn: a complete, explicit world state behind every frame. We turn one into a data instrument. GEST-Engine, our production-grade open-source system, deterministically executes Graphs of Events in Space and Time (GESTs), whether procedurally generated or derived from text, into videos of synchronized multi-actor scenarios, recording ground truth as it renders: 3D entity and camera state, pairwise

Read source article
arXiv cs.CVResearch

CreatiParser: Generative Image Parsing of Raster Graphic Designs into Editable Layers

arXiv:2604.19632v2 Announce Type: replace Abstract: Graphic design images consist of multiple editable layers, such as text, background, and decorative elements, while most generative models produce rasterized outputs without explicit layer structures, limiting downstream editing. Existing graphic design parsing methods typically rely on multi-stage pipelines combining layout prediction, matting, and inpainting, which suffer from error accumulation and limited controllability. We propose a hybri

Read source article
arXiv cs.CVResearch

You Only Gaussian Once: Controllable 3D Gaussian Splatting for Ultra-Densely Sampled Scenes

arXiv:2604.21400v4 Announce Type: replace Abstract: 3D Gaussian Splatting (3DGS) has revolutionized neural rendering, yet existing methods remain predominantly research prototypes ill-suited for production-level deployment. We identify a critical "Industry-Academia Gap" hindering real-world application: unpredictable resource consumption from heuristic Gaussian growth, the "sparsity shield" of current benchmarks that rewards hallucination over physical fidelity, and severe multi-sensor data poll

Read source article
arXiv cs.CVResearch

INFANiTE: Implicit Neural representation for high-resolution Fetal brain spatio-temporal Atlas learNing from clinical Thick-slicE MRI

arXiv:2605.09977v3 Announce Type: replace Abstract: Spatio-temporal fetal brain atlases are important for characterizing normative neurodevelopment and identifying congenital anomalies. However, existing atlas construction pipelines necessitate days for slice-to-volume reconstruction (SVR) to generate high-resolution 3D brain volumes and several additional days for iterative volume registration, thereby rendering atlas construction from large-scale cohorts prohibitively impractical. We address t

Read source article
arXiv cs.CVResearch

Astra: a generalizable report generation foundation model for 3D computed tomography

arXiv:2605.31437v3 Announce Type: replace Abstract: Interpreting computed tomography (CT) requires review of hundreds of volumetric slices and remains time-intensive and expertise-dependent. Automated CT report generation offers a promising route to improving clinical efficiency, yet the field still lacks a generalizable CT report generation foundation model that supports multi-region reporting and remains robust across external real-world cohorts. Intrinsic inconsistencies in reporting style an

Read source article
arXiv cs.CVResearch

Z-Reward: Beyond Scalar Rewards by Internalizing Reasoning into Score Distributions

arXiv:2606.09076v3 Announce Type: replace Abstract: Reward models are central to text-to-image post-training, but visual preference is subjective and better represented as a distribution over rubric scores than as a deterministic scalar. Existing scalar, score-token, and pairwise reward models over-compress uncertainty and fine-grained score differences, while reasoning-based generative rewards provide stronger judgments but are costly to deploy and difficult to use as direct optimization signal

Read source article
arXiv cs.CVResearch

VISA: VLM-Guided Instance Semantic Auditing for 3D Occupancy World Models

arXiv:2606.13460v2 Announce Type: replace Abstract: Semantic 3D occupancy provides a voxelized world state for autonomous driving and robot decision making, but object and rare-class errors can affect free-space interpretation, collision checking, and temporal state propagation. We show that a common VLM strategy, aligning 3D voxel or object features with crop-caption embeddings, improves text-space similarity without reliably improving closed-set occupancy mIoU. Motivated by this mismatch, we p

Read source article
arXiv cs.CVResearch

Structure-Semantic Co-optimized Latent Diffusion Model for Fast Visual Anagram Synthesis

arXiv:2606.16241v3 Announce Type: replace Abstract: Visual anagram is an intriguing form of art creation wherein a single image presents different conceptual interpretations under transformations such as flipping or rotation. Recent work has achieved visual anagram synthesis by leveraging pretrained text-to-image (T2I) diffusion models, yet still suffers from several key limitations including computational inefficiency, suboptimal aesthetic quality, and weak semantic fidelity and expressiveness.

Read source article
arXiv cs.LGResearch

The Limits and Potentials of Local SGD for Distributed Heterogeneous Learning with Intermittent Communication

arXiv:2405.11667v2 Announce Type: replace Abstract: Local SGD is a popular optimization method in distributed learning, often outperforming other algorithms in practice, including mini-batch SGD. Despite this success, theoretically proving the dominance of local SGD in settings with reasonable data heterogeneity has been difficult, creating a significant gap between theory and practice. In this paper, we provide new lower bounds for local SGD under existing first-order data heterogeneity assumpt

Read source article
arXiv cs.LGResearch

RAFP: Identifying LLM Lineages via Rare-Region Fingerprints

arXiv:2505.12682v2 Announce Type: replace Abstract: Large language models (LLMs) are increasingly released under restricted licenses, creating a growing need for robust model ownership verification. Existing fingerprinting methods are often fragile under downstream finetuning, require invasive training modifications, or fail in black-box settings. We introduce RAFP, a robust framework for identifying LLM lineages via rare-region fingerprints. Our key insight is that downstream finetuning primari

Read source article
arXiv cs.LGResearch

Variational Mixture of Graph Neural Experts for Alzheimer's Disease Recognition across Frequency Bands in EEG Brain Networks

arXiv:2510.11917v3 Announce Type: replace Abstract: Dementia disorders such as Alzheimer's disease (AD) and frontotemporal dementia (FTD) exhibit overlapping electrophysiological signatures in EEG that challenge accurate diagnosis. Existing EEG-based methods are limited by full-band frequency analysis, which hinders the precise differentiation of dementia subtypes and severity stages. To address this limitation, we propose a Variational Mixture of Graph Neural Experts (VMoGE) framework that inte

Read source article
arXiv cs.LGResearch

When and Why Does Multi-Agent Debate Fail and Does It Really Underperform?

arXiv:2510.20963v2 Announce Type: replace Abstract: Multi-agent debate (MAD) was proposed as a promising approach for ensembling the wisdom of multiple large language models (LLMs) to improve reasoning and provide effective supervision to superhuman LLMs. However, increasing empirical evidence suggests that MAD may not outperform or even significantly underperform single-agent approaches (SA), raising doubts about the benefits of MAD. In this work, we investigate this issue by analyzing the ince

Read source article
arXiv cs.LGResearch

Auditable Context-Aware HFMD Forecasting with Structured LLM Agents

arXiv:2511.23276v2 Announce Type: replace Abstract: Effective HFMD surveillance requires forecasts capturing both time-series patterns and contextual drivers such as school calendars, weather, and policy or surveillance reports. In clinical settings, forecasts must be trusted and actionable; thus, beyond point accuracy, decision-makers require concise, auditable explanations of why risk is expected to rise or fall. Classical models (e.g., ARIMA and Prophet) and foundation models (e.g., Chronos,

Read source article
arXiv cs.LGResearch

AnySleep: a channel-agnostic deep learning system for high-resolution sleep staging in multi-center cohorts

arXiv:2512.14461v2 Announce Type: replace Abstract: Sleep is essential for health, yet studying its dynamics requires manual sleep staging, a labor-intensive step in research and clinical care. Across centers, polysomnography (PSG) recordings are traditionally scored in 30-s epochs for pragmatic, not physiological, reasons and vary in electrode count, montage, and subject characteristics. These constraints challenge harmonized multi-center studies and the discovery of robust biomarkers on shorte

Read source article
arXiv cs.LGResearch

A Replicate-and-Quantize Strategy for Plug-and-Play Load Balancing of Sparse Mixture-of-Experts LLMs

arXiv:2602.19938v2 Announce Type: replace Abstract: Sparse Mixture-of-Experts (SMoE) architectures are increasingly used to scale large language models efficiently, delivering strong accuracy under fixed compute budgets. However, SMoE models often suffer from severe load imbalance across experts, where a small subset of experts receives most tokens while others are underutilized. Prior work has focused mainly on training-time solutions such as routing regularization or auxiliary losses, leaving

Read source article
arXiv cs.LGResearch

Understanding LoRA as Knowledge Memory: An Empirical Analysis

arXiv:2603.01097v4 Announce Type: replace Abstract: Continuous knowledge updating for pre-trained large language models (LLMs) is increasingly necessary yet remains challenging. Although inference-time methods like In-Context Learning (ICL) and Retrieval-Augmented Generation (RAG) are popular, they face constraints in context budgets, costs, and retrieval fragmentation. Departing from these context-dependent paradigms, this work investigates a parametric approach using Low-Rank Adaptation (LoRA)

Read source article
arXiv cs.LGResearch

MUSA-PINN: Multi-scale Weak-form Physics-Informed Neural Networks for Fluid Flow in Complex Geometries

arXiv:2603.08465v3 Announce Type: replace Abstract: While Physics-Informed Neural Networks (PINNs) offer a mesh-free approach to solving fluid-flow PDEs, standard point-wise residual minimization suffers from convergence pathologies in topologically complex domains like Triply Periodic Minimal Surfaces (TPMS). The locality bias of point-wise constraints fails to propagate global information through tortuous channels, causing unstable gradients and conservation violations. To address this, we pro

Read source article
arXiv cs.LGResearch

ScatterPrism: convergence for generative simulation and inverse problems in particle and nuclear physics

arXiv:2604.01313v3 Announce Type: replace Abstract: High-fidelity simulations and complex inverse problems, such as detector modeling and unfolding, are computationally intensive bottlenecks across subatomic physics, yet essential for accurate physical interpretation. While Conditional Flow Matching (CFM) offers a robust acceleration approach, we demonstrate its standard training loss is fundamentally misleading. Specifically, utilizing a Jefferson Lab Nuclear Physics (NP) kinematic dataset ($\g

Read source article
arXiv cs.LGResearch

Beyond Logit Adjustment: A Residual Decomposition Framework for Long-Tailed Reranking

arXiv:2604.01506v2 Announce Type: replace Abstract: Long-tailed classification, where a small number of frequent classes dominate many rare ones, remains challenging because models systematically favor frequent classes at inference time. Existing post-hoc methods such as logit adjustment address this by adding a fixed classwise offset to the base-model logits. However, the correction required to restore the relative ranking of two classes need not be constant across inputs, and a fixed offset ca

Read source article
arXiv cs.LGResearch

NetForge RL: A Multi-Agent Simulation Environment for Cyber Defense with Durative Actions

arXiv:2604.09523v3 Announce Type: replace Abstract: Training reinforcement-learning agents for cyber defense requires an environment that reflects the operational setting: noisy, partial observations, several defenders coordinating across a network, and an adaptive adversary realized through self-play. We present NetForge RL, a multi-agent environment for this setting on procedurally generated enterprise and operational-technology (OT) networks. A red agent compromises hosts with partial observa

Read source article
arXiv cs.LGResearch

Inference-Time Machine Unlearning via Gated Activation Redirection

arXiv:2605.12765v3 Announce Type: replace Abstract: Large Language Models memorize vast amounts of training data, raising concerns regarding privacy, copyright infringement, and safety. Machine unlearning seeks to remove the influence of a targeted forget set while preserving model performance, ideally approximating a model retrained from scratch without the forget set. Existing approaches aim to achieve this by updating model parameters via gradient-based methods. However, these updates are com

Read source article
arXiv cs.LGResearch

A Geometric Approach to Constrained Online Learning

arXiv:2605.21107v2 Announce Type: replace Abstract: We study constrained online convex optimization with adversarial time-varying constraints. At each round the learner acts before observing the loss and constraint, and is compared with the best fixed action satisfying all constraints in hindsight. The goal is to obtain minimax-optimal regret while controlling cumulative constraint violation (CCV). Prior algorithms achieved $O(\log T)$ regret with $O(\sqrt{T\log T})$ CCV for strongly convex loss

Read source article
arXiv cs.LGResearch

AMUSE: Anytime Muon with Stable Gradient Evaluation

arXiv:2605.22432v2 Announce Type: replace Abstract: Modern deep learning commonly relies on AdamW with prescribed learning rate schedules, but recent works challenge both components: Schedule-Free optimization removes explicit schedules via iterate averaging, and Muon improves the update geometry by orthogonalizing momentum for matrix parameters. Despite Muon's strong empirical performance, its underlying mechanism remains partially understood. We study Muon through the river-valley loss landsca

Read source article
arXiv cs.LGResearch

Hierarchical Synthetic Tabular Data Generation: A Hybrid Top-Down and Bottom-Up Framework

arXiv:2605.28198v2 Announce Type: replace Abstract: Existing approaches for synthetic tabular data generation are based on either purely generative models or LLMs, both of which struggle with data heterogeneity, logical consistency, rare-event coverage, and robustness in low-data regimes. In this paper, we propose a hierarchical hybrid top-down and bottom-up (H-TDBU) framework that decouples semantic structures from stochastic texture. In the top-down path, structure-driven logical constraints a

Read source article
Hacker News AILLMs

Threading the AI in the UI for a personal app

Article URL: https://millfolio.app/blog/millwright-ui-layers/ Comments URL: https://news.ycombinator.com/item?id=48916095 Points: 2 # Comments: 0

Read source article
Hacker News AILLMs

Australia creates world-first Office of AI

Article URL: https://www.minister.industry.gov.au/t-ayres/media/ai-australias-interests Comments URL: https://news.ycombinator.com/item?id=48915705 Points: 1 # Comments: 0

Read source article
Hacker News AILLMs

AI Agents: Hype vs. Reality (2024)

Article URL: https://www.kadoa.com/blog/ai-agents-hype-vs-reality Comments URL: https://news.ycombinator.com/item?id=48915700 Points: 1 # Comments: 0

Read source article
The Guardian AIBusiness

AI may be the toughest challenge Anthony Albanese faces this term. Guardrails are urgently needed | Peter Lewis

<p>Coherent decision-making and internal accountability are critical to meeting this manic moment</p><ul><li><p><a href="https://www.theguardian.com/technology/2026/jul/14/anthony-albanese-promises-fast-track-approvals-for-datacentres-to-shore-up-ai-investment">Anthony Albanese promises fast-track approvals for datacentres to shore up AI investment</a></p></li></ul><p>The University of Sydney was the natural setting for Anthony Albanese to lay out his vision for how Australia should confront the

Read source article
AI may be the toughest challenge Anthony Albanese faces this term. Guardrails are urgently needed | Peter Lewis
Hacker News LLMLLMs

Ptolemaic and LLM Architecture Compared (In 3D)

Article URL: https://celestial-llm-mirror.vercel.app/ Comments URL: https://news.ycombinator.com/item?id=48915548 Points: 1 # Comments: 0

Read source article
Hacker News AILLMs

White House launches AI cybersecurity clearinghouse

Article URL: https://www.cnn.com/2026/07/14/tech/ai-cybersecurity-clearing-house-white-house Comments URL: https://news.ycombinator.com/item?id=48915488 Points: 1 # Comments: 0

Read source article