AiAnyTool - Best AI Tools Directory and Artificial Intelligence Software Hub Logo
Loading theme toggle
Real-Time Coverage

AI News Today

Live

32030 stories from 30+ sources, refreshed continuously.

arXiv cs.CVResearch

Mitigating Modality and Language-Style Gaps for Zero-Shot Video Moment Retrieval

arXiv:2607.19027v1 Announce Type: new Abstract: Zero-shot video moment retrieval aims to overcome the limitations of traditional approaches that require large-scale datasets annotated with text and its relevant temporal spans. Despite advances in pre-trained vision-language models and multimodal large language models, existing ZMR methods still heavily depend on query-to-video content similarity, making them vulnerable to modality and language-style gaps. These gaps lead to unreliable span propo

Read source article
arXiv cs.CVResearch

CoGoal3D: Collaborative 3D Object Detection with 3D-Aware Fusion and Refinement

arXiv:2607.19036v1 Announce Type: new Abstract: V2X collaborative object detection features overcoming the limitations of single-vehicle systems by aggregating environmental features from multiple collaborative agents. However, existing mainstream V2X perception methods mainly focus on 2D BEV object detection. When 3D detection task is concerned, inferior results are obtained because they ignore the 3D spatial misalignment caused by differing height and attitude among the collaborators. In this

Read source article
arXiv cs.CVResearch

FilmWorld: Agentic Novel-to-Film Generation through Dynamic Cinematic World Modeling

arXiv:2607.19038v1 Announce Type: new Abstract: Translating novels into films poses a grand challenge for generative artificial intelligence, requiring conversion of abstract literary prose into long-form, multi-scene visual narratives. While current video generation models excel at short, single-scene clips within narrow temporal and spatial contexts, novel-to-film generation operates in a more complex regime, demanding long-duration content across diverse scenes with dynamically evolving entit

Read source article
arXiv cs.CVResearch

Now You See the Hate: Adaptive View Retrieval for Hidden Hateful Illusions

arXiv:2607.19061v1 Announce Type: new Abstract: Hateful optical illusions expose a serious gap in current multimodal safety systems. On original-view hateful illusions, previous work shows that six moderation classifiers achieve at most 20.9 to 24.5% accuracy and nine state-of-the-art VLMs remain at or below 10.2% with illusion-aware prompting, leaving most hidden hate undetected. We formulate hidden hateful illusion detection as a perceptual retrieval problem and propose Adaptive View Retrieval

Read source article
arXiv cs.CVResearch

Mage-Flow: An Efficient Native-Resolution Foundation Model for Image Generation and Editing

arXiv:2607.19064v1 Announce Type: new Abstract: Large-scale visual generators are increasingly capable but costly to train, fine-tune, and deploy. We introduce Mage-Flow, a compact 4B-scale generative stack for efficient text-to-image generation and instruction-based image editing. The stack is built from two co-designed components: Mage-VAE, a lightweight high-fidelity latent tokenizer, and a Native-Resolution Multimodal Diffusion Transformer trained with rectified flow matching. Mage-VAE uses

Read source article
arXiv cs.CVResearch

Delineate Anything v2: A Global Foundation Model for Field Delineation

arXiv:2607.19069v1 Announce Type: new Abstract: Accurate agricultural field boundary delineation at large scale is a foundational task for food security, supply chain transparency, and carbon accounting. While vision foundation models like SAM show remarkable zero-shot capabilities, they frequently fail in geospatial domains due to topological complexity, cropland texturing patterns, and a lack of physical scale awareness. In this work, we introduce Delineate Anything v2, a globally scalable fou

Read source article
arXiv cs.CVResearch

Context-structured Video Anomaly Detection with Large Vision-Language Models

arXiv:2607.19077v1 Announce Type: new Abstract: Training video anomaly detectors is challenging due to the difficulty and cost of annotating diverse and rare abnormal events. Although recent large vision-language models enable training-free inference, existing approaches mostly rely on holistic inference over sampled video and may miss context-specific anomaly cues. In this paper, we present CSI-VAD, a training-free video anomaly detector that identifies abnormal events across diverse contexts.

Read source article
arXiv cs.LGResearch

Beyond Single-Dimensional Compression: The Compound Sparsity Frontier of Large Language Models

arXiv:2607.18280v1 Announce Type: new Abstract: Large language models (LLMs) are often compressed through static parameter pruning or dynamic token-level computation, yet aggressive sparsification can trigger rapid performance degradation beyond an essential sparsity boundary. This work asks \emph{whether combining these two mechanisms can delay such degradation by distributing the compression burden}. We study a minimalist compound sparsity framework that first applies low-rank approximation an

Read source article
arXiv cs.LGResearch

Compressing What Matters: Neuron Importance Meets Data-Aware Low Rank Approximation for Language Model Compression

arXiv:2607.18284v1 Announce Type: new Abstract: To excel at their domain large language models are comprised of billions of parameters. Yet this comes at the cost of huge memory requirements restricting their applicability in resource-constrained environments. To address the problem of neural network (NN) compression Singular Value Decomposition (SVD) has played a key role as a fundamental component for matrix compression through decomposition. To minimize compression error and to maximize the e

Read source article
arXiv cs.LGResearch

Edge-Efficient Transformer for End-to-End RF Spectrum Monitoring

arXiv:2607.18285v1 Announce Type: new Abstract: We present E-SpecFormer (Edge Spectrum monitoring Transformer) for end-to-end automatic modulation and covert channel (CC) recognition. We introduce LiTAN (Linear Tanh Attention Network), a Softmax- and LayerNorm-free attention mechanism that reduces complexity while increasing accuracy in RF tasks. E-SpecFormer is parameterized in four scalable variants (Nano, Small, Medium, Large) to accommodate diverse hardware constraints. Using the RadioML2018

Read source article
arXiv cs.LGResearch

BearingNAS: Obtaining In-Sensor Intelligent Fault Diagnosis Systems for Bearings Using a Laptop

arXiv:2607.18287v1 Announce Type: new Abstract: This paper introduces BearingNAS, a Hardware-Aware Neural Architecture Search (HW-NAS) framework designed to shift the intelligence directly onto the sensor die via in-sensor processing. BearingNAS frames the search as a constrained optimization problem targeting extreme micro-budgets (4 to 8 kiB of RAM and 16 to 32 kiB of Flash). To eliminate the reliance on expensive discrete GPUs, we propose a lightweight, derivative-free search strategy paired

Read source article
arXiv cs.LGResearch

SechKAN: Kolmogorov-Arnold Networks with Hyperbolic Secant Functions

arXiv:2607.18290v1 Announce Type: new Abstract: In recent years, Kolmogorov-Arnold Networks (KANs) have attracted increasing attention due to their effectiveness in machine learning and scientific computing tasks, offering a new paradigm for neural network design. In this paper, we present SechKAN, a KAN architecture based on hyperbolic secant (sech) functions. The hyperbolic secant basis is used for its smooth bell-shaped form, localized responses, and stable gradients. We employ 1D linear tran

Read source article
arXiv cs.LGResearch

The Information Shadow: Measuring Structural Limits on What Language Models Can Learn

arXiv:2607.18305v1 Announce Type: new Abstract: Some limits on what language models know are not gaps in data coverage but structural properties of learning from text. We introduce the information shadow: the region of phenomena that a text-trained learner cannot acquire regardless of scale, comprising (I) structures language cannot express, (II) functions that are statistically non-identifiable from the training distribution, and (III) functions that are representable but unreachable by gradien

Read source article
arXiv cs.LGResearch

Gradient-Energy Guided Block-Wise Perturbations for Sharpness-Aware Minimization

arXiv:2607.18306v1 Announce Type: new Abstract: Sharpness-Aware Minimization (SAM) improves generalization by minimizing the worst-case loss in a local parameter neighborhood. Standard SAM implicitly allocates its global perturbation budget across parameter blocks according to instantaneous minibatch gradient norms. Such an allocation can be noisy and may not reflect the sensitivity that blocks accumulate throughout training. We propose Gradient-Energy Adaptive Radius SAM (GEAR-SAM), which maint

Read source article
arXiv cs.LGResearch

Agentic Calibration of Grey-Box Simulation Models: An LLM-Driven Alternative

arXiv:2607.18308v1 Announce Type: new Abstract: Calibration of grey-box simulation models is a constrained optimization problem in which model evaluations are expensive, the parameter space can be high-dimensional, and the search must respect plausibility constraints. Although the simulation code is fully available to the analyst, the joint effect of multiple parameters remains difficult to predict analytically. Classical optimizers such as Nelder--Mead (NM) are simple to deploy but sample-ineff

Read source article
arXiv cs.LGResearch

Cost Accounting for Reactive Computational Graphs: Exhaustive Sweeps, Sequential Mutation, and the Backward-Locality Gap

arXiv:2607.18323v1 Announce Type: new Abstract: Exhaustive site-by-site interventions on a neural network's computational graph -- activation-patching sweeps, circuit-discovery searches, systematic ablation studies -- mutate the graph at every candidate site, and their cost is dominated by recomputation after each mutation. On a reactive graph engine whose invalidation provably touches exactly the downstream cone of a mutated node, we give a complete cost accounting for such workloads. First, th

Read source article
arXiv cs.LGResearch

Dynamic Loss Balancing for Joint SOH and RUL Prediction of Lithium-Ion Batteries via a Rotary SOH-Injected Prior Battery Transformer

arXiv:2607.18329v1 Announce Type: new Abstract: The deployment of reliable lithium-ion battery management systems is crucial for accelerating electrification, yet the joint prognosis of State of Health (SOH) and Remaining Useful Life (RUL) remains severely hindered by task heteroscedasticity. Conventional multi-task learning frameworks fail to balance the bounded, low-variance noise of SOH estimation with the unbounded, nonlinearly expanding uncertainty of long-term RUL predictions. Here, we pre

Read source article
arXiv cs.LGResearch

Federated Lightweight Fine-Tuning

arXiv:2607.18343v1 Announce Type: new Abstract: Federated fine-tuning is bottlenecked by communication: FedAvg and pseudo-gradient schemes transmit a payload that scales with the model, and gradient compression shrinks it by only a constant factor. We take a different lever. Mapping networks generate a network's weights from a small trainable latent through a frozen affine projection; because the map is shared and affine, averaging latents is exactly averaging the generated weights. We turn this

Read source article
arXiv cs.LGResearch

An Analysis of Residual-Stream Geometry Across Transformer Depth

arXiv:2607.18348v1 Announce Type: new Abstract: We propose a transition-centred geometric analysis of transformer residual streams. Relative displacement measures how \emph{far} representations move between consecutive layers, and orthogonal Procrustes analysis separates each transition into a rigid rotation and a non-rigid residual. Across six instruction-tuned models, on code generation and cross-lingual translation, these measurements reveal reproducible depth regularities. Relative displacem

Read source article
arXiv cs.LGResearch

Multi-layer MIMO Relay as Deep Physical Neural Networks: Power Amplifiers as Activation Functions

arXiv:2607.18354v1 Announce Type: new Abstract: Wireless physical neural networks (WPNNs) embed neural computation directly into analog hardware, offering lower energy consumption and latency than conventional digital implementations. In this paper, we propose a deep WPNN in which nonlinear activations are realized by a multi-hop multiple-input multiple-output (MIMO) relay network, in which each relay implements a trainable complex linear gain and bias, followed by the power amplifier's intrinsi

Read source article
arXiv cs.LGResearch

Physical Self-Supervised Learning: IMU Sensing without Manual Labels

arXiv:2607.18361v1 Announce Type: new Abstract: Deep neural networks have become a promising approach for IMU-based sensing, but their scalability is fundamentally limited by costly labeled data and poor robustness to heterogeneous devices, placements, and users. Existing unsupervised and self-supervised methods reduce but do not remove this dependence, still requiring labeled data for domain adaptation and largely ignoring known physical structure. We propose physical self-supervised learning,

Read source article
arXiv cs.LGResearch

A Controlled Study of Attention-Only Transformers

arXiv:2607.18363v1 Announce Type: new Abstract: Feed-forward networks hold two thirds of a transformer's non-embedding parameters, yet the architecture has not received a necessity test that controls parameters, compute, and depth at once. We pretrain attention-only decoder transformers (Simple Attention Networks, SANs) against standard transformers matched separately for parameter count, training FLOPs, and depth (2 to 48 layers), for up to 105B tokens at 6M to 87M parameters. Deleting feed-for

Read source article
arXiv cs.LGResearch

Intelligence from Learnable Novelty

arXiv:2607.18433v1 Announce Type: new Abstract: Intelligence appears under different names in different fields: as data compression in statistics and machine learning, as universal computation in dynamical systems, and as adaptive behavior in agents. Each field carries its own objective, and the two most influential drives often fail in mirror image: novelty search, which seeks surprise, is transfixed by a noisy television screen, while the free-energy principle, which avoids surprise, is most c

Read source article
arXiv cs.LGResearch

Automated Data Engineering and Feature Selection for the Case Study of Warpage Detection in Fused Deposition Modeling

arXiv:2607.18515v1 Announce Type: new Abstract: This study contributes toward development of an Automated Data Processing (ADP) framework designed to evaluate and reinforce optimal machine learning model-feature combinations for predictive tasks in fused deposition modeling (FDM) process datasets. The methodology is centered around a reinforcement learning-inspired policy updating mechanism, where multiple machine learning models are trained on both full feature sets and feature subsets selected

Read source article
arXiv cs.LGResearch

Operational Proto-Introspection in Looped Language Models: Process-Quality Taps, Executable Branching, and the Readout-Control Boundary

arXiv:2607.18553v1 Announce Type: new Abstract: Can a language model read the quality of ongoing computation, and can an external intervention turn that readout into better outcomes? We test both questions in a frozen 2.6B looped transformer, Ouro-RLTT. On GSM8K, a strict pre-answer probe excludes the answer region and gold value yet predicts eventual success: hidden states plus length and log-probability shortcuts reach AUROC 0.797, versus 0.731 for the shortcuts alone (incremental +0.066; task

Read source article
arXiv cs.LGResearch

Stale but Stable: Staleness-Adaptive Trust Regions for Stabilizing Asynchronous Reinforcement Learning

arXiv:2607.18722v1 Announce Type: new Abstract: Asynchronous reinforcement learning improves throughput by decoupling rollout generation from optimization, but staleness is an inevitable byproduct compounded by policy lag, engine delays, and mixture-of-experts routing. From a trust-region perspective, this mismatch is critical: training-inference divergence governs approximation error in finite-horizon bounds, whereas PPO clipping only gates sampled outward updates, acting as a sampled surrogate

Read source article
arXiv cs.LGResearch

ConceptCF: Concept-based Counterfactuals for the Explainability of Time Series

arXiv:2607.18748v1 Announce Type: new Abstract: This paper proposes ConceptCF, a method for counterfactual generation that operates on human-interpretable concepts. In high-stakes domains such as healthcare and predictive maintenance, artificial intelligence models can increase efficiency and safety. Explainability is key to ensure these models rely on causal relationships rather than spurious correlations. Counterfactual explanations identify minimal modifications that would change a model's pr

Read source article
Dev.to

Running Qwen3 Through the ExecuTorch MLX Delegate: Up to 4.52x Faster on M1 Max

<p>Hello, everyone.</p> <p>There are now many ways to run an LLM on a Mac, but exporting a PyTorch model for Apple Silicon and executing it in a lightweight runtime is still an evolving path. How much faster is it, and does 4-bit quantization change the output?</p> <p>Today, I am looking at ExecuTorch's experimental MLX delegate, released in May 2026. It enables PyTorch models to run on Apple Silicon GPUs. I use ExecuTorch 1.3.1 to run Qwen3-0.6B and compare it with PyTorch MPS.</p> <p>The short

Read source article
Dev.to

OpenAI ships Codex into Claude Code — two commands, or four?

<p>A rare thing happened on March 30, 2026: OpenAI shipped first-party tooling straight into a competitor's terminal. The result is <code>codex-plugin-cc</code>, and it changes how a two-agent workflow fits in one window.</p> <h2> What codex-plugin-cc is, and why OpenAI built it for Claude Code </h2> <p><code>codex-plugin-cc</code> is an official OpenAI plugin that runs the Codex coding agent from inside Anthropic's Claude Code CLI. It lives at <a href="https://github.com/openai/codex-plugin-cc"

Read source article
Dev.to

Six failure cases to test before shipping an AI workflow

<p>An AI workflow is easy to demonstrate and harder to finish. The happy path can look convincing while retries duplicate side effects, approvals are skipped, or completion is reported without evidence.</p> <p>This tutorial turns a workflow idea into a small contract and six acceptance tests. The examples are deliberately model-agnostic: they apply to an OCR review flow, a support-ticket agent, an API automation, or a research assistant.</p> <h2> Start with an observable contract </h2> <p>Before

Read source article
Dev.to

You shouldn't have to learn three config formats to use MCP servers

<p>If you've set up MCP servers across Claude Desktop, Claude Code, and Cursor, you've hit this: each client wants the same servers in a different file, in a slightly different place, and you end up copy-pasting JSON three times and getting the paths wrong.</p> <p>The MCP roadmap itself lists config portability as an open problem — "setting up an MCP server in one client means starting from scratch in another."</p> <h2> The three places nobody remembers </h2> <ul> <li> <strong>Claude Desktop</st

Read source article
Dev.to

Ebook Reviewer Wanted: Help Me Find What's Gone Stale

<p>Technical books have a shelf life that nobody prints on the cover.</p> <p>A cookbook from 2018 still works. A history book from 2018 still works. A Laravel book from 2018 will teach you a middleware pattern that was replaced twice, recommend a package whose maintainer archived it, and show you a test suite in a syntax that no longer runs.</p> <p>I've written a series of ebooks on PHP, OOP, SOLID, design patterns, Laravel conventions, testing with Pest, application architecture, and AI-assiste

Read source article
Hacker News AILLMs

If AI writes everything, why should anyone trust you?

Article URL: https://www.adgully.com/post/18212/if-ai-writes-everything-why-should-anyone-trust-you Comments URL: https://news.ycombinator.com/item?id=49001606 Points: 2 # Comments: 0

Read source article
OpenClaw Commits

refactor(config): config-surface reduction tranche 3 — product consol…

<pre style='white-space:pre-wrap;width:81ex'>refactor(config): config-surface reduction tranche 3 — product consolidations (review request) (#111527) * refactor(config): consolidate media model lists * refactor(config): unify memory configuration * refactor(config): consolidate TTS ownership * refactor(config): move typing policy to agents * refactor(config): retire product-level config surfaces * refactor(config): share scoped tool policy type * chore(config): refresh generated baselines * fix(

Read source article
OpenClaw Commits

fix(agents): emit diagnostic when sessions_yield parks without contin…

<pre style='white-space:pre-wrap;width:81ex'>fix(agents): emit diagnostic when sessions_yield parks without continuation evidence (#100146) * fix(agents): emit diagnostic when sessions_yield parks without continuation evidence (#100146) Add a shared hasYieldContinuationEvidence predicate in incomplete-turn.ts and emit a user-visible diagnostic payload in terminal-resolution.ts when a yielded turn has no same-turn continuation source (accepted spawn, async tool, messaging delivery, cron add). Sim

Read source article