←── back to feed
/topics/arxiv-ai-agents-papers-september-18
arXiv AI agents papers September 18
86 items●1 sources●updated 2d ago●trend 0
On September 18, 2026, arXiv published 20 machine learning papers spanning diverse topics including query suggestion, LLM compression, diffusion models, federated learning, attention mechanisms, and Bayesian optimization. The papers address efficiency, privacy, interpretability, and scalability challenges across language models, neural networks, and domain-specific applications.
- Intent-Driven Query Suggestion Framework uses dual-stage optimization with intent-aware diversity rewards for follow-up query generation
- Layer-wise curriculum learning enables efficient LLM compression through progressive knowledge transfer from teacher to student models
- Block parallelism (BP) reduces distributed attention communication overhead for long-context diffusion language model training
- Probe guidance method achieves state-of-the-art performance on continuous diffusion language models without additional inference passes
- Stiefel Attention constrains query/key projections to Riemannian manifold with O(d)-equivariant Riemannian Adam optimizer
[BLG]blog/rss86
Generative Query Suggestion via Intent Coverage and Query-Level Credit Assignment
Layer-wise Curriculum Learning for Efficient LLM Compression
Block Parallelism For Efficient Distributed Long-Context Diffusion Language Model Training
Randomized SVD Approximations for Spectral Co-Clustering of Word-Document Matrices
Radio-Frequency Convolutional Neural Networks
Learning-Induced Dynamical Transition in Recurrent Neural Networks
Personalized Federated Hierarchical Gaussian Processes for Privacy-Preserving Modeling of Heterogeneous Distributed Systems
How to Guide Your Language Flow
Smart Insole Human Activity Recognition for Continuous Monitoring in Elderly Care
Stiefel Attention: When the Geometry of Transformer Projection Matrices Dominates Optimizer Choice---and When It Does Not
Machine-Learning Assessment of the Predictive Value of Inflammatory Biomarkers for Cognitive Impairment in an Older Hispanic Adult Cohort
FCx: An algorithm for finding Feasible Counterfactual Explanations
Improving Offline Goal-Conditioned Reinforcement Learning via Selective Reward Stimulation
Bayesian Optimization with Rich Auxiliary Information via LLMs
Sharpness-Aware Minimization (SAM) Improves Classification Accuracy of Bacterial Raman Spectral Data Enabling Portable Diagnostics
Enhanced Agriculture-informed Neural Network by Domain Knowledge
Search at the Cost of Sampling: Nearly-Instant Latent Space Bayesian Optimization
Sample Count Is Not Enough: Candidate-Generation Strategy Shapes the Energy and Performance of LLM Test-Time Scaling
LSTM-UT and Recurrent-Depth Transformers on Cellular Automata
Compressed Active Subspaces for Scalable Bayesian Inference
FedFIbOS: Fisher Importance based Optimal Submodelling for Heterogeneous Federated Learning
The Complexity Kink: A Prompt-Side Structural Complexity Index for Code-Generation Reliability
A Policy Profile for Croissant: Refusal as a Property of the Dataset
CoRe: Coherence and Relational Alignment for Multivariate Time Series Forecasting
Conservation Buys Stability and Factoring Buys Counterfactuals in Physical World Models
Opinion Dynamics-based Coalition Formation for Federated Learning in Heterogeneous IoT Systems
Odds-Ratio Thompson Sampling: A Specification and Design Guide for Contrast-Based Multi-Armed Bandits
Learn Your Own Thoughts: Abstract Token Curriculum
Alliance Beats Isolation: Unifying Heterogeneous Allied Datasets Improves Classifier Performance
OceanMoE: Structured Conditional Sparse Computation for Long-Horizon Multivariate Ocean Forecasting
PhyRestore: Physics-Structured Latent-Factor Restoration
Learning-Based Reconstruction of Optical Properties in Bilayered Media from Single-distance Time-Resolved Reflectance Measurements
DeliveryGym: An RL Environment for Long-Horizon Embodied Agent Planning with Adaptive Curriculum
Beyond Flattened Tokens: Structure-Preserving EEG Decoding with Reusable TriDim Blocks
Expected Hypervolume Maximization for Multiobjective Optimization under Uncertainties
Pretrained Medical Representations for the Practical Screening of Drug Repositioning Candidates
AURA: Adaptive Uncertainty-Routed Analysis for Email Threat Detection
Regularized Emphatic Temporal-Difference Learning: Stability under Constant Stepsizes
BioPhys-Bridge: A Benchmark for Interdisciplinary Scientific Reasoning in Physics-Grounded Biological Research
What Do We Expect from LLMs? Mapping the Design of LLM Benchmarks
Position: It is Time to Virtualize Foundation Models with a Self-evolving Operating System Layer
What Do Current Systematic Generalization Tasks Miss? A Reasoning-Centered Analysis
Characterizing Web Search by Conversational LLM Agents: From Search Decisions and Strategies to Results and Responses
Do AI Agents Understand Computer Architecture?
MAGS: Multi-agent Auto-formalization Guarantees Safety for Agentic Outputs
Closed-World Resolution Against Tool Hallucination in LLM Agents
The syntax and semantics of goals
Compositional Reasoning in Language Models under Reinforcement Learning Post-Training
Safety Beyond the Interface: Detecting Harm via Latent States in Large Language Models
QVAC Genesis III: A Large-Scale, High-Quality Open Synthetic STEM Corpus for Efficient Language Model Pre-Training
LLM-as-an-Improver: Turning Verification into Better Candidates
An Architecture for Long-Horizon Agents: Levels, Ticks and Cascaded Intelligence
EconSkills: Studying Skill Transfer and Retrieval for Web Agents on Live Economic Data
A Unified Evaluation Framework for Trustworthy Large Language Models, Agentic AI, and Multimodal Systems
Self Improvement via Fast Tree-search
When Hiring Becomes Agent-Mediated: Evaluating Access and Recurrence in Two-Agent R\'esum\'e Screening
Agentic AI Networking for Heterogeneous Unmanned Aerial Systems in Low-Altitude Wireless Networks
Continual Enterprise World Model Discovery in Dynamic Systems
SIMLIFE: Pattern Understanding for Long-Horizon Human-Agent Partnership
From Intent to Action: Benchmarking LLM Safety in Vehicle Voice Command Authorization
Reach or Solve? Attributing Agentic RL Gains with Checkpoint Handoffs
ScientistTwo: Pioneering the Human Knowledge Frontier with Autonomous AI
Replan, Repair, or Edit? A Unified Empirical Evaluation of Travel Agents for Itinerary Revision under Resource Disruptions
When2Think: Learning Difficulty-Aware Length Control for Efficient Hybrid Reasoning Models
FINSKILLOPS: A Self-Evolving Multi-Agent System for SEC Filing QA
LearnActCoder: Role-Aware Error Memory for Adaptive Clinical Coding Agents
AutoData: Agentic Search for Pre-training Data Selection
Rethinking Multi-Agent Collaboration: When More Is Less
TorchCraft: Unified binder design by inverting an all-atom structure predictor
Integrating knowledge from case reports: a medical ontology based multimodal information system with structured summary
Contagion on the Trading Floor: How Adversarial Signals Spread in Multi-Agent Trading Systems
Steering Equilibrium Selection in Regularized Self-Play via the Reference Policy
MetaRTL: Meta-path Attention Enhanced Relational Table Learning
A Dual-Process Perspective on Nudge Susceptibility in LLM-Based GUI Agents
Constraint-Safe Graph-Context Scoring for Stable Point-Feature Labels Under Text-Width and Accessibility-Inspired Profiles
Reproducibility is not construct validity: LLM measurement of institutionally situated communication
Physical knowledge on historical data matters more than enforcing physical constraints on the forecast
TRACE: Accountable Agentic Retrieval for Source Discovery in Digital Archives
From "Who Is This User?" to "What Does This Purchase Mean?": A Deployed Pipeline for Semantic User Profiling at Bank Scale
Beyond Depth Truncation: Controlled Evaluation of Depth Utilization in Recursive Language Models
MaSCoD: A Multi-Agent Framework for Structural-Context-Guided Candidate Causal Graph Generation
Not All AI Agents Are Equal: Characterizing Resource and Performance Dynamics
Neuro-Symbolic Agentic AI for Networked Low-Altitude UAVs
Customizable and Jointly Optimized Route Planning: A Deep Architecture Enabling Differentiable Shortest-Path Search
E-AVI: Evidence-Grounded Multimodal Assessment for Automated Video Interviews
Geopolitical Divisions Across Languages in Large Language Models