←── back to feed
/topics/arxiv-ai-agents-papers-september-15
arXiv AI agents papers September 15
79 items●1 sources●updated 6d ago●trend 0
On September 15, 2026, arXiv published 20 papers on AI agents spanning foundation models, policy frameworks, professional task automation, physical robotics, scientific research, and domain-specific applications. The papers address core challenges including long-horizon execution, cost efficiency, reliability assurance, and governance of agentic systems across healthcare, supply chain, air traffic, and molecular design domains.
- ZGCM-1: 7B open foundation model with 256K context, combining internal thinking with external tool use for math and agentic search
- Generalized Agent Iteration framework formalizes recursive self-improvement and iterative policy improvement with theoretical properties
- Vibe Patenting testbed shows LLM judge-guided revision improves patent-draft quality versus unguided iteration
- Asclepius benchmarks long-horizon clinical agents on 8-hour emergency-department shifts under time and resource pressure
- AutoTailor meta-framework automatically constructs compact MCP API tool sets from web trajectories, reducing cost and latency
- Carbon-aware routing distributes function-calling queries across edge-cloud architecture to reduce LLM energy use and emissions
[BLG]blog/rss79
ZGCM-1: A Fully Open and Extremely Efficient Foundation Model for Math and Agentic Search
Generalized Agent Iteration: One Formal Framework for Iterative Policy Improvement and Recursive Self-Improvement
Vibe Patenting: Evaluating LLM Judges for Professional Patent-Drafting Agents
Toward Self-Adaptive Physical AI: Can LLM Agents Manage Long-Horizon Physical Tasks?
LabAgent: Customize Any Research Hubs for Scientific Discoveries Using AI Agents
TimeThink: Eliciting Compositional Reasoning in Timeseries Large Language Models
Root-Cause Attribution Is a Search Problem: Continual Search for Long-Horizon Agent Failures
Governing at Machine Speed: An Adaptive Intelligence Architecture for Real-Time AI Policy Enforcement
OrchSLM: Probing the Dynamics of Small Language Model Orchestration
Grounded Adjudication of Variations across Extracted TimeLines (GAVEL): Comparing Clinical Timelines Against Their Case Reports
Token Efficient Task Execution via Application Behavior Modeling for Web Agents
Fraglingo: Molecular Design via Attachment-Aware Autoregressive Fragment Generation
From Legal Text to AI-specific Risk Sources: A Systematic Analysis of the EU AI Act's High-Risk Requirements
Asclepius: An Adaptive Harness for Long-Horizon Clinical Agents
AutoTailor: Automatic, User-Aligned Capability Selection and Adaptation for Web Agents
Toward a Decision-Assurance Layer for AI-Assisted Flight Planning in Air Traffic Management
Carbon-Aware Routing for Function Calling in Edge-Cloud LLM Systems
A Hybrid Agentic AI Framework for Intelligent Supply Chain Analytics
Planning or Learning: Reliability and Cost in Multi-Asset Maintenance
Causal multi-modal AI for personalized chemosensitivity prediction
How User-AI Mistreatment Occurs and Matters in Conversational Systems?
FLoKD: Adaptive Knowledge Distillation for Federated Low-Rank LLM over Wireless Networks
$\tau$-Elicitation: Benchmarking multi-turn entity extraction in voice agents
Identity Is More Than Recall: A Benchmark for Persistent Identity in Deployed AI Agents
Solar Intelligence
Safety as a Constraint: Fine-Tuning a LLM Recommender to Explain Itself
FedV-KGQA in Practice: Design Lessons and an Interactive Prototype
Cost Characterization of Vertically Partitioned Federated Knowledge Graphs
GeoSkill:Experience-Driven Hierarchical Skill Learning with Collaborative Revision forGeospatialAgents
Enhancing Event Candidate Acquisition for Event Linking
Recoverability as a System Primitive for Long-Horizon AI Agents
Windowed A-K-MDP
Drift-Constrained Optimization: Only Direction Matters in Fine-Tuning Instruct Models
Degraded but Not Entirely Ineffective: PE-Based Deformable Graph Neural Networks
JaxAHT: A JAX-Based Library for Ad Hoc Teamwork
MANAS-2: Constrained Reconstruction for EEG Foundation Models
IBBench-Light: A Paired Evaluation of Task-Conditioned Responses to External Directives
Trustworthy Agentic AI: A Comprehensive Cybersecurity and Systems Survey on Threat Landscapes, Defense Architectures, and Open Challenges
How Many Thoughts Can a Vector Hold? The Capacity of Reasoning by Superposition
Positioning manuscripts in the scientific landscape with agentic AI
Homeostatic Continual Learning
Surprising Effectiveness of Self-Demonstrations in Enhancing Schema-Ontology Mapping with LLMs
Partition Scores Are Not System Scores: Deployment-Fidelity Gaps in Decomposed Algorithm Selection
Do Not Restart: Residual Completion for Stateful Agent Handoffs
LLM-Enhanced Multi-Agent Reinforcement Learning for Unified Electric Vehicles-Charging Station-Grid Optimization in Public Charging Systems
Bypass Observation: A Conceptual Design of a Non-Intrusive Layer-Wise Semantic Extraction Architecture
ViperQ: Order Flow Pattern Recognition via Auction Market Theory for Reinforcement Learning Trading
UniCAR-RL: Seeing Better before Thinking Deeper in Visual Mathematics
ClinAgent: A ReAct-Based Agent for Conversational Access to Clinical Trial Information
Token Merging for Multilingual Speech Recognition: A Systematic Study Across Model Scale and Fine-Tuning
PhysMent: An Interactive Approach For LLM Reasoning In Physics Problems
Lexical Prompt Compression for Large Language Models: A Training-Free, Deterministic Pipeline with Empirical Pareto Analysis Across Eleven Task Categories
TestHallVQA: Exploring LVLMs' Document-Level Reasoning under Redundant Contexts from Scientific Exams
Clinical Reasoning Under a Partially Observed Objective in Cone Beam CT Report Generation
RFCLLM: Evaluating LLMs' Reasoning Ability of Network Protocol State Machines
CVSS-X: A Multilingual Speech-to-Speech Translation Corpus for 28 Languages
Causal Analysis and Mitigation of Spurious Onsets in Full-Duplex Speech LLMs
Hindsight Bias in Clinical Temporal Reasoning: How Future Data Exposure Affects Large Language Model Judgment
A Hybrid Hierarchical 1D-CNN-BiLSTM Framework for Extractive Summarization of Biomedical and Clinical Text
From Token Probabilities to Semantic Constraints: Towards Declarative Probabilistic Evaluation of Language Models
Harmfulness Propagation Dynamics: Layer-wise Trajectories of Adversarial Intent in Large Language Models
Domain-Specific Jargon in Large Language Models: A Comparative Analysis between General-Purpose and Specialist Models
Toward Complete Hospital Discharge Summarization with Abstract Meaning Representation
Same Patient, Different Order: Action-Level Reliability of Clinical LLM Agents Under Repeated Runs
In the Blind: Building Pseudo-References for MT Evaluation
The University of Melbourne WMT 2026 CreoleMT Submission: A Domain-Balanced Approach to Low-Resource Pacific Creole Machine Translation
LayerRoute: Adaptive Layer-Skipping with LoRA-Preserved Quality for Efficient LLM Inference
Not all Negation Cues are Equal: Affixal Negations Yield Better Negation Understanding
When Edit Localization Amplifies Relative Selection Bias: Gradient Geometry, Target Mismatch, and Importance Weighting
Scaling Hindi Quantum Natural Language Processing through Automatic Pregroup Supertagging
PolicyMem: Geometric Policy Memory for LLM Governance
ForeSight: Enhancing Risk Monitoring via Early Safety Signal Distillation
Inside VLM Chart Reading: Tracing Value Reading from Vertical Bar Charts Across Space and Depth
HyperProve: Answer-Guided Hypergraph Expansion for Multi-Hop Question Answering
SyRHM: Symbolic-Language-Enhanced Reasoning with Associative Retrieval for Zero-shot Harmful Meme Detection
Understanding the Limits of Agentic ICD Coding
DARE: Dialectical Agentic Reasoning for Structured Knowledge Fact Checking
When Consistency Does Not Mean Reliability: Evaluating Local LLM Judges Against Human Ratings
Sweet Talkers: How Query Formulation Shapes Sycophancy in Romantic Relationship Advice