←── back to feed
/topics/arxiv-ai-agents-papers-september-15

arXiv AI agents papers September 15

79 items1 sourcesupdated 6d agotrend 0

On September 15, 2026, arXiv published 20 papers on AI agents spanning foundation models, policy frameworks, professional task automation, physical robotics, scientific research, and domain-specific applications. The papers address core challenges including long-horizon execution, cost efficiency, reliability assurance, and governance of agentic systems across healthcare, supply chain, air traffic, and molecular design domains.

  • ZGCM-1: 7B open foundation model with 256K context, combining internal thinking with external tool use for math and agentic search
  • Generalized Agent Iteration framework formalizes recursive self-improvement and iterative policy improvement with theoretical properties
  • Vibe Patenting testbed shows LLM judge-guided revision improves patent-draft quality versus unguided iteration
  • Asclepius benchmarks long-horizon clinical agents on 8-hour emergency-department shifts under time and resource pressure
  • AutoTailor meta-framework automatically constructs compact MCP API tool sets from web trajectories, reducing cost and latency
  • Carbon-aware routing distributes function-calling queries across edge-cloud architecture to reduce LLM energy use and emissions
[BLG]blog/rss79
ZGCM-1: A Fully Open and Extremely Efficient Foundation Model for Math and Agentic Search
arXiv cs.AI · Jiyan He, Guang Liang, Hao Liu, Haoxiang Guan, Jinbo Sun, Junyi Guo, Wenjun Feng, Yantai Xie, Yifei Shen, Bin Shao, Chuyang Wei, Kai Chen, Kexin Zhou, Minghang Zhu, Shuxin Zheng, Tie-Yan Liu, Taine Zhao, Wenhui Zhu, Xueyin Xu, Xiaoqing Zhang, Yatao Li, Yuxuan Ren · 6d
Generalized Agent Iteration: One Formal Framework for Iterative Policy Improvement and Recursive Self-Improvement
arXiv cs.AI · Hongyao Tang, Yi Ma, Pengyi Li, Yifu Yuan · 6d
Vibe Patenting: Evaluating LLM Judges for Professional Patent-Drafting Agents
arXiv cs.AI · Toshiaki Koike-Akino, Vlad Blaykhman, Ye Wang, Jing Liu, Gene V. Vinokur · 6d
Toward Self-Adaptive Physical AI: Can LLM Agents Manage Long-Horizon Physical Tasks?
arXiv cs.AI · Varun Kaushik, Yayun Tan, Xiaofan Yu · 6d
LabAgent: Customize Any Research Hubs for Scientific Discoveries Using AI Agents
arXiv cs.AI · Lei Liu, Yikun Zhang, Jialin Chen, Wanjia Zhao, Rex Ying, Wengong Jin, Hua Xu, James Zou, Tianyu Liu, Hongyu Zhao · 6d
TimeThink: Eliciting Compositional Reasoning in Timeseries Large Language Models
arXiv cs.AI · Sudarshan Regmi, Arvind Pillai, Yu Yvonne Wu, Yuliang Chen, Bibek Panthi, Tess Z. Griffin, Michael V. Heinz, Lisa Marsch, Nicholas C. Jacobson, Andrew Campbell · 6d
Root-Cause Attribution Is a Search Problem: Continual Search for Long-Horizon Agent Failures
arXiv cs.AI · Harsh Raj, David Lee, Anas Mahmoud, Renxiong Wang, Razvan-Gabriel Dumitru, Chenguang Wang, Tong Zhao, Yunzhong He, Darvin Yi, Vipul Gupta · 6d
Governing at Machine Speed: An Adaptive Intelligence Architecture for Real-Time AI Policy Enforcement
arXiv cs.AI · Sandeep Bokkasam, B. Durgalakshmi · 6d
OrchSLM: Probing the Dynamics of Small Language Model Orchestration
arXiv cs.AI · Chengxi Zhang, Yu Yao · 6d
Grounded Adjudication of Variations across Extracted TimeLines (GAVEL): Comparing Clinical Timelines Against Their Case Reports
arXiv cs.AI · Jack Cummins, Sayantan Kumar, Ketan Tamirisa, Jeremy C. Weiss · 6d
Token Efficient Task Execution via Application Behavior Modeling for Web Agents
arXiv cs.AI · Alexandru Ianta, Eleni Stroulia · 6d
Fraglingo: Molecular Design via Attachment-Aware Autoregressive Fragment Generation
arXiv cs.AI · Thao Nguyen, Jeonghwan Kim, Zhenhailong Wang, Heng Ji · 6d
From Legal Text to AI-specific Risk Sources: A Systematic Analysis of the EU AI Act's High-Risk Requirements
arXiv cs.AI · Ronald Schnitzer, Mike Auer, Rumpa Choudhury, Andreas Hapfelmeier, Maximilian Hoeving, Isabelle Painter, Josiane Xavier Parreira, Sonja Zillner · 6d
Asclepius: An Adaptive Harness for Long-Horizon Clinical Agents
arXiv cs.AI · Grace Chang Yuan, Xiaoman Zhang, Sung Eun Kim, Luyang Luo, Pranav Rajpurkar · 6d
AutoTailor: Automatic, User-Aligned Capability Selection and Adaptation for Web Agents
arXiv cs.AI · Xinyun Cao, Adriana Szekeres, Fazle Elahi Faisal · 6d
Toward a Decision-Assurance Layer for AI-Assisted Flight Planning in Air Traffic Management
arXiv cs.AI · Alexandre Barreto (George Mason University), Shou Matsumoto (George Mason University), Jorge Valverde-Rebaza (Tecnol\'ogico de Monterrey), Cleiton Ataide (DECEA: Department of Airspace Control), Paulo Costa (George Mason University) · 6d
Carbon-Aware Routing for Function Calling in Edge-Cloud LLM Systems
arXiv cs.AI · Aikaterini Maria Panteleaki, Varatheepan Paramanayakam, Spyros Tragoudas, Iraklis Anagnostopoulos · 6d
A Hybrid Agentic AI Framework for Intelligent Supply Chain Analytics
arXiv cs.AI · Xian Yeow Lee, Teppei Inoue, Haiyan Wang, Chetan Gupta · 6d
Planning or Learning: Reliability and Cost in Multi-Asset Maintenance
arXiv cs.AI · Xian Yeow Lee, Chandrasekar Venkatraman, Ahmed Farahat · 6d
Causal multi-modal AI for personalized chemosensitivity prediction
arXiv cs.AI · Dhruva Biswas, Jeroen Berrevoets, Alec McClean, Linus Bao, Jungkyu Park, Ken G. Zeng, Joseph Cappadona, Cerise Tang, Chuwen Liu, Bartosz Machura, Yin Wu, Valerie Speirs, Hatem Soliman, Rohit Bhargava, Sheheryar Kabraji, Thaer Khoury, David Page, Brian Piening, Carlo Bifulco, Claudia Meurs, Pieter Westenend, Sylvie Chabaud, Jerome Lemonnier, Paul H. Cottu, Florence Dalenc, Fabrice Andre, Frederique Madeleine Penault-Llorca, Thomas Bachelot, Frederick Howard, Francisco J. Esteva, Kevin Kalinsky, Lajos Pusztai, Jan Witowski, Krzysztof J. Geras · 6d
How User-AI Mistreatment Occurs and Matters in Conversational Systems?
arXiv cs.AI · Fanqi Zeng, Sadid A. Hasan, Chaocheng He · 6d
FLoKD: Adaptive Knowledge Distillation for Federated Low-Rank LLM over Wireless Networks
arXiv cs.AI · Xinlu Zhang, Na Yan, Yang Su, Yansha Deng, Toktam Mahmoodi · 6d
$\tau$-Elicitation: Benchmarking multi-turn entity extraction in voice agents
arXiv cs.AI · Soham Ray, Victor Barres · 6d
Identity Is More Than Recall: A Benchmark for Persistent Identity in Deployed AI Agents
arXiv cs.AI · Zhenyu Zhao, Roy Zhao · 6d
Solar Intelligence
arXiv cs.AI · Jyotsna Singh · 6d
Safety as a Constraint: Fine-Tuning a LLM Recommender to Explain Itself
arXiv cs.AI · Jiashu He, Emma Yanyang Kong, JJ Tan, David Fagnan · 6d
FedV-KGQA in Practice: Design Lessons and an Interactive Prototype
arXiv cs.AI · Md Saikat Islam Khan Bappy, Oshani Seneviratne · 6d
Cost Characterization of Vertically Partitioned Federated Knowledge Graphs
arXiv cs.AI · Md Saikat Islam Khan Bappy, Oshani Seneviratne · 6d
GeoSkill:Experience-Driven Hierarchical Skill Learning with Collaborative Revision forGeospatialAgents
arXiv cs.AI · Han Luo, Xian Xu, Yinhe Liu, Yanfei Zhong · 6d
Enhancing Event Candidate Acquisition for Event Linking
arXiv cs.AI · Ziyang Zhang, Yinan Liu, Boyi Xue, Yingxuan Huang, Bin Wang, Xiaochun Yang · 6d
Recoverability as a System Primitive for Long-Horizon AI Agents
arXiv cs.AI · Zhihui Zhang, Wei Liu · 6d
Windowed A-K-MDP
arXiv cs.AI · Xiangwen Yang, Frankie Cho, Iadine Chades · 6d
Drift-Constrained Optimization: Only Direction Matters in Fine-Tuning Instruct Models
arXiv cs.AI · Fei Yuan, Changjiang Gao, Yilei Tu, Yifeng Liu, Shujian Huang, Yu Qiao · 6d
Degraded but Not Entirely Ineffective: PE-Based Deformable Graph Neural Networks
arXiv cs.AI · Jinhua Wu, Xinliang Zhang · 6d
JaxAHT: A JAX-Based Library for Ad Hoc Teamwork
arXiv cs.AI · Caroline Wang, Rolando Fernandez, Zelal Su Mustafaoglu, Montek Kundan, Jiaxun Cui, Lingyun Xiao, Zhihan Wang, Di Yang Shi, Aditya Madhan, Johnny Liu, Arrasy Rahman, Peter Stone · 6d
MANAS-2: Constrained Reconstruction for EEG Foundation Models
arXiv cs.AI · Arvasu Kulkarni, Aditya Ray Mishra, Mahir Jain, Parshva Runwal, Lakshya Saini, Siddharth Panwar, Sandeep Singh · 6d
IBBench-Light: A Paired Evaluation of Task-Conditioned Responses to External Directives
arXiv cs.AI · Kainan Zhou, Gangzhen Qian, Zhaoyi Li, Hang Xiao · 6d
Trustworthy Agentic AI: A Comprehensive Cybersecurity and Systems Survey on Threat Landscapes, Defense Architectures, and Open Challenges
arXiv cs.AI · Seyedakbar Mostafavi · 6d
How Many Thoughts Can a Vector Hold? The Capacity of Reasoning by Superposition
arXiv cs.AI · Hongyu Gu, Chang Liu, Jingwen Fu · 6d
Positioning manuscripts in the scientific landscape with agentic AI
arXiv cs.AI · Jiawen Chen, Zichen Zhang, Bingxuan Li, Quan Sun, Yiyan Zhang, Edric Tam, Jinjie Lin, Didong Li, Yun Li, Bingxin Zhao · 6d
Homeostatic Continual Learning
arXiv cs.AI · Yue Jin · 6d
Surprising Effectiveness of Self-Demonstrations in Enhancing Schema-Ontology Mapping with LLMs
arXiv cs.AI · Siddhesh Thombre, Manasi Patwardhan, Sunita Sarawagi · 6d
Partition Scores Are Not System Scores: Deployment-Fidelity Gaps in Decomposed Algorithm Selection
arXiv cs.AI · Jiachen Zhang, Yu Tang, Li Zhu · 6d
Do Not Restart: Residual Completion for Stateful Agent Handoffs
arXiv cs.AI · Runzhi Deng, Yiming Zhong, Fang Zhao, Pan Zhou · 6d
LLM-Enhanced Multi-Agent Reinforcement Learning for Unified Electric Vehicles-Charging Station-Grid Optimization in Public Charging Systems
arXiv cs.AI · Yang Zhang, Lindong Xie, Chongyu Wang, Gaojunjie Li, Siqi Bu, Edward Chung · 6d
Bypass Observation: A Conceptual Design of a Non-Intrusive Layer-Wise Semantic Extraction Architecture
arXiv cs.AI · Haibin Tong, Jiang Yu · 6d
ViperQ: Order Flow Pattern Recognition via Auction Market Theory for Reinforcement Learning Trading
arXiv cs.AI · Asser Moustafa, Rares-Mihail Neagu, Jugal Kalita · 6d
UniCAR-RL: Seeing Better before Thinking Deeper in Visual Mathematics
arXiv cs.AI · Yuzhe Li, Hao Yan, Hao Wang, Xingchen Liu, Ya-Qi Yu, Jihao Wu, Minghui Liao, Wei Chen, Yuliang Liu · 6d
ClinAgent: A ReAct-Based Agent for Conversational Access to Clinical Trial Information
arXiv cs.AI · Antonino Vaccarella, Riccardo Cantini, Domenico Talia, Paolo Trunfio, Marianna Talia, Rosamaria Lappano, Marcello Maggiolini · 6d
Token Merging for Multilingual Speech Recognition: A Systematic Study Across Model Scale and Fine-Tuning
arXiv cs.CL · Dylan Luke Holyoak · 6d
PhysMent: An Interactive Approach For LLM Reasoning In Physics Problems
arXiv cs.CL · Joseph Chan, Utkarsh Jha, Xiyin Yang, Abhinav Jarajapu, Anik Sahai, Eddie Hu, Robin Jeshua Deepak, Stefano Saravalle, Aditya Shah · 6d
Lexical Prompt Compression for Large Language Models: A Training-Free, Deterministic Pipeline with Empirical Pareto Analysis Across Eleven Task Categories
arXiv cs.CL · Shamin Chokshi · 6d
TestHallVQA: Exploring LVLMs' Document-Level Reasoning under Redundant Contexts from Scientific Exams
arXiv cs.CL · Yongqi Yu, Yu Zhang · 6d
Clinical Reasoning Under a Partially Observed Objective in Cone Beam CT Report Generation
arXiv cs.CL · Ajo Babu George, Govind Arun, Sidharth N Krishna, Uma Ranjan · 6d
RFCLLM: Evaluating LLMs' Reasoning Ability of Network Protocol State Machines
arXiv cs.CL · Anqi Chen, Dan Goldwasser, Cristina Nita-Rotaru · 6d
CVSS-X: A Multilingual Speech-to-Speech Translation Corpus for 28 Languages
arXiv cs.CL · Lucas Rafael Stefanel Gris, Alef Iury Siqueira Ferreira, Frederico Santos de Oliveira, Augusto Seben da Rosa, Alexandre Costa Ferro Filho, Arlindo Rodrigues Galv\~ao Filho, Anderson da Silva Soares · 6d
Causal Analysis and Mitigation of Spurious Onsets in Full-Duplex Speech LLMs
arXiv cs.CL · Kento Nishi · 6d
Hindsight Bias in Clinical Temporal Reasoning: How Future Data Exposure Affects Large Language Model Judgment
arXiv cs.CL · Misaki Matsuura, Sayantan Kumar, Ojas Kadam, Jeremy C. Weiss · 6d
A Hybrid Hierarchical 1D-CNN-BiLSTM Framework for Extractive Summarization of Biomedical and Clinical Text
arXiv cs.CL · Saad Bin Ather, Muhammad Saif, Ali Hassan Khan, Manzer Abbas, Hajra Waheed · 6d
From Token Probabilities to Semantic Constraints: Towards Declarative Probabilistic Evaluation of Language Models
arXiv cs.CL · Kyle Richardson, Cullen Anderson, Pranav Balakrishnan, Takuto Ban, Daksha Ladia, Ankita Gupta, Marisa Hudspeth · 6d
Harmfulness Propagation Dynamics: Layer-wise Trajectories of Adversarial Intent in Large Language Models
arXiv cs.CL · Noor Islam S. Mohammad, Ulu\u{g} Bayaz{\i}t · 6d
Domain-Specific Jargon in Large Language Models: A Comparative Analysis between General-Purpose and Specialist Models
arXiv cs.CL · Darin Keng, Zhewei Sun · 6d
Toward Complete Hospital Discharge Summarization with Abstract Meaning Representation
arXiv cs.CL · Paul Landes, Sitara Rao, Aaron Jeremy Chaise, Barbara Di Eugenio · 6d
Same Patient, Different Order: Action-Level Reliability of Clinical LLM Agents Under Repeated Runs
arXiv cs.CL · Rohith Reddy Bellibatlu, Manpreet Singh, Zhoutian Han, Wenbin Zhang · 6d
In the Blind: Building Pseudo-References for MT Evaluation
arXiv cs.CL · Diptesh Kanojia, Chi-kiu Lo, Archchana Sindhujan, Samuel Larkin, Greg Hanneman, Alon Lavie · 6d
The University of Melbourne WMT 2026 CreoleMT Submission: A Domain-Balanced Approach to Low-Resource Pacific Creole Machine Translation
arXiv cs.CL · Rapha\"el Merx, Nick Thieberger, Ekaterina Vylomova · 6d
LayerRoute: Adaptive Layer-Skipping with LoRA-Preserved Quality for Efficient LLM Inference
arXiv cs.CL · Prateek Kumar Sikdar · 6d
Not all Negation Cues are Equal: Affixal Negations Yield Better Negation Understanding
arXiv cs.CL · Tian Tan, Eduardo Blanco · 6d
When Edit Localization Amplifies Relative Selection Bias: Gradient Geometry, Target Mismatch, and Importance Weighting
arXiv cs.CL · Shengwei Zhang, Haoda Dai, Yifei Li, Yuheng Song · 6d
Scaling Hindi Quantum Natural Language Processing through Automatic Pregroup Supertagging
arXiv cs.CL · Gautami Sanjay Naik, Krishna Bhatia, Mithun Paul Saint-Germain, H Aswath Babu · 6d
PolicyMem: Geometric Policy Memory for LLM Governance
arXiv cs.CL · Yuanchen Bei, Zhengzhang Chen, Yanjun Zhao, Haoyu Wang, Hanghang Tong, Haifeng Chen · 6d
ForeSight: Enhancing Risk Monitoring via Early Safety Signal Distillation
arXiv cs.CL · Hanling Wang, Chenlong Wei, Ling Xu, Hanyan Niu, Qi Cao, Shizhou Huang, Yang Yang, Xiaohui Zhu, Yao Zhu · 6d
Inside VLM Chart Reading: Tracing Value Reading from Vertical Bar Charts Across Space and Depth
arXiv cs.CL · Tianhao Niu, Qingfu Zhu, Wanxiang Che · 6d
HyperProve: Answer-Guided Hypergraph Expansion for Multi-Hop Question Answering
arXiv cs.CL · An Nguyen Phu, Dung Nguyen Quang, Luu Hieu An, Linh Ngo Van, Trung Le, Thien Huu Nguyen · 6d
SyRHM: Symbolic-Language-Enhanced Reasoning with Associative Retrieval for Zero-shot Harmful Meme Detection
arXiv cs.CL · Hanling Wang, Chenlong Wei, Yingjuan Li, Di Wu, Yuchao Zhang, Xiaohui Zhu, Yao Zhu · 6d
Understanding the Limits of Agentic ICD Coding
arXiv cs.CL · Chong Yock Eng, Yushi Cao, Yiming Chen, Kezhi Mao, Hongchao Jiang · 6d
DARE: Dialectical Agentic Reasoning for Structured Knowledge Fact Checking
arXiv cs.CL · Yifei Li, Xiaohan Zheng, Wentao Qian, Liansheng Zhuang · 6d
When Consistency Does Not Mean Reliability: Evaluating Local LLM Judges Against Human Ratings
arXiv cs.CL · Aakash Kumar Tiwari · 6d
Sweet Talkers: How Query Formulation Shapes Sycophancy in Romantic Relationship Advice
arXiv cs.CL · Helena Choi, Edric Castel Hao, Karl Bautista, Francis Gabriel Magleo, Renzo Panti, Danielle Beatrice Olalia · 6d