←── back to feed
/topics/arxiv-ai-agent-systems-papers-july-16

arXiv AI agent systems papers July 16

71 items1 sourcesupdated 21d agotrend 0

On July 16, 2026, arXiv published 20 papers on AI agent systems spanning data provenance, robot deployment, reasoning verification, memory management, safety, and human-AI collaboration. Topics range from technical frameworks for agent harnesses and self-improvement to practical challenges like root-cause analysis, insurance underwriting, and user reliance on AI advice.

  • OriginBlame enables record- and token-level data provenance to resolve author removal requests into precise forget sets for model unlearning.
  • SPINE framework automates bimanual robot deployment with multi-agent workflows, reducing expert calibration overhead for embodied AI.
  • Interventional grounding audits test whether LLM chain-of-thought reasoning genuinely depends on stated premises via predicate substitution.
  • Survey on self-improving agents frames modern systems as foundation models coupled with prompts, memory, tools, and control logic.
  • AI-native insurance framework maps agent risk states (autonomy, authority, governance) to event probabilities and loss severities for agentic deployments.
  • Study finds AI advice suppresses humans' willingness to say 'I don't know' even when advice is wrong and accuracy is incentivized (N=3,132).
[BLG]blog/rss71
OriginBlame: Record- and Token-Level Data Provenance for AI Training Datasets
arXiv cs.AI · Haolin Xue · 21d
SPINE: Bridging the Cyber-Physical Gap with Agentic AI
arXiv cs.AI · Minkyu Ham, Dongho Kim, Chan Lee, Jiayi Wang, Min Jun Kim, Yixi Zhang, Guo Ye, Jihai Zhao, Soyeon Park, Han Liu · 21d
Interventional Grounding Audits: Black-Box Premise-Dependency Tests for LLM Chain-of-Thought via Predicate Substitution
arXiv cs.AI · Hironao Nakamura · 21d
Probabilistic Extension of Neuro-Symbolic AGI Robots based on Belnap's Typed Intensional FOL
arXiv cs.AI · Zoran Majkic · 21d
Self-Improvements in Modern Agentic Systems: A Survey
arXiv cs.AI · Zhe Ren, Yimeng Chen, Dandan Guo, Guowei Rong, Tonghui Li, R. B. Xiong, Qingfeng Lan, Wenyi Wang, Li Nanbo, Yibo Yang, Mingchen Zhuge, J\"urgen Schmidhuber · 21d
Improving Molecular Property Prediction in Small Language Models Using Graph-based Tools
arXiv cs.AI · Konstantinos Bougiatiotis, Dimitrios Kelesis, Georgios Paliouras · 21d
Oracle Agent Memory as an Enterprise Memory Substrate for Long-Horizon AI Agents
arXiv cs.AI · Richmond Alake, Cesare Bernardis, Paul Cayet, Luca Engel, Damien Hilloulin, Sungpack Hong, Allen Hosler, Nickolas Kavantzas, Ingo Kossyk, Son Le, Rhicheek Patra, Kartik Talamadupula, Valentin Venzin · 21d
Learning Safe Agent Behaviour from Human Preferences and Justifications via World Models
arXiv cs.AI · Ilias Kazantzidis, Timothy J. Norman, Yali Du, Christopher T. Freeman · 21d
CayleyR: Solving the TopSpin puzzle via cycle intersection
arXiv cs.AI · Yuri Baramykov · 21d
Networked Intelligence: Active Shared Context Graphs for Human-AI Team Science
arXiv cs.AI · Sutanay Choudhury, Jeffrey J. Czajka, Lummy M. O. Monteiro, Erin Bredeweg, Jason McDermott, Katherine Wolf, Alex Beliaev, Josh Elmore, Paul Piehowski, Kylee Tate, Yuqian Gao, Aivett Bilbao, Kelly Stratton, Scott Baker, Jaydeep P. Bardhan, Kristin Burnum Johnson, Chris Oehmen, Robert Rallo · 21d
AI-Native Insurance for Agentic AI: Pricing, Underwriting, and End-to-End Automation
arXiv cs.AI · Quanyan Zhu · 21d
Cost-Optimal Foundation Model Deployment Portfolio for Transportation Management
arXiv cs.AI · Xi Cheng, Ke Liu, Siyuan Feng, Jane Lin, H. Oliver Gao · 21d
Harness Handbook: Making Evolving Agent Harnesses Readable,Navigable, and Editable
arXiv cs.AI · Ruhan Wang, Yucheng Shi, Zongxia Li, Zhongzhi Li, Yue Yu, Junyao Yang, Kishan Panaganti, Haitao Mi, Dongruo Zhou, Leoweiliang · 21d
Theory-Level Autoformalization: From Isolated Statements to Unified Formal Knowledge Bases
arXiv cs.AI · Marcus J. Min, Mike He, Zhaoyu Li, Zixuan Yi, Sharad Malik, Aarti Gupta, Xujie Si, Osbert Bastani · 21d
EZSMT Version 3, Matured
arXiv cs.AI · Yuliya Lierler · 21d
Set-shifting Behavioral Test for Harnessed Agents
arXiv cs.AI · Ziwei Ye · 21d
LAPO: Leave-One-Turn Attribution for Self-Generated Process Rewards in Multi-Turn Search Reasoning
arXiv cs.AI · Qiang Zhu, Jiajun Wu · 21d
How Far Can Root Cause Analysis Go on Real-World Telemetry Data?
arXiv cs.AI · Athira Gopal, Ashwanth Krishnan · 21d
Multi-Agent Collaborative Reasoning with Tool-Augmented Evidence for Urban Region Profiling
arXiv cs.AI · Xixuan Hao, Yutian Jiang, Jiabo Liu, Yihang Yang, Guangyin Jin, Song Gao, Yuxuan Liang · 21d
AI advice suppresses people's willingness to say "I don't know", even when the advice is wrong and accuracy is incentivized
arXiv cs.AI · Chiara Marcoccia, Walter Quattrociocchi, Valerio Capraro · 21d
SAFETY SENTRY: Context-Aware Human Intervention via EXECUTE-ASK-REFUSE Routing
arXiv cs.AI · Tianyu Chen, Chujia Hu, Wenjie Wang · 21d
Automatic Ordinary Differential Equations Discovery For Biological Systems Using Large Language Model Powered Agentic System
arXiv cs.AI · David Krongauz, Arad Zulti, Eran Segal, Teddy Lazebnik · 21d
STOCKTAKE: Measuring the Gap Between Perception and Action in LLM Agents with a Fair Oracle
arXiv cs.AI · Sagar Deb, Ashwanth Krishnan · 21d
UESF-Bench: Benchmarking and Probing for Unified Embodied Seeking and Following
arXiv cs.AI · Kun Yu, Jianhua Yang, Yixiang Chen, Changwei Wang, Hongyuan Yu, Yan Huang, Fushuo Huo, Ya Jing, Zhumin Chen, Keji He · 21d
Explaining Reinforcement Learning Agents via Inductive Logic Programming
arXiv cs.AI · Celeste Veronese, Edoardo Zorzi, Daniele Meli, Alessandro Farinelli · 21d
When Bots Join the Team: Bot Adoption and the Institutional Fabric of Open-Source Software Projects
arXiv cs.AI · Yongren Shi, Wenyi Gong · 21d
AgentCompass: A Unified Evaluation Infrastructure for Agent Capabilities
arXiv cs.AI · Zichen Ding, Jiaye Ge, Shufan Jiang, Kai Chen, Mo Li, Qingqiu Li, Zehao Li, Zonglin Li, Tiaohao Liang, Shudong Liu, Zerun Ma, Zixing Shang, Wenhui Tian, Zun Wang, Liwei Wu, Zhenyu Wu, Jun Xu, Bowen Yang, Dingbo Yuan, Qi Zhang, Songyang Zhang, Peiheng Zhou, Dongsheng Zhu · 21d
CAVA: Canonical Action Verification and Attestation for Runtime Governance of Agentic AI Systems
arXiv cs.AI · Zexun Wang · 21d
Experience Memory Graph: One-Shot Error Correction for Agents
arXiv cs.AI · Wenjun Wang, Yuchen Fang, Fengrui Liu, Zibo Liang, Kai Zheng · 21d
AIMO Interpretability Challenge
arXiv cs.AI · Michal \v{S}tef\'anik, Philipp Mondorf, Andreas Waldis, Qianying Liu, Chuan Yang, Michal Spiegel, Josef Kucha\v{r}, Marek Kadl\v{c}\'ik, Adam Vawda-Oomerjee, Chaoran Liu, Simon Frieder, Barbara Plank, Fazl Barez, Pontus Stenetorp · 21d
A Self-Evolving Agent for Longitudinal Personal Health Management
arXiv cs.AI · Haoran Li, Jiebi Deng, Tong Jin, Jinghong Han, Yuxin Wang, Zexin Wang, Qingyi Si, Weikang Gong, Xiahai Zhuang, Jia You, Wei Cheng, Jianfeng Feng, Hongcheng Guo · 21d
Do Agent Optimizers Compound? A Continual-Learning Evaluation on Terminal-Bench 2.0
arXiv cs.AI · Wenxiao Wang, Priyatham Kattakinda, Soheil Feizi · 21d
AI-accelerated End-to-End Framework for Rapid Professional Upskilling
arXiv cs.AI · Tam Nguyen, Hung Nguyen, Robert Ogburn · 21d
Earthquaker-AI: A Retrieval-Augmented Generation Framework with Rubric-Based Assessment for Primary School Earthquake Education
arXiv cs.AI · Xanthi Kokkinou, Chaido Mizeli, Nafsika Koulaxidou, Marina Delianidi, Konstantinos Diamantaras · 21d
Deep Interaction: An Efficient Human-AI Interaction Method for Large Reasoning Models
arXiv cs.AI · Hefeng Zhou, Jinxuan Zhang, Jiong Lou, Yuxin Liu, Chaochao Lu, Jingjing Qu, Jie Li · 21d
FixItFlow: Automated Troubleshooting Guide Generation from Cloud Incidents
arXiv cs.AI · Srihari Unnikrishnan, Jaskaran Singh Walia, Drishti Goel, Supriyo Ghosh · 21d
Ask Before You Diagnose: Safe-Psych, a Sequential Evaluation Benchmark for LLMs in Psychiatry
arXiv cs.AI · Oriana Presacan, Andreea Grama, Larisa Irimin\u{a}, Alireza Nik, Jaya Ojha, Vajira Thambawita, Ciprian I. B\u{a}cil\u{a}, Bogdan Ionescu, Michael A. Riegler · 21d
Designing Safety-Constrained LLM Systems for Public Health Information Access
arXiv cs.AI · Ben Torkian, Jun Zhou · 21d
Safeguard-Conditioned Uplift: Measuring Utility-Risk Frontiers for Dual-Use Biology Assistants
arXiv cs.AI · Dipesh Tharu Mahato · 21d
Final Authority in AI Governance: Frontier-Provider Sovereignty and Action-Centered Deployer Governance
arXiv cs.AI · Zexun Wang · 21d
LessonBench-V1: A Benchmark Dataset for Evaluating AI Lesson Generation Agents
arXiv cs.AI · Ravidu Suien Rammuni Silva, Ahmad Lotfi, Isibor Kennedy Ihianle, Golnaz Shahtahmassebi, Jordan J. Bird · 21d
Beyond Backbone Backpropagation: A Decoupled Strategy for Efficient Transfer Learning
arXiv cs.AI · Daniel Vila-Cruz, Laura Mor\'an-Fern\'andez, Ver\'onica Bol\'on-Canedo · 21d
The Perplexity Trap: When Patent Law Makes Human Writing Look Like AI
arXiv cs.AI · Anubhab Banerjee · 21d
Federated Explainable Artificial Intelligence: Roles, Architectures, Evaluation, and Open Challenges
arXiv cs.AI · Masoume Gholizade, Fabrizio Ruffini, Pietro Ducange, Francesco Marcelloni · 21d
Uncertainty-Aware Sequential Decision Rules for Event-Triggered LLM Invocation in Streaming Systems
arXiv cs.AI · Zhaohui Wang · 21d
Autonomous UAV Route Planning for Coverage Maximization in Environmental Monitoring: A Systematic Literature Review
arXiv cs.AI · Sebastian Jouannet-Contreras, Carola Figueroa-Flores · 21d
Compaction as Epistemic Failure: How Agentic LLM Tools Fabricate Confirmed Results from Killed Processes
arXiv cs.AI · Hiroki Tamba · 21d
HRO: Hierarchical Room-to-Object Framework for Zero-Shot Object Goal Navigation with Large Language Models
arXiv cs.AI · Luyuan Jia, Yinfeng Yu · 21d
When is the combined load identifiable from a stress-intensity profile? A coupled forward-inverse study on SIFBench finite-element data
arXiv cs.AI · Giansalvo Cirrincione, Filippo Grassia · 21d
The Entanglement Wall: Activation-Space Probes as Risk Detectors, Not Context Adjudicators
arXiv cs.AI · Dominik Schwarz · 21d
Automatic Differentiation from Scratch: How PyTorch Computes Gradients in Physics-Informed Neural Networks
arXiv cs.LG · Abdeladhim Tahimi · 21d
What Your Model Threw Away and Why You'll Want It Back: Masking, Fingerprinting, and Privacy from Discarded Geometry
arXiv cs.LG · Zachary P. Bradshaw · 21d
Targeted Recovery of Weight-Space Mechanisms From Neural Networks
arXiv cs.LG · Antoine Vigouroux, Lee Sharkey · 21d
TSSM: Triaxial State Space Model for Global Station Weather Forecasting with Temporal-Variable-Historical Modeling
arXiv cs.LG · Songru Yang, Zili Liu, Tao Han, Ben Fei, Fenghua Ling, Lei Bai, Chang Liu, Xiangyang Ji, Zhenwei Shi, Zhengxia Zou · 21d
Disentangling Knowledge States with Ability and Proficiency Modeling for Knowledge Tracing
arXiv cs.LG · Duantengchuan Li, Yingqian Bi, Jinsong Chen, Rui Zhang, Mingwen Tong · 21d
STKAN: Kolmogorov-Arnold Networks for Spatio-Temporal Forecasting
arXiv cs.LG · Sicong Lai, Yuehong Hu, Siru Zhong, Si Qiao, Yuxuan Liang, Guangyin Jin · 21d
A Hybrid Mamba for Audio-Visual Navigation
arXiv cs.LG · Yi Wang, Yinfeng Yu · 21d
CoDiffGRN: Rethinking Gene Regulatory Network Inference via the BEELINE-KGC Benchmark and Co-evolutionary Discrete Diffusion
arXiv cs.LG · Jiaze Song, Runhao Zhao, Minghao Xu, Bin Cui, Wentao Zhang · 21d
ShortOPD: Recovering Pruned LLMs with Short-to-Long On-Policy Distillation
arXiv cs.LG · Qingyu Zhang, Qianhao Yuan, Hongyu Lin, Yaojie Lu, Xianpei Han, Le Sun, Xiang Li, Ming Xu, Jiarui Li, Xiuyin Zhao · 21d
HEDGEHOG: Hierarchical Evaluation of Drug Generators Through Rigorous Filtration
arXiv cs.LG · Daria A. Ryabchenko (Ligand Pro, Moscow, Russia, Skolkovo Institute of Science and Technology, Artificial Intelligence Center, Moscow, Russia), Pavel Gurevich (Ligand Pro, Moscow, Russia, Skolkovo Institute of Science and Technology, Artificial Intelligence Center, Moscow, Russia), Shamil Kadyrov (Ligand Pro, Moscow, Russia), Daria Frolova (Ligand Pro, Moscow, Russia, Skolkovo Institute of Science and Technology, Artificial Intelligence Center, Moscow, Russia), Kseniia Fedisheva (Ligand Pro, Moscow, Russia), Sergei A. Nikolenko (Ligand Pro, Moscow, Russia), Alexander Shapeev (Ligand Pro, Moscow, Russia, Skolkovo Institute of Science and Technology, Artificial Intelligence Center, Moscow, Russia), Marina A. Pak (Ligand Pro, Moscow, Russia) · 21d
SteinGate: Tail-Sensitive Safe Reinforcement Learning via Stein Discrepancy
arXiv cs.LG · Yassine Chemingui, Chenhua Fan, Honghao Wei, Janardhan Rao Doppa · 21d
Concurrent Image Understanding and Generation: Self-Correcting Coupled Markov Jump Processes
arXiv cs.LG · Minh-Quan Le, Armand Comas, Alexandros Lattas, Stylianos Moschoglou, Pedro V\'elez, Amit Raj, Aaron Germuth, Thabo Beeler, Dimitris Samaras, Di Qiu · 21d
EMAGN: Efficient Multi-Attention Graph Network via Learned Clustering for Scalable Traffic Forecasting
arXiv cs.LG · Mingxing Xu, Rakesh Chowdary Machineni, Ke Liu, Xi Cheng, Chengqi Lu, Xin Hu, Lyuhao Chen, Xiangyu Li, Junwei You, Oliver Gao · 21d
Reassessing Muon for Matrix Factorization
arXiv cs.LG · Ali Parviz, Gal Mishne, Alex Cloninger · 21d
Deconstructing Actor-Critic: A Large-scale Empirical Study of Design Components for Practitioners
arXiv cs.LG · Haseeb Shah, Lingwei Zhu, Adam White, Martha White · 21d
Tabular Foundation Models for Discrete Choice Estimation
arXiv cs.LG · Liu Liu, Dan Zhang · 21d
Accuracy-Preserving Stability Regularization for Large-Scale Retail Demand Forecasting
arXiv cs.LG · Jize Li, Jiani He, Dishu Yang, Dingyan Shang, Jingjing Liu, Shiqi Huang · 21d
Agora: Collective and Permissionless Internet-Scale Pretraining of Large Language Models
arXiv cs.LG · Gil Avraham, Violetta Shevchenko, Hadi Mohaghegh Dolatabadi, Karol Pajak, James Snewin, Harry Xi, Rodney O'Donnell, Thalaiyasingam Ajanthan, Sameera Ramasinghe, Chamin Hewa Koneputugodage, Shamane Siriwardhana, Alexander Long · 21d
Weight Feedback Computes the Jacobian Transpose Locally in Modern Deep Networks
arXiv cs.LG · Junlong Shen, Xingyu Li · 21d
Where Should RL Post-Training Compute Go? Model Size, Search, Learning, and Feedback
arXiv cs.LG · Patrick Wilhelm, Odej Kao · 21d
Self-Improving is Often Sudden: Enlightenment-style Finetuning for Large-Scale Models
arXiv cs.LG · Jing-Xiao Liao, Tianwei Zhang, Yu-Hao Jiang, Feifei Zhang, Hang-Cheng Dong, Feng-Lei Fan · 21d