←── back to feed
/topics/arxiv-ai-agent-systems-and-reasoning-papers-july-20

arXiv AI agent systems and reasoning papers July 20

26 items1 sourcesupdated 17d agotrend 0

On July 20, 2026, arXiv released 20 papers on AI agent systems and reasoning, spanning medical diagnosis frameworks, causal reasoning auditing, healthcare-specialized models, local voice assistants, multi-agent math reasoning, coding agents, reinforcement learning explainability, mobile GUI safety, knowledge graph completion, and workflow generation. The papers address core challenges in agent design: balancing cost and accuracy, ensuring transparency and auditability, scaling to complex environments, and improving controllability and safety.

  • GraphDx uses Medical Diagnosis Knowledge Graphs to balance diagnostic accuracy against resource costs in sequential diagnosis.
  • Causal-Audit proposes explicit, auditable graph-based reasoning for context-free causal question answering with verifiable reasoning paths.
  • Cura 1T is a healthcare-specialized LLM trained via human-gated self-evolution for patient consultation, clinical reasoning, and EHR tool use.
  • AnovaX is a local-first desktop voice assistant running entirely on-device with LLM planning, typed executors, and adaptive recovery.
  • Reviewer precision alone does not guarantee critique uptake; broadcast-style peer discussion outperforms planner-executor-reviewer pipelines on harder math problems.
  • DSWorld introduces Data Science World Models to predict environment state transitions and reduce trial-and-error in autonomous data science workflows.
[BLG]blog/rss26
GraphDx: A Cost-Aware Knowledge-Enhanced Multi-Agent Framework for Sequential Diagnosis
arXiv cs.AI · Shaoting Tan, Ning Liu, Yuntao Du, Shuyue Wei, Wu Shuai, Qian Li, Yanyu Xu, Wei Zhang, Lizhen Cui, Haitao Yuan · 17d
Causal-Audit: Explicit and Auditable Graph-based Reasoning via Target-Aware Causal Chain Construction
arXiv cs.AI · Su Lan, Xuefei Yin, Yanming Zhu, Alan Wee-Chung Liew · 17d
Cura 1T: Specialized Model for Agentic Healthcare
arXiv cs.AI · actAVA AI, :, Haolin Chen, Leon Qi, Steve Brown, Deon Metelski, Tao Xia, Joonyul Lee, Qixuan Wang, Kevin Riley, Frank Wang, Weiran Yao · 17d
AnovaX: A Local, Multi-Agent Voice Assistant with LLM Planning, Typed Executors, and Adaptive Recovery
arXiv cs.AI · Raunak B Sinha · 17d
Precise but Uncoupled: Reviewer Precision Does Not Guarantee Critique Uptake in Multi-Agent Math Reasoning
arXiv cs.AI · Chih-Hsuan Yang, Jingyan Jiang, Vikram Vasudevan, Cheng-Hau Yang, Huihuo Zheng, Le Chen, Eliu A. Huerta, Venkatram Vishwanath, Ian T. Foster, Rajeev Thakur · 17d
Do Coding Agents Need Executable World Models, Simplification, and Verification to Solve ARC-AGI-3?
arXiv cs.AI · Sergey Rodionov · 17d
From Black Box to Executable Logic: Explainable Reinforcement Learning through Prolog Expert Systems
arXiv cs.AI · Eduardo C. Garrido-Merch\'an · 17d
A Critical Analysis of Trustworthy AI Tools, Mark Frameworks, and the Implementation Chasms
arXiv cs.AI · Michael Papademas, Xenia Ziouvelou, Kostas Karpouzis, Vangelis Karkaletsis · 17d
SeerGuard: A Safety Framework for Mobile GUI Agents via World Model Prediction
arXiv cs.AI · Xue Yu, Bo Yuan, Pengshuai Yang, Kailin Zhao, Hong Hu, Junlan Feng · 17d
MGDT: MLLM-Guided Diffusion Transformer with Relation-Adaptive Mixture-of-Experts for Multimodal Knowledge Graph Completion
arXiv cs.AI · Xu Hou, Meiyu Liang, Wei Huang, Yawen Li, Zhe Xue, Wu Liu, Guanhua Ye, Lei Shi, Kangkang Lu · 17d
ToolVerse: Unlocking Massive Environments and Long-Horizon Tasks for Agentic Reinforcement Learning
arXiv cs.AI · Shuaiyu Zhou, Fengpeng Yue, Zengjie Hu, Yuanzhe Shen, Chenyang Zhang, feng hong, Cao Liu, Ke Zeng · 17d
Behavioral Controllability of Agentic Models for Information Extraction: From Fixed Workflows to Reflective Agents
arXiv cs.AI · Lujia Zhang, Xingzhou Chen, Hongwei Feng · 17d
NeurOWL: An LLM-Based Neural-symbolic Framework for Incomplete OWL Ontology Reasoning
arXiv cs.AI · Hui Yang, Jiaoyan Chen, Yiping Song, Renate Schmidt, Wen Zhang · 17d
AgentFAIR: A Multi-Agent Collaborative Framework for FAIRness Evaluation of Geospatial Datasets
arXiv cs.AI · Ming Chen, Pranav Pai · 17d
Knowledge-Centric Agents for Workflow Generation
arXiv cs.AI · Zhendong Li, Lei Sun, Ruibo Ming, He Zhang, Danda Pani Paudel, Luc Van Gool, Jinjin Gu · 17d
DSWorld: A Data Science World Model for Efficient Autonomous Agents
arXiv cs.AI · Zherui Yang, Fan Liu, Hao Liu · 17d
A Formally Grounded ODRL Evaluator: Implementation and Comparison
arXiv cs.AI · Jaime Osvaldo Salas, Paolo Pareti, Adeel Aslam, Christopher Maidens, George Konstantinidis · 17d
SciForge: An AI-Native, Multimodal Workbench for Scientific Discovery
arXiv cs.AI · SciForge Team, Zhangyang Gao, Minghao Fang, Yifei Liu, Hanhui Yang, Xinyu Gu, Shixiang Tang, Siqi Sun, Lei Bai, Cheng Tan, Mengdi Liu, Hao Wu, Shuizhou Chen · 17d
Harmonizing AI Safety Thresholds
arXiv cs.AI · Wilber Sean Anterola, Matthew Ball, Luis F. Lafuerza, Markov Grey · 17d
CRAFT: Clustering Rubrics to Diagnose Weak LLM Capabilities and Generate Targeted Fine-Tuning Data
arXiv cs.AI · Vipul Gupta, Zihao Wang, Razvan-Gabriel Dumitru, MohammadHossein Rezaei, Aakash Sabharwal, Yunzhong He · 17d
AV-JEPA: Extending LeJEPA to Audio-Visual Self-Supervised Learning
arXiv cs.AI · Benjamin Robson, Santeri Mentu, Wenshuai Zhao, Arno Solin · 17d
Coercion and Deception in AI-to-AI Management: An Agentic Benchmark of Unprompted Escalation
arXiv cs.AI · Jasmine Brazilek, Maheep Chaudhary, Zoe Lu, Miles Tidmarsh · 17d
Design-Based Supervised Learning with Noisy Human Labels
arXiv cs.AI · Robert Chew, Matthew R. Williams · 17d
LLM-Driven AutoML for Cross-Lingual Handwritten OCR: Closed-Loop Neural Architecture Search with GPT-5, GPT-4o, and Claude Sonnet 4
arXiv cs.AI · Mobina Kashaniyan, Amirhossein Ghassemi, Nasser Mozayani · 17d
Evolutionary Algorithm-Guided LLMs for Physics-Informed Neural Network Design
arXiv cs.AI · Xu Yang, Mingyang Yu, Jing Xu, Keqian Li · 17d
Ask Twice, Look Twice: Prompt Echoing Resolves the Question-First Paradox in Vision-Language Models
arXiv cs.AI · Rakshanda Hassan Abhinandan, John Galeotti, Deva Ramanan, Gautam Rajendrakumar Gare · 17d