←── back to feed
/topics/arxiv-ai-agent-systems-and-reasoning-papers-july-20
arXiv AI agent systems and reasoning papers July 20
26 items●1 sources●updated 17d ago●trend 0
On July 20, 2026, arXiv released 20 papers on AI agent systems and reasoning, spanning medical diagnosis frameworks, causal reasoning auditing, healthcare-specialized models, local voice assistants, multi-agent math reasoning, coding agents, reinforcement learning explainability, mobile GUI safety, knowledge graph completion, and workflow generation. The papers address core challenges in agent design: balancing cost and accuracy, ensuring transparency and auditability, scaling to complex environments, and improving controllability and safety.
- GraphDx uses Medical Diagnosis Knowledge Graphs to balance diagnostic accuracy against resource costs in sequential diagnosis.
- Causal-Audit proposes explicit, auditable graph-based reasoning for context-free causal question answering with verifiable reasoning paths.
- Cura 1T is a healthcare-specialized LLM trained via human-gated self-evolution for patient consultation, clinical reasoning, and EHR tool use.
- AnovaX is a local-first desktop voice assistant running entirely on-device with LLM planning, typed executors, and adaptive recovery.
- Reviewer precision alone does not guarantee critique uptake; broadcast-style peer discussion outperforms planner-executor-reviewer pipelines on harder math problems.
- DSWorld introduces Data Science World Models to predict environment state transitions and reduce trial-and-error in autonomous data science workflows.
[BLG]blog/rss26
GraphDx: A Cost-Aware Knowledge-Enhanced Multi-Agent Framework for Sequential Diagnosis
Causal-Audit: Explicit and Auditable Graph-based Reasoning via Target-Aware Causal Chain Construction
Cura 1T: Specialized Model for Agentic Healthcare
AnovaX: A Local, Multi-Agent Voice Assistant with LLM Planning, Typed Executors, and Adaptive Recovery
Precise but Uncoupled: Reviewer Precision Does Not Guarantee Critique Uptake in Multi-Agent Math Reasoning
Do Coding Agents Need Executable World Models, Simplification, and Verification to Solve ARC-AGI-3?
From Black Box to Executable Logic: Explainable Reinforcement Learning through Prolog Expert Systems
A Critical Analysis of Trustworthy AI Tools, Mark Frameworks, and the Implementation Chasms
SeerGuard: A Safety Framework for Mobile GUI Agents via World Model Prediction
MGDT: MLLM-Guided Diffusion Transformer with Relation-Adaptive Mixture-of-Experts for Multimodal Knowledge Graph Completion
ToolVerse: Unlocking Massive Environments and Long-Horizon Tasks for Agentic Reinforcement Learning
Behavioral Controllability of Agentic Models for Information Extraction: From Fixed Workflows to Reflective Agents
NeurOWL: An LLM-Based Neural-symbolic Framework for Incomplete OWL Ontology Reasoning
AgentFAIR: A Multi-Agent Collaborative Framework for FAIRness Evaluation of Geospatial Datasets
Knowledge-Centric Agents for Workflow Generation
DSWorld: A Data Science World Model for Efficient Autonomous Agents
A Formally Grounded ODRL Evaluator: Implementation and Comparison
SciForge: An AI-Native, Multimodal Workbench for Scientific Discovery
Harmonizing AI Safety Thresholds
CRAFT: Clustering Rubrics to Diagnose Weak LLM Capabilities and Generate Targeted Fine-Tuning Data
AV-JEPA: Extending LeJEPA to Audio-Visual Self-Supervised Learning
Coercion and Deception in AI-to-AI Management: An Agentic Benchmark of Unprompted Escalation
Design-Based Supervised Learning with Noisy Human Labels
LLM-Driven AutoML for Cross-Lingual Handwritten OCR: Closed-Loop Neural Architecture Search with GPT-5, GPT-4o, and Claude Sonnet 4
Evolutionary Algorithm-Guided LLMs for Physics-Informed Neural Network Design
Ask Twice, Look Twice: Prompt Echoing Resolves the Question-First Paradox in Vision-Language Models