←── back to feed
/topics/arxiv-ai-agent-research-papers-august-5
arXiv AI agent research papers August 5
30 items●1 sources●updated 1d ago●trend 1
On August 5, 2026, arXiv published 20 papers on AI agent research spanning medical reasoning, time-series forecasting, CAD generation, tool-use parameter optimization, multi-agent economics, reinforcement learning, memory management, and adversarial testing. The papers address core agent challenges: grounding decisions in evidence, handling long-horizon interactions, maintaining consistency under pressure, and improving reasoning beyond average-case performance.
- MedPIC-Bench evaluates whether models use patient-specific information to apply medication-safety rules correctly, not just recall drug-risk associations.
- CastFSR framework adds explicit context identification and temporal constraint validation to LLM-based time-series forecasting.
- TraceCAD preserves repair history and failed outcomes as persistent state to improve LLM-based CAD agent correction loops.
- OM-GRPO decouples reward estimation from policy optimization in label-free reinforcement learning to prevent answer-token reinforcement shortcuts.
- TumorBoard multi-agent system coordinates radiology, pathology, and molecular diagnosis with auditable claim-evidence ledgers and adversarial critic oversight.
- The Agent Operating System (AOS) proposes a reference architecture for distributed agentic systems independent of specific framework implementations.
[BLG]blog/rss30
Evaluating Counterfactual Sensitivity to Patient Information in Medication-Safety Reasoning
CastFSR: A Fast--Slow--Reflect Agentic Reasoning Framework for Context-Aware Time Series Forecasting
TraceCAD: Trace-Guided Repair for Agentic CAD Generation
Getting the Parameters Right: A Difficulty-Graded Benchmark and Probe-Guided Training for LLM Tool Calls
AI Agent Economics: Can Autonomous Economic Behavior Emerge among AI Agents under Minimal External Conditions?
Don't Peek at the Answer: Outcome-Masked Group Relative Policy Optimization for Label-Free RLVR
Beyond Average Performance: Dynamic Instance Clustering and Specialized Algorithm Design in LLM-Assisted Evolutionary Search
Verifiable Memory: Learning Unified Memory Management with Local and Global Verifiers for Large Language Model Agents
Spatial proteomics guided by H&E-based AI reveals recurrence-risk niches in triple-negative breast cancer
UniGD: A Unified Generative-Discriminative Framework for Industrial Retrieval
Evidence-Grounded Multimodal Knowledge Graph Construction for Multi-Lecture Educational Reasoning
Adversarial Stress Testing of Role-Playing Language Agents using Multi-Agent Evaluation
Surrogate Substitution Preserves PHI Detectability: A Multi-Detector Equivalence Study
Diversity is Not Ambiguity: Toward Accurate and Efficient Ambiguity Detection for Open-Domain QA
TumorBoard: Evidence-Grounded Multi-Agent Decision Support for Longitudinal Neuro-Oncology
When Refusal Looks Safe: The Refusal-Cue Shortcut in Safety Guard Models
The Agent Operating System (AOS): A Reference Operating Architecture for Distributed Agentic Systems
Reachability Is Not Realization: Tracing the Sources of LLM Benchmark Gains
UniNav: A Unified World-Action Diffusion Model for Visual Navigation
One Knob to Rule Them All: A Unified Optimal Transport View of Cold-Start Active Learning
TaskPress: Query-Agnostic KV Cache Compression via Task-Guided Pruning
AgentPanel: Toward a New Paradigm for Human--AI Collaboration in Exploring Scientific Questions
DocTrace: Towards Traceable Long Document VQA via Hierarchical Evidence Graph Reasoning
Distractor-Aware Truncation: Disentangling Context-Length Effects from Signal Loss in Long-Context LLM Benchmarks
SeaSlides: Semantic Abstraction Layer for Agentic Slide Generation
Screenshots or Tools? Eliciting Tool Use and Managing Multimodal Context in Hybrid GUI-MCP Computer-Use Agents
Long-term Traffic Scene Prediction via Polynomial Representations in Autonomous Driving
Traceable Multi-Agent System for Knowledge-Based Forecasting
MMLongBench-Doc-V2: A Corrected-Annotation, Semantics-Aware Revision of MMLongBench-Doc
Towards Robust Tool Use in Agents via Experience-Driven Adaptive Guidance