←── back to feed
/topics/arxiv-ai-agent-systems-papers-july-16
arXiv AI agent systems papers July 16
71 items●1 sources●updated 21d ago●trend 0
On July 16, 2026, arXiv published 20 papers on AI agent systems spanning data provenance, robot deployment, reasoning verification, memory management, safety, and human-AI collaboration. Topics range from technical frameworks for agent harnesses and self-improvement to practical challenges like root-cause analysis, insurance underwriting, and user reliance on AI advice.
- OriginBlame enables record- and token-level data provenance to resolve author removal requests into precise forget sets for model unlearning.
- SPINE framework automates bimanual robot deployment with multi-agent workflows, reducing expert calibration overhead for embodied AI.
- Interventional grounding audits test whether LLM chain-of-thought reasoning genuinely depends on stated premises via predicate substitution.
- Survey on self-improving agents frames modern systems as foundation models coupled with prompts, memory, tools, and control logic.
- AI-native insurance framework maps agent risk states (autonomy, authority, governance) to event probabilities and loss severities for agentic deployments.
- Study finds AI advice suppresses humans' willingness to say 'I don't know' even when advice is wrong and accuracy is incentivized (N=3,132).
[BLG]blog/rss71
OriginBlame: Record- and Token-Level Data Provenance for AI Training Datasets
SPINE: Bridging the Cyber-Physical Gap with Agentic AI
Interventional Grounding Audits: Black-Box Premise-Dependency Tests for LLM Chain-of-Thought via Predicate Substitution
Probabilistic Extension of Neuro-Symbolic AGI Robots based on Belnap's Typed Intensional FOL
Self-Improvements in Modern Agentic Systems: A Survey
Improving Molecular Property Prediction in Small Language Models Using Graph-based Tools
Oracle Agent Memory as an Enterprise Memory Substrate for Long-Horizon AI Agents
Learning Safe Agent Behaviour from Human Preferences and Justifications via World Models
CayleyR: Solving the TopSpin puzzle via cycle intersection
Networked Intelligence: Active Shared Context Graphs for Human-AI Team Science
AI-Native Insurance for Agentic AI: Pricing, Underwriting, and End-to-End Automation
Cost-Optimal Foundation Model Deployment Portfolio for Transportation Management
Harness Handbook: Making Evolving Agent Harnesses Readable,Navigable, and Editable
Theory-Level Autoformalization: From Isolated Statements to Unified Formal Knowledge Bases
EZSMT Version 3, Matured
Set-shifting Behavioral Test for Harnessed Agents
LAPO: Leave-One-Turn Attribution for Self-Generated Process Rewards in Multi-Turn Search Reasoning
How Far Can Root Cause Analysis Go on Real-World Telemetry Data?
Multi-Agent Collaborative Reasoning with Tool-Augmented Evidence for Urban Region Profiling
AI advice suppresses people's willingness to say "I don't know", even when the advice is wrong and accuracy is incentivized
SAFETY SENTRY: Context-Aware Human Intervention via EXECUTE-ASK-REFUSE Routing
Automatic Ordinary Differential Equations Discovery For Biological Systems Using Large Language Model Powered Agentic System
STOCKTAKE: Measuring the Gap Between Perception and Action in LLM Agents with a Fair Oracle
UESF-Bench: Benchmarking and Probing for Unified Embodied Seeking and Following
Explaining Reinforcement Learning Agents via Inductive Logic Programming
When Bots Join the Team: Bot Adoption and the Institutional Fabric of Open-Source Software Projects
AgentCompass: A Unified Evaluation Infrastructure for Agent Capabilities
CAVA: Canonical Action Verification and Attestation for Runtime Governance of Agentic AI Systems
Experience Memory Graph: One-Shot Error Correction for Agents
AIMO Interpretability Challenge
A Self-Evolving Agent for Longitudinal Personal Health Management
Do Agent Optimizers Compound? A Continual-Learning Evaluation on Terminal-Bench 2.0
AI-accelerated End-to-End Framework for Rapid Professional Upskilling
Earthquaker-AI: A Retrieval-Augmented Generation Framework with Rubric-Based Assessment for Primary School Earthquake Education
Deep Interaction: An Efficient Human-AI Interaction Method for Large Reasoning Models
FixItFlow: Automated Troubleshooting Guide Generation from Cloud Incidents
Ask Before You Diagnose: Safe-Psych, a Sequential Evaluation Benchmark for LLMs in Psychiatry
Designing Safety-Constrained LLM Systems for Public Health Information Access
Safeguard-Conditioned Uplift: Measuring Utility-Risk Frontiers for Dual-Use Biology Assistants
Final Authority in AI Governance: Frontier-Provider Sovereignty and Action-Centered Deployer Governance
LessonBench-V1: A Benchmark Dataset for Evaluating AI Lesson Generation Agents
Beyond Backbone Backpropagation: A Decoupled Strategy for Efficient Transfer Learning
The Perplexity Trap: When Patent Law Makes Human Writing Look Like AI
Federated Explainable Artificial Intelligence: Roles, Architectures, Evaluation, and Open Challenges
Uncertainty-Aware Sequential Decision Rules for Event-Triggered LLM Invocation in Streaming Systems
Autonomous UAV Route Planning for Coverage Maximization in Environmental Monitoring: A Systematic Literature Review
Compaction as Epistemic Failure: How Agentic LLM Tools Fabricate Confirmed Results from Killed Processes
HRO: Hierarchical Room-to-Object Framework for Zero-Shot Object Goal Navigation with Large Language Models
When is the combined load identifiable from a stress-intensity profile? A coupled forward-inverse study on SIFBench finite-element data
The Entanglement Wall: Activation-Space Probes as Risk Detectors, Not Context Adjudicators
Automatic Differentiation from Scratch: How PyTorch Computes Gradients in Physics-Informed Neural Networks
What Your Model Threw Away and Why You'll Want It Back: Masking, Fingerprinting, and Privacy from Discarded Geometry
Targeted Recovery of Weight-Space Mechanisms From Neural Networks
TSSM: Triaxial State Space Model for Global Station Weather Forecasting with Temporal-Variable-Historical Modeling
Disentangling Knowledge States with Ability and Proficiency Modeling for Knowledge Tracing
STKAN: Kolmogorov-Arnold Networks for Spatio-Temporal Forecasting
A Hybrid Mamba for Audio-Visual Navigation
CoDiffGRN: Rethinking Gene Regulatory Network Inference via the BEELINE-KGC Benchmark and Co-evolutionary Discrete Diffusion
ShortOPD: Recovering Pruned LLMs with Short-to-Long On-Policy Distillation
HEDGEHOG: Hierarchical Evaluation of Drug Generators Through Rigorous Filtration
SteinGate: Tail-Sensitive Safe Reinforcement Learning via Stein Discrepancy
Concurrent Image Understanding and Generation: Self-Correcting Coupled Markov Jump Processes
EMAGN: Efficient Multi-Attention Graph Network via Learned Clustering for Scalable Traffic Forecasting
Reassessing Muon for Matrix Factorization
Deconstructing Actor-Critic: A Large-scale Empirical Study of Design Components for Practitioners
Tabular Foundation Models for Discrete Choice Estimation
Accuracy-Preserving Stability Regularization for Large-Scale Retail Demand Forecasting
Agora: Collective and Permissionless Internet-Scale Pretraining of Large Language Models
Weight Feedback Computes the Jacobian Transpose Locally in Modern Deep Networks
Where Should RL Post-Training Compute Go? Model Size, Search, Learning, and Feedback
Self-Improving is Often Sudden: Enlightenment-style Finetuning for Large-Scale Models