←── back to feed
/topics/arxiv-ai-and-ml-research-papers-september-3

arXiv AI and ML research papers September 3

75 items1 sourcesupdated 18d agotrend 0

On September 3, 2026, arXiv published 20 AI and ML research papers spanning evaluation robustness, agent architectures, model compression, and domain-specific applications. Key topics include measuring evaluation awareness in frontier LLMs, persistent-memory agent failures, post-training ternarization of Qwen3-4B, and benchmarks for statistical problem formulation and multi-hop document reasoning.

  • EvalDetectBench introduced to measure evaluation awareness in frontier language models using Inspect-compatible evaluations
  • Memory Trust Gap study on Qwen3 (0.6B–8B) shows stale facts override current evidence without warning as model capability increases
  • Qwen3-4B ternarized to 1.58-bit using KOTMS rotation and E2M-ATQ with 16-bit activations; effective bit accounting and perplexity evaluated
  • DocHop benchmark tests multimodal LLMs on integrated chart–context reasoning across information-dense documents
  • HeadWiseKV framework compresses residual global KV caches in hybrid language models without training, reducing long-context inference memory
[BLG]blog/rss75
EvalDetectBench: A Benchmark for Measuring Evaluation Awareness in Frontier Language Models
arXiv cs.AI · Xinning Li, Kemunto Ochwang'i, Aryasomayajula Ram Bharadwaj, Alexandra Souly, Robert Kirk · 18d
Meta-ethics and AI: exploring the novel meta-ethical questions in the era of AI
arXiv cs.AI · Shang Lu · 18d
When Can a Machine Trust a Statute? A Survival Certificate for Machine-Extracted Legal Logic
arXiv cs.AI · Surya Saka · 18d
When Does Information Sharing Improve Decentralized Discovery? Aggregation, Independent Rescue, and Equilibrium Selection
arXiv cs.AI · Yohei Nakajima · 18d
Induction and Inquiry via Probabilistic Reasoning over Language and Code
arXiv cs.AI · Wasu Top Piriyakulkij, Sam Acquaviva, Cassidy Langenfeld, Joshua Tenenbaum, Kevin Ellis · 18d
Architecting Conversational Data Systems for Stateless LLM APIs: The Hydration Proxy Pattern
arXiv cs.AI · Joseph Axisa · 18d
SSAKG 2.0: An Open-Source Package for Structural Associative Sequence Memory and Context-Based Retrieval
arXiv cs.AI · Przemys{\l}aw Stok{\l}osa, Janusz A. Starzyk, Pawe{\l} Raif · 18d
The Memory Trust Gap: Capability-Dependent Failures in Persistent-Memory Agents
arXiv cs.AI · Jundong Hu, Shekar Ramachandran · 18d
Belief-Calibrated Optimization: An Explicit World Model for Agentic Optimization
arXiv cs.AI · Yuhan Chen, Zhihua Tian, Mahavir Dabas, Charith Peris, Rahul Gupta, Ming Jin, Feiyang Kang, Siyuan Zhang, Nan Wang, Ruoxi Jia · 18d
Epistemic Sybil Resistance: Multiplying AI Agents Without Multiplying Evidence
arXiv cs.AI · Marc Bara · 18d
The Ceiling Is in the Channel: Auditing Learner Gaps and Measurement Frontiers in Clinical Prediction
arXiv cs.AI · Sayeed Shafayet Chowdhury, Nusrat Jahan, Snehasis Mukhopadhyay, Shiaofen Fang, Vijay R. Ramakrishnan · 18d
Looped Transformers under the Jacobian Lens: Does the Global Workspace Survive Recurrence?
arXiv cs.AI · Wenlong Wang, Fergal Reid · 18d
Post-Training Ternarization of Qwen3-4B Capability, Effective Bit Budget, Storage Compression, and Deployment
arXiv cs.AI · Anirudh Malik, M Sparsh Mehra, Poojith Devan · 18d
Benchmarking Language Models for Statistical Problem Formulation
arXiv cs.AI · Chen Wang, Junzhe Zhao, Xin Cong, Wanlu Deng, Ke Deng · 18d
When Agents Implement Systems: A Case Study in Defects, Detection, and Evaluation Rigor
arXiv cs.AI · Phanindra Reddy Madduru · 18d
ClaimReceipt: Verifying Evidence Sufficiency and Coverage in Agent Evaluations
arXiv cs.AI · Peiying Zhu, Sidi Chang · 18d
HeadWiseKV: Budgeted Per-Head Cache Residency for Hybrid Long-Context Language Models
arXiv cs.AI · Renjie Xie, Juncheng Yang, Aoting Hu, Mingxi Zhang, Liyao Wu, Zheheng Hong, Wei Xu · 18d
Monitoring Web Agents Without Internal Signals: Observable Trajectories and Key-Step Supervision
arXiv cs.AI · Sitong Pan, Yipeng Shen, Yilin Lu, Caiwen Ding, Lu Cheng, Qianwen Wang · 18d
DocHop: Benchmarking Out-of-domain Multi-hop Reasoning in Information-Dense Documents
arXiv cs.AI · Zhuoran Yu, Le Thien Phuc Nguyen, Jaden Park, Xinyi Gu, Zexue He, Soochahn Lee, Rogerio Feris, Yong Jae Lee · 18d
MineTRACE: An Evidence-Grounded Interactive Reasoning System for Mineral Prospectivity
arXiv cs.AI · Yiran Zhang, Jinwen Liu, Daniel Su, Yisu Chen, Qiang Sun, Chris Gonzalez, Eun-Jung Holden, Marco Fiorentini, Wei Liu, Yihao Ding · 18d
ToolGate: An Executable Acceptance Pipeline for Tool-Dependent Scientific Benchmark Construction
arXiv cs.AI · Ke Zhang, Yankang Liu, Roya Zandi, Maziar Raissi · 18d
CHIME: Credit-Aware Hierarchical Memory Evolution for Long-Horizon Agentic Planning
arXiv cs.AI · Yongshi Ye, Tian Lan, Feihu Jiang, Muyang Ye, Bin Zhu, Qianghuai Jia, Longyue Wang, Zhao Xu, Weihua Luo, Xiaodong Shi · 18d
Beyond Outcome Gaps: Process-Aware Fairness Diagnosis for LLM-based Multi-Agent Decision Systems
arXiv cs.AI · Yiran Zhao, Lu Zhou, Liming Fang, Yufei Chen, Jiafei Wu, Zhe Liu, Xiaogang Xu · 18d
MASkills: Continual Skills Optimization for Multi-Agent LLM Systems
arXiv cs.AI · Huaiyuan Yao, Xiaoou Liu, Charles Fleming, Tianlong Chen, Hua Wei · 18d
READY or Not: Reliable Enterprise Agent Deployment
arXiv cs.AI · Veronica Chatrath (Christy), Bryan Zhu (Christy), Jingxuan Fan (Christy), George Pu (Christy), Soham Dinesh Tiwari (Christy), Soham Dan (Christy), Ryan Young (Christy), Yuan (Christy), Li, Yuang Yao, Apaar Shanker, Minglai Yang, Daniel Yue Zhang, Yunzhong He, Ying Liu, Chenguang Wang, Zhijun Yin, Yuan Xue · 18d
Semantic Signal-Assisted Inspection and Recovery Allocation in Reverse Logistics
arXiv cs.AI · Jiani He, Dingyan Shang, Yihua Xu, Shiqi Huang, Yan Lyu, Jize Li, Shangjing Tang · 18d
Beyond Context Windows: Persistent Discovery Context for Data-Centric Agents
arXiv cs.AI · Jalal Mahmud · 18d
EmoStance: Response-Side Affective-Orientation Control for Empathetic Response Generation via Emoji Weak Supervision
arXiv cs.AI · Ziyuan Jin, Yuxuan Ge, Zheng Tian · 18d
FUSE: An Evaluating Framework for Dangerous Capabilities of LLMs
arXiv cs.AI · Zhengyi Jin, Ru Zhang, Xiao Chen, Xinbo Liu, Jiaxuan Lin, Jia Huang, Jianyi Liu, Zhen Yang · 18d
Examining the Vulnerability of Multi-Agent Medical Systems to Human Interventions for Clinical Reasoning
arXiv cs.AI · Benjamin C Liu, Dillon Mehta, Rishi Malhotra, Adam Zobian, Yong Ying Tan, Samir Chopra, Daniella Rand, Natalie Pang, Abhiram Gudimella, Kevin Zhu · 18d
ASCII Attack: Recontextualising Harmful Requests as Artistic Critique in Large Language Models
arXiv cs.AI · Da Cheng Gu, Yifei Dong, Xinghao Yang, Yongshun Gong, Wei Liu · 18d
PEARL: Path-Entity Aligned Relational Learning with Contextual Subgraphs for Inductive Knowledge Graph Completion
arXiv cs.AI · Yunchi Yang, Longlong Li, Cunquan Qu · 18d
SkillGLoW: Procedural-Family Skill Consolidation for Self-Improving Agents on Long-Horizon Task Streams
arXiv cs.AI · Ao Yan, Xin Zhang, Jiawei Du, Joey Tianyi Zhou · 18d
PhoenixNest-Video: Evidence-Grounded Multimodal Agent Framework for Automated Video Interview Assessment
arXiv cs.AI · Fan Yuxuan, Huang Miaojun, Zhang Haimei, Wu Jingshen, Liu Hao · 18d
PGPO: Potential-Guided Policy Optimization for Multi-Turn Agentic Tasks
arXiv cs.AI · Yuyao Zheng, Haipeng Sun, Junwei Bao, Lemao Liu, Hongfei Jiang, Yang Song, Dejing Dou · 18d
Propose to Learn, Learn to Propose: Evaluability-Aware Assistance under Bounded Rationality
arXiv cs.AI · Yifan Zhu, Sammie Katt, Samuel Kaski · 18d
Task-Level Natural Language Priors as Learning Signals for Low-Resource LLM Training
arXiv cs.AI · Jian Gao, Xiao Zhang, Xun Zhu, Miao Li, Ji Wu · 18d
LLM-as-a-Judge Is Not an Oracle: Why Self-Improving Agents Need Deterministic Guardrails
arXiv cs.AI · Vansh Wahi · 18d
APEx: Distillation of Agent Procedural Experience for Adaptive Deep Research Question Answering
arXiv cs.AI · Jie Ding, Rui Sun, Xinyuan Zhang, Zeyu Zhang, Xin Liu · 18d
Codebook Agent: Amortized Topology Design for LLM Multi-Agent Systems
arXiv cs.AI · Jinxi Yu, Yubei Li, Eric Hanchen Jiang, Zhi Zhang, Dong Liu, Wenxiao Zhao, Levina Li, Kai-Wei Chang, Ying Nian Wu · 18d
CoMerge: Conflict-Driven Preference Optimization for Multi-Task Model Merging
arXiv cs.AI · Mingjie Zheng, Zihao Chen, Wenqing Chen, Weile Yuan, Zhixuan Chu, Jianxing Yu, Zibin Zheng · 18d
SCX Router: Streaming Zero-Shot Model Selection with a Decoder-KV Classifier and a Real-World Task Ontology
arXiv cs.AI · Ihor Stepanov, Aleksandr Smechov, Mykhailo Shtopko, Dmytro Vodianytskyi, Oleksandr Lukashov · 18d
Improving Evaluation Realism with Inference-Time Compute and Deployment Scaffolds
arXiv cs.AI · Axel Ahlqvist, Richard Guan, Juan-Pablo Rivera, Adeline Kassler, Dmitrii Troitskii, Alexandra Souly, Kai Fronsdal, Robert Kirk, John Hughes · 18d
SALA: Semantic-Aware Logical Alignment for Complex Reasoning in In-Context Learning
arXiv cs.AI · Zhao Ji, Wenqing Chen, Zhixuan Chu, Jianxing Yu, Jingping Liu, Shanhe Zhao, Zibin Zheng · 18d
Diagnosing with Insights: Structured Analysis of Agent Failures via Behavioral Abstractions
arXiv cs.AI · Jiayi Bi, Yanjie Gao, Yuanmin Xie, Liqun Li, Tianyin Xu, Fan Yang, Mao Yang · 18d
Contrastive Explanations in Quantitative Bipolar Argumentation Frameworks
arXiv cs.AI · Xiang Yin, Nico Potyka, Antonio Rago, Francesca Toni · 18d
UTP-Bench: Uncertainty-aware Travel Planning Benchmark
arXiv cs.AI · Etcharla Revanth Rao, Priyanshu Karmakar, Shubhojit Mallick, Manish Gupta, Shreya Ghosh, Abhik Jana · 18d
CivBench: A Long-Horizon Benchmark for Tool-Mediated Agents in Civilization VI
arXiv cs.AI · Austin Tudor David Andrews, Liam Wilkinson, Jamie Heagerty, Harry Coppock, Jakob Nicolaus Foerster, Rui Ponte Costa · 18d
Collective creativity in hybrid societies
arXiv cs.AI · Mason Youngblood, Katie Mudd, Manuel Anglada-Tort, Cameron Jones, Elena Miu, Diana Omigie, Margaret Schedel · 18d
Loom: Weaving Diagnostic Strands into Free-Text Consensus via Embedding-Space Reweighting
arXiv cs.AI · Ron Begleiter, Katya Egert Berg, Gilad Saban, Gil Shabat · 18d
PRO-Step: Step-level Process Reward Optimization for Retrieval-Augmented Generation
arXiv cs.CL · MinKeon Kim, Namjun Lee, Jaekwang Kim · 18d
Learning Evidence Sufficiency Boundaries for Selective Answering in Grounded Multi-Hop QA
arXiv cs.CL · Haruto Sato, Yuki Tanaka, Ren Nakamura, Aoi Kobayashi, Mei Ito · 18d
SpeakPay: Domain-Adaptive LoRA Fine-Tuning of Whisper for Low-Resource Nepali Financial Speech Recognition
arXiv cs.CL · Biraj Subedi · 18d
MemeCULT-1K: Benchmarking South Asian Cultural Context and Humor Understanding of Multimodal Models
arXiv cs.CL · Tawsif Tashwar Dipto, Mehedi Ahamed, Radib Bin Kabir, Mueeze Al Mushabbir, Mohammed Saidul Islam, Mir Rayat Imtiaz Hossain, Md Tahmid Rahman Laskar, Sabbir Ahmed · 18d
VakyArth: Evaluating Pragmatic Competence in LLMs across Indic Languages
arXiv cs.CL · Usneek Singh, Poorvaja Veera Balaji Kumar, Parth Nanda, Anand Madhusoodanan, Geyang Guo, Wei Xu, Junyi Jessy L · 18d
Disentangling Statistical Preemption from Entrenchment in Language Models' Avoidance of Overgeneralization
arXiv cs.CL · Yixuan Wang, Freda Shi, Kanishka Misra · 18d
How Do Prompt Variations Affect Energy Consumption in On-Device LLMs?
arXiv cs.CL · Wei Hu, Xiaolong Tu, Dawei Chen, Yitao Chen, Kyungtae Han, Haoxin Wang · 18d
TalkFa: A Unified Benchmark for Farsi Dialogue Generation and Understanding
arXiv cs.CL · Neda Jamshidi, Kamyar Zeinalipour, Fahimeh Akbari, Monica Bianchini, Marco Maggini, Marco Gori · 18d
AVERT: Audio-Verified Adjudication for Spoken Dialogue State Tracking
arXiv cs.CL · Chunggi Lee, Hanspeter Pfister · 18d
Interpretable Symptom Vectors for Depression in a Large Language Model
arXiv cs.CL · Fangyi Zhu, Ajay Subramanian, Allison Constant, Camille Wang, Ravish Gupta, Corey J. Keller · 18d
Candidate Generation and Definition-Guided Verification for Sentence-Level Depression Symptom Recognition
arXiv cs.CL · Weiming Li, Catarina Barata, Miguel Constante, Joao Sanches · 18d
Cite or Decline: A Strict Course-Grounded Chatbot for STEM Lecture Videos
arXiv cs.CL · S M Masrur Ahmed, Jaspal Subhlok · 18d
Thinking effort aligns between humans and reasoning models in abductive reasoning
arXiv cs.CL · Henry Arthur · 18d
GAPS: Dimension-Level Gates for Conditional Activation Steering
arXiv cs.CL · Moghis Fereidouni, Muhammad Umair Haider, Hassan Sajjad, A. B. Siddique · 18d
Grounded, Compute-Efficient LLM Policy Agents for Energy-Poverty Equity in Physically-Constrained Peer-to-Peer Energy Markets
arXiv cs.CL · Kunal Jadhav, Siddhesh More · 18d
Sparse Readout Prism: Explaining Logit-Lens Scores in Features Instead of Tokens
arXiv cs.CL · Matteo He, William F. Shen, Xinchi Qiu, Nicholas D. Lane · 18d
NS-Copilot: An LLM-Driven Agent System for Autonomous Neuroscience Analysis
arXiv cs.CL · Wuche Liu, Yiran Qiao, Linlin Hou, Rui Yang, Shusen Pu, Song Wang, Jing Ma · 18d
How Output Format Confounds Data Quality and Capability in Instruction Tuning
arXiv cs.CL · Chengguang Gan, Hanjun Wei, Yunhao Liang, Qinghao Zhang, Shiwen Ni, Zhixi Cai · 18d
A Tri-Agent Framework for Evaluating and Aligning Question Clarification Capabilities of Large Language Models
arXiv cs.CL · Yikai Zhao, Saurabh Pandey, Pradeep Kumar Misra · 18d
HyGRAIL: Cost-Aware and Evidence-Grounded Scientific Hypothesis Discovery over Knowledge Graphs
arXiv cs.CL · Yihang Sun, Zhihan Zhu, Zhiyuan Jiang, Jingyi Ge, Zixuan Li, Jiaxuan You · 18d
IDEEA: training-free Input-Dependent stEEring via Activation cluster matching
arXiv cs.CL · Zheng Wang, Muchen Li, Renjie Liao, Yan Leng · 18d
Selective Knowledge Edit Reversal via Gated Singular Vector Shrinkage
arXiv cs.CL · Weifeng Jiang, Ruirui Chen, Qianren Mao, Junnan Liu, Qili Zhang, Kwok-Yan Lam · 18d
Predict, Don't Iterate: Efficient Adaptive-Length Infilling for Diffusion Language Models
arXiv cs.CL · Haobo Xu, Sirui Chen, Yuanchen Bei, Lingjie Chen, Yuchen Yan, Dongqi Fu, Jingrui He, Hanghang Tong · 18d
text2ql: Multi-Target Natural Language Querying via a Language-Agnostic Intermediate Representation
arXiv cs.CL · Ritesh Kumar · 18d
AI agents reshape consensus formation in human groups
arXiv cs.CL · Lin Chen, Ziyi Liu, Xia Hu, Yong Li · 18d