←── back to feed
/topics/arxiv-ai-and-ml-research-papers-september-3
arXiv AI and ML research papers September 3
75 items●1 sources●updated 18d ago●trend 0
On September 3, 2026, arXiv published 20 AI and ML research papers spanning evaluation robustness, agent architectures, model compression, and domain-specific applications. Key topics include measuring evaluation awareness in frontier LLMs, persistent-memory agent failures, post-training ternarization of Qwen3-4B, and benchmarks for statistical problem formulation and multi-hop document reasoning.
- EvalDetectBench introduced to measure evaluation awareness in frontier language models using Inspect-compatible evaluations
- Memory Trust Gap study on Qwen3 (0.6B–8B) shows stale facts override current evidence without warning as model capability increases
- Qwen3-4B ternarized to 1.58-bit using KOTMS rotation and E2M-ATQ with 16-bit activations; effective bit accounting and perplexity evaluated
- DocHop benchmark tests multimodal LLMs on integrated chart–context reasoning across information-dense documents
- HeadWiseKV framework compresses residual global KV caches in hybrid language models without training, reducing long-context inference memory
[BLG]blog/rss75
EvalDetectBench: A Benchmark for Measuring Evaluation Awareness in Frontier Language Models
Meta-ethics and AI: exploring the novel meta-ethical questions in the era of AI
When Can a Machine Trust a Statute? A Survival Certificate for Machine-Extracted Legal Logic
When Does Information Sharing Improve Decentralized Discovery? Aggregation, Independent Rescue, and Equilibrium Selection
Induction and Inquiry via Probabilistic Reasoning over Language and Code
Architecting Conversational Data Systems for Stateless LLM APIs: The Hydration Proxy Pattern
SSAKG 2.0: An Open-Source Package for Structural Associative Sequence Memory and Context-Based Retrieval
The Memory Trust Gap: Capability-Dependent Failures in Persistent-Memory Agents
Belief-Calibrated Optimization: An Explicit World Model for Agentic Optimization
Epistemic Sybil Resistance: Multiplying AI Agents Without Multiplying Evidence
The Ceiling Is in the Channel: Auditing Learner Gaps and Measurement Frontiers in Clinical Prediction
Looped Transformers under the Jacobian Lens: Does the Global Workspace Survive Recurrence?
Post-Training Ternarization of Qwen3-4B Capability, Effective Bit Budget, Storage Compression, and Deployment
Benchmarking Language Models for Statistical Problem Formulation
When Agents Implement Systems: A Case Study in Defects, Detection, and Evaluation Rigor
ClaimReceipt: Verifying Evidence Sufficiency and Coverage in Agent Evaluations
HeadWiseKV: Budgeted Per-Head Cache Residency for Hybrid Long-Context Language Models
Monitoring Web Agents Without Internal Signals: Observable Trajectories and Key-Step Supervision
DocHop: Benchmarking Out-of-domain Multi-hop Reasoning in Information-Dense Documents
MineTRACE: An Evidence-Grounded Interactive Reasoning System for Mineral Prospectivity
ToolGate: An Executable Acceptance Pipeline for Tool-Dependent Scientific Benchmark Construction
CHIME: Credit-Aware Hierarchical Memory Evolution for Long-Horizon Agentic Planning
Beyond Outcome Gaps: Process-Aware Fairness Diagnosis for LLM-based Multi-Agent Decision Systems
MASkills: Continual Skills Optimization for Multi-Agent LLM Systems
READY or Not: Reliable Enterprise Agent Deployment
Semantic Signal-Assisted Inspection and Recovery Allocation in Reverse Logistics
Beyond Context Windows: Persistent Discovery Context for Data-Centric Agents
EmoStance: Response-Side Affective-Orientation Control for Empathetic Response Generation via Emoji Weak Supervision
FUSE: An Evaluating Framework for Dangerous Capabilities of LLMs
Examining the Vulnerability of Multi-Agent Medical Systems to Human Interventions for Clinical Reasoning
ASCII Attack: Recontextualising Harmful Requests as Artistic Critique in Large Language Models
PEARL: Path-Entity Aligned Relational Learning with Contextual Subgraphs for Inductive Knowledge Graph Completion
SkillGLoW: Procedural-Family Skill Consolidation for Self-Improving Agents on Long-Horizon Task Streams
PhoenixNest-Video: Evidence-Grounded Multimodal Agent Framework for Automated Video Interview Assessment
PGPO: Potential-Guided Policy Optimization for Multi-Turn Agentic Tasks
Propose to Learn, Learn to Propose: Evaluability-Aware Assistance under Bounded Rationality
Task-Level Natural Language Priors as Learning Signals for Low-Resource LLM Training
LLM-as-a-Judge Is Not an Oracle: Why Self-Improving Agents Need Deterministic Guardrails
APEx: Distillation of Agent Procedural Experience for Adaptive Deep Research Question Answering
Codebook Agent: Amortized Topology Design for LLM Multi-Agent Systems
CoMerge: Conflict-Driven Preference Optimization for Multi-Task Model Merging
SCX Router: Streaming Zero-Shot Model Selection with a Decoder-KV Classifier and a Real-World Task Ontology
Improving Evaluation Realism with Inference-Time Compute and Deployment Scaffolds
SALA: Semantic-Aware Logical Alignment for Complex Reasoning in In-Context Learning
Diagnosing with Insights: Structured Analysis of Agent Failures via Behavioral Abstractions
Contrastive Explanations in Quantitative Bipolar Argumentation Frameworks
UTP-Bench: Uncertainty-aware Travel Planning Benchmark
CivBench: A Long-Horizon Benchmark for Tool-Mediated Agents in Civilization VI
Collective creativity in hybrid societies
Loom: Weaving Diagnostic Strands into Free-Text Consensus via Embedding-Space Reweighting
PRO-Step: Step-level Process Reward Optimization for Retrieval-Augmented Generation
Learning Evidence Sufficiency Boundaries for Selective Answering in Grounded Multi-Hop QA
SpeakPay: Domain-Adaptive LoRA Fine-Tuning of Whisper for Low-Resource Nepali Financial Speech Recognition
MemeCULT-1K: Benchmarking South Asian Cultural Context and Humor Understanding of Multimodal Models
VakyArth: Evaluating Pragmatic Competence in LLMs across Indic Languages
Disentangling Statistical Preemption from Entrenchment in Language Models' Avoidance of Overgeneralization
How Do Prompt Variations Affect Energy Consumption in On-Device LLMs?
TalkFa: A Unified Benchmark for Farsi Dialogue Generation and Understanding
AVERT: Audio-Verified Adjudication for Spoken Dialogue State Tracking
Interpretable Symptom Vectors for Depression in a Large Language Model
Candidate Generation and Definition-Guided Verification for Sentence-Level Depression Symptom Recognition
Cite or Decline: A Strict Course-Grounded Chatbot for STEM Lecture Videos
Thinking effort aligns between humans and reasoning models in abductive reasoning
GAPS: Dimension-Level Gates for Conditional Activation Steering
Grounded, Compute-Efficient LLM Policy Agents for Energy-Poverty Equity in Physically-Constrained Peer-to-Peer Energy Markets
Sparse Readout Prism: Explaining Logit-Lens Scores in Features Instead of Tokens
NS-Copilot: An LLM-Driven Agent System for Autonomous Neuroscience Analysis
How Output Format Confounds Data Quality and Capability in Instruction Tuning
A Tri-Agent Framework for Evaluating and Aligning Question Clarification Capabilities of Large Language Models
HyGRAIL: Cost-Aware and Evidence-Grounded Scientific Hypothesis Discovery over Knowledge Graphs
IDEEA: training-free Input-Dependent stEEring via Activation cluster matching
Selective Knowledge Edit Reversal via Gated Singular Vector Shrinkage
Predict, Don't Iterate: Efficient Adaptive-Length Infilling for Diffusion Language Models
text2ql: Multi-Target Natural Language Querying via a Language-Agnostic Intermediate Representation
AI agents reshape consensus formation in human groups