←── back to feed
/topics/arxiv-ai-and-ml-research-papers-september-4
arXiv AI and ML research papers September 4
135 items●1 sources●updated 15d ago●trend 0
On September 4, 2026, arXiv published 20 papers spanning computational linguistics, speech recognition, and legal AI. Key advances include methods for optimizing LLM agent prompts (HARNESSEVO), efficient retrieval-augmented generation (R²Adapter), speculative decoding for faster inference (AdaptiveSpec), and domain-adapted biomedical models (DRET), alongside new benchmarks for misinformation detection (BharatGather), legal issue identification (LexIssue), and document parsing (Jina-OCR-v1).
- HARNESSEVO decomposes LLM agent harness into four separately evolvable slots: role, task-strategy, tool/format-rules, and reflection/control.
- R²Adapter routes queries between text and graph RAG strategies to balance reasoning capability and inference latency.
- AdaptiveSpec enables training-free per-step lossy speculative decoding without fixed token budgets or strict verification rules.
- DRET injects biomedical domain knowledge into lightweight DistilBERT via priority-based embedding transfer.
- Jina-OCR-v1 combines 3B mixture-of-experts decoder with FastMTP speculative decoding (K=3 steps) for efficient document parsing on low-budget GPUs.
- BharatGather dataset targets misinformation detection in Indian public events with socio-cultural context.
[BLG]blog/rss135
Where Does Harness-Optimization Value Live? Localized Gains and the Budget-Splitting Trap in Self-Evolving LLM Agents
Bounded Personas Match Retrieval on Classification but Not Regression for a Frozen Agent
Counterexamples as Feedback for Agent Self-Correction
Probe Generalization as Subspace Selection for OOD Deception Detection
R$^{2}$Adapter: A Routing and Rewriting Adapter for Efficient Hybrid RAG
BharatGather: A Culturally-Informed Benchmark Dataset for Misinformation and Fake News Detection in Indian Public Events
PiPMRE: A Pipeline Based on Language Model for Medical Relation Extraction
Margins, Not Windows: Training-Free Per-Step Lossy Speculative Decoding
Distilled Rapid Embedding Transfer (DRET): Parameter-Efficient Biomedical Domain Adaptation via Priority-Based Embedding Transfer
Contamination Inflates Scores but Rarely Reorders Large Language Model Leaderboards
Dual-Form ASR: Semantics-Aware Inverse Text Normalization for Chinese Speech Recognition
RL-ADA: A World-Feedback Framework for Adversarially Robust Enterprise Dialogue Agents
Listen to the Latents: Self-Correcting Speech Recognition in Large Audio Language Models Through Hidden-State Interactions
Judging LLM-as-a-Judge: Concerning Rubric Artifacts in LLM-based Automated Text Generation Evaluation
LexIssue: Benchmarking Legal Issue Identification in Chinese Civil Litigation
Unifying Conformal Language Tasks with In-Context Ensembles
SHELF: A Synthetic Harness for Multi-Task Bibliographic Benchmarking
Large Language Models in Resolving Contextual Knowledge Conflicts
No country for old linguists: LLM-brain alignment underdetermines neural computation
Jina-OCR-v1: Efficient Document Parsing with Speculative Decoding and Dense Verifiable Rewards
MemoryLACE: Memory Lifecycle-Aware Consolidation and Evidence Retrieval
LLMs Learn Better In-Context from Rules than from Examples
SWIM: Student Writing Simulation via Proficiency-Conditioned Generation
The Analyst in the Prompt: Role, Retrieval, and Memory Biases in LLM Financial Analysis
Counterfactual Fairness Audits of Multi-Step Clinical LLM Agents Require a Measured Per-Action Instability Floor
SGD-KV: Summarization Guided KV Cache Compression
What Else Needs Fixing? Exploring Cost-Effective Test-Time Compute for Revision Propagation in Artifacts Generated Through Conversation
Contextual Tamil Spelling and Grammar Correction Using Progressively Fine-Tuned Sequence-to-Sequence Transformers
PACE: Towards Surfacing Hidden Conflicts in User Requests
Decoupling Turn-Taking from Semantics: A Decoupled Data Approach for Finite-State-Machine-Based Full-Duplex Dialogue
How Perturbations Propagate: A Multi-Level Analysis of Robustness in Large Language Models
Less Is Moral: A CHARMing Framework for Moral Foundations Detection in Endorsement Behaviour
FPCO-Dialog: A Multi-Turn False-Premise Benchmark for Correction and Cooperation in Vision-Language Models
Accountable AI with Grounded, Faithful, Consistent, Actionable Rationales: A Case Study in Clinical Trial Matching with VERDICT
FrameBench:A Language Understanding Benchmark Based on Frame Semantics
Chiaroscuro for Emotions: A Contrastive Emotion Benchmark Grounded in Appraisal Theory
TabScope: Question-Adaptive Scope Selection for Table Question Answering
To What Extent Do Large Language Models Understand Bangla Idioms?
Lngram v2: Latent N-Gram Memory with Interpretable Discrete Representations
Random Attention: Rethinking KV Cache Eviction for Efficient Reasoning
Decoupled Analysis-Judging: An Automated Creativity Evaluator Using LLMs in Complex Multi-step Creativity Tasks
When Retrieval Helps: Selective Retrieval for Single-Turn Mental-Health QA
When Users Don't Ask: Benchmarking Context-Driven Memory Retrieval in Conversational Agents
Pattern Over-Generalization of Knowledge Graph Embedding
Building and Evaluating Fixed-Voice Thai TTS from Synthetic Speech
The Attention Triangle in Audio-Video Models
KC-Bench: A Dynamic Interactive Benchmark for Evaluating Knowledge Conflicts in LLM Agents
A computable representation of the physical laboratory enables verifiable workflows
Analysis of Prompt Engineering for Drug Toxicity Prediction
Synthetic Semantic Supervision for Contrastive Code Representation Learning in Small Transformers: An Empirical Study
Counterfactual Routing Using Integer Programming with Constraint Generation
Artificial Intelligence for Energy Optimization in Data Centers
Proactive Service Agents: A Unified Decision Framework, Methods, and Evaluation
SimSkill: A Lifelong Learning AI Agent for Autonomous Mastery of Traffic Simulation
Rethinking World Models for Safety-Critical Embodied Systems
DNative-Twin: Decision Graphs and Digital Twins for Reconstructable Agentic Decisions
Transfiver: Human-AI Co-Inference through a Shared Editable State
Govern the Model, Not Only the Data: Storage, Circulation, and Learning in Creative AI
SVG-Score: Human-Aligned Evaluation of Text-to-SVG Generation
CauseCollab: Causal Unified and Modality-Agnostic Network for Heterogeneous Collaborative Perception
Semantic Bayesian World Models
Adapting to Evolving Requirements: Agentic AI for Retail Supply Chain Operations
Bioinfoysis Technical Report
STAIR (STructure Aware Information Retriever): A novel dataset and LLM based retriever for document structure augmentation
Xiaomi-TabLDM: A Tabular Foundation Model Technical Report
Inferring Affective Consciousness in an Artificial Agent: A Case Study
Lose the Order, Keep the Hierarchy: Deordering HTN Plans
Value-Preserving Architectures for Agentic AI Systems
Speak for Me: Giving LLMs the Situational Awareness to Participate in a Meeting
Towards Numerical TOHTN Planning with SMT-based HTN-SAT Encoding
More Criticism Does Not Make a Better Review: EquiReview-R
FiMI Banking: A Sovereign Model for Indian Retail Banking
Interface-Induced Trajectory Censoring
Common-Witness Certificates and Sharp Feature Bounds for Counterfactual Image Auditing
Occupancy-based Quantile Risk Control
A Closed-Form Formula for Consistent Lipschitz Regression on Metric Spaces with Sparse Neural Network Realizations
ALRA: Adaptive Local Relational Alignment for Logit-Based Pre-training Distillation of Autoregressive Language Models
Towards a Statistical Understanding of Mixture-of-Experts
Towards Scaling Reinforcement Learning to Massive Populations: Learning Mean-Field Representations
The Geometry of Ignorance: LLMs Know When to Temper Bayesian Priors
Statistical Feature Augmentation for Anomaly Detection in Dynamic Graphs
Tail-Likelihood Reinforcement Learning
TRACE: Spatiotemporal Contact Memory Graph Network Simulator for Granular Dynamics
No-Regret Bayesian Optimization with Finite-Library Input-Warped Kernels
Evaluating Graph Neural Networks for Change-Criticality Classification in Maritime Navigation Charts
Causal Foundation Models
Discrete Gromov-Wasserstein Duality: Algorithms and Isomorphism Testing
Spectral characteristics of autoencoder parameters as a vector representation of data
Correlated initialization of deep residual networks
Q-Edge: Symmetry-Reduced Quantum Simulation of Structured Extreme Dependence
High-Dimensional Learning Dynamics of Attention-Indexed Models
Cooperative Multi-Task Semantic Communication for Joint Classification and Regression Tasks
Reliable Selection of Heterogeneous Treatment Effect Estimators
Active learning for data-driven reduced models of parametric differential systems with Bayesian operator inference
Deep networks learn to parse uniform-depth context-free languages from local statistics
Semiparametric Inference for Counterfactual Regression under Intervention-Driven Shift
Online Learning of Functional Principal Component Analysis for Multidimensional Functional Data
Data-efficient Kernel Methods for Learning Hamiltonian Systems
Distributional Treatment Effect Transportability across Heterogeneous Sites
Entropy-Generated Attention Beyond Softmax and Entmax: Kaniadakis and Reciprocal-Symmetric Abe Operators
Algebraic Invariants of Lightning Self-Attention
Equation Recast for Canonical Operator Learning Across Parametric PDEs
From Euclidean to Graph-Structured Data: A Survey of Collaborative Learning
Modern Transformers Are Implicit Hybrids: From Functional Differentiation to Principled Hybrid Architecture Design
Mesh-Native Physics-Informed Graph Surrogates for TCAD-in-the-Loop Design Space Exploration
Verify Before You Distill: Prompt-Level Teacher Gating for On-Policy Distillation
ObserverBench: Testing Mechanistic Estimates for Intervention and Control
Learnable composition for neural operators
LeanStream: A Speculate-and-Refine Streaming Framework for Efficient on-Device LLM Inference
The Gradient Does Not See Rank: Rank-Indifference in Matrix-CODI on ProsQA
Distilling deep optical flow stereo methods to retrieve dense three-dimensional wind fields
Scaling Laws, Tabular Data and Actuarial Ratemaking Models
Kernel Reboot: Breaking the Boundaries of Neural Tangent Kernels for Neural Fields
Routing Is Not Enough: Diagnosing Intra-Adapter Subspace Contention in MoE+LoRA Fine-Tuning
Frontier LLMs are effective batch optimizers: Assessing reasoning models in continuous and discrete settings
Portable Causal Fairness Across Synthetic Data Generator Families
Language-encoded network topology enables large language models to reason about complex networks
The 2026 PNPL Competition: Word Classification and Efficient Cross-Subject Generalisation in LibriBrain100
B2B Customer Conversion Prediction: A Document Representation, Graph Theory, and CatBoost Driven Methodology
FlowBalance: Verifier-Grounded Self-Improvement from On-Policy Reasoning Experience
Selective Hypergraph Refinement for Frozen Graph Clustering
Latent Energy Action Planning with World Models
Geometry-Aware Graph Construction via Adaptive Spectral Bandwidth Control
Risk and Anomaly Identification for Distribution Network Optimal Operation Based on Reinforcement Learning and Uncertainty Quantification
DE-Venus: A Data-Efficient RLVR Framework for Large Language Models
A Large Open Multi-Energy Corpus of Soil Compaction Tests, with Machine-Learning Baselines
Gradients Know What Outcomes Don't: Unlocking Reinforcement Learning for LLM Reasoning with Gradient-Aligned Rewards
From Zero to Hero: An Open LLM Ecosystem for Armenian
Time Without Timesteps: Simulating Coupled Dynamical Systems via Self-Consistency
SimpleDesign: A Joint Model for Protein Sequence and Structure Codesign
RecurTrace: Adaptive Latent Reasoning with Loop-Time Memory
TIGPO: Temporal Instance-Graph Policy Optimization for Long-Horizon LLM Agents
Inferred Generative-Process Diversity Predicts Correlated Failure Across Language Models
TraveL: Transformer-based Multi-view Path Distributional Representation Learning
It's the Problem, Not the Path: Budget and Difficulty Confounds in LLM Reasoning Trajectories