←── back to feed
/topics/arxiv-ai-and-ml-research-papers-september-7
arXiv AI and ML research papers September 7
161 items●1 sources●updated 13d ago●trend 0
On September 7, 2026, arXiv published 20 AI/ML research papers spanning foundation models for finance and recruitment, agentic systems for evaluation and kernel generation, LLM faithfulness and reasoning benchmarks, and applications in power systems, XR networks, and biomarkers. Key contributions include EXAONE Finance (a financial time-series foundation model), Harbor Adapters (unified infrastructure for 80+ agentic benchmarks), and multiple studies on LLM agent behavior, safety, and reliability in real-world deployments.
- EXAONE Finance: financial time-series foundation model addressing quadratic self-attention costs and missing-data assumptions in general TSFMs
- Harbor Adapters: unified evaluation infrastructure porting 80+ benchmarks; evaluated 8 models across 54 benchmarks spanning capability tiers
- Iris-mini and Iris-pro: search agents at 35B-A3B and 397B-A17B scales trained on multi-hop entity-graph tasks with no string-matching shortcuts
- MaxKernel: multi-agent system for TPU kernel generation with human-in-the-loop and fully autonomous paradigms using LLM and compiler feedback
- HarvestBench: first benchmark quantifying LLM agent willingness to pay to avoid harming animals in farm-simulation gridworld environment
[BLG]blog/rss161
EXAONE Forecast for Finance
From Matching Models to Recruiting Agents: A Systematized Narrative Review of AI Recruitment Systems, Evaluation, and Governance
Harbor Adapters and Harbor-Index: Infrastructure and a Curated Meta-Dataset for Large-Scale Agentic Evaluation
Data-Optimized Contingency Screening: A Machine Learning Approach to Power System Security
Iris: Climbing to the Search Frontier
A Removal Based Approach to Improve LLM Faithfulness at Test-Time
Why Better Models Can Create Riskier Systems: Evidence from LLM Agents in Financial Markets
Corporate Language Model (CLM): Transforming Tacit and Fragmented Enterprise Knowledge into a Sovereign, Auditable, and Executable Corporate Intelligence Layer
HarvestBench: Measuring Whether LLM Agents Will Pay to Avoid Killing Animals
PerfReasoning: How Well Do LLMs Reason on Hardware Performance?
When Quantization Breaks Memory: Recurrent-State Write-Back in Low-Precision Temporal Inference
ResLearn-XR: Residual Learning for Network Traffic and Quality-of-Experience-Aware Modeling in Extended Reality
Rethinking Indirect Prompt Injection as a Test-Time Search Problem
BioSync: Transformer-Based Cross-Modal Fusion for a Multimodal Physiological Digital Biomarker
What Does Multi-Harness RL Learn? Credit Assignment and Portability in Coding Agents
MaxKernel: Agentic Kernel Generation for TPUs
Towards a universal language of concepts: A survey
Data-Driven Discovery of Composition-Dependent Constitutive Models for Hyperelasticity and Viscoelasticity of Digital Materials
From Answers to Interpretations: Rethinking Ambiguity-Induced Aleatoric Uncertainty Estimation in LLMs
IPGeoAI: Transformer-Based Geolocation with LLM Semantic Fusion
Reducing Hallucinated Transcripts in Whisper via Hallucination Space Projection
La Agente \'Optima: Towards Agentic Self-Driving Laboratories
Extremely Sparse Supervision Incentivizes Reasoning Ability
Does the Selected Object Reach the Reader? Auditing Identity Handoffs in Grounded Language-Model Pipelines
$\tau^\tau$-Bench: An Environment for End-To-End, Realistic Agent Construction
Leveraging Imperfect Restoration for Data Availability Attack
SiLR: Structure-Preserving Admission and Process Reward for LLM Tool Agents
A Cost-Aware Agentic Architecture for NL-to-SQL over Nested Enterprise Schemas, with a New Benchmark
Continual Graph Memory for Adaptive Recommendation under Intent Drift
Harness-agnostic detection and immunization of reward hacking in self-evolving language models
ERPBench: Evaluating LLM Agents for Enterprise Decision-Making Across Competitive Market Ecologies
Train What You Deploy:Token-Faithful Post-Training of a Production Coding
Predicting Spatiotemporal Mobile Sensing-Based PM2.5 Concentrations Using Low-Rank Adapted Spatially Attentive Graph Neural Network
SQL-Zero: Self-Evolving Text-to-SQL
Model Retirement Creates Reproducibility Risk in Biomedical AI Publications
FinalityBench: An Effect-Level Benchmark for Agent Decisions Under Delayed and Conflicting Financial Finality
PLUME: Parameter-Efficient Personalization of Large Language Models via Low-Rank User Modulation in Shared Subspaces
Aplaud: Adaptive Personalized Low-Rank Decomposition for User-Specific LLM
DCFA: Dual-view Causal-inspired Attribution for Failure Reasoning in LLM-based Multi-agent Systems
Shadow Queries for Private Retrieval in Vector Databases
Diffusion Language Models for Mobile Edge Agentic AI: Foundations, Applications, and Challenges
DODR: Deterministic Operator-Driven Reasoning in Latent Space
ProtLingo: Efficient Protein Language Modeling via Conditional Memory and Expert Routing
Whose record is this? Diagnosing and authorizing record use in personalized multimodal models
Hierarchical Possession-Aware Graph Pointer Network for Pass Receiver Selection
MedFlow: Class-Aware Multi-Scale Generation for Medical Time-Series Synthesis
When Financial Fine-tuning Fails: A Three-Level Detectability Analysis of Numerical Hallucination in Domain-Adapted Language Models
CPR-IE:A Compression-Prediction-Resource Intelligence Efficiency Metric
Long Horizon Transformer Quantile Fault Prediction for Multi Site Industrial Predictive Maintenance
ElderBench: Benchmarking Autonomous Mobile Agents for Older Adults
How Much Does Corpus Choice Change Dependency-Distance Estimates?
Memory as transformation: LETHE, a self-referential gan-inspired architecture
Evidence Integration in Large Language Models
MedProb: Probing Internal Representations of Vision-Language Models for Medical Question Answering
Adapting from Downturns: Prediction of Long-Term Conversational-Skill Development in Mental-Health Crisis Counselors
VERGE: Verification-Enhanced Refinement for Grounded Extraction of Early-Onset Colorectal Cancer Symptoms in Clinical Notes
You Really Didn't Get That? Benchmarking Social Pragmatic Inference for Indirect and Playful Chinese Online Comments
Evaluation of Phonetic Encoding Algorithms on Transcription Datasets
The Anatomy of an ASR Hallucination
A Systematic Evaluation of Cross-Lingual Consistency Enhancement Methods in Multilingual Language Models
What Attention Recalls and Recurrence Controls in Hybrid Language Models
GRACE: Graph-Grounded Reflective Agent Copilot Engine for Expert-in-the-Loop Knowledge Expansion
TRILOGUE: A Trilingual Spoken Dialogue Fact-Checking Benchmark with Evidence and Paired Audio
Shared circuits predict whether LLMs generalize across formats in arithmetic reasoning
Safety for Whom? Boundary-Aware Self-Distillation for Controlled LLM Safety Refusal
Patterns of Priming in Production: Lexical, Semantic and Structural Alignment in Language Model Generation
Cultural Misalignment in Large Language Models: Detection, Measurement, and Mitigation Through Targeted Fine-Tuning
Towards Understanding Pause Token Fine-Tuning Dynamics: A Mode Retention Perspective
LentEx: Generalizable Latent Entity Extraction via Synthetic Data and Instruction-Tuned LLMs
Scale-QLoRA: Code-Invariant Adapter Merging for Native 4-bit Microscaling LLMs
A Calibrated Reflection Approach for Enhancing Confidence Estimation in LLMs
Rhythms of Work: Multi-Scale Interpretation of Human Behavioral Traces for Workplace Agents
When Do Internal Probes Beat Reading the Answer? Miscalibrated Readouts and Behavior-Concealed Knowledge in Language Models
PetQA: Benchmarking Veterinary Knowledge and Clinical Reasoning
Tracing Audio Grounding and Answer Selection in Audio LLMs
CAGE: Coherence-Aware Graph Encoding for Retrieval-Augmented Generation
ConsensusBench: Benchmark of Consensus Nodes for LLM Reasoning via Outcome Reward Densifying
Choosing the Right Language Mode at Inference Time for Multilingual Reliability
Controlling and Assessing Appropriate Persona Use in LLM-based Dialogue Generation
How Do Language Models Represent and Use Phonological Information for Allomorph Selection?
Refuse without Refusal: A Structural Analysis of Safety-Tuning Responses for Reducing False Refusals in Language Models
Knowing What Not to Answer: Selective Non-Compliance in Vision-Language Models
Beneath the Surface of Chains-of-Thought: A Mechanistic Interpretation of Reasoning Operations in LLMs
Vectorizing Classical Tamil: Representation Learning for Verse-Commentary Pairs
Can Activation Steering Capture Multidimensional Authorship Style?
Recurrence Is Not Enough: Causally Validating Multilingual SAE Translation Features in Gemma 2 and 3
A Systematic Comparison of Multilingual Interpretability Methods Reveals Anisotropy-Driven Failures
Reinforcement Learning for improving Large Language Models' Catalan text simplification capabilities
Generating Constructive Feedback on Stories via Reinforcement Learning
On Epistemic Diversity in Large Language Models
MABPD: Multi-Agent Bias Probing & Detection via Structured Argument Debate
MMTClinic: Multimodal, Multilingual Time Series Question Answering and Reasoning Benchmark for Clinical Domain
CC-Mediation: Evaluating Large Language Models for Cross-Cultural Conflict Mediation
Cache-Aware Joint Router Adaptation for Memory-Efficient MoE Inference
RefactorPlatform: An Open-Source Harness for Controlled Evaluation of Repository-Scale Refactoring Agents
Discourse Dependency: A Continuous Criterion for Translation Difficulty
BIT.UA at BioASQ 14B: Modular Retrieval with pg_textsearch and Qdrant, and Agent-Based Answer Generation
MoirfEolas and Cr\'iochScore: Developing Resources for and the Evaluation of Tokenization Alignment with Irish Morphology
Leveraging Low-Level Symbolic Competences for Unsupervised Grounding in Hallucination Detection
How do LLMs Evaluate Perceived Moral Agency? Investigating Moral Decision-Making in Human-Artificial Agents Interactions
Minimax Lower Bound for Estimating Diffusion-based Local Intrinsic Dimension
An Analysis of Self-supervised Pre-training with Dependent Samples
FluxDisco: Symbolic Regression for Stoichiometric Dynamical Systems via Monte Carlo Graph Search
PAC-Bayesian Reconstruction Guarantees for Time Series Variational Autoencoders
Self-Supervised Pretraining of Molecular Graph Encoders with LeJEPA
Towards AI-Driven Nanomedicine Discovery: A Benchmark and Multimodal Learning Framework for Nano Self-Assembly Prediction
A Constraint-Aware Generative Framework for Synthetic Origin-Destination Demand in Logistics Networks
Optimal Stratified Allocation for Rare-Event Onset Forecasting in Dependent Sequences
Interpretability for Turing Machines
Faster Learning under Relaxed Local Differential Privacy
Reconciling Universal and Uniform Learning with $Q$-Aggregation
Confounding-Valid Conformal Inference for Counterfactual KPIs in Wireless Networks
Learning from VAE Errors to support ECG-based Differential Diagnosis of Myocardial Scar
A Gaussian process model for chemoinformatics with application to the hazard classification of organic solvents
High-dimensional censored MIDAS logistic regression for corporate survival forecasting
GyroSwin: 5D Surrogates for Gyrokinetic Plasma Turbulence Simulations
On the Asymptotic Inadmissibility of Double Machine Learning Estimators Under Structure-Agnostic Models
Spectral-Target Physical Latent Structuring for JEPA-Style World Models
ProToMEx: Rapid, Interpretable Explanations via Structured Representations
A Data Fusion Framework for Grounding Aerospace Surrogate Model via Experimental Wind-Tunnel Observations
Quantum-Assisted Memory-Efficient Training for Parameter-Intensive Wi-Fi-Based Human Activity Recognition
Evaluating Large Language Models for Forced Outage Risk Prediction: Benefits and Comparison to Machine Learning
BER-PEF: Unified Human Mobility Predictability Evaluation via Bayes Error Rate Estimation
Data-Driven Learning of Unknown Nonlinear Differential Equations Using Functional Analysis
Modular Deep Recurrent Neural Network: Application to Quadrotors
SharedSAE: One Feature Dictionary Across Language Models
A Quantum Variational Approach to Prototypical Recurrent Unit
On the Abundance of Critical Points of the t-SNE Energy
Disentangling Attention in Deep Operator Learning: A Controlled Study of Data-Driven and Physics-Informed Architectures
REFINE: LLM Refinement over Budgeted Text-Attributed Graphs for Personalized Medical Concept Representation
Beyond a Universal Forecasting Selector: Demand-Conditioned Model Selection across Demand Patterns and Horizons
A Repeated-Measurement Study for Cultural Analytics of English Song Lyrics Using Five Large Language Models
Conformity Breaks Conformal Prediction
When Load-Balancing Goes Too Far: Expert Pruning in Over-Dispersed Mixture-of-Experts Models
On-board ML for Trace Gas detection in Imaging Spectroscopy data
Nested Inductive Bias Framework for SPD Manifold Learning
Hakken: Predicting future discoveries to fill the gaps in today's knowledge
An Energy-Based Conservative-Dissipative Latent Neural Evolution Operator for Magnetization Dynamics
Distilled Continuous Diffusion Language Models Can Write Code in Few Steps---or One
Mitra-v2 Technical Report
Fast Surrogate Modeling of Excitable and Oscillatory FitzHugh-Nagumo Dynamics with Parametric Neural Operators
Training-Free Halving of Activated Experts in Fine-Grained Mixture-of-Experts Models
Optimizer Memory Schedules for Outscaling the Overtraining Axis
Representation Redundancy and Structural Complexity in Finite-Field Inversion
GNN-Guided Graph Coarsening and Adaptive QUBO Penalties for the Capacitated Vehicle Routing Problem with Time Windows on a Quantum Annealer
Too Rare to Learn: Prescribed Cyclone Tracks Degrade a Bay of Bengal Ocean Emulator
SMILE: Bridging Continuous Optimization and Discrete Symbolic Recovery
WEECFP-SuRGE: Wide Embedded Extended Connectivity Fingerprint with Substructure Rotary Graph-distance Encoding
Simulation-free Unbalanced Dynamic Optimal Transport with General Growth Penalty
Locating and Steering Refusal Beyond Attention
Training Large Language Models for Small-Molecule Design with Synthetic Task Scaling
A Fairness Audit of the Duckworth-Lewis-Stern Method: Format-Specific and Gender-Differential Bias, with an Interpretable Calibration Layer for Cricket Target Revision
Resilience Beyond Stationary Client Unavailability: Unlocking Efficient and Unbiased Federated Learning
A Robust Watermark-based Fingerprint Framework for GNNs Ownership Verification
Persistent Teacher Anchoring for Tool-Using Agents
Dynamic Heterogeneous Graph Representation Learning: A Survey
Learning-Augmented Algorithms: Guarantees, Construction Mechanisms, and System-Level Implications
How Faithful Is Attribution for Sales Forecasting? A Counterfactual Study
Federated Attack Campaign Detection via Contrastive Encoding of Threat Indicators in Gradient Updates
Communication-Efficient Personalized Federated Learning via Layer-Wise Multi-Threshold Random Sketching
PACE: Propagation-Aware Collaborative Correction for One-Shot Personalized Federated Graph Learning