←── back to feed
/topics/arxiv-ai-and-ml-research-papers-september-7

arXiv AI and ML research papers September 7

161 items1 sourcesupdated 13d agotrend 0

On September 7, 2026, arXiv published 20 AI/ML research papers spanning foundation models for finance and recruitment, agentic systems for evaluation and kernel generation, LLM faithfulness and reasoning benchmarks, and applications in power systems, XR networks, and biomarkers. Key contributions include EXAONE Finance (a financial time-series foundation model), Harbor Adapters (unified infrastructure for 80+ agentic benchmarks), and multiple studies on LLM agent behavior, safety, and reliability in real-world deployments.

  • EXAONE Finance: financial time-series foundation model addressing quadratic self-attention costs and missing-data assumptions in general TSFMs
  • Harbor Adapters: unified evaluation infrastructure porting 80+ benchmarks; evaluated 8 models across 54 benchmarks spanning capability tiers
  • Iris-mini and Iris-pro: search agents at 35B-A3B and 397B-A17B scales trained on multi-hop entity-graph tasks with no string-matching shortcuts
  • MaxKernel: multi-agent system for TPU kernel generation with human-in-the-loop and fully autonomous paradigms using LLM and compiler feedback
  • HarvestBench: first benchmark quantifying LLM agent willingness to pay to avoid harming animals in farm-simulation gridworld environment
[BLG]blog/rss161
EXAONE Forecast for Finance
arXiv cs.AI · Seunghan Lee, Jaehoon Lee, Jun Seo, Tae Yoon Lim, Dongwan Kang, Hwanil Choi, Minjae Kim, Sungdong Yoo, Junhyeok Kang, Sangjun Han, Soonyoung Lee, Wonbin Ahn · 14d
From Matching Models to Recruiting Agents: A Systematized Narrative Review of AI Recruitment Systems, Evaluation, and Governance
arXiv cs.AI · Ziyi Zhao, Guanzheng Wei · 14d
Harbor Adapters and Harbor-Index: Infrastructure and a Curated Meta-Dataset for Large-Scale Agentic Evaluation
arXiv cs.AI · Lin Shi (Audrey), Haowei Lin (Audrey), Zixuan Zhu (Audrey), Xiaoyue Zhou (Audrey), Xiang Li (Audrey), Xiangning Lin (Audrey), Yaxuan Deng (Audrey), Han Xu (Audrey), Yuangang Li (Audrey), Shanda Li (Audrey), Zizhao Chen (Audrey), Hanwen Xing (Audrey), Harsh Raj (Audrey), Bo Chen (Audrey), Quan Shi (Audrey), Steven Dillmann (Audrey), Yipeng Gao (Audrey), Puneesh Khanna (Audrey), Ruofan Lu (Audrey), Chao Beyond Zhou (Audrey), Michael Yang (Audrey), Robert Zhang (Audrey), Siyuan Chai (Audrey), Jiayu Chang (Audrey), Yizhao Chen (Audrey), Xiaokun Chen (Audrey), Yiwei Dai (Audrey), Wenting Yang (Audrey), Hange Liu (Audrey), Minghao Liu (Audrey), Zihan Wang (Audrey), Adnan El Assadi (Audrey), Benedikt Stroebl (Audrey), E. Kelly Buchanan (Audrey), Han Meng (Audrey), Junwei He (Audrey), Longxuan Yu (Audrey), Radin Shayanfar (Audrey), Yukyung Lee (Audrey), Zhikang Dong (Audrey), Allen G Hart (Audrey), Anjiang Wei (Audrey), Anurag Kashyap (Audrey), Arpandeep Khatua (Audrey), Audrey Jixin Zheng (Audrey), Chengrui Ma (Audrey), David Heineman (Audrey), Dubing Chen (Audrey), Hai-Anh Trinh (Audrey), Haishuo Fang (Audrey), Hefan Zhang (Audrey), Hui Shen (Audrey), Issa Sugiura (Audrey), Jiankai Sun (Audrey), Jiechao Gao (Audrey), Junhong Lin (Audrey), Junnan Li (Audrey), Kai Yang (Audrey), Lei Hsiung (Audrey), Maoyu Wang (Audrey), Mengze Tang (Audrey), Nabil Omi (Audrey), Negin Raoof (Audrey), Nicholas Edwards (Audrey), Octavia Guo (Audrey), Orfeas Menis Mastromichalakis (Audrey), Pengliang Ji (Audrey), Przemys{\l}aw Hejman (Audrey), Qi Qi (Audrey), Qunshu Lin (Audrey), Richard Zhuang (Audrey), Rui Yang (Audrey), Ruichen Zheng (Audrey), Ryan Marten (Audrey), Shaghayegh Fazliani (Audrey), Shizheng Hou (Audrey), Sicong Jiang (Audrey), Sijie Li (Audrey), Song Bian (Audrey), Terry Yue Zhuo (Audrey), Tianqing Wu (Audrey), Tom Tang (Audrey), Wanjia Zhao (Audrey), Weihao Xuan (Audrey), Wenhua Liang (Audrey), Xian Liu (Audrey), Xin Lan (Audrey), Xuan Zhang (Audrey), Xuandong Zhao (Audrey), Yanchuan Tang (Audrey), Yifan Jiang (Audrey), Yijiang Li (Audrey), Yitong Guan (Audrey), Yizhi Li (Audrey), Yonghui Liu (Audrey), Yuheng Tang (Audrey), Yujun (Audrey), Mao, Yunfei Zhao, Yuxin Wang, Yuxuan Tang, Zhenheng Tang, Zhifei Li, Ziruo Wang, Ziyu She, Kaiyuan Liu, Iheb Chaabane, Yuxin Tang, Xiangyi Li, Andy Konwinski, Boxuan Li, Leon Liangyu Chen, Alex Dimakis, Nicholas Carlini, Soroush Vosoughi, Di He, Etash Guha, Benjamin Feuer, Mike Merrill, Ludwig Schmidt, Alex Shaw · 14d
Data-Optimized Contingency Screening: A Machine Learning Approach to Power System Security
arXiv cs.AI · Joshua Salako, Folajimi Osikomaiya, Olakorede Olamiju · 14d
Iris: Climbing to the Search Frontier
arXiv cs.AI · Ziyuan Liu, Hengqi Liu, Zichuan Wang, Yang Qin, Jiachen Liang, Xu Chu, Shaowei Chen, Yuantao Gu, Mu Chuan · 14d
A Removal Based Approach to Improve LLM Faithfulness at Test-Time
arXiv cs.AI · Qinglan Luo, S M A Nahian, John Guttag, S. Mazdak Abulnaga, Katie Matton · 14d
Why Better Models Can Create Riskier Systems: Evidence from LLM Agents in Financial Markets
arXiv cs.AI · Jillian Ross, Eric So, Zoe De Simone, Charles Pozniak, Andrew W. Lo · 14d
Corporate Language Model (CLM): Transforming Tacit and Fragmented Enterprise Knowledge into a Sovereign, Auditable, and Executable Corporate Intelligence Layer
arXiv cs.AI · Fabricio C. Avini, Guilherme Trez · 14d
HarvestBench: Measuring Whether LLM Agents Will Pay to Avoid Killing Animals
arXiv cs.AI · Jasmine Brazilek, Miles Tidmarsh, Matthias Endres, Anshuman Singh, Jeremiah Miller · 14d
PerfReasoning: How Well Do LLMs Reason on Hardware Performance?
arXiv cs.AI · Dan Zhao, Karthikeyan Sankaralingam, Christos Kozyrakis, Qijing Huang · 14d
When Quantization Breaks Memory: Recurrent-State Write-Back in Low-Precision Temporal Inference
arXiv cs.AI · Ismail Erbas, Xavier Intes, Vikas Pandey · 14d
ResLearn-XR: Residual Learning for Network Traffic and Quality-of-Experience-Aware Modeling in Extended Reality
arXiv cs.AI · Yoga Suhas Kuruba Manjunath, Jie Gao, Lian Zhao · 14d
Rethinking Indirect Prompt Injection as a Test-Time Search Problem
arXiv cs.AI · Duong M. Nguyen, Joon Sik Kim, Blazej Manczak, Vaikkunth Mugunthan · 14d
BioSync: Transformer-Based Cross-Modal Fusion for a Multimodal Physiological Digital Biomarker
arXiv cs.AI · Seyed Mahmoud Sajjadi Mohammadabadi · 14d
What Does Multi-Harness RL Learn? Credit Assignment and Portability in Coding Agents
arXiv cs.AI · Chenqian Le, Jiayi Cheng, Qijia He, Runhao Li, Yinghao Li, Xupeng Chen · 14d
MaxKernel: Agentic Kernel Generation for TPUs
arXiv cs.AI · Shangkun Wang, Nina Cai, Charles Hoong, Julian Walker, Gerson Kroiz, George Vanica, Deepak Patil, Andi Gavrilescu, Hassan Sipra, Sethu Sankaran · 14d
Towards a universal language of concepts: A survey
arXiv cs.AI · Aishni Parab · 14d
Data-Driven Discovery of Composition-Dependent Constitutive Models for Hyperelasticity and Viscoelasticity of Digital Materials
arXiv cs.AI · Josu\'e Garc\'ia-\'Avila (Department of Mechanical Engineering, Columbia University, New York City, USA), Beijun Shen (Department of Mechanical Engineering, Columbia University, New York City, USA), Manuel K. Rausch (Department of Aerospace Engineering and Engineering Mechanics, University of Texas at Austin, Austin, USA, Department of Biomedical Engineering, University of Texas at Austin, Austin, USA, Department of Mechanical Engineering, University of Texas at Austin, Austin, USA), Mary C. Boyce (Department of Mechanical Engineering, Columbia University, New York City, USA), Adri\'an Buganza-Tepole (Department of Mechanical Engineering, Columbia University, New York City, USA) · 14d
From Answers to Interpretations: Rethinking Ambiguity-Induced Aleatoric Uncertainty Estimation in LLMs
arXiv cs.AI · Omer Nahum, Niv Nayman, Jonathan Fhima, Alon Zolfi, Jeremy Levy, Shai Mazor, Paolo Favaro · 14d
IPGeoAI: Transformer-Based Geolocation with LLM Semantic Fusion
arXiv cs.AI · Avinash Kadimisetty, Andy Jinqing Yu, Philip Favaloro, Wenlong Liu, Xiaolu Xiong · 14d
Reducing Hallucinated Transcripts in Whisper via Hallucination Space Projection
arXiv cs.AI · Maryam Abbasihafshejani, Murtuza Jadliwala · 14d
La Agente \'Optima: Towards Agentic Self-Driving Laboratories
arXiv cs.AI · Marcel M\"uller, Jiaru Bai, Willi Gottstein, Abhijoy Mandal, Mohammad Nazeri, Elia Savino, Yanlin Fang, Sujoy Das, Sergio Pablo Garc\'ia Carrillo, Yeonghun Kang, Juan B. P\'erez-S\'anchez, Simone Pilon, Martin Fitzner, Timothy No\"el, Frank Gu, Varinia Bernales, Al\'an Aspuru-Guzik · 14d
Extremely Sparse Supervision Incentivizes Reasoning Ability
arXiv cs.AI · Zhishuai Liu, Xingzi Xu, Mehmet Saygin Seyfioglu, Pan Xu, Karim Bouyarmane · 14d
Does the Selected Object Reach the Reader? Auditing Identity Handoffs in Grounded Language-Model Pipelines
arXiv cs.AI · Siddharth Vohra, Runmin Jiang, Xiaomo Li, Min Xu · 14d
$\tau^\tau$-Bench: An Environment for End-To-End, Realistic Agent Construction
arXiv cs.AI · Quan Shi, Keshav Dhandhania, Karthik Narasimhan, Victor Barres · 14d
Leveraging Imperfect Restoration for Data Availability Attack
arXiv cs.AI · Yi Huang, Jeremy Styborski, Mingzhi Lyu, Fan Wang, Adams Kong · 14d
SiLR: Structure-Preserving Admission and Process Reward for LLM Tool Agents
arXiv cs.AI · Chenyu Zhou, Qiliang Jiang, Shuning Wu, Xu Zhou · 14d
A Cost-Aware Agentic Architecture for NL-to-SQL over Nested Enterprise Schemas, with a New Benchmark
arXiv cs.AI · Yoga Sri Varshan Varadharajan, Ajay Yadav, Ritesh Goru, Prateek Chaudhury, Constantine Caramanis, Prateek Jain, Divyateja Pasupuleti, Sunil Kumar Pandey · 14d
Continual Graph Memory for Adaptive Recommendation under Intent Drift
arXiv cs.AI · Hao Nguyen Ngoc, Tung Nguyen, Nguyen Thi Hanh, Hoang Thai Dinh, Nguyen Xuan Tung · 14d
Harness-agnostic detection and immunization of reward hacking in self-evolving language models
arXiv cs.AI · Rongxin Yang, Yang Liu, Shang Luo, Haoxuan Jia, Chongyang Zhang, Hao Zheng, Yingguang Yang, Yulin Huang, Jianshen Zhang, Yongzhi Qi, Kefu Xu, Congjing Ran, Bin Chong · 14d
ERPBench: Evaluating LLM Agents for Enterprise Decision-Making Across Competitive Market Ecologies
arXiv cs.AI · Xinran Zhang, Pengrui Lu, Lyumanshan Ye, Pengfei Liu · 14d
Train What You Deploy:Token-Faithful Post-Training of a Production Coding
arXiv cs.AI · Cheng Li, Jiexiong Liu, Yixuan Chen, Chi Hong · 14d
Predicting Spatiotemporal Mobile Sensing-Based PM2.5 Concentrations Using Low-Rank Adapted Spatially Attentive Graph Neural Network
arXiv cs.AI · Om Chiddarwar, Priyanka Mandal, Praveen Kumar Chandaliya, Shriniwas Arkatkar · 14d
SQL-Zero: Self-Evolving Text-to-SQL
arXiv cs.AI · Daniel Machado Pedrozo, Julia Soares Dollis, Bryan Lincoln Marques de Oliveira, Vinicius Alboneti Aguiar, S\'avio Salvarino Teles de Oliveira, Telma Woerle de Lima Soares · 14d
Model Retirement Creates Reproducibility Risk in Biomedical AI Publications
arXiv cs.AI · Nathan Wolfrath, Meghan Conroy, Thomas Kosten, Dave Bell, Bhabishya Neupane, Jonah Kindel, Anjishnu Banerjee, Priya Deshpande, Bradley Taylor, Anai N. Kothari · 14d
FinalityBench: An Effect-Level Benchmark for Agent Decisions Under Delayed and Conflicting Financial Finality
arXiv cs.AI · Abhishek Sharma · 14d
PLUME: Parameter-Efficient Personalization of Large Language Models via Low-Rank User Modulation in Shared Subspaces
arXiv cs.AI · Xinyu Li, Hao Zhou, Jianfeng Zhu, Julina Maharjan, Ruixin Guo, Feodor Dragan, Ruoming Jin · 14d
Aplaud: Adaptive Personalized Low-Rank Decomposition for User-Specific LLM
arXiv cs.AI · Xinyu Li, Ruoming Jin, Jianfeng Zhu, Ruixin Guo, Zhi Liu · 14d
DCFA: Dual-view Causal-inspired Attribution for Failure Reasoning in LLM-based Multi-agent Systems
arXiv cs.AI · Zehao Wang, Lanjun Wang, Shilong Jin, Junjie Chen, Yanghua Xiao · 14d
Shadow Queries for Private Retrieval in Vector Databases
arXiv cs.AI · Xinguo Feng, Zhongkui Ma, Zihan Wang, Chuan Yan, Guowei Yang, Alsharif Abuadbba, Guangdong Bai · 14d
Diffusion Language Models for Mobile Edge Agentic AI: Foundations, Applications, and Challenges
arXiv cs.AI · Chenqi Li, Minghui Min, Dusit Niyato, Wei Ni · 14d
DODR: Deterministic Operator-Driven Reasoning in Latent Space
arXiv cs.AI · Weicai Huang (Beijing MQPat Technologies, Co., Ltd.) · 14d
ProtLingo: Efficient Protein Language Modeling via Conditional Memory and Expert Routing
arXiv cs.AI · Mingrui Li, Sixian Shen, Minzhang Li, Ruiyi Zhang, Kexin Zhang, Jiakai Zhang, Jingyi Yu · 14d
Whose record is this? Diagnosing and authorizing record use in personalized multimodal models
arXiv cs.AI · Xinyu Mao, Junsi Li, Chenyang Liu, Haoji Zhang, Ming Sun · 14d
Hierarchical Possession-Aware Graph Pointer Network for Pass Receiver Selection
arXiv cs.AI · Jingyi Wang, Da Li, Kaixin Wang, Zhangqin Huang · 14d
MedFlow: Class-Aware Multi-Scale Generation for Medical Time-Series Synthesis
arXiv cs.AI · Yanhao Huang, Shibo Feng, Wanjin Feng, Peilin Zhao, Chunyan Miao · 14d
When Financial Fine-tuning Fails: A Three-Level Detectability Analysis of Numerical Hallucination in Domain-Adapted Language Models
arXiv cs.AI · Xiaodong Li, Peiwei Liu · 14d
CPR-IE:A Compression-Prediction-Resource Intelligence Efficiency Metric
arXiv cs.AI · Xiantao Jiang · 14d
Long Horizon Transformer Quantile Fault Prediction for Multi Site Industrial Predictive Maintenance
arXiv cs.AI · David J Poland, Daniele Ravi, Na Helian · 14d
ElderBench: Benchmarking Autonomous Mobile Agents for Older Adults
arXiv cs.AI · Weide Zhan, Qumu Shaqu, Yuanqing Liu, Peng Zhang, Jiahao Liu, Kam Him Lam, Ning Gu, Zhan Hu, Tun Lu · 14d
How Much Does Corpus Choice Change Dependency-Distance Estimates?
arXiv cs.CL · Sirui Chen · 14d
Memory as transformation: LETHE, a self-referential gan-inspired architecture
arXiv cs.CL · Francesco Vitucci, Anthony Di Furia, Francesco Scagliola · 14d
Evidence Integration in Large Language Models
arXiv cs.CL · Sebastien Kawada, Manolis Kellis · 14d
MedProb: Probing Internal Representations of Vision-Language Models for Medical Question Answering
arXiv cs.CL · Erfan Nourbakhsh, Ke Yang, Anthony Rios · 14d
Adapting from Downturns: Prediction of Long-Term Conversational-Skill Development in Mental-Health Crisis Counselors
arXiv cs.CL · Vivian Nguyen, Lillian Lee, Elizabeth A. Olson, Cristian Danescu-Niculescu-Mizil · 14d
VERGE: Verification-Enhanced Refinement for Grounded Extraction of Early-Onset Colorectal Cancer Symptoms in Clinical Notes
arXiv cs.CL · Nikkie Hooman, Monarch Nigam, Amy E. Hughes, Rasmi G. Nair, Mehak Gupta · 14d
You Really Didn't Get That? Benchmarking Social Pragmatic Inference for Indirect and Playful Chinese Online Comments
arXiv cs.CL · Shiwei Hong, Junjie Ma, Emma Jiren Wang, Ethan Z. Rong, Siying Hu, Haichang Li, Ziying Wang, Zhicong Lu · 14d
Evaluation of Phonetic Encoding Algorithms on Transcription Datasets
arXiv cs.CL · Can \"Ozbey, Emre Kaplan, Berkin Deniz Kahya · 14d
The Anatomy of an ASR Hallucination
arXiv cs.CL · Hamees Sayed, Apoorv Singh, Kumar Aman, Akshat Mandloi · 14d
A Systematic Evaluation of Cross-Lingual Consistency Enhancement Methods in Multilingual Language Models
arXiv cs.CL · Jirui Qi, Mingyang Wang, Hinrich Sch\"utze, Raquel Fern\'andez, Arianna Bisazza · 14d
What Attention Recalls and Recurrence Controls in Hybrid Language Models
arXiv cs.CL · Kirill Afendulev, Alexey Dontsov, Elena Tutubalina, Anton Korznikov · 14d
GRACE: Graph-Grounded Reflective Agent Copilot Engine for Expert-in-the-Loop Knowledge Expansion
arXiv cs.CL · John Seon Keun Yi, Joshua R. Minot, Dokyun Lee · 14d
TRILOGUE: A Trilingual Spoken Dialogue Fact-Checking Benchmark with Evidence and Paired Audio
arXiv cs.CL · Chaewan Chun, Meruyert Aristombayeva, Jiyoung Choi, Mahjabin Nahar, Delvin Ce Zhang, Dongwon Lee · 14d
Shared circuits predict whether LLMs generalize across formats in arithmetic reasoning
arXiv cs.CL · Andrea Gregor de Varda, Sana Pandey, Pengrui Han, Jacob Andreas, Evelina Fedorenko · 14d
Safety for Whom? Boundary-Aware Self-Distillation for Controlled LLM Safety Refusal
arXiv cs.CL · Alejo L\'opez-\'Avila, Iker Garc\'ia-Ferrero, Jezabel Garcia, Antonio Tiene, Rom\'an Or\'us · 14d
Patterns of Priming in Production: Lexical, Semantic and Structural Alignment in Language Model Generation
arXiv cs.CL · Giulia Pucci, Ruizhe Li, Arabella Sinclair · 14d
Cultural Misalignment in Large Language Models: Detection, Measurement, and Mitigation Through Targeted Fine-Tuning
arXiv cs.CL · Antoni Czolgowski, Abel Iyasele · 14d
Towards Understanding Pause Token Fine-Tuning Dynamics: A Mode Retention Perspective
arXiv cs.CL · Jaehyeon Kim, Suhwan Kim, Nakyung Lee, Yeongoon Kim, Jimin Seo, Giho Lee, Jungwoo Lee · 14d
LentEx: Generalizable Latent Entity Extraction via Synthetic Data and Instruction-Tuned LLMs
arXiv cs.CL · Umesh Bodhwani, Yuan Ling, Cibi Chakravarthy Senthilkumar, Shujing Dong, Yarong Feng, Hongfei Li, Ayush Goyal · 14d
Scale-QLoRA: Code-Invariant Adapter Merging for Native 4-bit Microscaling LLMs
arXiv cs.CL · Tung-Ling Li, Jiale Huang, Lee-Chi Wang, Janaki Ram Gotei · 14d
A Calibrated Reflection Approach for Enhancing Confidence Estimation in LLMs
arXiv cs.CL · Umesh Bodhwani, Yuan Ling, Shujing Dong, Yarong Feng, Hongfei Li, Ayush Goyal · 14d
Rhythms of Work: Multi-Scale Interpretation of Human Behavioral Traces for Workplace Agents
arXiv cs.CL · Lin Ai, Scott Counts · 14d
When Do Internal Probes Beat Reading the Answer? Miscalibrated Readouts and Behavior-Concealed Knowledge in Language Models
arXiv cs.CL · Gnaneswar Villuri, Hashmath Shaik, Alex Doboli · 14d
PetQA: Benchmarking Veterinary Knowledge and Clinical Reasoning
arXiv cs.CL · Taegyun Kim, Youngwook Ham, Jungwook Rhim, Ju-Hyun An, Sungkyu Park, Kunwoo Park · 14d
Tracing Audio Grounding and Answer Selection in Audio LLMs
arXiv cs.CL · Hyebin Cho, Suho Yoo, Jihoo Jung, Joon Son Chung · 14d
CAGE: Coherence-Aware Graph Encoding for Retrieval-Augmented Generation
arXiv cs.CL · Tong Qi, Jingyu Wu, Youbing Yin, Spencer Hong, Daben Liu, Erin Babinsky · 14d
ConsensusBench: Benchmark of Consensus Nodes for LLM Reasoning via Outcome Reward Densifying
arXiv cs.CL · Shi-Qi Yan, Chao-Hong Tan, Qian Chen, Wen Wang, Xiangang Li, Zhen-Hua Ling · 14d
Choosing the Right Language Mode at Inference Time for Multilingual Reliability
arXiv cs.CL · Ekata Mitra, Ameeta Agrawal · 14d
Controlling and Assessing Appropriate Persona Use in LLM-based Dialogue Generation
arXiv cs.CL · Jongkyung Shin, Inkyu Lee, Chiehyeon Lim · 14d
How Do Language Models Represent and Use Phonological Information for Allomorph Selection?
arXiv cs.CL · Sangwoo Kim, Sangah Lee · 14d
Refuse without Refusal: A Structural Analysis of Safety-Tuning Responses for Reducing False Refusals in Language Models
arXiv cs.CL · Minji Kim, Hyounghun Kim · 14d
Knowing What Not to Answer: Selective Non-Compliance in Vision-Language Models
arXiv cs.CL · Minji Kim, Jihyoung Jang, Hyounghun Kim · 14d
Beneath the Surface of Chains-of-Thought: A Mechanistic Interpretation of Reasoning Operations in LLMs
arXiv cs.CL · Seogyeong Jeong, Jaehui Hwang, Dongyoon Han, Geonmo Gu, Alice Oh, Taekyung Kim · 14d
Vectorizing Classical Tamil: Representation Learning for Verse-Commentary Pairs
arXiv cs.CL · Amrit Gopinath, Sangeetha Sivanesan · 14d
Can Activation Steering Capture Multidimensional Authorship Style?
arXiv cs.CL · Hieu Tran, Calvin Bao, Marine Carpuat · 14d
Recurrence Is Not Enough: Causally Validating Multilingual SAE Translation Features in Gemma 2 and 3
arXiv cs.CL · Giang Son Nguyen, Nhi Ngoc-Yen Nguyen, Wray Buntine, Dung D. Le · 14d
A Systematic Comparison of Multilingual Interpretability Methods Reveals Anisotropy-Driven Failures
arXiv cs.CL · Oskar Holmstr\"om, Marcel Bollmann, Marco Kuhlmann · 14d
Reinforcement Learning for improving Large Language Models' Catalan text simplification capabilities
arXiv cs.CL · Arnau Ayguad\'e Domingo, Stefan Bott, Horacio Saggion · 14d
Generating Constructive Feedback on Stories via Reinforcement Learning
arXiv cs.CL · Maja Stahl, Timon Ziegenbein, Henning Wachsmuth · 14d
On Epistemic Diversity in Large Language Models
arXiv cs.CL · Elisabeth Kirsten, Nicole Kr\"amer, Muhammad Bilal Zafar · 14d
MABPD: Multi-Agent Bias Probing & Detection via Structured Argument Debate
arXiv cs.CL · Garvit Joshi (Graphic Era University, Dehradun, India), Stavya Dhyani (Graphic Era University, Dehradun, India), Jasmine (Graphic Era University, Dehradun, India), Arun Chauhan (Graphic Era University, Dehradun, India) · 14d
MMTClinic: Multimodal, Multilingual Time Series Question Answering and Reasoning Benchmark for Clinical Domain
arXiv cs.CL · Sourav Malakar, Harshit Nigam, Akash Ghosh, Sriparna Saha, Amlan Chakrabarti, Saptarsi Goswami, Priti Singh · 14d
CC-Mediation: Evaluating Large Language Models for Cross-Cultural Conflict Mediation
arXiv cs.CL · Suhyun Lee, Wenxuan Zhang, W. Quin Yow, Yang Deng · 14d
Cache-Aware Joint Router Adaptation for Memory-Efficient MoE Inference
arXiv cs.CL · Zhenhe Wu, Yaping Jin, Qinghua Xing, Hang Zhou, Wei He, Xianjie Wu, Xianfu Cheng, Jian Yang, Hanting Chen · 14d
RefactorPlatform: An Open-Source Harness for Controlled Evaluation of Repository-Scale Refactoring Agents
arXiv cs.CL · Aziz Ben Amor, Drish Mali, Mann Acharya, Vijayasri Iyer, S\'ebastien Brati\`eres · 14d
Discourse Dependency: A Continuous Criterion for Translation Difficulty
arXiv cs.CL · Ahrii Kim, Chanjun Park, Seong-heum Kim · 14d
BIT.UA at BioASQ 14B: Modular Retrieval with pg_textsearch and Qdrant, and Agent-Based Answer Generation
arXiv cs.CL · Andr\'e Ribeiro, R\'uben Garrido, Alexander Christiansen, Richard A. A. Jonker, S\'ergio Matos · 14d
MoirfEolas and Cr\'iochScore: Developing Resources for and the Evaluation of Tokenization Alignment with Irish Morphology
arXiv cs.CL · Jane Adkins, Abigail Walsh, Brian Davis, Elaine U\'i Dhonnchadha · 14d
Leveraging Low-Level Symbolic Competences for Unsupervised Grounding in Hallucination Detection
arXiv cs.CL · Renato Vukovic, Hsien-chin Lin, Carel van Niekerk, Benjamin Ruppik, Michael Heck, Shutong Feng, Nurul Lubis, Milica Gasic · 14d
How do LLMs Evaluate Perceived Moral Agency? Investigating Moral Decision-Making in Human-Artificial Agents Interactions
arXiv cs.CL · Fernanda Mansilla, Aloysius Tok, Bahia Guella\"i, Farah Benamara, Nancy F. Chen · 14d
Minimax Lower Bound for Estimating Diffusion-based Local Intrinsic Dimension
arXiv stat.ML · Jaehee Seo, Wontae Jeong, Jisu Kim · 14d
An Analysis of Self-supervised Pre-training with Dependent Samples
arXiv stat.ML · Maximilian Fleissner, Debarghya Ghoshdastidar, Samory Kpotufe · 14d
FluxDisco: Symbolic Regression for Stoichiometric Dynamical Systems via Monte Carlo Graph Search
arXiv stat.ML · Cassandra Durr (Lancaster University), Alvaro K\"ohn-Luque (University of Oslo), Chris Jewell (Lancaster University), Lloyd A. C. Chapman (Lancaster University) · 14d
PAC-Bayesian Reconstruction Guarantees for Time Series Variational Autoencoders
arXiv stat.ML · Chlo\'e Hashimoto-Cullen, Ghislain Agoua, Benjamin Guedj, Sylvain Le Corff · 14d
Self-Supervised Pretraining of Molecular Graph Encoders with LeJEPA
arXiv stat.ML · Micha{\l} Kulczykowski, Rafa{\l} {\L}ab\k{e}dzki · 14d
Towards AI-Driven Nanomedicine Discovery: A Benchmark and Multimodal Learning Framework for Nano Self-Assembly Prediction
arXiv stat.ML · Quan Hao, Mengyue Fan, Zifan Dong, Jianduo Zhao, Changhao Xiao, Shangqing Jiao, Hao Zhang, Yudong Wang, Fei Xia, Jigang Wang, Liguo Zhang, Chong Qiu · 14d
A Constraint-Aware Generative Framework for Synthetic Origin-Destination Demand in Logistics Networks
arXiv stat.ML · Leian Chen · 14d
Optimal Stratified Allocation for Rare-Event Onset Forecasting in Dependent Sequences
arXiv stat.ML · Jaskaran Singh · 14d
Interpretability for Turing Machines
arXiv stat.ML · Billy Snikkers, Rumi Salazar, Daniel Murfet, Will Troiani · 14d
Faster Learning under Relaxed Local Differential Privacy
arXiv stat.ML · Cristina Butucea, Huiyun Tang, Marie-Luce Taupin · 14d
Reconciling Universal and Uniform Learning with $Q$-Aggregation
arXiv stat.ML · Mikael M{\o}ller H{\o}gsgaard, Patrick Rebeschini, Tobias Wegel · 14d
Confounding-Valid Conformal Inference for Counterfactual KPIs in Wireless Networks
arXiv stat.ML · Abdessamed Qchohi, Jessica Moysen Cortes, Matteo Zecchin · 14d
Learning from VAE Errors to support ECG-based Differential Diagnosis of Myocardial Scar
arXiv stat.ML · Shayan Sharifi, Riccardo Treu, Ilaria Gandin, Federico Garoia, Marco Merlo, Giulia Cisotto · 14d
A Gaussian process model for chemoinformatics with application to the hazard classification of organic solvents
arXiv stat.ML · Arron Gosnell, Evangelos Evangelou · 14d
High-dimensional censored MIDAS logistic regression for corporate survival forecasting
arXiv stat.ML · Wei Miao, Jad Beyhum, Jonas Striaukas, Ingrid Van Keilegom · 14d
GyroSwin: 5D Surrogates for Gyrokinetic Plasma Turbulence Simulations
arXiv stat.ML · Fabian Paischer, Gianluca Galletti, William Hornsby, Paul Setinek, Lorenzo Zanisi, Naomi Carey, Stanislas Pamela, Johannes Brandstetter · 14d
On the Asymptotic Inadmissibility of Double Machine Learning Estimators Under Structure-Agnostic Models
arXiv stat.ML · Lin Liu, Rajarshi Mukherjee, James M Robins · 14d
Spectral-Target Physical Latent Structuring for JEPA-Style World Models
arXiv cs.LG · Penghao Zhu, Salvatore Penachio, Kaustav Mukherjee, Aneesh Jonelagadda · 14d
ProToMEx: Rapid, Interpretable Explanations via Structured Representations
arXiv cs.LG · Athina Georgara, Adarsh Valoor, Sarvapali D. Ramchurn · 14d
A Data Fusion Framework for Grounding Aerospace Surrogate Model via Experimental Wind-Tunnel Observations
arXiv cs.LG · Nitin Nagesh Kulkarni, Dheeraj Vemula, Yin Yu, Peter Lyu, Juan J. Alonso · 14d
Quantum-Assisted Memory-Efficient Training for Parameter-Intensive Wi-Fi-Based Human Activity Recognition
arXiv cs.LG · To Truong An, Jie Zhang, Guolin Yin, Junqing Zhang, Yanjiao Li, Trung Q. Duong, Simon L. Cotton · 14d
Evaluating Large Language Models for Forced Outage Risk Prediction: Benefits and Comparison to Machine Learning
arXiv cs.LG · Christos Petridis, Zoran Obradovic, Mladen Kezunovic · 14d
BER-PEF: Unified Human Mobility Predictability Evaluation via Bayes Error Rate Estimation
arXiv cs.LG · En Xu, Jingtao Ding, Zhiwen Yu, Yong Li · 14d
Data-Driven Learning of Unknown Nonlinear Differential Equations Using Functional Analysis
arXiv cs.LG · Seyyed Shaho Alaviani, Yongzhi Qu, Gregory W. Vogl · 14d
Modular Deep Recurrent Neural Network: Application to Quadrotors
arXiv cs.LG · Nima Mohajerin, Steven L. Waslander · 14d
SharedSAE: One Feature Dictionary Across Language Models
arXiv cs.LG · Daniil Ognev, C\'elian Vasson, Lijie Hu, Kentaro Inui, Benjamin Heinzerling · 14d
A Quantum Variational Approach to Prototypical Recurrent Unit
arXiv cs.LG · Mahyar Sadeghi Garjan, Tommaso Cesari, Michel Barbeau · 14d
On the Abundance of Critical Points of the t-SNE Energy
arXiv cs.LG · Nakul Haridas, Ryan Murray · 14d
Disentangling Attention in Deep Operator Learning: A Controlled Study of Data-Driven and Physics-Informed Architectures
arXiv cs.LG · Amar Alem Koric, Qibang Liu, Seid Koric · 14d
REFINE: LLM Refinement over Budgeted Text-Attributed Graphs for Personalized Medical Concept Representation
arXiv cs.LG · Mohsen Nayebi Kerdabadi, Arya Hadizadeh Moghaddam, Dongjie Wang, Zijun Yao · 14d
Beyond a Universal Forecasting Selector: Demand-Conditioned Model Selection across Demand Patterns and Horizons
arXiv cs.LG · Adolfo Gonz\'alez · 14d
A Repeated-Measurement Study for Cultural Analytics of English Song Lyrics Using Five Large Language Models
arXiv cs.LG · E. Cho Smith, Samuel Ho, Dawn Laux · 14d
Conformity Breaks Conformal Prediction
arXiv cs.LG · Yibo Hu, Hanyu Su · 14d
When Load-Balancing Goes Too Far: Expert Pruning in Over-Dispersed Mixture-of-Experts Models
arXiv cs.LG · Berkcan Kapusuzoglu, Connor Pryor, Sangwoo Cho, Supriyo Chakraborty, Shi-Xiong Zhang, Sambit Sahu, Milind Naphade · 14d
On-board ML for Trace Gas detection in Imaging Spectroscopy data
arXiv cs.LG · V\'it R\r{u}\v{z}i\v{c}ka, Adam Chlus, Andrew Thorpe, David R. Thompson · 14d
Nested Inductive Bias Framework for SPD Manifold Learning
arXiv cs.LG · Tushar Das · 14d
Hakken: Predicting future discoveries to fill the gaps in today's knowledge
arXiv cs.LG · Tarek R. Besold, Uchenna Akujuobi, Pablo Sanchez, Alessandra Toniato, Kana Maruyama, Jihun Choi, Samy Badreddine, Frederick Gifford, Daniel Evans-Yamamoto, Sucheendra K. Palaniappan, Miquel Ferrer, Kae Nagano, Iris Rossell, Tom Joy, Hatem ElShazly, Chrysa Iliopoulou, Christoph Wehner, Thiviyan Thanapalasingam, Susana Nunes, Pedro G. Cotovio, Peter Wurman, Peter Stone, Hiroaki Kitano, Michael Spranger · 14d
An Energy-Based Conservative-Dissipative Latent Neural Evolution Operator for Magnetization Dynamics
arXiv cs.LG · Sebastian Schaffer, Lukas Exl · 14d
Distilled Continuous Diffusion Language Models Can Write Code in Few Steps---or One
arXiv cs.LG · Fred Zhangzhi Peng, Kaiwen Zheng, Anru R. Zhang · 14d
Mitra-v2 Technical Report
arXiv cs.LG · Yefan Tao (Bernie), Xiyuan Zhang (Bernie), Xinyi Liu (Bernie), Boran Han (Bernie), Danielle Maddix (Bernie), Haoyang Fang (Bernie), Zhen Han (Bernie), Jiading Gai (Bernie), Xuanqing Liu (Bernie), Michael Bohlke-Schneider (Bernie), Yuyang (Bernie), Wang, Gerald Friedland, Kevan Mah, Chris Lee, Chris Kong · 14d
Fast Surrogate Modeling of Excitable and Oscillatory FitzHugh-Nagumo Dynamics with Parametric Neural Operators
arXiv cs.LG · Andrew Franck, Justin Li · 14d
Training-Free Halving of Activated Experts in Fine-Grained Mixture-of-Experts Models
arXiv cs.LG · Xing Chen, Hengshuai Yao · 14d
Optimizer Memory Schedules for Outscaling the Overtraining Axis
arXiv cs.LG · Katie Everett, Shikai Qiu · 14d
Representation Redundancy and Structural Complexity in Finite-Field Inversion
arXiv cs.LG · Zheng Zhang, Na Zhang · 14d
GNN-Guided Graph Coarsening and Adaptive QUBO Penalties for the Capacitated Vehicle Routing Problem with Time Windows on a Quantum Annealer
arXiv cs.LG · Youssef Kamel Rezk, Pawe{\l} Gora · 14d
Too Rare to Learn: Prescribed Cyclone Tracks Degrade a Bay of Bengal Ocean Emulator
arXiv cs.LG · Sumaiya Islam · 14d
SMILE: Bridging Continuous Optimization and Discrete Symbolic Recovery
arXiv cs.LG · Mansooreh Montazerin, Antonio Ortega, Ajitesh Srivastava · 14d
WEECFP-SuRGE: Wide Embedded Extended Connectivity Fingerprint with Substructure Rotary Graph-distance Encoding
arXiv cs.LG · Robert Epps · 14d
Simulation-free Unbalanced Dynamic Optimal Transport with General Growth Penalty
arXiv cs.LG · Junda Ying, Yuxuan Wang, Bowen Yang, Peijie Zhou, Lei Zhang · 14d
Locating and Steering Refusal Beyond Attention
arXiv cs.LG · Preethi Carmel Bosco, Gopalakrishnan Srinivasan · 14d
Training Large Language Models for Small-Molecule Design with Synthetic Task Scaling
arXiv cs.LG · Frank Hu, Shriram Chennakesavalu, Zichen Wang, Patricia Suriana, Bodhi Vani, Kirill Shmilovich, Kangway Chuang, Colin Grambow · 14d
A Fairness Audit of the Duckworth-Lewis-Stern Method: Format-Specific and Gender-Differential Bias, with an Interpretable Calibration Layer for Cricket Target Revision
arXiv cs.LG · Soumyadeep Roy · 14d
Resilience Beyond Stationary Client Unavailability: Unlocking Efficient and Unbiased Federated Learning
arXiv cs.LG · Ming Xiang, Stratis Ioannidis, Edmund Yeh, Carlee Joe-Wong, Lili Su · 14d
A Robust Watermark-based Fingerprint Framework for GNNs Ownership Verification
arXiv cs.LG · Han Zhang, Yan Wang, Guanfeng Liu, Pengfei Ding, Huaxiong Wang, Kwok-Yan Lam · 14d
Persistent Teacher Anchoring for Tool-Using Agents
arXiv cs.LG · Hyun Bin Park (Sogang University), Kyungho Song (University of Michigan, Ann Arbor), Sangmin Lee (Sogang University), Du-Seong Chang (Sogang University) · 14d
Dynamic Heterogeneous Graph Representation Learning: A Survey
arXiv cs.LG · Huan Liu, Pengfei Jiao, Jie Yin, Hongjiang Chen, Zhidong Zhao · 14d
Learning-Augmented Algorithms: Guarantees, Construction Mechanisms, and System-Level Implications
arXiv cs.LG · Hailiang Zhao, Peng Chen, Xueyan Tang, Jianwei Yin, Shuiguang Deng · 14d
How Faithful Is Attribution for Sales Forecasting? A Counterfactual Study
arXiv cs.LG · Glib Kechyn · 14d
Federated Attack Campaign Detection via Contrastive Encoding of Threat Indicators in Gradient Updates
arXiv cs.LG · Manuel R\"oder, Bibin Babu, Frank-Michael Schleif · 14d
Communication-Efficient Personalized Federated Learning via Layer-Wise Multi-Threshold Random Sketching
arXiv cs.LG · Xu Zhang, Xingyu Hou, Jiacheng Cheng, Kaiyuan Feng, Maoguo Gong · 14d
PACE: Propagation-Aware Collaborative Correction for One-Shot Personalized Federated Graph Learning
arXiv cs.LG · Ruizhe Huang, Chengran Li, Xiaochuan Shi · 14d