←── back to feed
/topics/arxiv-computational-linguistics-papers-september-21
arXiv computational linguistics papers September 21
39 items●1 sources●updated 20h ago●trend 2
[BLG]blog/rss39
Do small language models know what they don't know?
HERMES: Contrast-Aware Knowledge Graph Reasoning from Clinical Notes for Patient Outcome Prediction
TALON: A Temporally Aware Longitudinal Framework for Radiology Report Generation
From Discharge Notes to Patient Understanding: Persona-Grounded, Open-Ended Simulation of LLMs as Discharge Educators
Beyond WER: Entity and Disfluency Recall in Accented Conversational ASR
SAGE: Schema-Guided LLMs for Grant Review
Reviser: Revision-Capable Text Generation via Autoregressive Cursor Actions
Recursive Language Models Generalize Out of Domain
TatBLiMP: A Benchmark of Linguistic Minimal Pairs for Tatar
Transsion's Speaker-Attributed Multilingual ASR System for the MLC-SLM 2026 Challenge
Towards Secure Cloud-Native Computing: Unveiling Kubernetes Misconfigurations with Large Language Models
A Generative Grammar Underlying the Voynich Manuscript, the Pastiche Hypothesis: Evidence from Large Language Models
PhysioBench: A Unified Benchmark for Physiological Signal Question Answering
From Generation to Detection: Exploration of Discourse Driven Scenario based LLM Generated Fake News
Curriculum-Based Noise Adaptation for Phoneme-to-Text Reconstruction in Visual Speech Recognition
COAL-SQL: Coverage-Guided Augmentation and Failure-Driven Learning for Text-to-SQL Post-Training
VISPATH: Visual-Intent-Guided Path Reasoning for Multimodal Knowledge Graph Question Answering
Boosting Deepresearch and LongContext Ability with Self-Generated Deepresearch Rollouts Traces
Reading Less While Writing: A Closed-Form Bandwidth Dial for Streaming Multimodal Decoders
Rewarding Efficient Reasoning Improves Abstention on Underspecified Tasks in Reasoning Models
Reading Anxiety or Reading the Label? Comparing Fine-Tuned and Frontier Models for Anxiety Detection on Social Media
Enhancing Audio Reasoning via Semantic Summary Prediction
MME-Safety: A Fine-grained Benchmark for Safety Evaluation of MLLMs
Generative Artificial Intelligence Chatbots for Motivational Interviewing: A Scoping Review From System Design to Intervention Outcomes
$\mu^2$-Bench: A Multilingual Machine Unlearning Benchmark
Aligning with Lived Experience: Heterogeneous Benefits of Fine Tuning in Mental Health Support Generation
Geometry of Values: Task Vector Composition for Ethical Preference Alignment in Language Models
From Task Success to Productive Success: Evaluating Human-AI Collaboration by Quality and Cost
Scaling Forced Alignment to End-User Devices
CoLearn: An Agentic Tutor that Learns its Learner in a Human--AI Co-Learning Loop
Not All Irregularity Is Equal: Causally Isolating a Rare Failure Mode in Japanese Morphological Inflection
When Better Turns Do Not Make Better Agents: Diagnosing the Gap Between Next-Turn Metrics and Workflow Success
Hallucination-R1: Robustness-Oriented Paraphrase Generation for Factual Consistency
Beyond Reference-Based Evaluation: Reward Models for Meta-Evaluation of Grammatical Error Correction
When Does Reasoning Help in Machine Translation? A Hierarchical Analysis of LRM Reasoning Traces
How Many Humans Is a Judge Panel Worth?
From Memory to Behavior: A Behavior-Aware Role-Playing Framework for Social Media Influencers
Beyond Atomic Tokens: Factorizing Syllables for Language Model Pretraining
ArenaFlow: From Trajectory Ranking to Hierarchical Credit Propagation for Open-Ended Agent RL