←── back to feed
/topics/arxiv-computational-linguistics-papers-september-23
arXiv computational linguistics papers September 23
31 items●1 sources●updated 22h ago●trend 2
[BLG]blog/rss31
What Does 99% Accuracy Measure? A Reproducible Audit of Shortcut Learning in a Widely Used Fake News Corpus
Training a Language Model End-to-End in Rust: An Experience Report
Same Quantity, Different Answer: Numerical Representation Invariance in Language Models
A Computational Approach to Measuring Semantic Change in Sanskrit Literature
Retrieved-Span Training for Efficient Query-Focused Meeting Summarization on QMSum
From Tone to Trajectory: Continuous Sentiment and the Shape of Monetary Policy Communication
Peerify: Benchmarking Peer-Review Claim Verification
AIBuildAI-2.5: Efficient Autonomous AI Model Development Through LLM-Guided Tree Search
Prompt Breadth and Rollout Refresh Interact in On-Policy Distillation
Mitigating LLM Over-Refusal via Dynamic Semantic Routing Calibratione
FrontierMath Erd\H{o}s
LLM-Driven Training-free Location-Attribute Synergic Fusion: A Closed-Loop Paradigm for Dual-source Encrypted POIs and LULC Mapping
Self-Cleaning and Captured Anyway: One Measured Primitive for Error in a Store an Agent Writes to Itself, and What a Falling Score Actually Measures
LatentPort: Beyond KV Cache - Cross-Model Transfer of Recurrent Memory in Hybrid Language Models: A 4B-to-9B Hybrid-State Handoff Without Target Prefix Replay
MoM: Memory of Memory
ICDAR2026 Competition on Multimodal Reasoning over Documents in Multiple Domains
Graph-Based Inference for Feedback-Driven Word Deduction: A Scalable Framework for the Jotto Problem
ChainDoRA: Tensor-Train Factorized Weight-Decomposed Low-Rank Adaptation for Parameter-Efficient LLM Fine-Tuning
Understanding Reliability in LLM-based Human Behavior Simulation
ufakzeka-1: Building and Evaluating a 151M-Parameter Turkish Language Model from Scratch
Impact Is Not Invalidation: Ask About the Claim, Not the Diff
FinFIRST: Benchmarking Search Agents for Financial Information Retrieval, Sourcing and Traceability
FineWeb-CLaR: Culture, Language, and Region Annotations for Benchmark-Aligned Corpus Auditing
TelecomGPT-R1: Unified Post-Training for Reasoning Across Heterogeneous Telecom Tasks
Passes Alone, Fails Together: Benchmarking Semantic Coordination in Parallel LLM-Agent Development
Mining Legal Arguments in U.S. Corporate Case Law
Conduct Under Pressure: What Sixty Language Models Do When a User Pushes
Matryoshka attribution: Learning to attribute language model outputs to representations and weights
Compressing Long Context into Answer-Aligned Memory Embeddings for LLM Inference
Qwen3.8-Omni: Towards Native Omni-Modal Agents
From Utterances to Networks: Modelling Slang Adoption and Diffusion Across Subreddits