[HN]HN: RAG/bbc.com20h
Farhud memories: Baghdad's 1941 slaughter of the Jewsmarysminefnuf·▲ 4
[BSKY]@natolambert/bsky.app1d
My best guess is that for scaled RL most of the top Chinese AI labs are starting to use a lot of Huawei for inference and Nvidia for training (maybe not for weird architectures). As agent swarms, even more scaled post-training, etc becomes…@natolambert.bsky.social·▲ 21
[BLG]GitHub Trending — Python/github.com19h
mcncarl/yichen-skills[HN]HN: agents/dev.yuv.run19h
While the human's away, do the agents slip into foul play?alkisyuv·▲ 1
[BLG]arXiv cs.CL/arxiv.org19h
How Many Humans Is a Judge Panel Worth?Chao Li, Yingying Yu, Yunfeng Li
[BSKY]@emollick/bsky.app1d
Back in July, there was a report that got a lot of attention on BlueSky that Waymo was more dangerous than NYC for-hire vehicles. The analysis turned out to be wrong and the updated report finds they are much safer. Good that they updated …@emollick.bsky.social·▲ 134
[HN]HN: RAG/orca.radioglaciology.com21h
Open Radar Code Architecturetoomuchtodo·▲ 3
[BSKY]@emollick/bsky.app1d
It is ironic that the thing that is now most annoying about long-running agentic tasks with Large Language Models isn't coding or errors or hallucinations, but the fact that their language gets worse due to drift & cross-agent talk as a ta…@emollick.bsky.social·▲ 66
[BLG]arXiv cs.CL/arxiv.org19h
Beyond Atomic Tokens: Factorizing Syllables for Language Model PretrainingNghia Hieu Nguyen, Thai Bao Huynh, Binh-An Dinh-Le, Phu Gia Hoang, Dat Tien Nguyen, Kiet Van Nguyen, Ngan Luu-Thuy Nguyen
[BSKY]@hardmaru/bsky.app1d
日経新聞による、グーグルディープマインド東京を率いる全炳河(Heiga Zen)氏の素晴らしい特集記事。2018年、Heigaさんと2人でGoogle Brain東京チームを立ち上げた日々を懐かしく思います。当時は時差の厳しい深夜の会議をこなしながら、日本のAI研究の存在感を示すために必死でした。現在、彼がGDM Tokyoを率い、私が Sakana AI を起業して、東京のAIエコシステムがここまで大きく成長したことを本当に嬉しく思います!@hardmaru.bsky.social·▲ 9
[BLG]arXiv cs.CL/arxiv.org19h
When Better Turns Do Not Make Better Agents: Diagnosing the Gap Between Next-Turn Metrics and Workflow SuccessMd Tahmid Rahman Laskar, Xue-Yong Fu, Gundeep Singh, Karol Chang, Kevin Sanders, Shi Zong, Tania Habib, Julien Bouvier Tremblay, Shayna Gardiner, Harsh Saini, Matthias Lee, Elena Khasanova, Quinten McNamara, Shashi Bhushan TN
[HN]HN: GPT/github.com22h
Show HN: Lain, a structural code graph and agent coordinator for coding agentsspuentesp·▲ 3
[BLG]arXiv cs.CL/arxiv.org19h
From Task Success to Productive Success: Evaluating Human-AI Collaboration by Quality and CostSaki Imai, Mert \.Inan, Malihe Alikhani
[HN]HN: GPT/github.com21h
Deterministic grounding checks for LLM agents, no LLM in the hot pathtjsandhu·▲ 1
[HN]HN: GPT/github.com22h
Recursive Cognitive Optimization (RCO)sealvarezl·▲ 1
[BLG]GitHub Trending — Python/github.com19h
mem0ai/mem0[HN]HN: GPT/calmscroll.com21h
The Great Gatsbyamadeuspagel·▲ 2
[BLG]GitHub Trending — Python/github.com19h
NVIDIA/TensorRT-LLM[HN]HN: AI/theintercept.com21h
Flock Partnered with Nonprofit That Uses AI to Rally Public Supportdp-hackernews·▲ 5
[BLG]arXiv cs.CL/arxiv.org19h
Consistent Relexicalization of Clinical Documents using Graph-Based ApproachDipankar Das, Atri Mandal, Sandeep Singh, Tushar Shandhilya
[BLG]arXiv cs.CL/arxiv.org19h
Omni Demand Understanding: A Benchmark for Contextual User-Intent Inference in Multimodal InteractionQi Chen, Yunfei Chu, Haolin He, Yifan Yang, Zihan Liu, Yuxuan Wang, Ziyang Ma, Ruiyang Xu, Meng Gao, Yinsong Yan, Ling Wang, Hui Wang, Wen Huang, Yiheng Chen, Guanrou Yang, Qiuqiang Kong, Jin Xu, Xie Chen
[HN]HN: AI/cbsnews.com21h
Nvidia's Jensen Huang rejects AI extinction warnings as "doomsday narratives"AlexDragusin·▲ 7
[HN]HN: RAG/hollywoodreporter.com19h
Presley Gerber, Son of Cindy Crawford and Rande Gerber, Dies at 27andsoitis·▲ 1
[BLG]arXiv cs.CL/arxiv.org19h
Talking Past the Machine: Morality, Politeness, and Alignment in Human-AI DialogueMarina Mitiaeva, Lu Xiao
[HN]HN: GPT/github.com21h
Applesauce: Transparent Compression for Apple File System Compression (AFSC)colechristensen·▲ 1
[BLG]arXiv cs.CL/arxiv.org19h
From Memory to Behavior: A Behavior-Aware Role-Playing Framework for Social Media InfluencersJi-Lun Peng, Yi-Zhen Zhang, Chun-Nan Chou, Yun-Nung Chen
[BLG]arXiv cs.CL/arxiv.org19h
When Steering Fails in Latent Reasoning: A Latent-to-Language Transition GapGaoxiang Huang, Lei Qi
[HN]HN: GPT/paramrathour.github.io20h
Coding Theory: A Playful Introductionvismit2000·▲ 3
[BLG]arXiv cs.CL/arxiv.org19h
Not All Irregularity Is Equal: Causally Isolating a Rare Failure Mode in Japanese Morphological InflectionWen Zhang
[HN]HN: GPT/github.com21h
Jev – System-1 Agent Architecture Radar (open-source)noobplus·▲ 3
[BLG]arXiv cs.CL/arxiv.org19h
Analysing the Linearity of Linguistic Relations in Language Model Embedding SpacesVasudevan Nedumpozhimana, Fathima Thekkekara, John Kelleher
[HN]HN: GPT/github.com22h
Two similar AI judges fail together 7.7x more often than independence predictsLawless1987·▲ 2
[HN]HN: RAG/en.wikipedia.org19h
Rally Englishred369·▲ 1
[BLG]GitHub Trending — Daily/github.com19h
yynxxxxx/Codex-X[HN]HN: RAG/wikifarmer.com20h
Reviving Deserts with Syntropic Agroforestrylawrenceyan·▲ 2
[BLG]arXiv cs.CL/arxiv.org19h
MIRAGE: Multi-Perspective Creative Language Model Reasoning with Reinforcement Learning GuidanceArash Lagzian, Srinivas Anumasa, Dianbo Liu
[BLG]arXiv cs.CL/arxiv.org19h
Prediction Dynamics in Depth-Recurrent Language ModelsXinyue Luo, Fei Yu
[HN]HN: AI/openprompt.tech22h
OpenPrompt – one workspace for developers using multiple AI coding toolsharafernando·▲ 2
[HN]HN: AI/lemire.me20h
AI is breaking the academic sorting machinechmaynard·▲ 1
[BLG]GitHub Trending — Daily/github.com19h
higgsfield-ai/higgsfield[HN]HN: AI/artoftheproblem.com20h
Art of the Problem Launches $99 AI botbritcruise·▲ 3
[BLG]arXiv cs.CL/arxiv.org19h
$\mu^2$-Bench: A Multilingual Machine Unlearning BenchmarkKyomin Hwang, Hyeonjin Kim, Hyunho Lee, Yearim Kim, Yeji Song, Nojun Kwak
[HN]HN: GPT/github.com22h
Sub-15ms, non-autoregressive, local drop-in alternative to TypeSafe Jev0x1997·▲ 5
[BLG]arXiv cs.CL/arxiv.org19h
CoLearn: An Agentic Tutor that Learns its Learner in a Human--AI Co-Learning LoopKailai He, Zhihao Wu, Linhai Zhang, Runcong Zhao, Yulan He, Jiazheng Li
[HN]HN: RAG/pollresults.org22h
Show HN: Pollresults.org – free US election opinion poll APIronbenton·▲ 1
[BLG]arXiv cs.CL/arxiv.org19h
Chinese Competitive Debating Dataset and BenchmarkZongrui Yang, Haoyuan Li, Zhongsheng Wang, Zhirui Zeng, Pengqian Han, Yi Zhou, Yuting Wang, Jiamou Liu
[HN]HN: AI/allstead.dev21h
Why Bother with Lint Rules If AI Writes the Code?willio58·▲ 3
[BLG]arXiv cs.CL/arxiv.org19h
Enhancing Audio Reasoning via Semantic Summary PredictionFrancesco Bonzi, Pooneh Mousavi, Cem Subakan, Mirco Ravanelli
[HN]HN: GPT/github.com20h
Show HN: Ambits – agentic grep/rg tool will history trackingjoshLong145·▲ 3
[BLG]arXiv cs.CL/arxiv.org19h
Steering LLMs Responses Towards Moral Foundations on the Norwegian MFQ-30Hans Andersen, David Dichas
[BLG]GitHub Trending — Python/github.com19h
PenglongHuang/chinese-novelist-skill[HN]HN: GPT/github.com20h
Show HN: Hibi – An open-source obsidian alternativeryanamay·▲ 2
[BLG]arXiv cs.CL/arxiv.org19h
Reading Anxiety or Reading the Label? Comparing Fine-Tuned and Frontier Models for Anxiety Detection on Social MediaCris Huynh, Arlene Pham
[HN]HN: AI/theguardian.com19h
A Silicon Valley radical: Trump's AI whisperer pushing for limited regulationkuerbel·▲ 2
[BLG]arXiv cs.CL/arxiv.org19h
Beyond Reference-Based Evaluation: Reward Models for Meta-Evaluation of Grammatical Error CorrectionRuotian Wu, Bill E. Johnson, Gene Saunders, Osama Hamzeh, Ankit Vadehra, Pascal Poupart
[HN]HN: LLM/nytimes.com20h
Want to Entice New Residents? Offer Cash, for a Startlxm·▲ 1
[BLG]GitHub Trending — Daily/github.com19h
cloudflare/quiche[HN]HN: AI/bitu79.substack.com20h
AI Weekly Warns Firms on Google AI Studio Data Retention FraudCanusLupus79·▲ 21
[HN]HN: GPT/github.com22h
Show HN: Gdocs-me-up: a high-fidelity Google Docs exporterbehdad·▲ 3
[BLG]arXiv cs.CL/arxiv.org19h
ArenaFlow: From Trajectory Ranking to Hierarchical Credit Propagation for Open-Ended Agent RLQiang Zhang, Ruixue Ding, Fanrui Zhang, Xi Chen, Boli Chen, Shihang Wang, Yinfeng Huang, Yi Zheng, Pengjun Xie, Kaipeng Zhang, Jiawei Liu, Zheng-Jun Zha
[HN]HN: RAG/publish.obsidian.md20h
Relativistic Raytracingvismit2000·▲ 1
[BLG]arXiv cs.CL/arxiv.org19h
Aligning with Lived Experience: Heterogeneous Benefits of Fine Tuning in Mental Health Support GenerationMohit Chandra, Nabin Kim, Eli Min, Aamogh Sawant, Tanmay Sutar, Munmun De Choudhury
[BLG]arXiv cs.CL/arxiv.org19h
Reading Less While Writing: A Closed-Form Bandwidth Dial for Streaming Multimodal DecodersYasir Mehmood, Kashif Javed
[HN]HN: AI/proofatlas.ai22h
Fundamental Theorem of Calculus [Slop]throwyonion·▲ 1
[BLG]arXiv cs.CL/arxiv.org19h
Generative Artificial Intelligence Chatbots for Motivational Interviewing: A Scoping Review From System Design to Intervention OutcomesRunze Hu, Jingqi Kong, Yang Yang, Yihang Yang, Jingyao Liu, Haizhou Tang, Shanghang Zhang, Zheng Liu
[HN]HN: AI/wenbo.site19h
A Builder's Simple Perspective on the Future of the Job Markettracyhenry·▲ 1
[BLG]GitHub Trending — Daily/github.com19h
ruanyf/weekly[HN]HN: RAG/vsqrd.com22h
Show HN: Vsqrd – AWS for Biology Experimentsmisterchocolat·▲ 1
[HN]HN: AI/chabad.org21h
Can I Let My AI Agent Run on Shabbat?some-guy·▲ 30
[BLG]arXiv cs.CL/arxiv.org19h
Scaling Forced Alignment to End-User DevicesLawry Sorenson, Michael Crandall, Eric K. Ringger, Stephen D. Richardson
[HN]HN: GPT/github.com21h
ReBarUEFI: Resizable BAR for almost any UEFI systemnateb2022·▲ 2
[BLG]arXiv cs.CL/arxiv.org19h
MME-Safety: A Fine-grained Benchmark for Safety Evaluation of MLLMsYueming Lyu, Yilian Shi, Haoxiang Tan, Linzhuang Zou, Qihao Wang, Guihua Yu, Jie Qin, Xin Gao, Chenyang Si, Jing Dong, Caifeng Shan
[HN]HN: GPT/github.com20h
ZCode, embroiled in a controversy over stealing user code, is now open sourcelinzhangrun·▲ 3
[BLG]arXiv cs.CL/arxiv.org19h
Rewarding Efficient Reasoning Improves Abstention on Underspecified Tasks in Reasoning ModelsPolina Tsvilodub, Max H\"oth, Michael Franke, Bj\"orn Deiseroth, Carina Kauf
[BLG]arXiv cs.CL/arxiv.org19h
Geometry of Values: Task Vector Composition for Ethical Preference Alignment in Language ModelsUtkarsh Agarwal, Monojit Choudhury
[HN]HN: GPT/github.com21h
ZCode: Z.ai's coding agent harness. Powerful, intelligent, extensibledoppp·▲ 2
[BLG]arXiv cs.CL/arxiv.org19h
Hallucination-R1: Robustness-Oriented Paraphrase Generation for Factual ConsistencyWenhan Yu, Wenxin Wu, Hao Wang, Lei Sha
[HN]HN: GPT/github.com20h
Memoization application that you have used but might not knowvismit2000·▲ 1
[BLG]arXiv cs.CL/arxiv.org19h
When Does Reasoning Help in Machine Translation? A Hierarchical Analysis of LRM Reasoning TracesYuxiang Liu, Jiaming Luo, Eleftheria Briakou, Colin Cherry
[HN]HN: RAG/nytimes.com19h
NYTimes Journalist Regrets Long Ordeal Sending Daughter to Worse NY Schoolsdoctorpangloss·▲ 1
[BLG]arXiv cs.CL/arxiv.org19h
Evaluating In-Context Learning and Retrieval Strategies for Devanagari Post-OCR CorrectionAbhishek Bhandari, Gaurav Harit
[HN]HN: RAG/thesignalmemo.substack.com19h
I Investigated Coinbase a Year Ago For Forbes. The Bigger Story Now Is Its Powersindhya1·▲ 1
[HN]HN: GPT/github.com20h
Skills to Make the Perfect WebsiteNoamank·▲ 2
[BLG]arXiv cs.CL/arxiv.org19h
Benchmarking Gender Bias in Machine Translation Evaluation Metrics across OccupationsOrfeas Menis Mastromichalakis, Giorgos Filandrianos, Wafaa Mohammed, Giuseppe Attanasio, Chrysoula Zerva