Page 10 of 192

AllHigh signalRecent
7679 articles
arXiv cs.AI·

When the Tool Decides: LLM Agents Defer Blindly to Graph Neural Network Tools, and Stronger Backbones Defer More

LLM agents equipped with GNN tools fail to exercise judgment: they blindly adopt the GNN's predictions 97.6-99.2% of the time. This deference increases with model capability (Qwen2.5 0.5B-7B), creating a 'GNN parrot' that bypasses its own reasoning. Simple alternatives outperform the GNN at high homophily, yet the agent still defers.

AI AgentsBenchmarksReasoning
SIG
78
HYP
15
arXiv cs.CL·

LLMs Contain Multitudes: How Deployment Context Reshapes Model-Level Preferences and Values

Study of 1.2M decisions showing deployment context (Reddit vs news article) produces far larger variations in model preferences and values than prompt paraphrasing or temperature controls. Measured biases (Global North favoritism) and cardinal exchange rates between outcomes shift by factor 2.47 across contexts, questioning stability of model-level properties.

EvalsAI safetyAlignment
SIG
78
HYP
15
arXiv cs.CL·

OdysSim: Building Foundation Models for Human Behavior Simulation

OdysSim presents the largest systematic investigation of behavioral foundation models for human behavior simulation. Researchers propose SOUL, a taxonomy of 5 axes (CONV, SS, COG, ROLE, EVAL) unifying 62 datasets and 23 benchmark tasks. The open 8B OSim model ranks first/tied-first on 8/23 tasks, outperforming frontier models, with 93.2% reaction alignment vs 93.5% for real users.

BenchmarksReasoningReinforcement learning
SIG
78
HYP
25
arXiv cs.CL·

Does the Judge Prefer English? Evaluating Language-Switching Invariance in LLM-as-a-Judge

Judge-LS evaluates whether LLMs used as automatic judges exhibit language bias. On 419 LLMBar benchmark items transformed into English, Chinese, and mixed-language variants, models show 10.7–14.4% preference flips across languages, with highest accuracy in English. Translation-equivalent probes reveal no systematic English preference, though most are judged as ties.

EvalsBenchmarksAI safety
SIG
78
HYP
15
arXiv cs.AI·

VeriGeo: Controllable Geometry Question Generation with Numerical and Analytical Verification

VeriGeo generates controllable geometry problems via executable reasoning traces. An Author agent creates the problem and diagram per user constraints, a Solver agent produces the proof. A three-stage pipeline verifies numerical, analytical, and global consistency. Fine-tuning on 8.7k examples achieves best reported GeoQA performance and strong results on PGPS9K and MathVista-GPS.

ReasoningVisionBenchmarks
SIG
78
HYP
15
arXiv cs.AI·

Topical Phase Transitions in Artificial Intelligence Research: Large-Scale Evidence and an Early-Warning Signature for Emerging Topics

Analysis of 80,814 papers from 5 major AI conferences (2017-2025) reveals research topics advance through abrupt phase transitions, not gradually. LLMs dominant by 2025; diffusion models and vision-language models surged within 1-3 years. Early-warning signature flags reasoning, test-time compute, agentic AI, multimodal LLMs, RAG, and world models as topics to monitor 2026-2028.

BenchmarksPapersReasoning
SIG
78
HYP
25