Page 23 of 192

AllHigh signalRecent
7679 articles
arXiv cs.AI·

Tool-Augmented Agent for Closed-loop Optimization,Simulation,and Modeling Orchestration

COSMO-Agent, a tool-augmented RL framework, trains LLMs to orchestrate iterative CAD-CAE processes. The system learns to generate parametric geometry, solve simulations, and revise designs under multiple constraints. Industry-aligned dataset covering 25 component categories. Trained small LLMs outperform large open-source and closed-source models in feasibility and stability.

AI AgentsReinforcement learningTools
SIG
78
HYP
25
arXiv cs.AI·

ScenePilot: Controllable Boundary-Driven Critical Scenario Generation for Autonomous Driving

ScenePilot generates critical scenarios for autonomous driving testing via multi-objective reinforcement learning. The framework combines RSS-derived physical feasibility with an AV-risk predictor to target boundary-band scenarios: physically solvable yet causing failures. Results: +6.2 percentage points collision rate on SafeBench while preserving physical validity.

Reinforcement learningAI safetyEvals
SIG
78
HYP
15
arXiv cs.AI·

DeepWeb-Bench: A Deep Research Benchmark Demanding Massive Cross-Source Evidence and Long-Horizon Derivation

DeepWeb-Bench is a deep research benchmark evaluating 9 frontier models on tasks requiring massive evidence collection, cross-source reconciliation, and long-horizon multi-step derivation. Errors stem primarily from derivation and calibration (>70%), not retrieval (12-14%). Strong and weak models fail differently: incomplete derivation vs hallucinated precision.

BenchmarksReasoningAI Agents
SIG
78
HYP
25
arXiv cs.LG·

Conformal Selective Acting: Anytime-Valid Risk Control for RLVR-Trained LLMs

CSA (Conformal Selective Acting) is a deployment wrapper for RLVR-fine-tuned LLMs guaranteeing per-round risk control without pooling across deployments. Tested on 480 specialist streams and 10,300 Expert-Iteration rounds with LoRA, CSA maintains a Ville e-process per threshold and achieves selective-risk bound R_T^act ≤ α+O(N_T^{-1/2}) with anytime pathwise validity.

Reinforcement learningAI safetyEvals
SIG
78
HYP
15
arXiv cs.CL·

Mechanics of Bias and Reasoning: Interpreting the Impact of Chain-of-Thought Prompting on Gender Bias in LLMs

arXiv study on Chain-of-Thought (CoT) impact on gender bias in LLMs. Researchers combine benchmark evaluation, mechanistic interpretability, and reasoning chain analysis. Finding: CoT does not consistently reduce bias gaps; observed improvements stem from memorization rather than genuine understanding, with gender bias remaining embedded in hidden representations.

ReasoningAI safetyAlignment
SIG
78
HYP
15
arXiv cs.LG·

OmniISR: A Unified Framework for Centralized and Federated Learning via Intermediate Supervision and Regularization

OmniISR proposes a unified framework for centralized and federated learning via intermediate supervision and regularization. The framework uses mutual information to align internal covariate shifts and negative entropy to regularize overconfident predictions. O(1/sqrt(T)) convergence guaranteed theoretically; CL-FL gap reduced by 22.60% in experiments.

Reinforcement learningAlignmentPapers
SIG
78
HYP
15
arXiv cs.CL·

Retrieval-Augmented Long-Context Translation for Cultural Image Captioning: Gators submission for AmericasNLP 2026 shared task

Two-stage pipeline for captioning cultural images in Indigenous languages: Qwen2.5-VL generates Spanish intermediate caption, then Gemini 2.5 Flash produces target-language caption via retrieval-augmented prompting. Achieves 164.1% (Bribri), 131.7% (Guaraní), 122.6% (Orizaba Nahuatl) improvements over baseline. Overall winner of AmericasNLP 2026 shared task.

VisionRAGGemini
SIG
78
HYP
25
arXiv cs.CL·

Distributional Alignment as a Criterion for Designing Task Vectors in In-Context Learning

New d_NTP metric evaluates task vector quality in ICL by measuring alignment of next-token probability distributions. Linear Task Vector (LTV) method minimizes d_NTP via closed-form linear regression, improves accuracy by 9.2% across 8 benchmarks and 5 LLMs, reduces inference latency. Task vectors transferable across model scales (+6.4% for smaller model).

Prompt engineeringReasoningBenchmarks
SIG
78
HYP
15