Archives

June 2026

2731 articles

arXiv cs.CL·

Nemotron 3 Ultra: Open, Efficient Mixture-of-Experts Hybrid Mamba-Transformer Model for Agentic Reasoning

NVIDIA introduces Nemotron 3 Ultra, a 550B-parameter (55B active) Mamba-Transformer MoE hybrid model pre-trained on 20T tokens with 1M context length. Uses SFT, RL, and multi-teacher distillation. Achieves ~6x inference throughput of public LLMs with comparable accuracy. Base, post-trained, and quantized checkpoints, training data, and recipe open-sourced on HuggingFace.

AI AgentsReasoningOpen source
SIG
82
HYP
35
arXiv cs.CL·

Beyond Layer Importance in Layer-wise Sparsity: An Inter-Layer Perturbation-Absorption Perspective

Study on layer-wise redundancy in LLMs. Authors characterize how layers absorb or amplify perturbations during pruning: early layers amplify, middle and late layers absorb. They propose absorption-aware correction using a per-layer absorption coefficient, improving OWL and AlphaPruning by 7.13% perplexity reduction and 1.02% zero-shot accuracy boost at 70% sparsity.

PapersBenchmarksFine-tuning
SIG
78
HYP
15
arXiv cs.AI·

CogGuard: Cognitive and Operational Profiling for Proactive Warning in Edge Intelligent Services

CogGuard is a proactive-warning framework for edge intelligent services using offline LLMs to build cognitive and operational profiles, then online SLMs for real-time scoring. Achieves 48% reduction in profile construction time and 19% in distributed fine-tuning on heterogeneous clusters. Reduces prediction error by 15.4% vs strongest baseline on educational datasets.

ReasoningFine-tuningBenchmarks
SIG
72
HYP
18
arXiv cs.AI·

Fusion is not one-size-fits-all: Cross-Modal Representation Alignment for Time-to-Event Modeling

Multimodal fusion framework for time-to-event prediction (PE mortality, CVD outcomes) aligning CT and longitudinal EHR representations using foundation models. Four strategies tested (late fusion, contrastive alignment, cross-attention, co-attention) on 3,099–2,951 patients. Contrastive fusion improves concordance index by 1.5–5.4% vs unimodal baselines.

BenchmarksEmbeddingsVision
SIG
72
HYP
18
arXiv cs.AI·

Your Agent Has a Genome: Sequence-Level Behavioral Analysis and Runtime Governance of LLM-Powered Autonomous Agents

Base Sequence Analysis framework encodes LLM-powered autonomous agent behavior into symbolic sequences (X/E/P/V). Analysis of 347 production ReAct traces reveals P-X-P pattern reduces success by 10.4% and P-ratio negatively predicts success (r=-0.256). Governor runtime intervention system achieves +6.2% absolute success increase and 44% token reduction. Validated on 2,000 SWE-agent trajectories.

AI AgentsReasoningEvals
SIG
78
HYP
22
arXiv cs.AI·

Large Language Models as Optimizers: A Survey of Direct vs. Tool-Augmented Approaches and Their Performance Frontiers

Survey of LLMs as mathematical optimizers across three paradigms: direct optimization (iterative prompting), tool-augmented optimization (translating to formal specs), and tool-creating optimization (discovering reusable algorithms). Identifies critical reasoning gap and proposes trade-offs between future potential and auditability.

ReasoningAI AgentsTools
SIG
72
HYP
18
arXiv cs.LG·

Can Neural Networks Achieve Optimal Computational-statistical Tradeoff? An Analysis on Single-Index Model

Theoretical study demonstrating that neural networks trained with gradient-based methods can achieve optimal computational-statistical tradeoff for Gaussian single-index models. Proposed algorithm (two-layer network) achieves sample complexity Õ(d^{s*/2} ∨ d) matching SQ lower bounds, with extension to k-sparse case via weight perturbation technique.

PapersReasoningBenchmarks
SIG
78
HYP
15
arXiv cs.LG·

Unlocking Latent Dimensions: Exploring Representations of Large-Scale X-ray Scattering Data using Variational Autoencoders

Variational Autoencoder (C-VAE) trained on 1.5 million X-ray scattering images to learn low-dimensional representations. Model reveals organized clusters and generates controlled synthetic images. Deployed without retraining across two synchrotron facilities, outperforms DINOv3 in interpretability. Integrated into Latent Space Explorer (MLExchange).

VisionBenchmarksTools
SIG
72
HYP
18
arXiv cs.LG·

Phase-Localized Curation Does Not Help: A Negative Result on Per-Phase Metric Selection for Demonstration Filtering

Negative result on per-phase metric selection for demonstration filtering in robotic manipulation. Across three LIBERO pick-and-place tasks, phase-gated curation never outperforms global metrics (Task 1: 86.0 vs 92.0). Rank-aggregating defect signals across phases dilutes informative scores. Authors recommend identifying a single defect-informative metric over phase-based decomposition.

RoboticsReinforcement learningBenchmarks
SIG
72
HYP
15
arXiv cs.LG·

TriAdReview: Triangular Adversarial Review Architecture for Multi-Model Technical Document Generation

TriAdReview proposes a triangular adversarial architecture with two reviewer models (engineering and security perspectives) to improve technical document generation. Across 75 experiments, the triple model achieves +10.1% over baseline (26.2 vs 23.8/50, p<0.05), with strong gains on security audit (+27.6%), code generation (+20.8%), architecture design (+15.6%), but -7.5% degradation on requirements analysis.

Multi-agentCode generationBenchmarks
SIG
72
HYP
18
arXiv cs.CL·

Context Compression Is Not One Thing: Readable Symbolic Re-expression vs. Coherent Summary at Matched Budget

Telegraph English, a readable symbolic format, rewrites retrieved passages into structured entity-relation statements for context compression. On MuSiQue, TwoWiki, and HotpotQA, it outperforms three matched-budget baselines (deletion, truncation, sub-sampling) by 13–20 F1 points, and exceeds coherent prose summaries on the hardest dataset.

RAGReasoningBenchmarks
SIG
75
HYP
15
arXiv cs.LG·

GRASP: Gradient-Aligned Sequential Parameter Transfer for Memory-Efficient Multi-Source Learning

GRASP enables multi-source transfer learning with O(1) memory instead of O(K) by sequentially merging source models. Using parameter-wise gradient alignment and iterative fine-tuning, it achieves 93.5% mean accuracy on continual learning benchmarks (Yearbook, CLEAR-10/100) versus 71.7% for ensembles, while remaining production-deployable.

Fine-tuningReinforcement learningBenchmarks
SIG
78
HYP
25
arXiv cs.CL·

PACUTE: Phonology-, Affix-, and Character-level Understanding of Tokens for Filipino

PACUTE is a 4,600-task benchmark evaluating morphological understanding of Filipino in LLMs. The benchmark tests 6 compositional levels including infixation, reduplication, and diacritic distinctions. Open-weight models perform near chance on morpheme decomposition; frontier models recover affixes but remain far below ceilings on morphological composition tasks.

BenchmarksPapersReasoning
SIG
78
HYP
15
arXiv cs.LG·

M-CTX: Exact and Scalable Spatial Context Retrieval for Trajectory Analytics

M-CTX is a spatial context-retrieval framework for trajectory analytics. It replaces three brute-force stages (OSM range retrieval, SDF computation, moving-vessel neighbor lookup) with index-backed operators. On a 5.48M-anchor maritime corpus, it reduces context construction from 17 CPU-days to 1.8 hours (226x speedup), with exact reproduction of reference context.

BenchmarksInfrastructureOpen source
SIG
78
HYP
15
arXiv cs.LG·

StarOR: Synergizing Tree Search and Test-Time Reinforcement Learning for Optimization Modeling

StarOR synergizes Monte Carlo Tree Search with test-time reinforcement learning for optimization modeling. The framework decomposes modeling into four stages, refines a transient LoRA adapter via GRPO at each node, and employs an unsupervised multi-faceted reward system. Achieves state-of-the-art results across five optimization benchmarks with a 4B backbone.

ReasoningReinforcement learningFine-tuning
SIG
75
HYP
25
arXiv cs.LG·

Machine Learning and the Random Walk Puzzle: Forecasting the CAD/USD Exchange Rate with Expanding Window Evaluation and SHAP Interpretability

Study comparing 5 ML models (linear regression, random forest, gradient boosting, XGBoost, AdaBoost) to forecast monthly CAD/USD rate (2017-2026, 113 observations). Only linear regression statistically outperforms random walk (DM=3.06, p=0.0071). Random Forest achieves MAPE=1.17%. SHAP shows short lags (lag1-2) and rolling means dominate predictions.

BenchmarksEvalsPapers
SIG
72
HYP
15
arXiv cs.CL·

CHILLGuard: Towards Fine-Grained Chinese LLM Safety Guardrail with Scalable Data Construction and Model-aware Preference Alignment

CHILLGuard is a safety guardrail system for Chinese LLMs with fine-grained taxonomy (5 macro, 31 micro categories). Authors construct 405k training samples via RAG and prompt rewriting, plus 51k annotated test samples. Model achieves +15.92% F1 improvement over Qwen3Guard-8B-Strict using Direct Preference Optimization.

AI safetyAlignmentFine-tuning
SIG
78
HYP
25
arXiv cs.CL·

ESBMC-PLC: Formal Verification of IEC 61131-3 Ladder Diagram Programs Using SMT-Based Model Checking

ESBMC-PLC is the first open-source formal verifier with native support for IEC 61131-3 ladder diagrams (PLCopen XML format). The tool translates rungs to GOTO IR, models the PLC scan cycle, and verifies safety properties via SMT-based bounded model checking or k-induction. Evaluation on 13 benchmarks: 8 bugs detected, 7 unbounded k-induction proofs, all runs under 60ms.

AI safetyBenchmarksOpen source
SIG
78
HYP
15
arXiv cs.AI·

VGPT-RSI for RH-Adjacent Formal Progress: Boundary Certificates, Verified Finite Lagarias Inequalities, and Explicit Failure Localization

VGPT-RSI system applied to two RH-adjacent certification tasks: construction of formally verified RH-boundary certificates in Coq, and initiation of a formalized Lagarias route. Explicitly identifies unresolved mathematical obstructions (Lagarias equivalence, global tail theorem, reduction to extremal integers).

ReasoningPapersBenchmarks
SIG
72
HYP
15
Reddit r/LocalLLaMA·

HalBench: 29 OSS models tested on a custom built Sycophancy and Hallucination Benchmark, Qwen 3.6 and Gemma 4 scoring far above their weight! (While Meta keeps proving they forgot how to spend their money...)

HalBench v2.3 benchmarks 29 open-source models on sycophancy and hallucination across 3,076 audited questions with false premises. Qwen 3.6 (~27B) scores 36.6% pushback, outperforming all larger open models, GPT-5.4, and Gemini 3.1 Pro. Only Sonnet 4.6 and Grok exceed 50%. Phi-4 scores 2.3%.

BenchmarksOpen sourceEvals
SIG
72
HYP
35
Reddit r/MachineLearning·

Open weights are not enough: we need open training frameworks for research and better algorithms [P]

FeynRL, an open-source framework for RL post-training of LLMs and agents, aims to make training transparent and modifiable. The author argues open weights alone are insufficient: explicit training codebases separating algorithms from systems are needed. Framework supports SFT, DPO, multi-GPU and cluster setups.

Open sourceReinforcement learningCode generation
SIG
72
HYP
28