OpenAI public policy agenda
OpenAI publishes its public policy agenda for AI covering safety, youth protection, workforce transition, and global standards to ensure AI benefits society.
2731 articles
OpenAI publishes its public policy agenda for AI covering safety, youth protection, workforce transition, and global standards to ensure AI benefits society.
Meta adjusts its MCI (Model Capability Initiative) following internal employee criticism. The company strengthens protections regarding employee gesture tracking used to train its AI models.
Trump signed an executive order allowing AI companies to share their models. The proposed regulation remains voluntary and depends on tech giants' willingness to comply.
Spatial reasoning benchmark on LLMs using Sokoban under zero-shot conditions. ChatGPT, Qwen3.7-max, and Gemini 3.5-thinking pass; Gemini 3.5-flash, Qwen 3.6/3.7-plus, GLM-5, and Gemma4 fail. Strict formatting (UP/DOWN/LEFT/RIGHT only) prevents chain-of-thought cheating.
UK publishers can now opt out of having their content appear in Google's AI search results. This option adds to existing content control mechanisms.
Microsoft unveils Scout, an autonomous agent integrated into Microsoft 365 that can organize, coordinate, and execute tasks continuously to automate enterprise work.
Helvete-nano, a compact 2B model, has been released for unrestricted conversations and creative freedom.
A malware exploiting Minecraft's popularity has infected over 116,000 victims. Propagation is accelerating, requiring increased user vigilance.
Microsoft announces MAI-Thinking-1 and MAI family models at Build conference. New models with technical specifications detailed.
H Company (France) releases Holo3.1, VLM family fine-tuned on Qwen 3.5 for computer use agents. Models 0.8B to 35B-A3B, web/desktop/mobile support, native function-calling, multiple quantizations (BF16, FP8, Q4 GGUF). Apache 2.0 license.
DeepL releases a study on real-time voice translation in B2B context. The article presents key figures from this research on voice translation adoption as the next step in linguistic AI.
Mellum and Granite embedding models are now available on llama.cpp. Two pull requests add support for these models in the framework.
llama.cpp build b9455 with tensor-split achieves 70+ tokens/s on Qwen3.6-27B-UD-Q8_K_XL with 2x3090, matching vllm performance. MTP speculative decoding and flash-attention enabled. Context up to 262K tokens, prefill at 1400+ t/s.
Microsoft announces two on-device models at Build 2026: Aion 1.0 Instruct (efficient small model, open-weights, competes with Apple AFM-3B) and Aion 1.0 Plan (14B parameters, reasoning + tool-calling, 32K context, built into Windows). Aion 1.0 Plan enables local agentic workflows.
Microsoft launches Solara, an operating system dedicated to AI. The platform aims to optimize the execution of artificial intelligence workloads with a specialized architecture.
Study across 6,000 task-condition pairs shows multi-agent debate degrades generation (-1.6 to -15.5pp) via critique-induced confusion, yet improves error detection (+27.4pp F1). Adversarial separation with code-execution grounding and evidence-gated generation achieves +5.3pp on generative tasks.
GRZO is a zeroth-order optimizer for memory-efficient LLM fine-tuning. It draws one perturbation per mini-batch example and aggregates losses via group-relative normalization, increasing effective gradient directions from one to batch size at no additional forward cost. On Llama3-8B, GRZO achieves +3.0 accuracy over MeZO with 23% lower peak GPU memory.
IdiomX is a large-scale multilingual benchmark for idiom understanding, containing 190K+ contextualized examples across 12K+ idioms in English, Arabic, and French. The dataset includes idiomatic/literal usage labels and linguistic metadata. Four tasks evaluate idiom detection, retrieval, and interpretation.
Study of 'handoff debt': the rediscovery cost when a coding agent resumes an interrupted task. Across 75 tasks and 724 runs, structured notes reduce median agent events by 20–59% and tokens by 42–63% vs. repository-only takeover. Agent benchmarks should evaluate resumption efficiency, not just task resolution.
RRISE compresses randomized smoothing certification into a single forward pass via a learned surrogate, replacing up to 10⁴ Monte Carlo evaluations per query. Conformal calibration ensures conservative certified radii. On CIFAR-100 and Tiny ImageNet, 1.23–1.91× higher certified accuracy than prior offline-surrogate methods.
SNMPBB, first adaptation of nonmonotone Barzilai-Borwein methods to Symmetric NMF. Demonstrates 6× speedup over SymANLS on synthetic data. Extensions for graph clustering (Graph-SNMPBB) and large-scale problems (LAI-SNMPBB) with proven global convergence to first-order stationary points.
arXiv study on breast cancer recurrence prediction using multi-modal machine learning. Integrates treatment records, pathology reports, and clinician notes. Uses regex-based extraction and conflict reconciliation to recover tumor characteristics from free-text narratives. Shows multi-modal integration consistently improves predictive accuracy over single-modal methods.
RelGT-AC extends RelGT for autocomplete tasks on relational databases. The model combines column masking, a unified task head (classification/regression), and a TF-IDF encoder for free text. On 7 RelBench v2 tasks, it outperforms GraphSAGE on 3 regression tasks and gains +10 AUROC on text-heavy eligibility tasks.
Benchmark evaluating environmental attitudes across 31 LLMs (proprietary and open-weight). Models exhibit more progressive environmental positions than average German survey respondents, but show no systematic relationship with model origin, size, or release context. Detects prompting manipulation risks and sycophantic shifts.
arXiv paper on LLM representations: lexical overlap influences embeddings more than semantic content, persists across layers and architectures, and degrades performance in summarization and model editing tasks.
arXiv study showing LLMs struggle to infer user sociodemographics from single conversation histories. Disparities in advice (legal, medical, financial) are minimal but present. Conversation topics prove more predictive than sociodemographic data and affect LLM outputs unpredictably.
Researchers propose Bank of Values (BoV), replacing context-dependent value vectors with context-free vectors stored as sparse parameters in the last third of layers. On 135M and 780M models, BoV improves validation loss and performance across 21 benchmarks with reduced compute and memory.
Padyam2Gadyam: dataset of 600 classical Telugu poems (13th-17th centuries) with human-verified modern Telugu and English prose translations. Evaluation of 5 contemporary LLMs shows insufficient performance and limitations of existing MT evaluation metrics for this task.
Systematic audit of FOLIO and MALLS benchmarks reveals 39% and 36% errors in FOL formalizations respectively. Authors release corrected annotations and an LLM-based framework to guide manual relabeling, achieving 90% dataset accuracy by reviewing <24% of instances versus >70% for unguided review. Testing on Gemma 31B, Qwen3-30B, and GPT-4o-mini shows +9 to +22 percentage point accuracy gains.
ALAR (Adaptive Latent Agentic Reasoning) is a dual-mode framework alternating between compact latent reasoning and explicit chain-of-thought based on decision difficulty. Trained using agent actions as supervision, it reduces generated tokens by 43.6% in search and 84.6% in tool-use while maintaining accuracy.
Framework combining conformal prediction and collaborative filtering-style annotator representation to analyze LLM behavior against human annotators in content moderation. Introduces Ghost Prediction metric to quantify model-human divergences. Evaluation across 4 LLMs and 4 datasets shows larger models more confident on texts with no human alignment, revealing structural demographic bias.
Study on linguistic productivity in LLMs: models reproduce entrenchment (frequent structures) but fail to implement preemption (observed absence of structures). Large models handle contextual coercion but overgeneralize patterns never seen in training data.
EURO-5K is a 5K-sentence corpus for extracting reporting obligations from EU legislation (136 legislative acts). Comparison of fine-tuned BERT and LLMs (QLoRA): generic and legal BERT achieve similar 0.89 F1; legal pretraining helps mainly for parameter-efficient tuning. Convergence at 3K samples.
Method to predict best-of-N inference scaling gains without running the full procedure. Ridge predictor identifies 3 stable features (prompt-level agreement spread, label-assisted first-correct-sample position, completion-length variance) plus entropy, reaching Spearman ρ=0.90 correlation with actual gains across model families and math/reasoning tasks.
Locally deployed RAG-based academic advising system to help students select courses while respecting prerequisites. Combines LLM with retrieval from structured syllabus data, preserving data privacy.
TypewriterLM is a 7.24B parameter language model trained exclusively on English text predating 1913. Authors construct TypewriterCorpus (54B tokens), a cleaned historical corpus with leakage mitigation, and introduce lexically grounded instruction tuning to ground responses in historical documents. Three datasets and History-Event benchmark are released.
HDPO (Hint-Guided Diversified Policy Optimization) enhances LLM reasoning through reinforcement learning with verifiable rewards. The method prompts models to first generate multiple candidate solution outlines (hints), then select the most reliable one. Two stages: Cold Start for structured reasoning, then hint-guided RL to diversify and improve solution reliability.
New inference-time method DCO (Dynamic Contextual Orthogonalization) to reduce hallucinations in LLMs. Based on hypothesis that hallucinations manifest as orthogonal noise relative to semantic manifold of residual stream. Evaluated on Llama-3 (8B/70B) with improvements on XSum, NQ-Swap, IFEval and TriviaQA.
SEA-Embedding is an open and reproducible text-embedding pipeline for Southeast Asian languages trained exclusively on public data. The study examines three core factors: data composition, training objective, and base encoder initialization. Achieves state-of-the-art results on SEA-BED.
SkillDAG models inter-skill relationships as a typed directed graph for dynamic LLM agent skill selection at inference time. On ALFWorld and SkillsBench with MiniMax-M2.7, it achieves 67.1% success and 27.3% reward, exceeding Graph-of-Skills baselines by +12.8 and +8.6 points. The graph self-evolves during execution via a propose-then-commit protocol, accumulating structure across episodes.
CORE is a framework training MLLMs to detect multimodal fake news by identifying semantic or physical conflicts across modalities. Using an annotated corpus (CAC) with conflict factors, CORE generalizes to unseen manipulations in few-shot or zero-shot settings, outperforming existing methods.
DeltaMem organizes LLM agent experience memory into two residual trees: one stores goal-conditioned tasks as reusable skills, another stores scene-level environment knowledge. Each tree uses root nodes for generalized base experiences and delta nodes for variations, eliminating redundancy. An autonomous consolidation mechanism distills high-frequency paths into new root nodes.
arXiv paper proposing CLEAR, an optimal budget allocation method for LLM inference grounded in economic theory. Using a shifted-surge utility function and global shadow pricing, CLEAR performs rational abandonment and reallocates resources from insolvent to solvable queries. Results: 3x improvement in global accuracy vs uniform allocation under resource scarcity.
Study of representational geometry to understand how prompts reshape behavior in LLMs and VLMs. Nested decomposition framework testing translation, rigid transformation, scaling, affine and nonlinear maps on 3 LLMs, 3 VLMs and 6 datasets. Finding: cross-dimensional linear mixing (affine transformation) is the key mechanism for representational reorganization toward task structure.
Novel framework combining importance-aware news compression and process-level retrieval supervision for time series forecasting. A reward model estimates each article's forecasting utility for sequential fusion, while a PRM ranks supplementary-news candidates based on error profile. Experiments on finance, energy, traffic, and bitcoin benchmarks show improved accuracy and fewer refinement iterations.
DeskCraft is a desktop GUI benchmark for agents on long-horizon professional workflows (>50 steps) in design, video, audio, and 3D with human-agent collaboration. 18 agents tested on 538 tasks: GPT-5.4 reaches 31.6% on standard tasks and 27.6% on interactive tasks. Reveals persistent failures in proactive clarification and long-horizon workflow delivery.
EvoTrainer co-evolves LLM policies and training harnesses via empirical feedback for autonomous agentic RL. Tested on mathematical reasoning, competitive programming code generation, and software engineering, the system matches or exceeds human-engineered RL baselines, with largest gains on long-horizon agentic SWE tasks.
Multi-agent LLM systems lose up to 72% of issue-critical facts during deliberation, creating a 'deliberative illusion'. DelibTrace measures factual attrition and stance homogenization. Agents converge toward consensus while forgetting essential elements needed to interpret the problem.
Geometric study showing inter-LLM agreement on subjective evaluations does not reflect human alignment. Across 41 LLM judges and 8 Indic languages, models use 30-50% of human score range, with evaluation axis nearly orthogonal to humans (87-89° vs 78-81°). LLM-LLM agreement (r≈0.35) exceeds LLM-human (r≈0.27-0.32). Only post-hoc calibration improves all rubrics.
G²C-MT proposes graph-guided context selection for document-level machine translation. The system models discourse dependencies between paragraphs via a lightweight graph and uses depth-biased random walks to extract context paths. Tested on DeepSeek-V3, Gemini-2.5-Flash-lite, and Qwen-2.5/3, the approach outperforms baselines across multiple domains.
Regret Pre-training introduces a self-supervised framework based on LUPI using dual-view architecture generating Student (causal) and Teacher (future-conditioned) distributions. On OLMoE-1B-7B after 4B tokens, GlobalRegret and LocalRegret achieve 33.9% and 32.2% average accuracy vs 30.2% baseline, with 18.1pp gain on BoolQ. No additional parameters.
New arXiv study on editing factual opinions in LLMs. FOE benchmark with 261 public figures, 19 issue categories, 2,178 opinion records. Current editing methods fail to maintain consistency between edited opinion and model-generated evidence. Proposes Self-Generated Evidence-Aligned method for opinion-evidence alignment.
PhotoCraft introduces a hierarchical memory system for photo-search agents combining working, episodic, and semantic memory to maintain logical consistency across multi-step reasoning. Training-free approach tested on DISBench achieves up to 18.5% improvement in context-aware retrieval across MLLM backbones.
RL-based method for adaptive sampling control at test-time on LLMs. A lightweight controller trained with RL dynamically decides to stop or continue sampling, balancing answer correctness, latency, and computation cost. MDP formulation with Lagrangian interpretation. Outperforms ASC and ESC on trade-offs.
Internal Coherence Maximization (ICM) automatically generates examples to align AI models with diverse human values without extensive human supervision. Tested across four benchmarks, ICM matches gold label performance. Example coherence improves generalization even at constant accuracy, especially for underrepresented personas.
LEDE, an offline reinforcement learning framework, optimizes LLM inference by dynamically selecting exit layer and speculation length based on local sequence context. On Llama-2 and Llama-3, it achieves 2.0×–2.7× speedup over autoregressive decoding, +17% over static speculative baselines.
DMT-CBT models longitudinal therapeutic state evolution in CBT across multiple sessions. The framework maintains structured states, incorporates multimodal data and tool-augmented interventions. DMTCorpus, a synthetic multimodal dataset, demonstrates improved counseling fidelity and therapeutic alliance over post-hoc extraction approaches.
SenseJudge is a customizable judgment framework driven by human preferences for evaluating LLM responses. Paired with SenseBench, a benchmark derived from real-world multi-turn interactions, it outperforms existing methods in adapting to user preferences and achieves model ranking aligned with human judgment.
Factorial study of 4 open-source LLMs rating clinical decisions in type 2 diabetes pharmacotherapy. LLMs as AI raters score 74–78 points under rubric-free protocol vs 7.69–49.64 points under anchored Gold Rubric. Rubric amplifies discrimination between CDSS models (1.76–5.10×) and reveals behavioral variation suppressed without rubric.
MedCUA-Bench is an interactive benchmark for evaluating computer-use agents in clinical interfaces. It covers 18 medical scenarios across 10 domains with authentic interfaces. Best closed-source models reach 54.2% strict success, open-source agents average 2.5%, exposing a major gap with required reliability.
Two-stage framework (PRPF) for proactive mobile agents decouples perception (deciding when to intervene) from reasoning (how to assist). Lightweight perceptor gates false triggers and compresses context, activating the MLLM reasoner only when needed. Reduces false trigger rates and improves efficiency on ProactiveMobile benchmark.
Empirical study detecting natural experiments (implicit interventions) in real-world datasets using causal discovery and feature selection. Authors validate on synthetic data then evaluate 50+ real datasets, showing that treating data as interventional rather than observational improves model performance.
Method to distill Answer-Set Programming (ASP) rules from LLMs for Visual Question Answering (VQA). The approach uses VQA dataset examples to guide the LLM in extending an initial reasoning theory, with validation by the ASP solver. Demonstrates effectiveness across multiple datasets with few examples.
LEAP is an agentic framework enabling LLMs to generate mechanically verifiable formal proofs in Lean. The system decomposes complex problems into smaller units through iterative interaction with the Lean compiler. On 2025 Putnam Competition (12 problems), LEAP solves all 12; on Lean-IMO-Bench, it achieves 70% one-shot solve rate versus <10% for general-purpose LLMs.
ToolGate is a lightweight controller that decides whether to execute or skip tool calls proposed by vision-language agents. Across five benchmarks with Qwen3-VL, it reduces token cost to 64-69% of the ReAct baseline while preserving accuracy, and further improves accuracy by 1.65 points with matched-domain training.
TriEval is an LLM evaluation pipeline assessing bias, toxicity, and truthfulness simultaneously with minimal resources. Compatible with open-source and closed-source models, runs on standard laptop without GPU. Tested on Llama 3 8B, Mistral 7B, Gemma 2 9B, and Claude Haiku, revealing toxicity and truthfulness differences between models.
Researchers propose an agent economy where AI agents self-coordinate through auctions and payment exchanges without centralized control, inspired by Hayek's economic theory. This approach generates emergent multi-step reasoning strategies and outperforms baselines on five tasks including mathematical reasoning, financial research, and distributed-system optimization.
WRIT is a trajectory synthesis pipeline for multi-turn agent training. It generates complex tasks along two axes: number of write decisions and evidence burden per decision. With 2K synthesized trajectories, a 4B model outperforms GPT-5.1 no-think on τ²-bench while reducing inference-time token usage.
Fast-dLLM++ improves diffusion LLM inference by replacing homogeneous confidence token selection with Fréchet profile decoding. Training-free, it exploits heterogeneous confidence profiles to parallelize more tokens safely, achieving up to 37% higher throughput on GSM8K, MATH, HumanEval, and MBPP with LLaDA-8B while maintaining accuracy.
Philosophical paper examining whether AI chatbot outputs (e.g., Anthropic's Claude) produce meaningful language. Authors argue that standard human language theory already applies to LLMs without requiring anthropomorphic assumptions about intentions or mental states.
Unified framework for memory access and selection in long-context dialogue systems. Uses Bayes factor to quantify utility of historical turns by measuring likelihood improvement of reference responses. Outperforms embedding-based retrieval on four benchmarks, especially for preference-intensive tasks with changing user preferences.
Conditional hypothesis generation framework for LLM-based text analysis incorporating researcher-specified covariates. Addresses stratum imbalance and sign reversal through feature-covariate interactions and inverse-frequency reweighting. Validated on synthetic and real-world datasets in computational social science.
Framework for LLM agents operating under underspecified instructions. Proposes Information Gain Reward metric quantifying clarification utility via Bayesian belief updates. Validation on τ-Bench: +3.7% success rate vs no-clarification baseline, +0.3 interaction steps.
Dataset of 410,499 tropical species covering plants, aquatic, and pet animals. Integrates taxonomic identifiers (GBIF, NCBI, iNaturalist, etc.), adds cross-domain ontology, Chinese vernacular names (99.50% coverage) with explicit provenance, and CITES linkage. Deposited on Zenodo.
Two automated metrics assess LLM lexical misalignment: Lexical Alignment Score detects term overuse ('suggest', 'additionally', 'strategy'), Triangulated Preference Shift quantifies RLHF impact. Tested on 6 model families (Falcon, Gemma, Llama, Mistral, OLMo, Yi) via PubMed abstracts, no manual annotation required.
HyperPatch introduces a hypergraph-based sequential knowledge editing framework for LLMs. It addresses N-ary Structural Drift, where complex relation updates fracture relational integrity. On MQuAKE-CF and MQuAKE-T benchmarks, it achieves 96.24% and 21.06% relative gains in Hop-wise Accuracy over strongest baselines.
MemTrain introduces a self-supervised training framework to enhance context-memory capabilities of LLM agents. Two coupled proxy tasks on Wikipedia (masked entity reconstruction and intermediate memory recall) are jointly optimized using GRPO. Achieves gains up to 17.67 points on long-text QA and search-based QA benchmarks.
Study showing visual graphs help LLMs organize multi-hop reasoning better than text representations. Graph-structured mind maps guide models without direct answer hints, unlike flattened text graphs. Advantage persists after supervised fine-tuning and KL-based distillation.
AURA-Mem introduces a constant-size recurrent memory (4,224 bytes) for robot policies, with a learned gate that writes only when observations would change the next action. On LIBERO-Long with OpenVLA-OFT 7B, it matches baseline policy (0.233 success) while reducing memory writes by 7× and VRAM by 6,061× versus KV-cache.
Study of demographic bias in skin lesion classification using ResNet-based models. Three strategies evaluated: single-task, reinforcing multi-task, adversarial learning. Sex bias stems from data imbalance; age bias consistently favors younger groups. Mitigation strategies show partial effectiveness depending on data distribution.
Study on activation transfer between language models (Pythia-160M to Pythia-410M). A linear translation layer strongly aligns hidden states (cosine similarity 0.97), but injecting translated activations does not improve downstream performance at inference time. Negative result: offline representational alignment is insufficient for useful causal communication.
Comparative study of brain region importance (EEG) for cognitive workload prediction. Analysis across 4 public datasets: frontal electrodes outperform full-scalp baseline by 15-20% in relative rank position using fewer sensors. Fronto-central regions show most stable predictive utility.
Deterministic orchestration framework for validating fragmented ESG data (Scope 1-3) with temporal anomaly detection, imbalance-aware ensemble learning, and audit provenance tracing. Synthetic benchmark calibrated against GHG Protocol, PCAF, ISSB standards. Evaluation on classification, calibration, and provenance chain completeness metrics.
GATD (Geometry-Aware Tabular Diffusion) improves tabular synthesis by incorporating pairwise angles and lengths between columns as inputs and auxiliary targets. The MLP instantiation achieves SOTA with 3.5x fewer parameters on average, winning 8/10 on Shape, 7/10 on Trend, and 9/10 on downstream utility, reducing Shape and Trend errors by 27% and 20%.
Neural network pruning method using Marchenko–Pastur random-matrix theory with minimal post-pruning fine-tuning. On ImageNet-1k, ViT-B/16 achieves 83.41% top-1 with 59.81% MAC reduction after 3 distillation epochs; ResNet50 8:16 reaches 75.87% with 1.62× A40 speedup.
Theoretical paper on generalization bounds under distribution shift. Proposes framework quantifying extra risk from regime composition mismatch in Markov-switching environments, with exact decomposition and minimax lower bounds. Tested on 25 years of global equity indices.
Adaptive multifidelity algorithm for machine learning in quantum chemistry. Dynamically determines dataset composition by querying samples at each fidelity level. Reduces data generation costs by up to 30× versus single-fidelity methods and 5× versus standard MFML on coupled cluster energies and excitation energies.
An arXiv study analyzes 8 benchmarks for multivariate time series anomaly detection. A diagnostic framework shows 79-100% of anomalies are univariately detectable on 6 datasets. Cross-channel models provide no measurable gain. Current benchmarks fail to validate multi-channel modeling capabilities.
FiRe-OPD introduces fine-grained on-policy distillation combining trajectory filtering and soft token reweighting. Validated on AIME 2024 (+6.25 strong-to-weak) and Miner (+18.81 multi-teacher), the method outperforms recent token-level OPD approaches in stability and performance.
ML method for real-time binary road surface classification (dry/damp grip vs snow/ice slip) using production vehicle signals during cruising. Feature-based and end-to-end frameworks using wheel speeds, torques, longitudinal acceleration, steering angle, yaw rate. Validated on public-road data.
QUIVER enriches classical ML models with quantum views based on quantum Fisher information matrix extracted from variational quantum circuits. Tested on QM9 (molecular properties) and JetClass (LHC), the paradigm improves performance without requiring fault-tolerant quantum hardware.
Qift introduces a zero-free level set for W2A4/KV4 quantization ({±0.5, ±1.5}) leveraging Hadamard rotation. Training-free and codebook-free, it improves perplexity on LLaMA-2-7B and LLaMA-3.1-8B versus standard levels {-2,-1,0,+1}, narrowing the gap to W3A4 in mixed precision while keeping half transformer layers at two-bit.
Comparative study of encoder-only Transformer vs LSTM for streamflow prediction in ungauged basins. On NOAA NWM data, LSTM outperforms Transformer. Incorporating downstream information improves all models (+60% median NNSE). Recurrent memory proves better suited for this hydrologic inference task.
SpecFlow introduces a lightweight multimodal spatial reasoning framework representing intermediate visual thoughts in fixed-size discrete cosine space. Classifier-free guidance enables autoregressive textual thoughts to steer visual state updates without context expansion. Achieves competitive reasoning performance while reducing computation and KV cache costs by up to 2.1×.
BehaviorBench is a benchmark for evaluating personalized decision modeling from real-world behavioral traces. Built on 2,000 wallets with 141,445 belief-prediction instances and 1,485,972 trade-prediction instances, it tests whether generative models can adapt predictions to individual users without relying on simulated behavior.
ChatHealthAI aligns structured EHR representations from a pretrained EHR foundation model with a frozen LLM's semantic space via a task-aware resampler. The multimodal framework integrates longitudinal patient representations with refined clinical event descriptions, improving interpretable clinical reasoning while maintaining competitive predictive performance on the EHRSHOT benchmark.
Systematic review of hybrid architectures for wind power interval forecasting. Approaches combining deep learning, modal decomposition (VMD, EEMD), and statistical methods improve accuracy. Dominant strategy: two independent models (LSTM, ELM) for lower/upper bounds. Challenges: lack of standardized metrics, computational complexity, limited real-world validation.
Method to extract reasoning primitives from ReAct agent traces. Recurrent reasoning moves are clustered and converted into typed pseudo-tools. Induced libraries outperform the source agent: +44pp on RuleArena NBA, +30pp on MuSR, +22pp on NatPlan.
Study of three approaches to automatically generate enemy morphologies in video games based on collision data. The tested methods match or outperform an evolutionary baseline adapted from robotics morphology work.
Paper proposes modular reference architecture for Embedded Agent Systems on microcontrollers. Decouples On-Device Agents (compressed neural networks, rule-based logic) from Cloud-Augmented Agents (SLMs for reasoning). Includes cross-cutting Governance Layer for observability and safety across distributed device fleets.