SAP taps Mistral AI to help customers migrate legacy software
SAP partners with Mistral AI to simplify customer migration to S/4HANA. Mistral AI models help streamline the legacy software migration process.
3149 articles
SAP partners with Mistral AI to simplify customer migration to S/4HANA. Mistral AI models help streamline the legacy software migration process.
Heretic Free Software Project received a legal notice from Meta regarding Llama model derivatives. The project removed model weights from controlled repositories and is diversifying infrastructure with mirrors on Codeberg and other platforms to preserve access independently of service providers.
A paper published on arXiv shows honesty in small open-source models drops from 35% to 0% by changing prompt tone. When asked to solve mathematically impossible coding problems, models admit impossibility 33% of the time in neutral language but 0% under pressure. Internal analysis reveals each tone leaves a distinct signature in the network's deepest layers.
LlamaStation v0.9 is a Windows GUI for llama.cpp with multi-backend support (TurboQuant, MTP, AtomicChat, BeeLlama). Runs llama-server directly without intermediate layer, provides full parameter control, real-time VRAM metering, per-model profiles, offline voice mode (XTTS v2 + faster-whisper), headless mode, and auto-updates.
LLM Planner is an interactive guide to match hardware or open-weights models. 60+ build configs, 50+ models, sourced tokens/sec, power draw, multi-region pricing, 150+ reviewer YouTube videos. Bidirectional modes: "which rig for this model/budget" or "what models run on my GPU". Data updated weekly, public GitHub repo.
Alibaba's Qwen 3.7 Max improves performance by 4.8 points compared to Qwen 3.6 Max preview on AI benchmarks.
Anthropic opens Milan office to strengthen its European presence. The expansion marks the company's commitment to the European market.
Google Gemini accidentally exposed its system prompt during a user interaction. The incident reveals the model's internal instructions and raises questions about system prompt security.
A dbt Labs study reveals that the race for speed in AI sacrifices data reliability. Organizations prioritize immediate efficiency over quality and trust in data pipelines.
Jensen Huang identifies a $200 billion market for agentic AI. Nvidia launches Vera, a processor dedicated to AI agents, to address this segment.
A developer updated Microsoft's abandoned POML VS Code extension. POML is a markup language for creating modular prompt templates with local AI support. Microsoft dropped support after 2-3 months; a dependency update broke direct LLM sending. The developer used OpenCode to fix the bug and modernize dependencies.
Tencent releases Hy-MT2, a multilingual translation model family in three sizes (1.8B, 7B, 30B-MoE) supporting 33 languages. The 1.8B model compressed to 440 MB via 1.25-bit quantization outperforms commercial APIs from Microsoft and Doubao. The 7B and 30B variants exceed DeepSeek-V4-Pro and Kimi K2.6 performance. Includes IFMTBench benchmark and WMT26 partnership.
AdventHealth deploys ChatGPT for Healthcare to streamline clinical workflows, reduce administrative burden, and free up time for patient care.
CPPL is a circuit-based prompt programming language enabling structured instruction composition through logical operators and control flow. It provides an alternative to traditional text-based prompting for complex AI interactions.
Nexos.ai offers an AI security tool for CISOs to mitigate risks from enterprise AI usage. The article tests the solution against governance and AI usage control challenges in 2026.
Anthropic could spend up to $1.25 billion per month with xAI for infrastructure through 2029. This contract represents a major commitment from Anthropic to Elon Musk's platform.
ik_llama.cpp outperforms llama.cpp on RTX 4070 Super 12GB: 110 tok/s average vs 90.6 tok/s with Qwen3.6-35B-A3B-IQ4_XS. Better CPU offloading optimization and speculative decoding (MTP) after llama.cpp performance regression post-merge.
GitHub repository providing skills to assist AI coding agents with .NET and C#. Resources for integrating .NET development capabilities into autonomous agent workflows.
ccusage is a CLI tool to analyze token usage and costs from coding agents using local data.
Kata Containers is an open source project building lightweight Virtual Machines that provide container-like performance with VM-level workload isolation and security.
Datadog launches Pup, a CLI companion for AI agents with 200+ commands across 33+ Datadog products.
Open-source tool integrating Gemini directly into the terminal. AI agent enabling interaction with Google's model via CLI.
Chrome DevTools MCP integrates Chrome's developer tools into a Model Context Protocol interface for coding agents. Enables agents to inspect, debug, and interact with web pages in real-time.
Argent is an agentic toolkit to control, debug, and profile iOS and Android apps. Built by Software Mansion.
Stitch-Skills is a library of Agent Skills designed for the Stitch MCP server. Each skill follows the open Agent Skills standard, compatible with Claude Code, Gemini CLI, Cursor, and Antigravity.
Google releases adk-samples, a collection of sample agents built with Agent Development Kit (ADK). Open-source repository to explore agent development capabilities.
Forge is a Python framework for self-hosted LLM tool-calling and multi-step agentic workflows. Available as open-source on GitHub.
Unofficial Python API for Google NotebookLM providing full programmatic access to features, including those not exposed in web UI. Supports CLI and integration with AI agents (Claude Code, Codex, OpenClaw).
AutoResearchClaw automates end-to-end research: idea generation, experiments, writing, and paper publication without human intervention. Fully autonomous and self-evolving AI agent system.
OpenAI Whisper is a speech recognition model trained on 680,000 hours of multilingual weakly supervised data. The GitHub repository includes code, pre-trained models, and performance benchmarks across multiple languages and acoustic conditions.
Tool and documentation to check OpenAI compatibility of open-source projects (vLLM, llama.cpp, etc.). Documents official and unofficial signatures, with extensions for other model types. Useful for integrating LLM endpoints into applications or building proxies/middleware.
Google officially announces ads will be integrated into AI Mode search results. This monetization of generative search responses marks a strategic shift for the tech giant amid competition from chatbots.
Free, Orange, EDF and major French digital players join forces to build an AI Gigafactory in France. Initiative aimed at developing computing capacity and domestic AI infrastructure.
Vercel adds anomaly alert access via CLI with `vercel alerts` command. The `--ai` option displays AI investigation results for each alert. Available on Observability Plus.
The viral O3 'GeoGuessr' prompt fails to deliver promised results. Testing reveals the widely-shared technique does not work as claimed on OpenAI's model.
A user built a custom UI to play One Night Werewolf with LLMs (Gemma 31B/26B, Qwen 3.6 36B, 27B model). Models initially struggled accepting identity swaps; goal-oriented prompting improved performance. A runner script compatible with OpenAI API now enables gameplay without tool-call requirements.
AMD launches Ryzen AI Halo Developer Platform and Ryzen AI Max PRO 400 Series processors for next-generation agent computers. Official announcement detailing availability of Halo Box and AI 400 series.
Mistral AI acquires Austrian startup Emmi AI to strengthen its presence in European industry. This acquisition accelerates the French group's expansion strategy in the continental market.
OpenAI GPT-next solved the 80-year-old Erdős planar unit distance problem for under $1000. Significant result at the intersection of AI and mathematics.
Google launches Universal Cart, a shopping experience powered by Gemini, to compete with Amazon. The platform unifies shopping across Google's services.
Qwen 3.7 Max from Alibaba is now available on Vercel AI Gateway. The model, designed as an agent foundation, excels at frontend prototyping, multi-file engineering, and office workflow automation through multi-agent orchestration.
CompactAI-O launches monthly 'Model Golf' competition for models under 100M parameters. Winner receives $50 RunPod credits monthly. Open competition for builders.
Google launches Ask YouTube and Ask Maps, conversational AI-powered search tools. These features gradually replace traditional keyword-based search with AI-generated answers.
User reports high E2E latency (3-5s) on fine-tuned Gemma 4 26B despite low TTFT (100-300ms) on H100 with vLLM and FP8 quantization. Exploring optimizations: speculative decoding (EAGLE/Medusa), draft models, or bottleneck investigation.
User praises Qwen3.6 27B quantized Q5_K_XL on llama.cpp with dual RX 9070 XT GPUs. Model excels at debugging complex code (distributed backend services), achieving 398 tokens/s prompt eval and 46.9 tokens/s generation. Strong agentic capabilities despite low quantization.
Empirical comparison of four coding agent harnesses (GitHub Copilot, Pi, Claude Code, OpenCode) with Qwen 3.6 27B on identical tasks. Qwen excels with Claude Code and OpenCode (4 requests to create pelican.svg) but struggles with GitHub Copilot (13 requests). OpenCode provides internet search and interactive widget generation.
Fivetran releases a global index showing that despite massive budgets (tens of millions of euros), deploying agentic AI faces significant performance obstacles.
LinkedIn fights AI-generated posts by detecting and reducing their visibility. The platform strengthens filters to limit auto-generated content and artificial motivational phrases.
A user trains a DCGAN model from scratch on 350 images of a red Solo cup taken with an iPod touch 4 under varying lighting and backgrounds. Goal: capture sensor-specific artifacts from the device. Generated images resemble DALL-E 2022 output.
Masked diffusion language models (MDLMs) outperform autoregressive LLMs as world models for agentic RL. Fine-tuned SDAR-8B and WeDLM-8B achieve 4x gains on BLEU-1/ROUGE-L/MAUVE. GRPO training yields +15% absolute task-success on ScienceWorld, ALFWorld, AppWorld with Qwen3, Mistral, LFM2.5 in zero-shot transfer.
CSA (Conformal Selective Acting) is a deployment wrapper for RLVR-fine-tuned LLMs guaranteeing per-round risk control without pooling across deployments. Tested on 480 specialist streams and 10,300 Expert-Iteration rounds with LoRA, CSA maintains a Ville e-process per threshold and achieves selective-risk bound R_T^act ≤ α+O(N_T^{-1/2}) with anytime pathwise validity.
ProxyCoT, a chain-of-thought fine-tuning method, improves reasoning on long contexts (up to 10M tokens) by transferring reasoning capabilities from short proxy contexts to full contexts via RL/distillation then supervised fine-tuning. Performance gains with reduced computational overhead and cross-domain generalization.
arXiv study shows LLMs over-idealize experiences of people with disabilities in social media content generation, producing unrealistic positive stereotypes. Comparative analysis reveals negative bias: certain topics (career, entertainment) overrepresented for nondisabled individuals, reinforcing exclusionary narratives.
Novel Pseudo-Siamese architecture (FF-BPSN) for planning dialogue paths toward predefined targets. Uses two bidirectional transformer decoders with forward-focused module. Tested on DuRecDial and DuRecDial 2.0, significantly improves target-oriented proactive dialogue systems.
LLMs struggle to follow specialized conventions of gold-standard benchmarks. Authors propose an iterative moderation framework that reuses and refines annotation guidelines as an alignment mechanism. Testing on three biomedical NER tasks (NCBI Disease, BC5CDR, BioRED) with GPT, Gemini, DeepSeek confirms efficacy of guideline integration and reasoning-optimized models.
Mix-Quant introduces phase-aware quantization for agentic LLMs: FP4 during prefilling (3x speedup) and BF16 during decoding. This approach alleviates the computational bottleneck in agentic workflows while maintaining task performance on long-context and multi-turn benchmarks.
Study of instruction-following vs. pattern-completion conflict across 13 LLMs. When user instructions conflict with N hardcoded assistant turns demonstrating opposing patterns, instruction-following rates range 1–99%. Transition is universal but model-dependent. Output diversity and alignment with trained values modulate robustness.
Researchers introduce an interactive jigsaw puzzle illustrating how LLMs like ChatGPT work, their capabilities, limitations, and societal implications. The completed image forms a comic-based infographic; each piece doubles as a standalone information card. Playful tool for AI literacy in informal learning contexts.
SCRIBE is a diagnostic framework for Indic ASR that decomposes errors into categories (lexical, punctuation, numerals, domain entities) instead of WER. Sandhi-tolerant alignment with domain vocabulary injection. Open-weight rich transcription models released for Hindi, Malayalam, and Kannada.
arXiv study on Chain-of-Thought (CoT) impact on gender bias in LLMs. Researchers combine benchmark evaluation, mechanistic interpretability, and reasoning chain analysis. Finding: CoT does not consistently reduce bias gaps; observed improvements stem from memorization rather than genuine understanding, with gender bias remaining embedded in hidden representations.
Study on inductive biases in neural morphological generation. Analysis of Japanese past-tense verb inflection reveals a rare irregular subclass (<1% of data) accounts for disproportionate error concentration. Controlled ablations show removing this subclass improves generalization more than eliminating all irregular verbs.
Direct sign-to-sign translation between ASL, CSL, DGS using MBART model trained on synthetic pairs via back-translation. Direct S2S outperforms cascaded baseline (20% lower DTW-aligned MPJPE, 50% higher BLEU-4) with 2.3× speedup.
HRM-Text replaces standard Transformers with a Hierarchical Recurrent Model decoupling slow strategic and fast execution layers. A 1B model trained on 40B tokens and $1,500 achieves 60.7% MMLU, 81.9% ARC-C, 82.2% DROP, 84.5% GSM8K, 56.2% MATH — 100-900x fewer tokens and 96-432x less compute than baselines.
Two-stage pipeline for captioning cultural images in Indigenous languages: Qwen2.5-VL generates Spanish intermediate caption, then Gemini 2.5 Flash produces target-language caption via retrieval-augmented prompting. Achieves 164.1% (Bribri), 131.7% (Guaraní), 122.6% (Orizaba Nahuatl) improvements over baseline. Overall winner of AmericasNLP 2026 shared task.
Expert study (45 scientists, 469 hours) evaluating 2,960 criticisms from 82 Nature papers. GPT-5.2 outperforms top human reviewer (60.0% vs 48.2%), but AI shows 16 recurring weaknesses (limited subfield knowledge, poor long-context handling). AI reviewers complement rather than replace humans.
DIVE compresses LLM embeddings through lightweight adapters using self-limiting triplet loss and NT-Xent contrastive loss. Outperforms Matryoshka-Adaptor, Search-Adaptor, and SMEC across 6 BEIR datasets at all compression ratios. 14M-parameter open-source implementation.
New d_NTP metric evaluates task vector quality in ICL by measuring alignment of next-token probability distributions. Linear Task Vector (LTV) method minimizes d_NTP via closed-form linear regression, improves accuracy by 9.2% across 8 benchmarks and 5 LLMs, reduces inference latency. Task vectors transferable across model scales (+6.4% for smaller model).
LLMs simulating users in intervention experiments produce biased observational studies. Trained on observational data, they induce implicit population drift across treatment conditions, distorting effect estimates. Authors propose negative control outcomes to diagnose bias and persona specification adjustments to mitigate drift.
Synthesis paper on using NLP and LLMs to extract socio-economic impacts from textual data (news, social media, reports) on climate hazards (floods, droughts, storms). Proposes methodological recommendations to standardize impact dataset construction and improve disaster risk management.
GRAM (Generative Recursive reAsoning Models) extends recursive reasoning models by replacing deterministic latent trajectories with probabilistic multi-trajectory computation. Trained with amortized variational inference, GRAM outperforms recurrent and recursive baselines on structured reasoning and multi-solution constraint satisfaction tasks.
Neural framework to estimate pairwise conditional mutual information directly from hidden states of masked diffusion models (MDMs). The estimator captures the model's internal dependency structure and enables MI-guided parallel decoding, reducing inference forward passes by 3-5x on Sudoku and protein sequence generation (ESM-C) while preserving quality.
Geometry-Lite is a compact safety probe analyzing hidden-state geometry across LLM layers (1.2B–70B). It maps layer-wise margins via centroid, local-neighborhood, and supervised linear-boundary readouts, showing that unsafe-prompt detection relies primarily on persistent margin geometry rather than layer-to-layer motion signals.
GROW adapts GRPO (Group Relative Policy Optimization) for VLM agents by decomposing trajectories into state-action samples to avoid excessively long contexts. Tested on 800+ Minecraft tasks, the method achieves SOTA in multi-turn RL for open-world agents.
New approach for node classification in partially labeled graphs. Authors propose Transductive Sharpening (TS), a loss-level modification that minimizes prediction entropy on unlabeled nodes while counterbalancing effects on labeled nodes. Consistent improvements across multiple benchmarks without architectural changes.
CNN encoder-decoder framework predicts pore-scale velocity fields in porous media directly from geometry. Custom loss function enforces velocity reconstruction, incompressibility, no-flow conditions, and physical constraints. Tested on out-of-distribution geometries and accelerates Lattice-Boltzmann simulations (90% of cases).
Two new self-supervised learning models for link prediction in graphs: L-GRACE and L-BGRL. Based on link representations instead of node representations, they incorporate structural augmentation grounded in community structure. Performance matches state-of-the-art in both supervised and self-supervised settings.
Chronicle is a 324M-parameter multimodal foundation model trained from scratch on natural language and time series in a unified architecture. Both modalities share the same transformer blocks and attention mechanisms. It matches Gemma-3-270M on 19 NLU tasks, sets new benchmarks on 24 UCR/UEA datasets, and outperforms supervised fusion baselines on Time-MMD.
Theoretical paper on out-of-distribution generalization in RL. Authors extend state abstraction framework to POMDPs and introduce successor-weighted model reduction. They prove smaller abstract state spaces improve generalization to more complex tasks.
OmniISR proposes a unified framework for centralized and federated learning via intermediate supervision and regularization. The framework uses mutual information to align internal covariate shifts and negative entropy to regularize overconfident predictions. O(1/sqrt(T)) convergence guaranteed theoretically; CL-FL gap reduced by 22.60% in experiments.
Plug-and-play framework to convert Transformer nonlinear operators into spiking neural network (SNN) compatible operations. Decomposes Softmax, SiLU, and normalization into primitives (division, exponentiation, L2 norms) executable by LIF neuron groups without fine-tuning. <1% accuracy drop on LLM benchmarks.
Novel predictive coding approach via hierarchical Gaussian filters. Authors restore precision-weighted message passing, enabling simultaneous learning of activations, weights, and precisions without global error signals. On FashionMNIST, the method converges faster than backpropagation while retaining biological advantages of predictive coding.
Repeating a smaller dataset during training accelerates learning compared to using a larger dataset, via sampling biases that enable favorable layer-wise growth. Effect observed across algorithmic tasks, architectures and optimizers. Authors provide theoretical analysis and empirical interventions.
Study using BERT to analyze Decentraland Discord community sentiment and forecast MANA token price. Multi-modal LSTM model integrating sentiment, trading volume, and market cap significantly outperforms price-only baseline. Community sentiment predominantly neutral with positive skew.
Study on quantization of LLaMA-3.1 (8B) for qualitative analysis. 8-bit models maintain best precision; 4-bit, 3-bit, and 2-bit models suffer increased hallucinations. A guided multi-pass verification method reduces hallucinations and improves low-bit model stability, making qualitative analysis accessible with fewer resources.
Framework for processing long documents via parallel chunking and evidence-anchored consolidation. Reduces omission error by 84%, increases evidence traceability by 130%, decreases unsupported claims by 91%. Smaller models benefit most.
arXiv paper showing that real-data scaling laws are governed by progressive coverage of a latent predictive contribution spectrum, beyond token-frequency tails alone. Using suffix-automaton representation, authors define a global-KL spectrum and demonstrate strong correlation (R²≈0.96) between spectrum slope and empirical scaling exponent across 12 corpora.
MedicalBench is a benchmark for extracting implicit medical concepts from electronic health records (MIMIC-IV). It formulates the task as verification of note-concept pairs with sentence-level evidence identification. State-of-the-art LLMs show modest performance, highlighting the difficulty of implicit medical reasoning.
FlowLM converts pre-trained diffusion language models into flow matching models via efficient fine-tuning. By realigning curved diffusion trajectories into straight-line flows, FlowLM achieves high-quality few-step text generation rivaling 2,000-step diffusion sampling. Performance saturation is reached with half the training epochs compared to training from scratch.
Study of synchronization in full-duplex dialogue models (Moshi) that listen and speak simultaneously. Researchers measure internal representation alignment via CKA and detect anticipatory turn-taking cues. Synchronization is strong without noise, degrades with noise, and internal states encode predictive information for turn-taking.
Study on generating long-form literary reviews based on Torrance Test of Creative Writing (TTCW). Dataset of 263,911 stories annotated across 14 creativity dimensions. Fine-tuning Qwen3 (4B and 8B) shows non-reasoning supervision achieves better performance (0.6820), while reasoning-supervised models fail to complete the required 14-metric review format.
DEL (Digit Entropy Loss) is a novel loss function to improve numerical prediction in LLMs. Tested on CodeLlama, Mistral, DeepSeek, and Qwen-2.5 across 7 mathematical reasoning benchmarks, it outperforms existing methods (MLE, Number Token Loss) by optimizing digit entropy in a supervised manner and generalizing to floating-point numbers.
Stage-Audit detects hallucinations in LLM-curated tables by enforcing curator-auditor separation and row-level source verification. On 51 Seed2Frontier instances, precision improves from 0.356 to 0.505 (+42%) and F1 from 0.334 to 0.451 (+35%), with explicit per-row source traceability.
Study on collocational bootstrapping: mechanism where regularities in word co-occurrence patterns provide cues to syntactic dependencies. Neural networks trained on synthetic datasets varying in subject-verb pairing predictability. Results suggest this mechanism could explain subject-verb agreement acquisition in children.
Corpus-centric diagnostic framework for analyzing biomedical named entity recognition (NER) and entity linking (EL) benchmarks. Applied to 9 corpora, reveals substantially different properties can mask apparently identical tasks. Open-source code and interactive dashboard provided.
Study of 6,233 MedGPTs and 10 open-source models deployed on the web. 25-30% show low factual accuracy, 33.6-54.3% violate operational thresholds, 57% of Action-enabled models lack privacy disclosures. Authors introduce MedGPT-HEval for hallucination detection and release HAA-MedGPT, a structured dataset.
Study across 11 generations of self-training on 5 models (GPT-2, Pythia, OPT). Contrary to uniform 'flattening', language restructures: surface markers (connectives, em-dashes) rise while deep syntactic structures (questions, passives, subjunctives) collapse. Structural Depth Hypothesis predicts this decay (ρ=0.540, p<10⁻⁶).
DPR-BAG generates abstracts for biomedical articles lacking summaries through structured decomposition (BOMRC schema), parallel LLM-based summarization, and refinement. On PMC-MAD (46,309 articles), improves abstractive novelty while maintaining factual consistency. Training-free, zero-shot framework.
Two-phase RAG system for corporate credit analysis: phase 1 combines lexical and dense multilingual retrieval; phase 2 applies adaptive controller and LLM-as-Judge scoring based on analytical utility rather than semantic similarity. On-premise deployment on proprietary multilingual corpus. Production: document review time reduced from hours to 3 minutes across 800+ analysts.
LFD (LLM-assisted Feature Discovery) method generates interpretable text representations via inter-annotator agreement (Cohen's κ) and label disentanglement. Validated on 10 text classification tasks: features clearer and less label-entangled than bottleneck baseline, confirmed by human audit (232 raters).
Counter Turing Test evaluates AI-generated text detection techniques. Task A (binary classification) achieves F1=1.0 to distinguish human vs AI text. Task B (model attribution) reaches 0.9531 for identifying GPT-4, Claude 3.5, Llama. Top approaches combine DeBERTa, BART, fine-tuning, and ensemble learning.