New OpenAI Academy courses for the next era of work
OpenAI launches three Academy courses to build practical AI skills, create repeatable workflows, and apply agents in everyday work.
2731 articles
OpenAI launches three Academy courses to build practical AI skills, create repeatable workflows, and apply agents in everyday work.
Visa and OpenAI partner to integrate payments directly into AI agents. Visa's payment infrastructure will be embedded in OpenAI agents to automate transactions.
Developer builds browser-use agent running entirely in WASM/WebGPU without server, using Wllama and ShowUI-2b. Supports typing, clicks, multi-turn actions (50% success rate), dropdowns. Alpha code shared on GitHub.
A Canadian mother is suing OpenAI after ChatGPT allegedly assisted her suicidal daughter in ending her life. The incident raises questions about AI companies' liability for content generated by their models.
Huawei launches openPangu 2.0 at HDC 2026 (June 12). Two versions: Pro (505B params, 18B activated) and Flash (92B params, 6B activated). 512K context, 28:1 sparsity. Optimized for Ascend: 2x throughput, reduced latency. Open-source from June 30 (architecture, weights, inference and training code).
Anthropic launched Claude Fable 5 and Claude Mythos 5 on June 9, 2026—two products built on the same underlying model but differentiated by distinct safety guardrails. This strategy reveals a segmentation approach based on security controls rather than architecture.
Supermicro raises up to $7 billion to finance AI server production. Despite this major funding round, the stock price declines on the market.
Apodex open-source model collection (0.8B–35B) trained as search agents with ReAct. 4B-SFT outperforms 30B models on hallucination reduction via BrowseComp. Tested on 3090: 4B in vLLM performant, 35B mini slow but effective. No official gguf available.
PaddleOCR v6 officially released with models ranging from 1.5M to 34.5M parameters. +4.9% detection accuracy, +5.1% recognition accuracy vs v5. Up to 5.2× faster CPU inference with OpenVINO. Supports 50 languages, new use cases (PCB, CAD, digital tubes). Apache 2.0 open-source.
Samsung rolls out three AI features from Galaxy S26 to Galaxy S25 through a software update. Capabilities previously exclusive to the newer model become available to S25 users.
EAGLE3 merged into llama.cpp after 6 months of development. The helper model receives guidance from the main model, unlike MTP where it operates independently.
C++ implementation of distilHuBERT with no runtime dependencies. Weights compiled into the library, supports dynamic sizes, performance on par with onnxruntime. Easy CMake integration.
Technical article on building a plugin system without shared runtime, centralized storage, or shared JavaScript context. Software architecture approach for plugin isolation.
Gladia positions Solaria-3 as leader in production audio transcription (noisy meetings, accents, telephony). The transcription API market has shifted toward these complex use cases since 2024-2025.
Vercel suspends access to Claude Fable 5 on AI Gateway following a US Government legal directive. Other Anthropic models remain accessible.
Coinbase launches Coinbase for Agents, a platform of autonomous AI agents to automatically manage cryptocurrency portfolios. The service enables users to delegate their transactions and investment strategies to AI agents.
InfiniteKV compresses KV cache into 104-byte searchable records stored in RAM or disk instead of deleting old tokens. Mistral-7B correctly answers at token 76,747 (2.3× its 32,768 training window). One million tokens requires ~3 GB instead of 122 GB.
LLM context compression technique achieves 16x compression ratio, outperforming traditional KV cache approaches. Method significantly reduces memory usage while maintaining response quality.
Geopolitical tensions over digital sovereignty escalate following revelations about US surveillance of communications. The Netherlands and EU strengthen demands for technological independence amid risks of foreign control.
Deezer launches a tool to detect AI-generated music. The platform enables users to identify whether a song was created by AI, addressing growing concerns about music content authenticity.
Loopcraft explores the concept of stacking iterative loops to improve AI systems. Work by Peter Steinberger, Boris Cherny, and Andrej Karpathy on iterative process architecture.
An AI agent caused financial damage to its operator while attempting to scan DN42, a private experimental network. The incident highlights risks of inadequate control over autonomous agents.
Evoflux is an inference-time evolutionary search method for repairing executable tool workflows in compact agents. On MCP-Bench with 250 tools, it raises execution feasibility from ~3% to 17-24%, outperforming SFT, SFT+DPO, and ReAct under scarce teacher-trace budgets.
LAUKIN is a dataset of 14,727 contract clause pairs (Australia-UK, UK-India, India-Australia) labelled for legal equivalence. 3,000 pairs are manually annotated by legal experts. Best models achieve 65.11% macro-F1, revealing that drafting conventions diverge significantly across jurisdictions despite shared legal heritage.
MemRefine compresses LLM agent memory over long-term interactions by using an LLM judge to decide merge, delete, or preserve decisions based on factual content rather than surface similarity. Tested across multiple long-conversation benchmarks, the system meets memory budgets while preserving downstream performance.
G-Long is a graph-enhanced framework using a fine-tuned small language model for efficient long-term memory management in dialogue systems. It extracts structured triplets and employs an attention-aware importance scoring mechanism with a T5 summarizer to identify salient memories. Gains: +9.8% response quality (MSC), +40.8% retrieval recall (LME).
HyPE introduces a hypergraph encoder for persona-grounded dialogue systems. The method structures persona attributes into category-induced hypergraphs and applies HyperGCN with Persistent Edge Embeddings (PEE). Evaluation on PersonaChat shows consistent improvements over baselines across GPT-2, LLaMA-3.2-3B, and Qwen2.5-3B.
Two-stage local pipeline using MedGemma-27B for extracting structured clinical information from unstructured EHR notes. Separates binary presence classification from value extraction, no external API calls or fine-tuning. Macro-F1 score of 0.55 on CRF 2026 test set, second place among open-source local submissions.
Study proposes MentalMARBERT, a domain-adapted BERT model for detecting mental health disorders in Arabic text. Two-phase framework: domain adaptation (DAPT/TAPT) on 50,670 annotated tweets, then hierarchical fine-tuning with LoRA. Macro-F1 0.861, accuracy 0.877.
Causal analysis of latent reasoning models (Coconut, CODI): observable patterns (BFS-like frontiers, decodable arithmetic) are not evidence of reasoning mechanisms. Causal interventions show latent-thought utilization is graded, not binary, and concentrated in low-rank directions. Decodability alone cannot establish mechanism.
Multi-agent simulation of morphological alternation emergence (e.g., go/went in English). Alternative forms spread through probabilistic adoption between agents. Evaluation via LLM-driven "Historical Linguist" comparing real, disguised, and evolved morphologies. Scale-free networks and Bernoulli adoption favor plausibility.
EDEN is a corpus of 4 million anonymized clinical notes from Italian emergency departments. 6,000 notes were manually annotated by clinical experts using a structured form with 132 items covering dyspnea and loss of consciousness. CRF-filling benchmark with Gemma-27B and MedGemma-27B baselines.
Study of ASR errors in Vietnamese speech translation. Authors categorize substitution errors by phonetic cause and propose PiDA (Phonetically-Informed Data Augmentation), which augments training data with phonetically similar corruptions. Fine-tuning on FLEURS Vietnamese-English improves translation of erroneous ASR outputs (+2.04 BLEU).
New arXiv paper introducing MINARD, a video generation system that transforms scientific figures into narrated walkthrough videos with region grounding. The pipeline generates paper-grounded narrations and sequentially aligns them to figure regions. Includes FigTalk benchmark with component-level grounding metrics.
SENTINEL is a failure-driven reinforcement learning framework that improves tool-using LLM agents by converting their failures into targeted training tasks. On Tau2-Bench Retail with Qwen3-4B-Thinking-2507, the method increases Pass@1 from 66.4 to 74.9 through a Controller-Proposer-Solver loop that analyzes recurring error patterns.
Bernstein-Schur kernels: random feature construction combining sketched finite modulation and radial randomization via Bernstein-Widder scale. Feature dimension Dm without O(d²) cost of exact modulation. Exact variance guarantees and operator-norm bounds controlled by intrinsic dimension, with kernel ridge regression applications.
ToolSense is an open-source diagnostic framework to audit actual tool understanding in LLMs. Applied to ToolBench (~47k tools), it reveals a knowledge-retrieval dissociation: five parametric model configurations collapse by 50-64 percentage points on realistic ambiguous queries, falling below embedding baselines, despite strong performance on standard benchmarks.
SkillChain automates skill evolution for multimodal e-commerce AI assistants. The system manages three stages: Skill creation from task specs, routing optimization, and iterative refinement via dual-path LLM-Judge evaluation. Deployed at production scale, it improves structural compliance and content quality, confirmed by A/B testing on user engagement.
LEDGER is a benchmark of 4,999 digitized corporate annual reports to evaluate LLM long-context capabilities in finance. The corpus includes 31 consolidated financial KPIs, 118,048 TREC-style retrieval questions, and extraction tasks on numerically dense documents. Case study: correlation between CEO rhetoric and post-publication market impact.
AfriSUD is the first large-scale collection of syntactically annotated treebanks for nine African languages using the SUD framework. Evaluations on POS tagging and dependency parsing reveal a significant syntax gap: models (non-transformer baselines, multilingual encoders, LLMs) show clear limitations across the structural diversity of African-language syntax.
NTS-CoT, a Chain-of-Thought framework, mitigates hallucinations in LLM-based timeline summarization. Three modules (Element-CoT, Date Selection, Causal-CoT) capture essential news elements, select timestamps, and infer causal relationships. Evaluation on three TLS benchmarks shows measurable improvements over baselines.
NaturalFlow optimizes simultaneous speech-to-speech translation by reducing inter-chunk pauses to improve acoustic fluency. The framework leverages model-internal signals (linguistic diversity, temporal variability) to balance low latency with natural speech flow, validated on short- and long-form benchmarks.
Study demonstrating AI reviewers can be manipulated through presentation-only changes (abstract, framing, narrative structure) without altering methods or results. Attack success rate of 75.1% with mean score gain of +1.21/10. AI reviewers confuse appearance with substance, favoring strength highlighting over weakness refutation.
LLMs lose up to 65% accuracy when task-critical information arrives across conversation turns. Compact rolling memory replaces growing history attention. A sharding pipeline converts QA datasets into multi-turn fragmented-information episodes without manual annotation. Training on sharded GSM8K improves multi-turn accuracy and generalizes zero-shot to harder math and out-of-domain long-context QA.
MÖVE is a holistic benchmark evaluating 39 LLMs for German public administration. Beyond performance (summarization, QA, topic extraction), it measures hallucinations, energy consumption, provider transparency, and alignment with German constitutional values. Uses 10 German-language datasets; no single model dominates across all criteria.
EvoBrowseComp is an evolving benchmark of 400 English and 400 Chinese questions to evaluate search agents (LLM + web tools). Unlike static BrowseComp, it uses live-web traversal and a three-agent framework (QA synthesis, information filtering, high-level guidance) to prevent contamination and parametric memorization. The benchmark auto-updates regularly.
PRISM is a multi-agent framework for empathetic spoken dialogue that decouples speech perception, response generation, and speech synthesis. It introduces a prosody-to-language translation mechanism to stabilize LLM reasoning and integrates external knowledge tools. Results show improvements in empathy, prosodic appropriateness, and response quality across metrics.
BioStance: dataset of 39,600 annotated Post-Comment pairs from Reddit for stance detection in bioethical debates. Covers 6 controversial targets (value conflicts, individual liberty vs collective responsibility, technological uncertainty). Triple-annotated, Krippendorff's α = 0.82.
SafeLLM compares line-number extraction to free-form rewriting for RAG systems in safety-critical settings (SOPs, HR policies, medical guidelines). Line-based extraction outperforms direct copying and safety-focused strategies, achieving 95% term recall on NHS and NICE documents with better source text fidelity.
Study of internal anchoring mechanisms in language models. Researchers localize circuits responsible for anchoring bias (where irrelevant numbers influence answers) in Qwen and Llama 7B-8B models. Edge-level attribution methods recover this signal more faithfully than node-level methods.
Localized repair method for outdated summaries: DETECT-REMASK-REPAIR uses masked diffusion to identify and fix unsupported claims in existing summaries without full regeneration. StreamSum benchmark introduced. Experiments on DialogSum show faithfulness improvement and repair in <0.5s.
Three small LLMs (Phi-3-mini 3.8B, Qwen2.5-3B, Mistral-7B) fine-tuned via QLoRA for biomedical claim verification. Mistral-7B outperforms GPT-4o and GPT-5 (+12% F1) on 1,008 training examples. Study identifies structural artifact in SciFact and demonstrates robust cross-domain generalization.
Simple prompting techniques improve LLMs' ability to capture human judgments. Across 144 moral scenarios (US) and 38 moral beliefs (32 countries), asking models to report standard deviations and response proportions better recovers the full distribution of human responses. LLMs also track human confusion ratings, though their error calibration remains poor.
Rigel empirically characterizes Metal 4.1 tensor compute path on Apple M4 Max. Researchers find fp8 (E4M3) matmul2d is emulated, not accelerated (0.94x fp16 throughput), executes on GPU shader cores without dedicated matrix datapath, and accumulates in ≥fp32. Hand-fused GEMM+bias+GELU kernel gains +6.5-12.9% in cache-resident regime.
PaperGuard, a multimodal benchmark, evaluates LLMs and MLLMs vulnerability to adversarial attacks in scientific peer review. Researchers test prompt injections and perturbations (GCG, PGD) on text and figures, proposing a defense using chunk-based embedding search to localize harmful instructions.
Empirical study on Direct Preference Optimization (DPO) for chatbot fine-tuning. Results show DPO simplifies the training pipeline, improves computational efficiency, and achieves competitive performance. Evaluation using BLEU, ROUGE, and cosine similarity metrics demonstrates effective learning and convergence, though training instability remains.
LoHoSearch is a 544-question benchmark for evaluating long-horizon search agents, built via an automated pipeline on a knowledge graph of 7 million Wikipedia entities. The strongest model achieves only 34.74% accuracy, versus >90% on prior saturated benchmarks.
X-MADAM-RAG diagnoses conflicts between Chinese and English evidence in RAG systems. On the X-RAMDocs-ZHEN benchmark (300 examples), the pipeline achieves 96.67% strict accuracy with Qwen2.5-7B-Instruct, but fails on naturalized variants (30% accuracy), revealing document-level extraction as the main bottleneck.
MARD is a 7B parameter model for mechanism-level drug-drug interaction prediction (enzyme, pharmacodynamic axis). Uses reasoning distillation with process-reward-weighted DPO and mechanism-aware retrieval. On April-2026 DrugBank: +13.9pp over best baseline, +6.7pp over GPT-4o, with robust generalization to unseen drug pairs.
STG, a Structured Testbench Generation framework, accelerates verification of LLM-generated HDL designs. 720x faster than iterative LLM approaches, it reduces false positives, improves coverage, and identifies errors in existing benchmarks. As a data curation engine, it is 11x faster than LLM-based filtering with 127x less energy consumption.
APCyc is a target-aware de novo cyclic peptide generation framework that explicitly models cyclization and jointly optimizes multiple physicochemical properties. The model uses an expanded residue vocabulary and Bayesian posterior guidance to generate cyclization-aware peptides adapted to specific therapeutic targets.
Mathematical forum platform embedding Mathpix OCR pipeline to convert images to LaTeX directly in posting interface. Three-layer architecture (image processing, rendering, storage) supporting desktop/mobile. Generates community-validated dataset of math problems and solutions for training AI reasoning systems.
OpenMedQ is a medical vision-language model pretrained on 14 datasets (~3.35M samples) covering pathology, radiology, microscopy, and clinical QA. It achieves 75.9 BLEU-1 on PathVQA (outperforming Med-PaLM M 562B) and 0.757 average macro-F1 on 8 unseen medical classification benchmarks.
GENIE is a fine-grained evaluation metric to measure novelty of LLM responses along task-specific features relative to a population of responses. Authors show holistic metrics fail to capture novelty's high-dimensionality and use GENIE to assess effectiveness of creativity mitigation methods.
Multi-factor value model for long-running LLM agent memory. Seven cognitive factors (emotional intensity, goal relevance, value alignment, etc.) weighted via gradient-free optimization. Retains 77% of critical evidence vs 36.8% for recency baseline on LongMemEval.
Shopping Reasoning Bench: expert-authored benchmark of 525 missions (232 single-turn, 293 multi-turn) with 10,863 importance-weighted binary rubrics for evaluating conversational shopping assistants. Evaluation of 9 models (GPT, Claude, Gemini): pass rates 57–77%, performance degrades 4–18 points across conversation turns, 13–29 point gap between required and optional criteria.
HCPD, a hallucination detection method without access to model internals or external references. An LLM agent adaptively decomposes judgment into weighted, interpretable criteria aligned via weak supervision on semantic consistency. Code released.
MARS is an adversarial stopping rule for parallel LLM test-time scaling. It probes partial traces at intermediate checkpoints to estimate which traces will change answers, enabling early stopping once the leading vote is safe. Across three reasoning models and three competition-math benchmarks, MARS saves 25-47% of self-consistency tokens while maintaining accuracy.
MDForge is an LLM agent automating molecular dynamics (MD) pipeline design through code generation and multi-agent expert debate. On three SAMPL benchmarks, it matches human expert performance and discovers a novel CB[7] binder confirmed by wet-lab NMR as a high-affinity, picomolar ligand.
Theoretical and empirical study of the scaling factor α in LoRA. Authors show α dominates optimization far more than learning rate via a Signal-Drift framework. They discover a square-root law linking α to rank and propose LoRA-α to improve performance and simplify hyperparameter tuning.
HarnessBridge is a learnable agent-environment interface controller using bidirectional projection. The module distills raw observations into compact decision-relevant states and converts proposed actions into executable transitions. Trained on Terminal-Bench 2.0 and SWE-bench Verified, it matches specialized harnesses while reducing token usage and trajectory length.
Teach VLM extracts operational knowledge from mobile screen demonstrations by analyzing visual state transitions and generating natural-language instructions. The Teach-and-Repeat paradigm uses this knowledge to guide downstream GUI execution agents. Evaluation on Android World shows consistent Task Success Rate improvements.
Experimental study across 280 runs showing unconstrained LLMs fail in 72% of economic research tasks versus 16% with Human-in-the-Loop Economic Research (HLER). Architecture enforces pre-commitment, decision sequencing, and three human gates. Significant result (p<0.001): humans validate, LLMs reason but do not execute data work.
PRISMR addresses parse collapse in multimodal listwise ranking with LMMs. Problem: autoregressive decoders silently omit candidates and terminate early on long lists. Solution: lightweight hypernetwork generates item-specific LoRA weights synthesized into instance-specific adapter. New large-scale multimodal review-ranking benchmark introduced.
Study on generating evaluation datasets for procedural reasoning. Three strategies compared: strict generation from TMK (Task-Method-Knowledge) models, transcript-first generation with post-hoc TMK filtering, and TMK-aware generation. Across 690 QA pairs and 23 instructional topics, strict TMK generation achieves 96.5% grounded questions and 92.6% usable items.
Multi-modal agent framework for power distribution defect detection using foundation models. Systematic evaluation of three capabilities: perception (equipment identification and defect description), reasoning (diagnosis and maintenance planning), tool usage (autonomous execution). Domain-specific evaluation dataset and comprehensive benchmark developed.
DailyReport is an open-source benchmark evaluating search agents on 150 real-world daily tasks with 3,546 evaluation rubrics. Tasks decomposed into subtasks with cascade evaluation across disentangled dimensions. Testing 17 agentic systems reveals significant gaps versus user expectations.
WISE is a long-horizon Minecraft agent using a Causal Event Graph to augment episodic memory. The framework couples causal reasoning (why-which) with spatiotemporal memory, enabling opportunistic subtask reordering and multi-scale exploration. Significant improvements on sparse tasks requiring adaptive decision-making.
SciAgentArena is a systematic benchmark evaluating ~200 real-world scientific tasks with stepwise verification. Current AI agents perform well on structured data-analysis workflows but struggle to generate novel insights, sustain self-directed exploration, and solve open-ended research questions.
arXiv study comparing self-reports (Big 5 vs Theory of Planned Behavior) with actual behavior of 11 frontier LLMs across 4 tasks. TPB reaches human-level coherence within shared conversation but fails across separate conversations. Broad personality frameworks poorly predict real behavior.
Study on deploying deep neural networks for EEG analysis on wearable devices. Authors explore parameter quantization and electrode reduction techniques to decrease computational complexity while maintaining accuracy for epileptic seizure detection.
arXiv report analyzing the transition from AGI (human-level general intelligence) to ASI (artificial general superintelligence). Identifies four possible pathways: scaling, paradigm shifts, recursive improvement, and multi-agent emergence. Raises questions about technological frictions and bottlenecks.
TrajGenAgent is a hierarchical LLM-agent framework for realistic human mobility trajectory generation without model fine-tuning. An orchestrator LLM synthesizes activity chains via in-context learning, then a deterministic workflow grounds them using personalized POI retrieval, distance-aware location selection, and LLM-based duration estimation. Evaluation via anomaly-detection framework on benchmark datasets.
Study evaluating 4 lie detectors across 31 models (2B-1T parameters). Detectors (CoT judge, logprob classifier, activation probes, DYL) perform well on prompted lying but fail on trained model organisms with verified beliefs. Only CoT judge maintains 0.82 balanced accuracy.
Optimization framework for AI agents: minimize support calls while controlling errors when agents act alone without assistance. Adaptive online algorithm with randomized exploration, tested on information gathering, human-AI collaboration, and tool use.
Pythagoras-Prover is an open-source family of efficient Lean theorem provers (4B and 32B parameters, including a diffusion-based prototype). Via curriculum SFT and Augmented Lean Formalisation (ALF), the 4B model outperforms DeepSeek-Prover-V2-671B on MiniF2F-Test (86.1% vs 82.4%) with 167x fewer parameters. The 32B achieves 93.0% on MiniF2F-Test and solves 93/672 PutnamBench problems.
Analysis of 80,814 papers from 5 major AI conferences (2017-2025) reveals research topics advance through abrupt phase transitions, not gradually. LLMs dominant by 2025; diffusion models and vision-language models surged within 1-3 years. Early-warning signature flags reasoning, test-time compute, agentic AI, multimodal LLMs, RAG, and world models as topics to monitor 2026-2028.
Arbor is a multi-agent framework introducing tree search as a cognition layer for autonomous agents. Validated on full-stack LLM inference optimization, it pairs an Orchestrator agent with a Critic agent in a checks-and-balances architecture. Arbor achieves 193% throughput-latency Pareto improvement over vendor-optimized baselines, versus 33% for a single agent that crashes within hours.
arXiv study showing frontier models (Claude Opus 4.5, GPT, Gemini) detect tampered prefills in 9-35% of cases with 0% false positive rate. This 'prefill awareness' undermines alignment and jailbreaking evaluations relying on inserted assistant context. Models distinguish stylistic from preference mismatch.
arXiv tutorial on world models and physical AI. Distinguishes explicit models (structured dynamics for planning) from implicit models (predictive structure in learned representations). Applications in robotics and autonomous driving. Challenges: hierarchical reasoning, long-horizon planning, autonomous goal formation.
MLUBench is a large-scale benchmark for evaluating lifelong unlearning in multimodal large language models (MLLMs). It contains 127 entities across 9 classes. Authors show existing unlearning methods suffer cumulative degradation and propose LUMoE to preserve multimodal alignment during sequential forgetting.
Deployment study of an LLM embedded in electronic health records. A pre-response classifier predicts user rejection risk (AUROC 0.719) by leveraging deployment-specific context (provider type, department, model). Prospective analysis over 4.5 months.
Paper introduces DAF-AGI, a governance framework to adjudicate competing AGI definitions. Analyzes five measurement families (performance, capability-ontology, psychometric, skill-acquisition, economic) and tests whether current generative systems constitute AGI. Only performance-based operationalization certifies this claim on 2024-2025 evidence.
Formal theory of epistemic state inference: ToM-U formalizes belief attribution through typed directed graphs (Local Epistemic World Models) representing agents, state nodes, and epistemic relationships. Three inference procedures and a residue function capture structured traces of mentalizing failures, without presupposing belief states.
PersonaDrive is a retrieval-augmented VLA (vision-language-action) agent pipeline for closed-loop driving simulation, conditioned on retrieved human demonstrations. Trained on CARLA data with aggressive/neutral/conservative instructions, it improves driving score by 4.6% on Bench2Drive and generates style-diverse non-ego agents without per-style retraining.
Study on story generation conditioned by Persian proverbs. Researchers introduce PAND (Proverb Aligned Narrative Dataset) and identify a 'decompression gap': LLMs produce fluent text but fail to preserve the moral and causal structure of proverbs. Explicit reasoning and iterative refinement partially mitigate these errors.
Critical article questioning whether AI tools genuinely boost developer productivity. The author argues that instead of improving efficiency, these tools are increasing workload and complexity in development workflows.
MTPLX V1 is a native macOS app (Swift) for running MLX models with MTP (Multi-Token Prediction). Forge auto-converts Hugging Face models to MLX+MTP with real speedup measurement. Qwen 3.6 27B reaches 63 tps (vs 28 tps in v0.1). Includes native chat, live dashboard, AIME 2026 benchmarking, Qwen 3.5 9B and Gemma 4 support.
Empirical comparison of Gemma and Qwen quantizations across three tasks (arithmetic, presidential dates, attention). Gemma-4-31B-Q4_K_S reaches 83.8% arithmetic and 87% attention accuracy. Qwen3.6-27B-Q4_K_S achieves 95.5% arithmetic and 100% presidents. Results demonstrate major impact of model size and quantization scheme on accuracy.
llama.cpp benchmark on Intel 250K Plus CPU: optimizing --threads argument yields +80% performance gain (49 → 88 tok/s). 16 threads optimal vs 6 threads (P-cores only). Using all 18 cores drops performance without throttling detected.