GLM 5.2 Release Video [Made with GLM 5.2]
GLM 5.2 generates videos via Remotion, comparable to Fable but below Gemini 3.1 Pro. Server overload observed on OpenRouter with timeouts on long outputs.
AI video generation refers to the automatic creation of animated sequences from text, images, or audio inputs. Sora (OpenAI) is a notable example of a model that produces realistic clips from a plain text description.
GLM 5.2 generates videos via Remotion, comparable to Fable but below Gemini 3.1 Pro. Server overload observed on OpenRouter with timeouts on long outputs.
OpenMontage is an open-source, agentic video production system with 12 pipelines, 52 tools, and 500+ agent skills. Converts an AI coding assistant into a full video production studio.
OpenMontage is an open-source, agentic video production system with 12 pipelines, 52 tools, and 500+ agent skills. Converts an AI coding assistant into a full video production studio.
xAI makes Grok Imagine Video 1.5 accessible, its video generation model now capable of producing videos with synchronized audio.
Mel AI demonstrates video-native AI characters that talk, lip-sync, show facial reactions, and respond in real time to camera context. The system detects user environment and adapts responses accordingly. This approach moves beyond text-based Character AI (founded by former Google/LaMDA developers).
Mirage, a video world model from Microsoft Research, stores scene information in latent space instead of pixel-based point clouds. This cuts compute time and graphics memory while maintaining spatial consistency through long camera moves. Object tracking across segments remains unreliable.
New arXiv paper introducing MINARD, a video generation system that transforms scientific figures into narrated walkthrough videos with region grounding. The pipeline generates paper-grounded narrations and sequentially aligns them to figure regions. Includes FigTalk benchmark with component-level grounding metrics.
Adaptive video tokenisation method exploiting temporal redundancy in frozen tokeniser latent space via fixed threshold on per-position temporal-L1 differences. Latent Inpainting Transformer (LIT) reconstructs dropped positions. Single encoder + one LIT pass pipeline: 31× speedup over ElasticTok-CV, 2× over InfoTok on TokenBench and DAVIS benchmarks.
Google DeepMind releases Gemma 4 12B, a unified encoder-free multimodal model. The model processes text, images, and video in a single architecture, optimized for on-device inference.
Hugging Face introduces Room360, a video-to-3D spatial reconstruction platform. The tool converts video sequences into usable 3D models for immersive and architectural applications.
NVIDIA releases Nemotron 3.5 Content Safety, an open-source multimodal safety model detecting harmful content across text, image, and video. Customizable for global enterprises, it provides granular policy control for moderation across regions and use cases.
NVIDIA releases Cosmos, an open platform of world models, datasets, and tools enabling developers to build Physical AI for robots, autonomous vehicles, smart infrastructure, and more.
Pixelle-Video is an AI-powered automated short video generation engine. The GitHub project provides an end-to-end solution for creating short videos without manual intervention.
xAI releases grok-imagine-video-1.5-preview, an image-to-video model generating cinematic videos up to 720p from still images and text prompts. Multiple clips can be stitched together into longer scenes.
Grok Imagine Video 1.5 from xAI is now available on AI Gateway. The model generates video from input image with synchronized audio in single pass. Improvements: audio quality, prompt following, photorealism, character consistency across longer sequences, expanded reference image support for visual style control.
NVIDIA releases Cosmos 3, a collection of omnimodal world models (Nano 16B, Super 64B) capable of generating dynamic video, image, audio, and action commands from text, image, video, and action trajectory inputs. Available on Hugging Face for Physical AI applications.
GEM is a concept erasure framework for Rectified Flow Transformers. It bridges trajectory-based unlearning (Generative Flow Networks) and teacher-guided erasure, using geometric guidance signals to suppress unwanted concepts while preserving benign generation and preventing harmful content synthesis.
NVIDIA announces Cosmos 3 (video model), Nemotron 3 Ultra (optimized LLM), and RTX Spark. Jensen Huang claims a major win for the company.
Ethan He, Grok Imagine lead at xAI, discusses building the video generation model in 3 months, compares video generation to world models approach, and argues why video agent models represent the next frontier.
Nvidia launches three physical AI models at GTC Taipei: Cosmos 3 (world model), Alpamayo 2 Super (scaled-up autonomous driving model), and an open reference platform for humanoid robots.
Google fixes bugs in Gemini usage limits where a single Omni video consumed entire quotas. Ultra members now get twice as many video generations, failed requests are no longer charged, and Google plans increased transparency on usage.
Amazon MGM Studios and AWS launch a creators' fund and in-house AI platform called 'Project Nara'. Three animated series are in production with five-week timelines for pilots. Amazon claims the only end-to-end AI content ecosystem in the industry.
Diffusion models applied to basketball trajectory simulation conditioned on partial sketches of player movements. The model jointly refines all player trajectories, producing more natural simulations than autoregressive generation. Code and model fully open-sourced.
YouTube deploys automatic detection system to flag AI-generated or heavily AI-altered content starting May 2026. Labels will display more prominently: below player for long videos and as overlay on Shorts. Recommendations and monetization unaffected.
YouTube enforces visible labels to identify AI-generated videos. This measure aims to improve transparency and help users distinguish authentic content from synthetic content.
MoneyPrinterTurbo: open-source tool generating high-definition short videos with one click using AI LLMs. Automates video content creation.
MoneyPrinterTurbo: open-source tool generating HD short videos with one click using AI LLMs. Automates video content creation.
Autoregressive video diffusion models use quantized KV caches to reduce memory, but quantization creates an attention bias (Jensen bias) that degrades quality. Authors propose a per-attention-score correction computed from quantization step sizes, recovering quality lost with INT2 quantization while using 50% less memory than INT4.
Tail-Aware HiFloat4 applies W4A4 post-training quantization to the Wan2.2 text-to-video generation model. The method adapts ViDiT-Q using HiFloat4 format, quantizes transformer linear layers, preserves numerically sensitive modules in high precision, and introduces activation-tail-aware percentile calibration to reduce impact of rare outliers.
Creators produced a cinematic heist movie trailer by combining 4 AI models for $60. Demonstrates feasibility of low-cost AI video production.
A short film presented at Cannes cost $500k to produce, with $400k spent on AI compute. The ratio reveals the growing share of infrastructure costs in video generation and creative content production.
Meituan releases LongCat-Video-Avatar 1.5, an open-source framework for audio-driven human avatar video generation. Upgrades audio encoder from Wav2Vec2 to Whisper-Large, supports Audio-Text-to-Video and Video Continuation with 8-step inference. Human evaluation on 508 image-audio pairs across 6 scenarios and 2 languages.
xAI releases Grok Imagine Video 1.5, which animates still images into short clips with synchronized audio in a single pass. Guide on how to maximize output quality.
PRISM is a 10,372 instruction-code pair benchmark for evaluating programmatic video generation by LLMs. It proposes 4 metrics: code reliability, spatial coherence, visual complexity, and temporal density. Evaluation of 7 LLMs reveals a 41% execution-spatial gap: executable code does not guarantee spatially coherent output.
Google unveils Gemini Omni at its I/O 2026 conference, a video AI capable of mastering physics and maintaining character consistency in video generation.
ByteDance releases Lance, an open-source multimodal model with 3B active parameters. Supports image/video generation and editing in a single framework. Trained from scratch on 128 A100-GPU budget.
Hyperframes is a framework enabling AI agents to generate video content through HTML. Tool designed to automate video creation in agent workflows.
ViMax is an agentic video generation system integrating director, screenwriter, producer, and video generator roles. The GitHub project presents a multi-agent architecture for end-to-end video creation orchestration.
Focused Forcing optimizes KV caches in autoregressive video diffusion generation by selecting relevant historical frames per-frame and per-head. The method combines attention scores with diversity scores, achieving 1.48× end-to-end acceleration without training while improving visual quality and text alignment.
ANVIL is a multimodal generative system automating production of analogy-based instructional animations for computer science. Given a concept definition, it generates textual analogies, compiles them into structured visual screenplays, and produces executable manim code. Evaluation includes teacher studies and user adoption assessment.