Wan 3.0 debuts at #1 on the Artificial Analysis Video Editing Leaderboard, and is a close #2 in Text to Video with Audio Wan 3.0 is Alibaba's new all-in-one video generation and editing model, positioned as a single system for turning multimodal creative direction into video. It generates up to 30 seconds at 1080p with native audio and accepts text, images, video, audio, documents, and web pages as creative references. The same model supports Text to Video, Image to Video, reference-based generation, and instruction-led editing, including changes to visuals, plot, dialogue, and sound. In the Artificial Analysis Video Arena, Wan 3.0 ranks #1 in Video Editing with Audio, #2 in Text to Video with Audio, and #5 in Image to Video with Audio. Wan 3.0 marks a large generational improvement: against the most recent Wan 2.7 version on each leaderboard, it rises from #5 to #1 in Video Editing with Audio, #6 to #2 in Text to Video with Audio, and #12 to #5 in Image to Video with Audio. Wan 3.0 is available now in public preview through Alibaba Cloud Model Studio. Pricing starts at $0.05 per second for 480p, increasing to $0.10 for 720p and $0.20 for 1080p. Congratulations to @Alibaba_Wan and @alibaba_cloud on the release! See below for comparisons between Wan 3.0 and other leading models in the Artificial Analysis Video Arena 🧵
Text to Video (With Audio) Prompt: A cat staring at its own reflection in a toaster, paw tapping the chrome surface. The distorted cat reflection taps back. Audio: Paw taps, confused meow.
Image to Video (With Audio) Prompt: Dragon spreads its wings, lifts off, and flies across the sky with powerful wingbeats. Tail trails behind as it soars into the distance. Heavy wings flapping, deep roar, and rushing wind.
Video Editing (With Audio) Prompt: Restyle the cat's fisheye POV into a hand-drawn cartoon look while keeping the first-person motion.
One Alibaba model now covers text to video, image to video, and edit-by-instruction with native audio up to 30 seconds at 1080p, priced $0.05 to $0.20 per second, so a video pipeline no longer needs a separate system per stage.
postMeta has released Muse Spark 1.3, their fourth Muse Spark model release in five…
postMuse Spark 1.3 scores 64 on coding tasks at a fifth of Opus 5's cost per task
postInworld's newly released Realtime TTS-2 is the new #1 on the Artificial…Checking sign-in…
Loading comments…