AI Video Is Getting Good Enough To Make The Model Invisible
Generative AI video is getting so good these days that making impressive videos on their own is no longer enough. People now want AI video systems that can act as a producer, director, cinematographer, script writer and shot creator. To this point, Kling AI is opening early access to Kling 4.0, its newest video model, with native 30 second generation, support for up to 10 keyframes, multimodal reference material, stereo audio and video extensions reaching two minutes. While that might sound like a technical list of functions, it shows that the AI video race is moving from simply generating clips toward controlling production.
For creative teams, agencies and enterprises, this sort of control is important. A beautiful few-second clip has little value when the person in the next shot looks different, the product mutates, the logo becomes unreadable or the director cannot reproduce an action. Businesses need video systems that take direction, survive revisions and produce something usable without burning tokens and time through generation after generation.
With so many models now capable of producing high quality results, that leads to an important question for vendors in the market. What happens when AI video systems become good enough and interchangeable enough that users start choosing them the same way developers increasingly choose large language models, selecting a different model for each task rather than committing to one vendor?
Kling’s Latest Preview Puts More Control Into The Generation Process
This week, Kling shared a preview of its upcoming 4.0 launch that provides a good view of where this market is going. According to information provided for the release, users can feed the model reference images, video clips and subjects that help define characters, scenes and other elements. Kling says generated videos can be extended up to two minutes.
The model adds 10 bit HDR output at 4K and 1080p, supports a 21:9 format, generates two channel stereo sound and improves lip synchronization. Kling says it can retain multilingual text, emojis and logos through camera movement and scene transitions. A public release is planned for October.
Text prompts leave a large amount of interpretation to the model. That can be entertaining when someone is experimenting. It gets expensive when a client wants a particular sequence and expects revisions. Keyframes provide fixed visual moments that the model must connect. A creator can establish what should appear at several points, then generate the movement between them. The model still creates the footage, but it has fewer opportunities to wander.
That is especially important in commercial work, where creative freedom is rarely unlimited. Products need to look like the products being sold and characters need continuity. The opening and closing moments of an advertisement might already be decided and models need to respect that. Similarly, brand teams have rules for logos, typography and color which models need to also respect.
The Market Is Rapidly Iterating To Solve Similar Problems
Kling has plenty of company in the AI video generation market. ByteDance released Seedance 2.5 on July 31 with 30 second audio and video generation, multiple rounds of extension, richer reference inputs and editing controls. A single generation can accept up to 30 images, 10 videos and 10 audio clips, according to ByteDance. The model can edit content with timestamp level instructions and offers controls aimed at camera perspective, green screen work and complex production.
MiniMax is taking a related route with H3. Released July 31 and later made available with open weights, H3 accepts text, image, video and audio context and generates video with stereo sound at up to 2K resolution and 15 seconds in length. MiniMax is targeting commercial uses that include advertising, ecommerce, product design and branding, with attention to text rendering and motion transfer.
Runway’s Gen 4.5 takes detailed instructions for camera choreography, scene composition and event timing. Its current text to video and image to video modes generate sequences from two to 10 seconds. Runway charges 12 credits per second for Gen 4.5 generation.
Google’s Veo 3.1 focuses heavily on reference material and continuity. Creators can feed it images defining a scene, object or character, give it style references, control camera framing and movement, extend scenes and guide sequences using first and last frames.
While the implementations differ, creators and buyers are starting to evaluate these models with the same practical questions.
How long can the model maintain coherence? How closely does it follow direction? Can the creator anchor characters and products? What happens when the sequence gets complicated? How many attempts does it take to get usable footage?
One Project Could Use Several Video Models
A year or two ago, a shiny demo was impressive, but now the demo is merely just proof of a model’s minimal suitability.
Like all corners of the AI model markets, rapidly improving models create both opportunities and problems for vendors in the market. Improvements can make switching easier.
Suppose one model handles human motion extremely well but has trouble with package lettering. Another excels at product shots but costs too much when an agency needs 300 social variations. A third handles complex camera moves better. Why choose just one model when the best results might come from chaining several together?
That approach is already taking place in the software layer above the models. Adobe added Kling 3.0 and Kling 3.0 Omni to Firefly in April, putting an outside model inside Adobe’s creative environment rather than forcing creators to leave the workflow and work directly in Kling.
Magnific, formerly Freepik, is pushing in a similar direction. The company’s products increasingly center on keeping the creative workflow in one place rather than making users bounce among individual model providers making use of a range of video and image generation models.
Higgsfield’s Canvas product goes a step further by letting users chain multiple models into a single visual workflow. Its video workspace offers models from Kling, ByteDance, Google, OpenAI, Alibaba and others inside the same environment. API documentation allows users to programmatically use one model for stills, another for motion, a cheaper engine for drafts and a stronger model for the final output.
A creative team may not need to decide whether it is a Kling shop, a Veo shop or a Seedance shop. It may simply pick the model that works best for each step, then assemble those outputs inside a platform such as Adobe Firefly, Magnific or Higgsfield.
For the model vendors, that is both an opportunity and a threat. Distribution through these platforms can put their models in front of more users. It can also turn the underlying model into one option among many, selected for a particular shot or task rather than chosen as the foundation for an entire production workflow.
When users go directly to a model provider, the provider owns the experience and the customer relationship, but when several models appear inside the same creative environment, the model starts looking more like a selectable engine.
That puts pressure on vendors to develop capabilities that go beyond output quality and focus on cost and speed of generation, reliability and consistency of outputs, and precise motion and direction control.
Generating video is not the same as generating images or text since the little details matter. This means that the real cost of a usable clip depends on how many times somebody has to generate it. Imagine Model A costs twice as much per second as Model B but usually gets the required shot after three attempts. Model B needs 15. The cheap model is no longer cheap. This is not even factoring in the time to build the prompts and review the outputs.
To help with this iteration time, Kling is introducing Kling 4.0 Flash as a faster, lower cost option intended for frequent creation. We can expect rapid growth of smaller, flash-style models from others as they seek to reduce iteration costs and time.
As markets mature for AI, they are starting to look less like a contest to produce the sharpest demo or most clever output, and more like a collection of engines and coordination systems for production. We’re moving into a more sober phase of AI where flashy matters less and “does it work for me now” matters more.