Kling 3.0
Kuaishou's audio-visual model — Kling 3.0 shoots 3–15 second takes where characters actually speak: dialogue in five languages, authentic dialects, multi-character scenes, up to true 4K. Generate below.
- Failed generations auto-refunded
- Credit packs never expire
- Every model's price is public
- Cancel anytime in one click
- 3–15 s
- Single take length
- Up to 4K
- Three resolution tiers
- 5
- Spoken dialogue languages
- From 22 cr/s
- Kling 3.0 720P rate
Six takes, six full prompts
A moonlit 15-second chase, a tearful dinosaur reunion, Spanish street directions, a Cantonese office rant, a four-person family scene and a 4K lettering shot — every clip below comes from Kling's official Video 3.0 guide and ships the complete prompt that made it. Hit Remix to load any prompt straight into the generator above and make it yours.
- Kling 3.0
- Image to video
- 15s
- 1080P
Opening with an ultra-wide-angle medium-long shot tracking horizontally, the stabilizer moves low to the ground, with a highly contrasting romantic cinematic tone of cold blue night and silvery white starry sky, exuding a strong poetic realism and classical epic temperament. The protagonist is a young woman in a dark green long dress, running with all her might on the garden lawn illuminated by moonlight; her skirt billows in the wind forming surging dynamic curves, she clutches a small white flower in her right hand and lifts the hem of her dress with her left, breathing rapidly yet with a firm gaze. At the 4th second, the camera accelerates forward with her, and multiple men and women in old-era ball gowns break into the frame one after another from the left and right sides in the background, running alongside her—some try to approach, some turn back to shout, yet none truly touch her, implying pursuit and escape. At the 8th second, the camera gradually zooms in to a medium shot, pans to track forward in front of the protagonist and lifts slightly; she glances back briefly at a young male character behind her, their gazes meet for a split second, emotions erupt mid-run, and the woman and man join hands to run together. At the 12th second, the music and movement reach a climax; the camera moves forward close to her side face and fluttering hair, she releases the white flower and tosses it into the air, the flower drifting down in slow motion as the crowd behind brushes past it. In the final 3 seconds, the camera keeps moving forward, the woman and man break through the crowd and dash toward the starry sky at the end of the garden, their figures gradually taking over the center of the frame. The overall atmosphere is fiery, romantic and resolute, a burst of narrative about fate, choice and freedom.
- Kling 3.0
- Image to video
- 15s
- 1080P
This is a 15-second cinematic long take, a single unbroken shot with no edited transitions. The scene is set inside a tower of plaster statues dappled with light and shadow, surrounded by towering white plaster sculptures, evoking an air of mystery and oppression. The shot opens with the protagonist skidding to a halt at the center of the scene after a frantic run, chest heaving, expression dazed and helpless, fear glinting in their eyes. The camera orbits the protagonist in a smooth 360-degree pan. As the camera rotates, the protagonist glances anxiously around and shouts: "Alex! Alex where are you! Are you here?" A cute dinosaur cry then echoes in the background, and the camera pushes in over the protagonist's shoulder to their back— a small to medium-sized, adorable baby dinosaur steps out from behind a plaster pillar, letting out a sweet chirp. Startled by the sound, the protagonist whips around; catching sight of the dinosaur, they burst into tears instantly and rush forward without hesitation to clasp it tightly in their arms. The dinosaur nestles obediently against them. Sobbing, the protagonist strokes the dinosaur gently and trembles: "I found you! Thank God, I'm so scared!" The overall lighting and shadow boast a cinematic texture, with the emotion shifting from despair to an overwhelmingly touching reunion.
- Kling 3.0
- Image to video
- 10s
- 1080P
Sunlight fills the old streets of Madrid. In front of a street-side bakery, a Chinese female tourist and a male tourist wearing a gray hoodie walk toward the shop clerk, both wearing polite smiles. Female tourist (speaking slightly slowly, with an awkward accent, in Spanish): Disculpe, ¿dónde está la plaza mayor? A white-haired Spanish shop clerk (turning slightly and pointing forward, with a light and cheerful tone, in Spanish): Por allí, a dos calles. Muy cerca. The female tourist nods to express her thanks. The male tourist also nods in agreement and says (in Spanish): Muchas gracias. The shop clerk smiles and nods in response. The two tourists then turn and walk in the indicated direction.
- Kling 3.0
- Image to video
- 10s
- 1080P
In a high-rise office building, the man leaned back, wearing a tired, disdainful expression, and said in Cantonese: 「其实……我真系唔系好 buy 你呢个 logic 啰。成个 proposal 根本 align 唔到我哋个 core value。你个 flow 咁乱,点样去 convince 个 client 呀?不如你返去 re-think 下个 angle,听朝早我要见到个 final version。」
- Kling 3.0
- Image to video
- 10s
- 1080P
Home setting with a faint hum of the living room air conditioner in the background for a realistic daily vibe. Mom (softly, in a surprised tone): Wow, I didn't expect this plot at all. Dad (in a low voice, agreeing, in a calm tone): Yeah, it's totally unexpected. Never thought that would happen. Boy (in an excited tone): It's the best twist ever! Girl (nodding along, in an enthusiastic tone): I can't believe they did that!
- Kling 3.0
- Image to video
- 10s
- 4K
The camera remains fixed on the word "KLING" emblazoned on the baseball bat as the player swings and hits the ball.
One still, one finished commercial
Image-to-video with native-level text: a single product shot of a perfume bottle goes in as the start frame, and the prompt directs three camera moves plus three voiceover lines. Kling 3.0 keeps the gold lettering on the bottle crisp through every move and delivers the ad with the narration already in the soundtrack.

The prompt
By the window of a Parisian apartment, with soft French piano BGM in the background, the gilded afternoon sunlight filters through the shutters onto the perfume bottle, casting dappled light and shadow. The camera pans slowly in from the scattered rose petals, shifting focus to the faceted cut of the Kling perfume bottle. Voiceover (lazy French female voice, British accent, slow pace): Bathe in the golden hour. The camera orbits the perfume bottle in slow motion, capturing the play of light and shadow on the golden lettering and bottle body. Voiceover: Kling, a whisper of Parisian elegance. The camera pulls back and freezes on the complete scene—the Kling perfume bottle standing on a velvet pedestal, with Parisian buildings faintly visible outside the window. Voiceover: Wrap yourself in luxury with every breath.
What Is Kling 3.0?
Kling 3.0 is the video generation model behind Kling AI, built by Kuaishou. Its defining trait is that sound is not an afterthought: dialogue, ambience and music generate together with the picture, characters lip-sync their lines, and in multi-character scenes each person speaks exactly the line you assign them. Spoken dialogue works in five languages — Chinese, English, Japanese, Korean and Spanish — and the model renders authentic dialects and accents, from Cantonese to Indian English, when you tag them in the prompt.
The rest of the sheet is just as serious. Length is a continuous dial from 3 to 15 seconds at 720p, 1080p or true 4K. Image-to-video animates a still and preserves native-level text — signs, captions and lettering stay crisp instead of smearing, which is why the perfume and baseball samples above hold up frame by frame. Element references compose a video from 2–4 photos of one subject, locking its identity across camera moves, and image-to-video also accepts a start and end frame pair for controlled transitions.
On Molyin, Kling 3.0 runs inside the same AI video generator as every other model we host — one prompt box, one credit balance. Credits work the same everywhere; the pricing page has the numbers.
Kling 3.0 vs Seedance 2.0
Both are live on Molyin, and they're genuinely complementary — one prompt box, one credit balance, so the honest answer to "which one?" is: run your prompt on both and keep the winner.
| Spec | Kling 3.0 | Seedance 2.0 |
|---|---|---|
| Duration | 3–15s, continuous | 4–15s, continuous |
| Resolution | 720p · 1080p · 4K | 480p · 720p · 1080p · 4K |
| Aspect ratios | 3 — 16:9, 9:16, 1:1 | 6 |
| Audio | Native audio, toggleable — dialogue in 5 languages plus dialects and accents | Native audio generation |
| References | Element references — 2–4 images of one subject | Up to 9 images + 3 videos + 3 audio tracks |
| Price | 22 / 29 / 71 credits/s (720P / 1080P / 4K) | 11 / 28 / 65 / 135 credits/s by tier |
| Reaches its best in | Spoken performances, multilingual dialogue, on-screen text, single-subject consistency | 480p quick drafts, big multi-modal reference stacks, seed-controlled iteration |
Rates shown are current execution prices on Molyin.
Kling 3.0 vs models live on Molyin
| Model | Positioning | Release Date | On Molyin |
|---|---|---|---|
| Kling 3.0 | Native audio in 5 languages + element references, up to 4K | Feb 4, 2026 | Live — generate above |
| Seedance 2.5 | ByteDance flagship — 30s takes, directed second by second | Aug 7, 2026 (API) | Live — generate above |
| FLUX 3 | Black Forest Labs — storyboard keyframes + take continuation | Aug 7, 2026 (video API) | Live — generate above |
| Wan 3.0 | Alibaba's All-in-One model, up to 30s takes | Aug 6, 2026 | Live — generate above |
| MiniMax H3 | Omni-modal flagship, native 2K + stereo | Jul 31, 2026 · open weights announced | Live — generate above |
| HappyHorse | Expressive generation + instruction video editing | Jun 2026 (1.1) | Live — generate above |
| Seedance 2.0 family | Standard / Fast / Mini workhorses, up to 4K | Apr 2026 (API) | Live — generate above |
| Veo 3.1 | Google's cinematic model with audio | Oct 2025 | Live — generate above |
One subscription, one credit balance — every model available above is included.
Kling 3.0 full specifications
| Parameter | Value | Notes |
|---|---|---|
| Model | Kling 3.0 — Kuaishou's unified audio-visual model | The model behind Kling AI, generating picture and sound in one pass. |
| Generation modes | Text / Image / Elements to video | Three modes: prompt from scratch, animate a still, or compose from element reference photos. |
| Resolution tiers | 720P / 1080P / 4K | Draft at 720P, finish at 1080P, master at true 4K — the baseball sample above is a 4K output. |
| Clip length | 3–15 s | A continuous slider, not fixed presets — both 15-second samples above are single unbroken takes. |
| Aspect ratios | 16:9 · 9:16 · 1:1 | Widescreen, vertical and square; in image-to-video the output follows your start frame's ratio. |
| Audio | Native audio — optional toggle | Sound generates together with the picture and dialogue lip-syncs; switch it off when you only need the visuals. |
| Dialogue languages | Chinese · English · Japanese · Korean · Spanish | Plus authentic dialects and accents — Cantonese, Sichuanese, British or Indian English — tagged right in the prompt. |
| Element references | 2–4 images of one subject | Multi-angle photos of a character, product or scene lock its identity across camera moves. |
| Start & end frames | Supported in image-to-video | Pin both endpoints of the clip and the model animates the transition between them. |
| Seed control | Not supported | No fixed-seed reproduction — iterate by refining the prompt or remixing a result you like. |
| Credits per second | 22–71 cr/s | From 720P to 4K — full per-tier pricing in the table below; failed generations are refunded. |
Everything listed here is already supported on the site, and stays in sync with the model's latest official capabilities.
Kling 3.0 pricing on Molyin
Billed per second by resolution tier, from the same credit balance as everything else on Molyin. The numbers below are computed straight from our billing engine, and failed generations are refunded automatically.
| Model | Resolution | Credits/sec | 5s Video | 5s Video + 3s Reference |
|---|---|---|---|---|
| 720p | 22 | 110 credits | Reference images only | |
| 1080p | 29 | 145 credits | Reference images only | |
| 4k | 71 | 355 credits | Reference images only |
How to generate with Kling 3.0
Three steps from idea to a finished take
Pick a model and a mode
Choose text-to-video, image-to-video, or reference mode with 2–4 element photos.
Write the prompt
One concrete subject plus one clear motion, then put dialogue in quotes and name who says it — exactly like the samples on this page. Tag a language or dialect for authentic delivery.
Generate and export
Pick 720P, 1080P or 4K and any length from 3 to 15 seconds, preview in place, then download the original file.
The element reference formula
Kling 3.0's signature move: upload 2–4 photos of one subject and describe the scene around it — the model locks the subject's identity while everything else obeys the prompt. Straight from our full prompt guide.
[Elements: Image 1, Image 2…] + [Scene description] + keep [element] consistent
Reference Elements
2–4 images, cited in the prompt as "Image N"; upload order is the numbering order.
Scene Description
What happens in the new scene: action, environment, camera.
Consistency
One line of "stay consistent with the reference images" reinforces the appearance lock.
Kling 3.0 FAQ
What creators ask about Kuaishou's video model — answered from our own integration.
Start Creating with Molyin Today
Sign up free and turn your first idea into a cinematic video in minutes.