REQUIREMENTSBefore you start
Two ways to set up
Shoot today through an API, or on your own GPU without counting videos. Either way, the same shot list.
CLOUD API
Run on APIs
No hardware, from the day you sign up. You pay each provider's API fees for what you use.
- fal.ai API key
- Anthropic or OpenAI API key
LOCAL GPU
Run locally
Your own GPU and a subscription. No API fees, however many clips you render.
- GPU machine + ComfyUI
- Claude or ChatGPT subscription
01—TWO SETSTwo ways to run
API or local, two ways to run it
Where video renders, and the AI you chat with. Each can be set up two ways.
CLOUD API
Run on APIs
LOCAL GPU
Run locally
VIDEO—Rendering video
- Where video renders
- Run on APIs
fal.ai's cloud API
- Run locally
ComfyUI running on your own GPU
- Engines
- Run on APIs
MiniMax H3 via the fal.ai API (no quantization, 5–15 s per shot)
MiniMax Hailuo via the fal.ai API (6 / 10 s per shot)
- Run locally
MiniMax H3 on a local GPU (int8 build that fits in 12GB)
Seedance on a local GPU (slot only; workflow in preparation)
- What you need
- Run on APIs
A fal.ai account and API key
Just paste it into Organization settings → Generative AI (組織設定 → 生成 AI). Stored encrypted; only the last 4 characters are shown
- Run locally
A machine with an NVIDIA GPU, and ComfyUI
The full MiniMax H3 model set (about 93GB of storage) and custom nodes
- What you pay
- Run on APIs
fal.ai usage fees, billed straight to the key you registered
Pay per use, by resolution and length. MiniMax H3 defaults to 768P; at 2K it costs about twice as much per second (as of 2026-09-14)
- Run locally
No API fees
You pay only for the hardware and the electricity
- Waiting
- Run on APIs
No waiting in line for a GPU. Shots are generated in parallel
How long it takes depends on how busy fal.ai is
- Run locally
About 139 s for one 4-second shot (measured)
One at a time per GPU, in order. You see an estimated time before you start
- Not possible yet
- Run on APIs
Reference mode with 2 or more keyframes, and video → video (v2v)
Generating cast profile sheets and set sheets (both local GPU only for now)
- Run locally
Long shots depend on VRAM. With 12GB, 60 s won't render even at 480P
A length that can't render is stopped before you start
AI—Chat AI
- Chat AI
- Run on APIs
An Anthropic (Claude) or OpenAI (GPT) API key
Pay per use. Register it under Settings → API settings → AI integration (設定 → API 設定 → AI 連携)
- Run locally
The allowance of a Claude Pro / Max or ChatGPT Plus / Pro subscription
Konte's local worker (Node 18 or later) calls the official CLI (claude / codex) on your machine
- Other options
- Run on APIs
An OpenAI-compatible in-house gateway (LiteLLM / vLLM etc.)
- Run locally
A local LLM (Ollama / LM Studio). Can be registered without a key
02—COMMONEither way
Four things that hold either way
How the engine is chosen, and the AI that builds the shot list.
The AI that builds the shot list
Breaking the conversation into shots, writing up the prompt right before rendering, and reading your cast and sets are handled by Claude, running on Konte's server.
One engine per organization
The organization's owner picks the default render engine. If shots in the same video were rendered on different engines depending on who pressed the button, the picture wouldn't match.
Video and chat are chosen separately
You can mix them: video on your own GPU and chat on an API key, for example.
Work the shot list from your own AI (optional)
Connect Claude Code or Codex to Konte over MCP and you can ask it for everything from writing the Wiki and making cast, sets, props and shots to rendering (52 tools). It renders only when you ask it to, and deletions that can't be undone happen only on screen.
03—LOCAL SPECThe machine we measured
Your own GPU, as we measured it
Every number here was measured on this one machine. A different GPU gives different numbers.
- GPU
- NVIDIA GeForce RTX 4070 (12GB VRAM)
- CUDA
- 13.0
- System memory
- 24GB allocated (peaks at about 22GB generating 30 s at 480P)
- Speed-up
- SageAttention 2.2.0 (required by the fast-generation workflow)
- ComfyUI nodes
- ComfyUI-MiniMaxH3-Easy / ComfyUI-KJNodes / ComfyUI-VideoHelperSuite. For sheets, also ComfyUI-MiniMax-H3-Image-Studio
- Models
- MiniMax H3 itself (int8), text encoder (Qwen3-VL 32B), VAE, Turbo 8step LoRA, RealESRGAN. About 93GB of storage
ONE CUT — one 4-second shot
- Generate at 480P (Turbo 8step)84.2 s
- ESRGAN x2 → up to 1080p55.0 s
1920×1080, with sound139.2 s
| Resolution | 4 s | 15 s | 30 s | 60 s |
|---|---|---|---|---|
| 480P | 1 min 24 s | 6 min 11 s | 31 min 57 s | Out of VRAM |
| 360P | — | 2 min 47 s | 6 min 25 s | 52 min 47 s |
Generation only, without upscaling to 1080p. Konte uses these measurements to show an estimate before you start. — means not measured.
GET STARTEDStart here
Your film studio, starting today
Start by creating a single cast member. We're happy to talk it through.