Skip to content
Konte
Before you start
SH-02 · 6s · KONTE · MiniMax H3 on a local GPU

REQUIREMENTSBefore you start

Two ways to set up

Shoot today through an API, or on your own GPU without counting videos. Either way, the same shot list.

CLOUD API

Run on APIs

No hardware, from the day you sign up. You pay each provider's API fees for what you use.

  • fal.ai API key
  • Anthropic or OpenAI API key

LOCAL GPU

Run locally

Your own GPU and a subscription. No API fees, however many clips you render.

  • GPU machine + ComfyUI
  • Claude or ChatGPT subscription

01—TWO SETSTwo ways to run

API or local, two ways to run it

Where video renders, and the AI you chat with. Each can be set up two ways.

VIDEO—Rendering video

Where video renders
Run on APIs

fal.ai's cloud API

Run locally

ComfyUI running on your own GPU

Engines
Run on APIs

MiniMax H3 via the fal.ai API (no quantization, 5–15 s per shot)

MiniMax Hailuo via the fal.ai API (6 / 10 s per shot)

Run locally

MiniMax H3 on a local GPU (int8 build that fits in 12GB)

Seedance on a local GPU (slot only; workflow in preparation)

What you need
Run on APIs

A fal.ai account and API key

Just paste it into Organization settings → Generative AI (組織設定 → 生成 AI). Stored encrypted; only the last 4 characters are shown

Run locally

A machine with an NVIDIA GPU, and ComfyUI

The full MiniMax H3 model set (about 93GB of storage) and custom nodes

What you pay
Run on APIs

fal.ai usage fees, billed straight to the key you registered

Pay per use, by resolution and length. MiniMax H3 defaults to 768P; at 2K it costs about twice as much per second (as of 2026-09-14)

Run locally

No API fees

You pay only for the hardware and the electricity

Waiting
Run on APIs

No waiting in line for a GPU. Shots are generated in parallel

How long it takes depends on how busy fal.ai is

Run locally

About 139 s for one 4-second shot (measured)

One at a time per GPU, in order. You see an estimated time before you start

Not possible yet
Run on APIs

Reference mode with 2 or more keyframes, and video → video (v2v)

Generating cast profile sheets and set sheets (both local GPU only for now)

Run locally

Long shots depend on VRAM. With 12GB, 60 s won't render even at 480P

A length that can't render is stopped before you start

AI—Chat AI

Chat AI
Run on APIs

An Anthropic (Claude) or OpenAI (GPT) API key

Pay per use. Register it under Settings → API settings → AI integration (設定 → API 設定 → AI 連携)

Run locally

The allowance of a Claude Pro / Max or ChatGPT Plus / Pro subscription

Konte's local worker (Node 18 or later) calls the official CLI (claude / codex) on your machine

Other options
Run on APIs

An OpenAI-compatible in-house gateway (LiteLLM / vLLM etc.)

Run locally

A local LLM (Ollama / LM Studio). Can be registered without a key

02—COMMONEither way

Four things that hold either way

How the engine is chosen, and the AI that builds the shot list.

  • The AI that builds the shot list

    Breaking the conversation into shots, writing up the prompt right before rendering, and reading your cast and sets are handled by Claude, running on Konte's server.

  • One engine per organization

    The organization's owner picks the default render engine. If shots in the same video were rendered on different engines depending on who pressed the button, the picture wouldn't match.

  • Video and chat are chosen separately

    You can mix them: video on your own GPU and chat on an API key, for example.

  • Work the shot list from your own AI (optional)

    Connect Claude Code or Codex to Konte over MCP and you can ask it for everything from writing the Wiki and making cast, sets, props and shots to rendering (52 tools). It renders only when you ask it to, and deletions that can't be undone happen only on screen.

03—LOCAL SPECThe machine we measured

Your own GPU, as we measured it

Every number here was measured on this one machine. A different GPU gives different numbers.

GPU
NVIDIA GeForce RTX 4070 (12GB VRAM)
CUDA
13.0
System memory
24GB allocated (peaks at about 22GB generating 30 s at 480P)
Speed-up
SageAttention 2.2.0 (required by the fast-generation workflow)
ComfyUI nodes
ComfyUI-MiniMaxH3-Easy / ComfyUI-KJNodes / ComfyUI-VideoHelperSuite. For sheets, also ComfyUI-MiniMax-H3-Image-Studio
Models
MiniMax H3 itself (int8), text encoder (Qwen3-VL 32B), VAE, Turbo 8step LoRA, RealESRGAN. About 93GB of storage

ONE CUT — one 4-second shot

  1. Generate at 480P (Turbo 8step)84.2 s
  2. ESRGAN x2 → up to 1080p55.0 s

1920×1080, with sound139.2 s

Generation time, by length
Resolution4 s15 s30 s60 s
480P1 min 24 s6 min 11 s31 min 57 sOut of VRAM
360P—2 min 47 s6 min 25 s52 min 47 s

Generation only, without upscaling to 1080p. Konte uses these measurements to show an estimate before you start. — means not measured.

GET STARTEDStart here

Your film studio, starting today

Start by creating a single cast member. We're happy to talk it through.