Director's Handbook
/docs — THE COMPLETE GUIDE TO THE STUDIO AND AI GENERATION LIFECYCLE.
01. What Is Banana Split
Prompting is dead. Banana Split puts you in the director's chair — where your job is to have a vision, not memorize syntax.
Generative AI is incredibly powerful, but raw. To get high-fidelity results, you usually have to think like an engineer, configuring complex workflows and tuning model inputs by trial and error. Banana Split changes this. We've built a professional-grade monochrome studio that translates your creative intent into precise technical parameters. You dictate the scene, the camera behavior, the style, and the actors. The studio handles the pipeline execution, letting you focus on directing the action.
02. The Studio at a Glance
Before diving into the tools, orient yourself to the workspace layout. The interface is optimized to keep you in a continuous flow state where inputs, references, and renders live side by side.
2.1 The Split-Screen Workspace
The studio uses a resizable split-pane layout:
- Left Pane (Controls & Inputs): This area hosts the Director Chat by default, and dynamically displays the settings panels when you load Manual Tools.
- Right Pane (Universal Canvas): Dedicated entirely to rendering. Every image, video, upscale, and console log appears here inline as a live output tab.
A draggable vertical border divides the panes, allowing you to slide the ratio at any time. Shrink the controls to inspect high-resolution canvas results, or widen them when writing script sequences.
2.2 The Asset Library
The right edge of the screen contains the collapsible **Asset Library** sidebar. Widen this panel to explore your media archive:
- Uploads Tab: Houses your imported reference files, mood boards, and aesthetic profiles.
- Results Tab: Houses outputs generated in the studio, catalogued chronologically.
- Card Preview Mode: Filters out code and metadata files to display a visual media grid in native aspect ratios, displaying a pink play badge over video files. Row List mode displays files as text lists.
2.3 How Files Move Through the Studio
Everything in your Asset Library is reusable. Production follows a simple loop:
03. The Director
The Director is your production orchestrator. It is a reasoning agent that reads your creative directions, plans the steps required, and dispatches the execution tasks to the GPU.
3.1 How It Works
The Director processes instructions in a two-phase loop:
- Planning Phase: You write your creative direction in plain language. The Director reads the instructions, plans which models are needed, and drafts a collapsible task checklist.
- Execution Phase: You review the proposed production checklist. Clicking "Approve" dispatches the tasks sequentially. The Director locks the input panel, renders the prompt, and displays active loading states for each task.
🎬 You stay in control of the creative direction. The Director handles the technical translation.
3.2 Directing Effectively
To write instructions that the Director can interpret accurately, describe your scene by combining details about characters or aesthetics (using saved Identities), actions, environmental lighting, and camera behavior.
"A woman walking in the city."
"A cinematic shot of @elena walking through a rain-soaked Tokyo street at night, neon reflections on the pavement, slow tracking shot, moody and desaturated color grade."
💡 The more specific your creative intent, the less the Director has to guess — and the closer your first result will be to your vision.
3.3 Pinning Files & References
You link context to a scene direction by combining three elements inside the input panel:
- Asset Handles: Drag images or videos from the Asset Library sidebar directly into the chat area. They drop as handle text mentions (e.g. `@uploads/filename.png` or `@results/filename.mp4`), pinning the visuals as style or starting-frame references.
- Identity Handles: Type the direct name handle of a saved identity (e.g. `@elena` or `@retro-grade`) to apply trained face models, color configurations, or environments.
- Descriptive Direction: Explain the motion, story elements, and camera movements in natural prose surrounding your handles.
3.4 Batches & Chained Tasks
In typical AI pipelines, you must wait for an image, download it, and manually load it into a video engine. Banana Split removes this manual work by scheduling multi-task plans:
- Batches (Parallel Generation): You can generate multiple variations concurrently. Asking for:
Generate three variations of the opening scene, each using @marco, but vary the lighting: one golden hour, one overcast, one neon nighttriggers three tasks to run in parallel. - Chains (Sequential Pipelines): You can link tasks so the output of one step is piped automatically into the next. Directing:
Generate a portrait of @sara in the library, then take that output and animate a slow push-in camera move with a 4s durationcreates a sequential chain.
🎬 Batches save time. Chains save the entire manual pipeline. Together, they're how you produce at scale.
3.5 Your Results
Once task execution completes, the results are loaded as canvas tabs and stored in the library sidebar. You can download the assets using the Canvas overlay, drag them back into the chat input as context pins for subsequent scenes, or send them directly to a manual tool like the Upscaler.
3.6 Reference Media & Video Analysis
Instead of building visual pacing or camera behavior from scratch, you can feed reference media directly to the Director. The studio includes an automated **Visual Forensic & Video Analysis engine** that reads reference videos and images to extract stylistic parameters.
- Video Analysis: You can send any public TikTok, Instagram, YouTube, or direct video URL to the Director Chat. The engine transcribes any spoken audio, dissects the cinematography (lighting, camera angles, movement vectors), and compiles a complete **Reusable Production Blueprint**. You can then instruct the Director to "generate a scene matching that blueprint" to duplicate its stylistic DNA.
- Image Analysis: Dragging an uploaded image or generated result into the Director Chat and referencing it triggers a visual breakdown. The Director automatically runs a stylistic forensic analysis, cacheing the descriptive visual features locally. This cache is then read by downstream generation pipelines, preserving character identity, composition, and clothing details without sending the original image files to the generation servers.
🎬 Reference media analysis bridges the gap between existing content and AI production, giving you surgical control over pacing and cinematography.
04. Manual Tools
The Director is the recommended way to work in Banana Split — but sometimes you know exactly what you want and prefer to configure it yourself. Manual Tools give you direct access to each generation capability with full parameter control. The output workflow is identical: everything lands in your Asset Library and Canvas.
4.1 Image Generation
Generates still images from a text description, with full control over format and quality.
Use this when you have a clear image in mind and want to configure every output parameter yourself — aspect ratio, resolution, color palette, and batch size — without going through the planning phase.
4.2 Video Generation
Generates video from a text description or a starting image frame.
Use this when you want to create a moving scene and have a specific duration, motion feel, or starting image in mind. Supports both text-to-video and image-to-video depending on the model selected.
4.3 Motion Control
Transfers movement from a reference video onto a still character image.
Use this when you want a specific character — especially one defined in your Identities — to perform a movement from a source video, without filming anything new.
⚠️ For best results, the target character image should be a clean, well-lit, front-facing shot. Cluttered or partially cropped images reduce transfer accuracy.
4.4 Upscaling
Increases the resolution of an existing image, rebuilding fine detail in the process.
Use this as a finishing step — after you've confirmed a generated image is the one you want, run it through Upscaling to bring it to delivery quality.
4.5 Voice & Audio
Generates natural voiceovers or clones an existing voice for narration, dialogue, or dubbing.
05. Identities
An Identity is a reusable creative reference you define once and invoke across any generation. The Identity Studio is a general-purpose engine capable of saving and training identities for characters, products, aesthetics, compositions, backgrounds, clothes, and literally any possible element you want to train and keep consistent across your production.
5.1 What You Can Identify
Identities support a wide range of creative training inputs, ensuring visual consistency across every visual dimension:
- Characters: A saved facial profile, facial structure, actor appearance, or specific person.
- Products: A watch, shoe, signature prop, or object that must keep its precise geometry.
- Aesthetics: A visual style, artistic mood board, or color grading profile to bind tones and lighting.
- Compositions: Framing constraints (such as flatlays, symmetry, low-angles, or perspective).
- Backgrounds: A persistent set location, landscape, backdrop, or interior environment.
- Clothes / Wardrobe: A specific outfit, uniform, costume, or clothing style to carry across scenes.
- Literally Any Element: Any consistent visual concept, brand element, or texture you wish to train.
5.2 Creating an Identity
To create an Identity, open the **Identity Studio** from the left panel:
- Name the Identity: Choose a clean, unique handle (e.g. `@elena`, `@studio-red`, `@grunge-style`). Note that identities have no prefix (like `@identity/` or `@identities/`); you reference them in your directions simply using `@name`.
- Provide a Description: Detail the visual features, build, geometry, or styling constraints you are saving and training.
- Upload References: Upload 5 to 10 high-resolution reference images showing the subject or aesthetic from multiple angles and under uniform lighting.
💡 The quality of your references matters more than the quantity. Ten consistent, well-lit, clearly composed images of the same subject will outperform fifty mixed or cluttered ones. The Identity learns from what you show it — so show it the best version of what you want.
5.3 Using Identities in Production
Casting identities is handled by typing their direct handle names anywhere inside your scene direction or manual prompt (without any `@identities/` prefix):
The orchestration engine automatically binds the trained facial features, product geometries, or aesthetic parameters together, eliminating long descriptive prompts and reducing your token costs.
5.4 Technical Architecture & Privacy
Identities are constructed exclusively to help the Director Agent formulate highly coherent prompts and scenarios.
When you upload reference images (Visual Anchors) to an Identity, the studio performs a **local, private forensic analysis** using Google Gemini to extract a highly detailed aesthetic and stylistic text description. This text description (the "compiled aesthetic context") is saved as the reference point for the Director Agent.
🔒 Privacy Note: The source images uploaded to your Identities are processed strictly for local context compilation. The raw image files are never sent or transmitted to generative text-to-image or text-to-video AI models.
06. Models
Banana Split provides access to a curated set of best-in-class generation models. You do not need to manage their technical hosting — choose the model that excels at your target task.
Image Models
Nano Banana 2 / Edit
Best for: High-speed photo generations and subject adjustments.
Supported modes: Text-to-Image / Image-to-Image / Local Edit
Default settings: Resolution: 1K (supports up to 4K) | Aspect ratio: auto | Batch: 1
Z-Image Turbo / i2i
Best for: Sub-second prototyping and storyboard testing.
Supported modes: Text-to-Image / Image-to-Image
Default settings: Resolution: 1K | Aspect ratio: auto | Batch: 1
Seedream 5.0 Lite / Edit
Best for: Rendering typography, text labels, and graphic designs.
Supported modes: Text-to-Image / Image-to-Image
Default settings: Resolution: 1K (supports up to 4K) | Aspect ratio: auto | Batch: 1
Stable Diffusion 3.5 Large Turbo / SDXL
Best for: Cost-effective generation with custom dimensions.
Supported modes: Text-to-Image / Image-to-Image
Default settings: Resolution: 1K (supports up to 2K) | Aspect ratio: auto | Batch: 1
FLUX 1.1 Pro Ultra
Best for: Unprocessed realistic details and cinematic compositions.
Supported modes: Text-to-Image
Default settings: Resolution: 1K (supports up to 2K) | Aspect ratio: auto | Batch: 1
Recraft V4 / V4 Pro
Best for: Logo design and vectors using strict color palette boundaries.
Supported modes: Text-to-Image
Default settings: Resolution: 1K (supports up to 2K) | Aspect ratio: auto | Batch: 1
Higgsfield Soul
Best for: Tonal grading and character consistency filters.
Supported modes: Image-to-Image
Default settings: Resolution: 1K (supports up to 2K) | Aspect ratio: auto | Batch: 1
Video & Motion Models
Kling 3.0 Pro / 3.0 4K
Best for: Detailed 4K cinematic video clips and camera panning.
Supported modes: Text-to-Video / Image-to-Video
Default settings: Resolution: 1080p (supports up to 4K) | Duration: 5s (supports up to 10s) | Aspect: 16:9
Kling 2.6 / 3.0 Motion Control
Best for: Animating character images using reference video motions.
Supported modes: Video-to-Video / Image-to-Video
Default settings: Resolution: 720p | Duration: 5s (supports up to 10s) | Aspect: 16:9
Sora 2 Pro / Sora 2 Actor
Best for: Complex physical interactions with synchronized audio.
Supported modes: Text-to-Video / Image-to-Video / Actor Training
Default settings: Resolution: 1080p | Duration: 4s (supports up to 12s) | Aspect: 16:9
Seedance 2.0
Best for: Consistent character movement and stable visual rendering.
Supported modes: Text-to-Video / Image-to-Video
Default settings: Resolution: 1080p | Duration: 5s (supports up to 12s) | Aspect: 16:9
Audio & Voice Models
ElevenLabs TTS / SVC
Best for: Narration, voice cloning, and audio conversion.
Supported modes: Text-to-Speech / Voice Conversion
Default settings: Sample rate: 44.1kHz | Voice: Auto-assigned
OmniVoice TTS / Voice Clone
Best for: Zero-shot voice cloning in over 600 languages.
Supported modes: Text-to-Speech / Voice Clone
Default settings: Language: Auto-detect | Sample rate: 44.1kHz
07. Token Economy
Banana Split operates on a token-based system to standardize billing across different hardware models. Think of tokens as the fuel for your GPU generations.
7.1 What Tokens Are
Tokens are the unified credit system in Banana Split. Every image or video you generate consumes tokens based on the complexity and model runtimes of your request.
Tokens come bundled with your subscription plan each month. Refill purchase rates are calculated based on your subscription tier:
- Starter / Basic: $0.050 per token (~20 tokens per $1 USD)
- Creator / Pro: $0.045 per token (~22.2 tokens per $1 USD)
- Studio: $0.040 per token (~25 tokens per $1 USD)
7.2 How We Protect Your Tokens
We optimize prompt compilation to keep your token consumption low:
- Unified Planning: When you generate a batch or sequential plan using the Director, we do not charge you for multiple AI planning passes. The Director batches your direction into a single scheduling run, charging you only for the final outputs.
- Identity Shortcuts: Typing a direct handle name (e.g. `@elena`) replaces a long descriptive paragraph. This keeps prompt text input short, reducing token expenses on context processing.
🎬 We built these optimizations because we believe you should be spending your tokens on creative output — not on overhead.
7.3 What Different Actions Cost
Token costs scale dynamically depending on the hardware engine used and render complexity. The overall cost logic is simple:
- High resolutions and larger canvas dimensions require more tokens.
- Longer video files consume more tokens than short clips.
- Generating batches multiplies the base model cost per visual created.
For guidance on cost values, high-quality professional still images range from **3 to 6 tokens**, while professional 15-second cinematic video files require **around 150 tokens**.
7.4 Topping Up
To buy additional tokens, open **Account Settings → Token Refill**:
- Choose a preset size check: **$20**, **$100**, or **$200**.
- Click the "Other" preset card to unlock the input box for custom amounts, which has a minimum limit of **$5** and a maximum of **$200** per top-up.
- Review the live calculator to check the exact token count you will receive before confirming your payment.
- Expiration Policy: Additional tokens purchased through refills expire exactly **90 days** from the date of purchase.
Your token balance is displayed in the header at all times, ensuring you spend only what you choose to load. Monthly subscription tokens refresh each billing cycle, whereas top-up reserve tokens are consumed as spillover once monthly tokens are exhausted.
