Director's Handbook

/docs — THE COMPLETE GUIDE TO THE STUDIO AND AI GENERATION LIFECYCLE.

01. What Is Banana Split

Prompting is dead. Banana Split puts you in the director's chair — where your job is to have a vision, not memorize syntax.

Generative AI is incredibly powerful, but raw. To get high-fidelity results, you usually have to think like an engineer, configuring complex workflows and tuning model inputs by trial and error. Banana Split changes this. We've built a professional-grade monochrome studio that translates your creative intent into precise technical parameters. You dictate the scene, the camera behavior, the style, and the actors. The studio handles the pipeline execution, letting you focus on directing the action.

02. The Studio at a Glance

Before diving into the tools, orient yourself to the workspace layout. The interface is optimized to keep you in a continuous flow state where inputs, references, and renders live side by side.

2.1 The Split-Screen Workspace

The studio uses a resizable split-pane layout:

  • Left Pane (Controls & Inputs): This area hosts the Director Chat by default, and dynamically displays the settings panels when you load Manual Tools.
  • Right Pane (Universal Canvas): Dedicated entirely to rendering. Every image, video, upscale, and console log appears here inline as a live output tab.

A draggable vertical border divides the panes, allowing you to slide the ratio at any time. Shrink the controls to inspect high-resolution canvas results, or widen them when writing script sequences.

2.2 The Asset Library

The right edge of the screen contains the collapsible **Asset Library** sidebar. Widen this panel to explore your media archive:

  • Uploads Tab: Houses your imported reference files, mood boards, and aesthetic profiles.
  • Results Tab: Houses outputs generated in the studio, catalogued chronologically.
  • Card Preview Mode: Filters out code and metadata files to display a visual media grid in native aspect ratios, displaying a pink play badge over video files. Row List mode displays files as text lists.

2.3 How Files Move Through the Studio

Everything in your Asset Library is reusable. Production follows a simple loop:

The Production Cycle
1. Upload Reference Files OR Generate fresh assets
2. Files appear immediately in Asset Library tabs
3. Drag or type file handles directly into Chat or Manual Tools
4. Render runs → outputs save back to library as results

03. The Director

The Director is your production orchestrator. It is a reasoning agent that reads your creative directions, plans the steps required, and dispatches the execution tasks to the GPU.

3.1 How It Works

The Director processes instructions in a two-phase loop:

  • Planning Phase: You write your creative direction in plain language. The Director reads the instructions, plans which models are needed, and drafts a collapsible task checklist.
  • Execution Phase: You review the proposed production checklist. Clicking "Approve" dispatches the tasks sequentially. The Director locks the input panel, renders the prompt, and displays active loading states for each task.
🎬 You stay in control of the creative direction. The Director handles the technical translation.

3.2 Directing Effectively

To write instructions that the Director can interpret accurately, describe your scene by combining details about characters or aesthetics (using saved Identities), actions, environmental lighting, and camera behavior.

Vague Prompt (Avoid):
"A woman walking in the city."
Clear Direction (Recommended):
"A cinematic shot of @elena walking through a rain-soaked Tokyo street at night, neon reflections on the pavement, slow tracking shot, moody and desaturated color grade."
💡 The more specific your creative intent, the less the Director has to guess — and the closer your first result will be to your vision.

3.3 Pinning Files & References

You link context to a scene direction by combining three elements inside the input panel:

  • Asset Handles: Drag images or videos from the Asset Library sidebar directly into the chat area. They drop as handle text mentions (e.g. `@uploads/filename.png` or `@results/filename.mp4`), pinning the visuals as style or starting-frame references.
  • Identity Handles: Type the direct name handle of a saved identity (e.g. `@elena` or `@retro-grade`) to apply trained face models, color configurations, or environments.
  • Descriptive Direction: Explain the motion, story elements, and camera movements in natural prose surrounding your handles.

3.4 Batches & Chained Tasks

In typical AI pipelines, you must wait for an image, download it, and manually load it into a video engine. Banana Split removes this manual work by scheduling multi-task plans:

  • Batches (Parallel Generation): You can generate multiple variations concurrently. Asking for: Generate three variations of the opening scene, each using @marco, but vary the lighting: one golden hour, one overcast, one neon night triggers three tasks to run in parallel.
  • Chains (Sequential Pipelines): You can link tasks so the output of one step is piped automatically into the next. Directing: Generate a portrait of @sara in the library, then take that output and animate a slow push-in camera move with a 4s duration creates a sequential chain.
🎬 Batches save time. Chains save the entire manual pipeline. Together, they're how you produce at scale.

3.5 Your Results

Once task execution completes, the results are loaded as canvas tabs and stored in the library sidebar. You can download the assets using the Canvas overlay, drag them back into the chat input as context pins for subsequent scenes, or send them directly to a manual tool like the Upscaler.

3.6 Reference Media & Video Analysis

Instead of building visual pacing or camera behavior from scratch, you can feed reference media directly to the Director. The studio includes an automated **Visual Forensic & Video Analysis engine** that reads reference videos and images to extract stylistic parameters.

  • Video Analysis: You can send any public TikTok, Instagram, YouTube, or direct video URL to the Director Chat. The engine transcribes any spoken audio, dissects the cinematography (lighting, camera angles, movement vectors), and compiles a complete **Reusable Production Blueprint**. You can then instruct the Director to "generate a scene matching that blueprint" to duplicate its stylistic DNA.
  • Image Analysis: Dragging an uploaded image or generated result into the Director Chat and referencing it triggers a visual breakdown. The Director automatically runs a stylistic forensic analysis, cacheing the descriptive visual features locally. This cache is then read by downstream generation pipelines, preserving character identity, composition, and clothing details without sending the original image files to the generation servers.
🎬 Reference media analysis bridges the gap between existing content and AI production, giving you surgical control over pacing and cinematography.

04. Manual Tools

The Director is the recommended way to work in Banana Split — but sometimes you know exactly what you want and prefer to configure it yourself. Manual Tools give you direct access to each generation capability with full parameter control. The output workflow is identical: everything lands in your Asset Library and Canvas.

4.1 Image Generation

Generates still images from a text description, with full control over format and quality.

Use this when you have a clear image in mind and want to configure every output parameter yourself — aspect ratio, resolution, color palette, and batch size — without going through the planning phase.

Inputs: Prompt script (scene description) and an optional Color Palette array (hex values to lock color bounds).
Parameters: Aspect Ratio (selection from standard formats, e.g. 1:1, 16:9, 9:16), Resolution (1K, 2K, 4K), and Batch Size (number of parallel variations).
Output: Images saved to `@results/` and displayed in the Canvas.

4.2 Video Generation

Generates video from a text description or a starting image frame.

Use this when you want to create a moving scene and have a specific duration, motion feel, or starting image in mind. Supports both text-to-video and image-to-video depending on the model selected.

Inputs: Prompt description and an optional Reference Image frame (accepts `@uploads/` or `@results/` links for image-to-video).
Parameters: Duration (e.g. 4s, 8s, 12s, 20s), Motion Steerability (controls camera vectors, panning, and track speed), and Audio Sync (toggle to generate matching soundscapes).
Output: Video files saved to `@results/` and playable directly in the Canvas.

4.3 Motion Control

Transfers movement from a reference video onto a still character image.

Use this when you want a specific character — especially one defined in your Identities — to perform a movement from a source video, without filming anything new.

Inputs: Reference Video clip (holding the animation movement) and a Target Image (the static character profile). Accepts `@uploads/` or `@results/` handles.
Parameters: Motion extraction is handled automatically by mapping reference vectors onto the target image coordinates.
Output: Animated video file saved to `@results/`.
⚠️ For best results, the target character image should be a clean, well-lit, front-facing shot. Cluttered or partially cropped images reduce transfer accuracy.

4.4 Upscaling

Increases the resolution of an existing image, rebuilding fine detail in the process.

Use this as a finishing step — after you've confirmed a generated image is the one you want, run it through Upscaling to bring it to delivery quality.

Inputs: Target Image file. Accepts library handles or imports directly from active Canvas overlay buttons.
Parameters: Enhancement Strength (level of detail generation and sharpening) and Output Scale (4K or 8K).
Output: High-definition visual file saved to `@results/`.

4.5 Voice & Audio

Generates natural voiceovers or clones an existing voice for narration, dialogue, or dubbing.

Inputs: Script text (dialogue to speak) and an optional Voice Reference audio file (needed for Voice Clone mode).
Parameters: Mode (preset Text-to-Speech or Voice Clone) and Language (Zero-shot cloning support for 600+ languages).
Output: Narration audio file saved to `@results/`.

05. Identities

An Identity is a reusable creative reference you define once and invoke across any generation. The Identity Studio is a general-purpose engine capable of saving and training identities for characters, products, aesthetics, compositions, backgrounds, clothes, and literally any possible element you want to train and keep consistent across your production.

5.1 What You Can Identify

Identities support a wide range of creative training inputs, ensuring visual consistency across every visual dimension:

  • Characters: A saved facial profile, facial structure, actor appearance, or specific person.
  • Products: A watch, shoe, signature prop, or object that must keep its precise geometry.
  • Aesthetics: A visual style, artistic mood board, or color grading profile to bind tones and lighting.
  • Compositions: Framing constraints (such as flatlays, symmetry, low-angles, or perspective).
  • Backgrounds: A persistent set location, landscape, backdrop, or interior environment.
  • Clothes / Wardrobe: A specific outfit, uniform, costume, or clothing style to carry across scenes.
  • Literally Any Element: Any consistent visual concept, brand element, or texture you wish to train.

5.2 Creating an Identity

To create an Identity, open the **Identity Studio** from the left panel:

  1. Name the Identity: Choose a clean, unique handle (e.g. `@elena`, `@studio-red`, `@grunge-style`). Note that identities have no prefix (like `@identity/` or `@identities/`); you reference them in your directions simply using `@name`.
  2. Provide a Description: Detail the visual features, build, geometry, or styling constraints you are saving and training.
  3. Upload References: Upload 5 to 10 high-resolution reference images showing the subject or aesthetic from multiple angles and under uniform lighting.
💡 The quality of your references matters more than the quantity. Ten consistent, well-lit, clearly composed images of the same subject will outperform fifty mixed or cluttered ones. The Identity learns from what you show it — so show it the best version of what you want.

5.3 Using Identities in Production

Casting identities is handled by typing their direct handle names anywhere inside your scene direction or manual prompt (without any `@identities/` prefix):

"A product shot of @model-watch on the wrist of @elena, shot in the @studio-aesthetic with soft key light and dark background."

The orchestration engine automatically binds the trained facial features, product geometries, or aesthetic parameters together, eliminating long descriptive prompts and reducing your token costs.

5.4 Technical Architecture & Privacy

Identities are constructed exclusively to help the Director Agent formulate highly coherent prompts and scenarios.

When you upload reference images (Visual Anchors) to an Identity, the studio performs a **local, private forensic analysis** using Google Gemini to extract a highly detailed aesthetic and stylistic text description. This text description (the "compiled aesthetic context") is saved as the reference point for the Director Agent.

🔒 Privacy Note: The source images uploaded to your Identities are processed strictly for local context compilation. The raw image files are never sent or transmitted to generative text-to-image or text-to-video AI models.

06. Models

Banana Split provides access to a curated set of best-in-class generation models. You do not need to manage their technical hosting — choose the model that excels at your target task.

Image Models

Nano Banana 2 / Edit

Best for: High-speed photo generations and subject adjustments.

Supported modes: Text-to-Image / Image-to-Image / Local Edit

Default settings: Resolution: 1K (supports up to 4K) | Aspect ratio: auto | Batch: 1

Z-Image Turbo / i2i

Best for: Sub-second prototyping and storyboard testing.

Supported modes: Text-to-Image / Image-to-Image

Default settings: Resolution: 1K | Aspect ratio: auto | Batch: 1

Seedream 5.0 Lite / Edit

Best for: Rendering typography, text labels, and graphic designs.

Supported modes: Text-to-Image / Image-to-Image

Default settings: Resolution: 1K (supports up to 4K) | Aspect ratio: auto | Batch: 1

Stable Diffusion 3.5 Large Turbo / SDXL

Best for: Cost-effective generation with custom dimensions.

Supported modes: Text-to-Image / Image-to-Image

Default settings: Resolution: 1K (supports up to 2K) | Aspect ratio: auto | Batch: 1

FLUX 1.1 Pro Ultra

Best for: Unprocessed realistic details and cinematic compositions.

Supported modes: Text-to-Image

Default settings: Resolution: 1K (supports up to 2K) | Aspect ratio: auto | Batch: 1

Recraft V4 / V4 Pro

Best for: Logo design and vectors using strict color palette boundaries.

Supported modes: Text-to-Image

Default settings: Resolution: 1K (supports up to 2K) | Aspect ratio: auto | Batch: 1

Higgsfield Soul

Best for: Tonal grading and character consistency filters.

Supported modes: Image-to-Image

Default settings: Resolution: 1K (supports up to 2K) | Aspect ratio: auto | Batch: 1

Video & Motion Models

Kling 3.0 Pro / 3.0 4K

Best for: Detailed 4K cinematic video clips and camera panning.

Supported modes: Text-to-Video / Image-to-Video

Default settings: Resolution: 1080p (supports up to 4K) | Duration: 5s (supports up to 10s) | Aspect: 16:9

Kling 2.6 / 3.0 Motion Control

Best for: Animating character images using reference video motions.

Supported modes: Video-to-Video / Image-to-Video

Default settings: Resolution: 720p | Duration: 5s (supports up to 10s) | Aspect: 16:9

Sora 2 Pro / Sora 2 Actor

Best for: Complex physical interactions with synchronized audio.

Supported modes: Text-to-Video / Image-to-Video / Actor Training

Default settings: Resolution: 1080p | Duration: 4s (supports up to 12s) | Aspect: 16:9

Seedance 2.0

Best for: Consistent character movement and stable visual rendering.

Supported modes: Text-to-Video / Image-to-Video

Default settings: Resolution: 1080p | Duration: 5s (supports up to 12s) | Aspect: 16:9

Audio & Voice Models

ElevenLabs TTS / SVC

Best for: Narration, voice cloning, and audio conversion.

Supported modes: Text-to-Speech / Voice Conversion

Default settings: Sample rate: 44.1kHz | Voice: Auto-assigned

OmniVoice TTS / Voice Clone

Best for: Zero-shot voice cloning in over 600 languages.

Supported modes: Text-to-Speech / Voice Clone

Default settings: Language: Auto-detect | Sample rate: 44.1kHz

07. Token Economy

Banana Split operates on a token-based system to standardize billing across different hardware models. Think of tokens as the fuel for your GPU generations.

7.1 What Tokens Are

Tokens are the unified credit system in Banana Split. Every image or video you generate consumes tokens based on the complexity and model runtimes of your request.

Tokens come bundled with your subscription plan each month. Refill purchase rates are calculated based on your subscription tier:

  • Starter / Basic: $0.050 per token (~20 tokens per $1 USD)
  • Creator / Pro: $0.045 per token (~22.2 tokens per $1 USD)
  • Studio: $0.040 per token (~25 tokens per $1 USD)

7.2 How We Protect Your Tokens

We optimize prompt compilation to keep your token consumption low:

  • Unified Planning: When you generate a batch or sequential plan using the Director, we do not charge you for multiple AI planning passes. The Director batches your direction into a single scheduling run, charging you only for the final outputs.
  • Identity Shortcuts: Typing a direct handle name (e.g. `@elena`) replaces a long descriptive paragraph. This keeps prompt text input short, reducing token expenses on context processing.
🎬 We built these optimizations because we believe you should be spending your tokens on creative output — not on overhead.

7.3 What Different Actions Cost

Token costs scale dynamically depending on the hardware engine used and render complexity. The overall cost logic is simple:

  • High resolutions and larger canvas dimensions require more tokens.
  • Longer video files consume more tokens than short clips.
  • Generating batches multiplies the base model cost per visual created.

For guidance on cost values, high-quality professional still images range from **3 to 6 tokens**, while professional 15-second cinematic video files require **around 150 tokens**.

7.4 Topping Up

To buy additional tokens, open **Account Settings → Token Refill**:

  • Choose a preset size check: **$20**, **$100**, or **$200**.
  • Click the "Other" preset card to unlock the input box for custom amounts, which has a minimum limit of **$5** and a maximum of **$200** per top-up.
  • Review the live calculator to check the exact token count you will receive before confirming your payment.
  • Expiration Policy: Additional tokens purchased through refills expire exactly **90 days** from the date of purchase.

Your token balance is displayed in the header at all times, ensuring you spend only what you choose to load. Monthly subscription tokens refresh each billing cycle, whereas top-up reserve tokens are consumed as spillover once monthly tokens are exhausted.

© 2026 Banana Split Studio