ChatGPT Claude Codex CLI Cohere Cursor DeepSeek Gemini GitHub Copilot GLM Grok Kimi Llama MiniMax Mistral OpenAI opencode Skill

generate-nanobanana

Generate and edit images/video with Google's Gemini media models (Nano Banana 2/Pro, Gemini Omni Flash), with cost-approval gates, reference-image support, and a prompt/output log per call.

LLM Mart · 0 points · 26 views 17 listing impressions 0 install-command copies

#image-generation

Virus-scanned Reviewed automatically before listing.

Full trust report

Download sickn33-agentic-awesome-skills-skills_generate-nanobanana-1f67c44.zip · 11 KB
Part of sickn33/agentic-awesome-skills — 427 skills
This skill couldn't be refreshed from GitHub on the last check — you're seeing the last imported snapshot.

Install

skills CLI npx skills add https://github.com/sickn33/agentic-awesome-skills/tree/main/skills/generate-nanobanana
Claude Code claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install sickn33-agentic-awesome-skills@llmmart
Git git clone https://github.com/sickn33/agentic-awesome-skills.git

The skills CLI installs just this skill, for any of its supported agents. Claude Code installs the whole sickn33/agentic-awesome-skills collection as a plugin from our marketplace. Git is the plain clone.

Skill manifest

Generate Nanobanana

Overview

generate-nanobanana calls Google's Gemini media models directly through the Gemini API — no third-party routing layer — to generate and edit images and video. It routes each request to the right model tier (draft, standard, quality, or video), loads real reference images instead of relying on text descriptions, gates every paid call behind explicit user approval, and writes a JSON sidecar next to every output recording the exact prompt, model, and cost. It registers a single /generate command.

This skill adapts the workflow (model routing, reference-image handling, sidecar logging) from AntonioCardenas/generate-nanobanana. The actual request shapes in references/ were independently verified against the live Gemini API docs rather than copied from that upstream repo, whose examples predate Google's migration to the Interactions API and use stale, non-functional request methods. Model IDs, request contracts, and pricing all change on Google's own schedule — re-verify against the docs linked from each reference file before relying on this skill in a new session.

When to Use This Skill

  • Use when the user asks to generate, create, or make an image or video, or wants a thumbnail.
  • Use when the user wants to animate a still image, or says "generate on brand" or "generate from reference".
  • Use when the user wants to link or import a folder of reference images (logos, faces, product shots) for reuse across generations.
  • Use when the user invokes /generate or /generate frf <set>, even without naming a specific model.

How It Works

Step 1: Route to a model

Pick the model for the job and read its reference file under references/ before calling anything — each file holds the current, verified request shape for that model.

Task Model Model ID Reference
Image (draft) Nano Banana 2 Lite gemini-3.1-flash-lite-image references/gemini-3.1-flash-lite-image.md
Image (standard) Nano Banana 2 gemini-3.1-flash-image references/gemini-3.1-flash-image.md
Image (quality, multi-image fusion) Nano Banana Pro gemini-3-pro-image references/gemini-3-pro-image.md
Video Gemini Omni Flash gemini-omni-flash-preview references/gemini-omni-flash-preview.md

All four models are called through the Interactions API (client.interactions.create(...), REST POST /v1beta/interactions) — see each reference file for the exact shape, including reference-image input and, for video, large-output retrieval. Every call is billable; see Step 3.

Draft on Nano Banana 2 Lite first and rerun the picked favorite on Nano Banana 2 or Pro; reserve Pro for heavy multi-image fusion, character-consistent series, or dense on-image text.

Step 2: Load references

Pull real reference images from generations/refs/, or from a named reference set when the request says "on brand" or invokes /generate frf <set>. Never substitute a text description for a reference image (logo, face, brand mark) that already exists — stop and ask if a named reference is missing instead of approximating it.

Reference sets are registered by importing (copying files into generations/refs/<set>/, a snapshot) or linking (recording the source path in generations/refs/sets.json, read live at generation time). A set may carry a style.md whose contents are prepended verbatim to every prompt generated from that set.

Step 3: Generate

Call the Gemini API per the model's reference file. Every generation — image or video — is billable and requires an explicit approval gate: quote the current per-unit price from the live pricing page for the selected model and get explicit user go-ahead before that specific call. One approval covers exactly one call; a rerun needs its own. Run generations one at a time, never in parallel, so approval and cost tracking stay accurate.

No model in this skill documents a seed or reproducibility parameter — do not promise an identical re-roll. For "same image but change X" requests, reuse the exact original prompt and reference images (from the sidecar log) and change only the requested delta; for video, chain edits via previous_interaction_id where supported (see the Omni Flash reference).

Step 4: Verify and log

Confirm the generated file is on disk and non-empty, then write a matching .json sidecar next to it (see Examples) recording the exact model ID, prompt, references used, response id, cost, and timestamp. Never log a generation whose file isn't there, and never write a sidecar for a failed or safety-blocked call.

Examples

Example 1: On-brand thumbnail from a linked reference set

User: generate a thumbnail on brand for the new pricing page

The skill resolves the brand reference set from generations/refs/sets.json, prepends its style.md (if present), picks the relevant reference images (e.g. the logo and a style shot), quotes the current Nano Banana 2 Lite price and gets approval, then saves the result to generations/pricing_page_thumbnail_<timestamp>.png with a sidecar.

Example 2: Sidecar log written beside an output

{
  "model": "gemini-3.1-flash-lite-image",
  "prompt": "the exact prompt sent",
  "reference_images": ["generations/refs/brand/logo_dark.png"],
  "reference_set": "brand",
  "response_id": "v1_...",
  "params": { "aspect_ratio": "16:9", "image_size": "1K" },
  "cost": "{price quoted from the live pricing page before running}",
  "created": "2026-07-31T14:20:00Z",
  "approved_by_user": true
}

Best Practices

  • ✅ Quote the current price and get explicit approval before every paid generation — image or video, not just video. A quote is not approval, and each rerun needs its own.
  • ✅ Use real reference images for faces, logos, and brand marks instead of describing them in text.
  • ✅ Read the model's reference file in references/ before calling it — model IDs and request shapes have already changed once in this skill's lifetime (Interactions API migration, gemini-3-pro-image-preview shutdown).
  • ❌ Don't generate "on brand" from an empty or nonexistent reference set — bootstrap the folder and stop until it has at least one real image.
  • ❌ Don't claim a generation is exactly reproducible — no model here documents a seed parameter. Reuse the exact prompt and references instead of promising identical output.
  • ❌ Don't run generations in parallel or reconstruct a prompt from memory when the original's sidecar still has the exact text.

Limitations

  • Covers Google Gemini models only; there is no multi-provider routing to other image/video generators.
  • Requires a Google AI Studio API key (GEMINI_API_KEY) and, outside Antigravity's native tool fallback, the google-genai Python package.
  • No model documents a seed or reproducibility guarantee; reruns are best-effort via the saved prompt and references, not identical output.
  • Model IDs and pricing are Google's to change; the reference files carry the model IDs verified at the time this skill was last updated, and each links to the live docs to re-verify against.
  • This skill does not replace environment-specific validation, testing, or expert review of generated assets.
  • Stop and ask for clarification if a required reference image, permission, or the API key is missing.

Security & Safety Notes

  • Network — Generation and file-transfer calls go to generativelanguage.googleapis.com; checking current docs or pricing contacts ai.google.dev, and an explicitly approved package install contacts the configured PyPI index. Never send prompts or reference media to any other endpoint.
  • Secrets — GEMINI_API_KEY is only ever read from the environment or a workspace .env the user already set up; it is never logged, printed, or written into a sidecar, prompt, or committed file. The skill never creates or edits .env, .env.example, or .gitignore itself.
  • File writes — skill-authored project outputs are confined to the workspace's generations/ folder (including generations/refs/, REST request/response files, and sets.json); nothing is written outside the current project except an explicitly approved package installation in its selected environment.
  • Package installs — only the official google-genai PyPI package, and only when missing; never installed silently or alongside any other package.
  • Cost — every call spends real money against the user's Google AI Studio billing; that, plus filesystem writes, is why this skill is risk: critical rather than safe.
  • Treat any change that would add a new network endpoint, a new package install, or a write outside generations/ as a design decision for the user to approve, not something to do quietly.

Common Pitfalls

  • Problem: Requesting "on brand" generation before any reference images exist. Solution: Create generations/refs/<name>/, tell the user its path, and wait for at least one image before generating.
  • Problem: Varying an existing image by re-describing it from memory. Solution: Read the original's sidecar for its exact prompt and references, and change only the requested delta.
  • Problem: Running an image or video generation without a cost quote. Solution: Always quote the current per-unit price from the live pricing page and get explicit approval before submitting any paid call.
  • Problem: Calling a model ID from memory instead of the reference file. Solution: Model IDs shift (e.g. gemini-3-pro-image-preview was shut down and replaced by gemini-3-pro-image) — always read references/<model>.md first.

Related Skills

  • @image-generator - Nano Banana Pro image generation and editing without the multi-model routing, reference-set library, or cost-gate workflow.
  • @nanobanana-ppt-skills - AI-powered PPT generation with document analysis and styled images.
  • @2slides-ppt-generator - Presentation generation via 2slides API.
Files (agentic-awesome-skills)
  • references
    • gemini-3-pro-image.md 3.7 KB
      # Gemini 3 Pro Image (`gemini-3-pro-image`)
      
      ## Overview
      Nano Banana Pro is the flagship model for highest-quality rendering, complex multi-image fusion, character consistency, and sharp on-image typography.
      
      > **Model ID note**: the earlier `gemini-3-pro-image-preview` was deprecated 2026-05-28 and shut down 2026-06-25. `gemini-3-pro-image` is the current generally-available (GA) replacement. Re-verify against [ai.google.dev/gemini-api/docs/image-generation](https://ai.google.dev/gemini-api/docs/image-generation) before relying on this ID, since Google rotates preview/GA model names on its own schedule.
      
      ## Model Specification
      - **Model ID**: `gemini-3-pro-image`
      - **API**: Interactions API (`client.interactions.create`) — this model does not use the older `generate_content` method.
      - **Primary Use**: Premium graphics, multi-image fusion, dense on-image text, complex composite scenes.
      - **Cost**: Billable per call. Quote the current price from the live [pricing page](https://ai.google.dev/gemini-api/docs/pricing) and get explicit user approval before every generation — see the skill's cost-approval rule.
      - **Reference images**: Up to 14 supported as additional `image` input parts.
      - **Reproducibility**: No `seed` parameter is documented for this model. Treat every generation as non-deterministic; for "same image but change X" requests, reuse the exact original prompt and reference images rather than promising an identical re-roll.
      
      ## Request Shape
      
      ### Python SDK (`google-genai`, Interactions API)
      ```python
      from google import genai
      import base64
      
      client = genai.Client()
      
      interaction = client.interactions.create(
          model="gemini-3-pro-image",
          input="A high-end editorial magazine cover featuring a futuristic electric car with legible headline text 'THE FUTURE OF MOBILITY'",
          response_format={
              "type": "image",
              "aspect_ratio": "3:4",
              "image_size": "2K",
          },
      )
      
      with open("generations/cover.png", "wb") as f:
          f.write(base64.b64decode(interaction.output_image.data))
      ```
      
      ### Multi-Image Reference & Fusion
      Nano Banana Pro supports up to 14 reference images for composition and style fusion:
      ```python
      from google import genai
      import base64
      
      client = genai.Client()
      
      with open("generations/refs/brand/character.png", "rb") as f:
          subject_bytes = f.read()
      with open("generations/refs/brand/logo.png", "rb") as f:
          logo_bytes = f.read()
      
      interaction = client.interactions.create(
          model="gemini-3-pro-image",
          input=[
              {"type": "text", "text": "Combine the character from the first image and place the logo from the second image on their jacket in a retro synthwave city"},
              {"type": "image", "data": base64.b64encode(subject_bytes).decode("utf-8"), "mime_type": "image/png"},
              {"type": "image", "data": base64.b64encode(logo_bytes).decode("utf-8"), "mime_type": "image/png"},
          ],
          response_format={"type": "image", "aspect_ratio": "16:9", "image_size": "2K"},
      )
      ```
      
      ### REST API (`curl`)
      ```bash
      mkdir -p generations
      cat > generations/pro_image_request.json << 'EOF'
      {
        "model": "gemini-3-pro-image",
        "input": [
          {"type": "text", "text": "A high-end editorial magazine cover featuring a futuristic electric car with legible headline text 'THE FUTURE OF MOBILITY'"}
        ],
        "response_format": {
          "type": "image",
          "aspect_ratio": "3:4",
          "image_size": "2K"
        }
      }
      EOF
      
      curl -s -X POST \
        "https://generativelanguage.googleapis.com/v1beta/interactions" \
        -H "x-goog-api-key: $GEMINI_API_KEY" \
        -H "Content-Type: application/json" \
        -d @generations/pro_image_request.json > generations/pro_image_response.json
      ```
      
      The response's `output_image.data` field holds the base64-encoded image bytes; decode and write them to the target file.
      
    • gemini-3.1-flash-image.md 3 KB
      # Gemini 3.1 Flash Image (`gemini-3.1-flash-image`)
      
      ## Overview
      Nano Banana 2 is the standard production model for image generation. It balances crisp detail, accurate style adherence, and high speed for most finished work.
      
      ## Model Specification
      - **Model ID**: `gemini-3.1-flash-image`
      - **API**: Interactions API (`client.interactions.create`) — this model does not use the older `generate_content` method.
      - **Primary Use**: Production image generation, brand assets, social media graphics.
      - **Cost**: Billable per call. Quote the current price from the live [pricing page](https://ai.google.dev/gemini-api/docs/pricing) and get explicit user approval before every generation — see the skill's cost-approval rule.
      - **Reference images**: Up to 14 supported as additional `image` input parts.
      - **Reproducibility**: No `seed` parameter is documented for this model. Treat every generation as non-deterministic; for "same image but change X" requests, reuse the exact original prompt and reference images rather than promising an identical re-roll.
      
      ## Request Shape
      
      ### Python SDK (`google-genai`, Interactions API)
      ```python
      from google import genai
      import base64
      
      client = genai.Client()
      
      interaction = client.interactions.create(
          model="gemini-3.1-flash-image",
          input="A sleek modern product advertisement for wireless headphones on a clean marble table, studio lighting",
          response_format={
              "type": "image",
              "aspect_ratio": "16:9",
              "image_size": "2K",
          },
      )
      
      with open("generations/headphones.png", "wb") as f:
          f.write(base64.b64decode(interaction.output_image.data))
      ```
      
      ### Reference Image Input
      ```python
      from google import genai
      import base64
      
      client = genai.Client()
      
      with open("generations/refs/brand/style_sample.png", "rb") as f:
          style_bytes = f.read()
      
      interaction = client.interactions.create(
          model="gemini-3.1-flash-image",
          input=[
              {"type": "text", "text": "Generate a pricing page banner adhering to the color scheme and lighting of this style reference"},
              {"type": "image", "data": base64.b64encode(style_bytes).decode("utf-8"), "mime_type": "image/png"},
          ],
          response_format={"type": "image", "aspect_ratio": "16:9", "image_size": "2K"},
      )
      ```
      
      ### REST API (`curl`)
      ```bash
      mkdir -p generations
      cat > generations/flash_image_request.json << 'EOF'
      {
        "model": "gemini-3.1-flash-image",
        "input": [
          {"type": "text", "text": "A sleek modern product advertisement for wireless headphones on a clean marble table, studio lighting"}
        ],
        "response_format": {
          "type": "image",
          "aspect_ratio": "16:9",
          "image_size": "2K"
        }
      }
      EOF
      
      curl -s -X POST \
        "https://generativelanguage.googleapis.com/v1beta/interactions" \
        -H "x-goog-api-key: $GEMINI_API_KEY" \
        -H "Content-Type: application/json" \
        -d @generations/flash_image_request.json > generations/flash_image_response.json
      ```
      
      The response's `output_image.data` field holds the base64-encoded image bytes; decode and write them to the target file.
      
    • gemini-3.1-flash-lite-image.md 3 KB
      # Gemini 3.1 Flash Lite Image (`gemini-3.1-flash-lite-image`)
      
      ## Overview
      Nano Banana 2 Lite is Google's fastest and cheapest Gemini image model — the draft tier for rapid concept exploration and quick visual iteration before promoting a picked result to a higher tier.
      
      ## Model Specification
      - **Model ID**: `gemini-3.1-flash-lite-image`
      - **API**: Interactions API (`client.interactions.create`) — this model does not use the older `generate_content` method.
      - **Primary Use**: Image drafts, rapid prototyping, thumbnail concepts.
      - **Cost**: Billable per call. Quote the current price from the live [pricing page](https://ai.google.dev/gemini-api/docs/pricing) and get explicit user approval before every generation — see the skill's cost-approval rule.
      - **Reference images**: Up to 14 supported as additional `image` input parts.
      - **Reproducibility**: No `seed` parameter is documented for this model. Treat every generation as non-deterministic; for "same image but change X" requests, reuse the exact original prompt and reference images rather than promising an identical re-roll.
      
      ## Request Shape
      
      ### Python SDK (`google-genai`, Interactions API)
      ```python
      from google import genai
      import base64
      
      client = genai.Client()
      
      interaction = client.interactions.create(
          model="gemini-3.1-flash-lite-image",
          input="A futuristic city skyline at sunset, cyberpunk aesthetic, high detail",
          response_format={
              "type": "image",
              "aspect_ratio": "16:9",
              "image_size": "1K",
          },
      )
      
      with open("generations/output.png", "wb") as f:
          f.write(base64.b64decode(interaction.output_image.data))
      ```
      
      ### Reference Image Input
      Pass reference images as additional `input` parts (base64-encoded), alongside the text prompt:
      ```python
      from google import genai
      import base64
      
      client = genai.Client()
      
      with open("generations/refs/brand/logo.png", "rb") as f:
          logo_bytes = f.read()
      
      interaction = client.interactions.create(
          model="gemini-3.1-flash-lite-image",
          input=[
              {"type": "text", "text": "Incorporate this logo style into a draft banner for summer sale"},
              {"type": "image", "data": base64.b64encode(logo_bytes).decode("utf-8"), "mime_type": "image/png"},
          ],
          response_format={"type": "image", "aspect_ratio": "16:9"},
      )
      ```
      
      ### REST API (`curl`)
      ```bash
      mkdir -p generations
      cat > generations/lite_image_request.json << 'EOF'
      {
        "model": "gemini-3.1-flash-lite-image",
        "input": [
          {"type": "text", "text": "A futuristic city skyline at sunset, cyberpunk aesthetic, high detail"}
        ],
        "response_format": {
          "type": "image",
          "aspect_ratio": "16:9",
          "image_size": "1K"
        }
      }
      EOF
      
      curl -s -X POST \
        "https://generativelanguage.googleapis.com/v1beta/interactions" \
        -H "x-goog-api-key: $GEMINI_API_KEY" \
        -H "Content-Type: application/json" \
        -d @generations/lite_image_request.json > generations/lite_image_response.json
      ```
      
      The response's `output_image.data` field holds the base64-encoded image bytes; decode and write them to the target file.
      
    • gemini-omni-flash-preview.md 5.8 KB
      # Gemini Omni Flash Video (`gemini-omni-flash-preview`)
      
      ## Overview
      Gemini Omni Flash generates and edits video. It supports text-to-video, image-to-video, subject-reference video, stateful multi-turn video editing, and editing a user's own uploaded video. **Every paid run requires explicit user cost approval before execution — see the skill's cost-approval rule.**
      
      ## Model Specification
      - **Model ID**: `gemini-omni-flash-preview`
      - **API**: Interactions API (`client.interactions.create`) — this model does not use the older `generate_videos` or `:predictLongRunning` methods.
      - **Primary Use**: Text-to-video, image-to-video, subject-reference video, video editing.
      - **Cost**: Billable per call, priced per output. Quote the current price from the live [pricing page](https://ai.google.dev/gemini-api/docs/pricing) and get explicit user approval before submitting — one approval covers exactly one run.
      - **Aspect ratios**: `16:9`, `9:16` documented for aspect-ratio-controlled requests.
      - **Reproducibility**: No `seed` parameter is documented for this model. Treat every generation as non-deterministic.
      
      ## Request Shape
      
      ### Text-to-Video (Python SDK, Interactions API)
      ```python
      import base64
      from google import genai
      
      client = genai.Client()
      
      # Quote cost and wait for explicit user approval before running!
      interaction = client.interactions.create(
          model="gemini-omni-flash-preview",
          input="A marble rolling fast on a chain reaction style track, continuous smooth shot.",
      )
      with open("generations/marble.mp4", "wb") as f:
          f.write(base64.b64decode(interaction.output_video.data))
      ```
      
      ### Control Aspect Ratio
      ```python
      interaction = client.interactions.create(
          model="gemini-omni-flash-preview",
          input="A futuristic city with neon lights and flying cars, cyberpunk style",
          response_format={
              "type": "video",  # optional
              "aspect_ratio": "9:16",  # supported: "9:16", "16:9"
          },
      )
      ```
      
      ### Image-to-Video
      Pass a reference image and instructions as separate `input` parts:
      ```python
      import base64
      from google import genai
      
      client = genai.Client()
      
      with open("generations/refs/start_frame.png", "rb") as f:
          frame_bytes = f.read()
      
      interaction = client.interactions.create(
          model="gemini-omni-flash-preview",
          input=[
              {"type": "image", "data": base64.b64encode(frame_bytes).decode("utf-8"), "mime_type": "image/png"},
              {"type": "text", "text": "The scene animates smoothly as the character steps forward into the misty forest."},
          ],
      )
      with open("generations/forest.mp4", "wb") as f:
          f.write(base64.b64decode(interaction.output_video.data))
      ```
      
      ### Subject Reference (multiple reference images)
      ```python
      interaction = client.interactions.create(
          model="gemini-omni-flash-preview",
          input=[
              {"type": "image", "data": cat_b64, "mime_type": "image/png"},
              {"type": "image", "data": yarn_b64, "mime_type": "image/png"},
              {"type": "text", "text": "A cat playfully batting at a ball of yarn."},
          ],
      )
      ```
      
      ### Stateful Multi-Turn Video Editing
      Chain an edit onto a prior generation with `previous_interaction_id` — this is the closest thing this model offers to controlled reruns, not a seed:
      ```python
      # Turn 1: generate
      res1 = client.interactions.create(model="gemini-omni-flash-preview", input="A woman playing violin outdoors.")
      
      # Turn 2: edit the previous result
      res2 = client.interactions.create(
          model="gemini-omni-flash-preview",
          previous_interaction_id=res1.id,
          input="Make the violin invisible.",
      )
      with open("generations/violin.mp4", "wb") as f:
          f.write(base64.b64decode(res2.output_video.data))
      ```
      
      ### Editing a User's Own Uploaded Video
      ```python
      import time
      from google import genai
      
      client = genai.Client()
      
      video_file = client.files.upload(file="Video.mp4")
      while video_file.state == "PROCESSING":
          time.sleep(10)
          video_file = client.files.get(name=video_file.name)
      if video_file.state == "FAILED":
          raise ValueError(video_file.state)
      
      interaction = client.interactions.create(
          model="gemini-omni-flash-preview",
          input=[
              {"type": "document", "uri": video_file.uri},
              {"type": "text", "text": "When the person touches the mirror, make the mirror ripple beautifully like liquid, and the person's arm turns into reflective mirror material"},
          ],
      )
      with open("generations/mirror.mp4", "wb") as f:
          f.write(base64.b64decode(interaction.output_video.data))
      ```
      
      ### Large Outputs: Retrieve via URI Instead of Inline Base64
      For outputs too large for inline base64, request `delivery: "uri"` and poll the Files API until `ACTIVE`:
      ```python
      import time
      from google import genai
      
      client = genai.Client()
      
      interaction = client.interactions.create(
          model="gemini-omni-flash-preview",
          input="A beautiful sunset over a calm ocean.",
          response_format={"type": "video", "delivery": "uri"},
      )
      
      video_output = interaction.output_video
      file_name = video_output.uri.split("/")[-1]
      
      while True:
          f_info = client.files.get(name=f"files/{file_name}")
          if f_info.state.name == "ACTIVE":
              break
          if f_info.state.name == "FAILED":
              raise RuntimeError("Generation failed.")
          time.sleep(5)
      
      video_bytes = client.files.download(file=video_output.uri)
      with open("generations/output.mp4", "wb") as f:
          f.write(video_bytes)
      ```
      
      ### REST API (`curl`)
      ```bash
      curl -X POST "https://generativelanguage.googleapis.com/v1beta/interactions" \
        -H "x-goog-api-key: $GEMINI_API_KEY" \
        -H "Content-Type: application/json" \
        -d '{
           "model": "gemini-omni-flash-preview",
           "input": "A marble rolling fast on a chain reaction style track, continuous smooth shot."
          }'
      ```
      
      The response's `output_video.data` field holds base64-encoded video bytes (or `output_video.uri` when `delivery: "uri"` was requested); decode/download and write to the target file. The response envelope also includes an `id` field — log it so multi-turn edits can chain via `previous_interaction_id`.
      
  • SKILL.md 10.8 KB
    ---
    name: generate-nanobanana
    description: "Generate and edit images/video with Google's Gemini media models (Nano Banana 2/Pro, Gemini Omni Flash), with cost-approval gates, reference-image support, and a prompt/output log per call."
    category: media
    risk: critical
    source: community
    source_repo: AntonioCardenas/generate-nanobanana
    source_type: community
    date_added: "2026-08-04"
    author: antonio
    tags: [nanobanana, gemini, google-ai-studio, image-generation, video-generation]
    tools: [claude, cursor, gemini, codex, antigravity]
    license: "MIT"
    license_source: "https://github.com/AntonioCardenas/generate-nanobanana/blob/main/LICENSE"
    ---
    
    # Generate Nanobanana
    
    ## Overview
    
    `generate-nanobanana` calls Google's Gemini media models directly through the Gemini API — no third-party routing layer — to generate and edit images and video. It routes each request to the right model tier (draft, standard, quality, or video), loads real reference images instead of relying on text descriptions, gates every paid call behind explicit user approval, and writes a JSON sidecar next to every output recording the exact prompt, model, and cost. It registers a single `/generate` command.
    
    This skill adapts the workflow (model routing, reference-image handling, sidecar logging) from [AntonioCardenas/generate-nanobanana](https://github.com/AntonioCardenas/generate-nanobanana). The actual request shapes in `references/` were independently verified against the live [Gemini API docs](https://ai.google.dev/gemini-api/docs/image-generation) rather than copied from that upstream repo, whose examples predate Google's migration to the Interactions API and use stale, non-functional request methods. Model IDs, request contracts, and pricing all change on Google's own schedule — re-verify against the docs linked from each reference file before relying on this skill in a new session.
    
    ## When to Use This Skill
    
    - Use when the user asks to generate, create, or make an image or video, or wants a thumbnail.
    - Use when the user wants to animate a still image, or says "generate on brand" or "generate from reference".
    - Use when the user wants to link or import a folder of reference images (logos, faces, product shots) for reuse across generations.
    - Use when the user invokes `/generate` or `/generate frf <set>`, even without naming a specific model.
    
    ## How It Works
    
    ### Step 1: Route to a model
    
    Pick the model for the job and read its reference file under [`references/`](references/) before calling anything — each file holds the current, verified request shape for that model.
    
    | Task | Model | Model ID | Reference |
    | --- | --- | --- | --- |
    | Image (draft) | Nano Banana 2 Lite | `gemini-3.1-flash-lite-image` | [`references/gemini-3.1-flash-lite-image.md`](references/gemini-3.1-flash-lite-image.md) |
    | Image (standard) | Nano Banana 2 | `gemini-3.1-flash-image` | [`references/gemini-3.1-flash-image.md`](references/gemini-3.1-flash-image.md) |
    | Image (quality, multi-image fusion) | Nano Banana Pro | `gemini-3-pro-image` | [`references/gemini-3-pro-image.md`](references/gemini-3-pro-image.md) |
    | Video | Gemini Omni Flash | `gemini-omni-flash-preview` | [`references/gemini-omni-flash-preview.md`](references/gemini-omni-flash-preview.md) |
    
    All four models are called through the **Interactions API** (`client.interactions.create(...)`, REST `POST /v1beta/interactions`) — see each reference file for the exact shape, including reference-image input and, for video, large-output retrieval. Every call is billable; see Step 3.
    
    Draft on Nano Banana 2 Lite first and rerun the picked favorite on Nano Banana 2 or Pro; reserve Pro for heavy multi-image fusion, character-consistent series, or dense on-image text.
    
    ### Step 2: Load references
    
    Pull real reference images from `generations/refs/`, or from a named reference set when the request says "on brand" or invokes `/generate frf <set>`. Never substitute a text description for a reference image (logo, face, brand mark) that already exists — stop and ask if a named reference is missing instead of approximating it.
    
    Reference sets are registered by **importing** (copying files into `generations/refs/<set>/`, a snapshot) or **linking** (recording the source path in `generations/refs/sets.json`, read live at generation time). A set may carry a `style.md` whose contents are prepended verbatim to every prompt generated from that set.
    
    ### Step 3: Generate
    
    Call the Gemini API per the model's reference file. **Every generation — image or video — is billable and requires an explicit approval gate**: quote the current per-unit price from the live [pricing page](https://ai.google.dev/gemini-api/docs/pricing) for the selected model and get explicit user go-ahead before that specific call. One approval covers exactly one call; a rerun needs its own. Run generations one at a time, never in parallel, so approval and cost tracking stay accurate.
    
    No model in this skill documents a `seed` or reproducibility parameter — do not promise an identical re-roll. For "same image but change X" requests, reuse the exact original prompt and reference images (from the sidecar log) and change only the requested delta; for video, chain edits via `previous_interaction_id` where supported (see the Omni Flash reference).
    
    ### Step 4: Verify and log
    
    Confirm the generated file is on disk and non-empty, then write a matching `.json` sidecar next to it (see Examples) recording the exact model ID, prompt, references used, response `id`, cost, and timestamp. Never log a generation whose file isn't there, and never write a sidecar for a failed or safety-blocked call.
    
    ## Examples
    
    ### Example 1: On-brand thumbnail from a linked reference set
    
    ```
    User: generate a thumbnail on brand for the new pricing page
    ```
    
    The skill resolves the `brand` reference set from `generations/refs/sets.json`, prepends its `style.md` (if present), picks the relevant reference images (e.g. the logo and a style shot), quotes the current Nano Banana 2 Lite price and gets approval, then saves the result to `generations/pricing_page_thumbnail_<timestamp>.png` with a sidecar.
    
    ### Example 2: Sidecar log written beside an output
    
    ```json
    {
      "model": "gemini-3.1-flash-lite-image",
      "prompt": "the exact prompt sent",
      "reference_images": ["generations/refs/brand/logo_dark.png"],
      "reference_set": "brand",
      "response_id": "v1_...",
      "params": { "aspect_ratio": "16:9", "image_size": "1K" },
      "cost": "{price quoted from the live pricing page before running}",
      "created": "2026-07-31T14:20:00Z",
      "approved_by_user": true
    }
    ```
    
    ## Best Practices
    
    - ✅ Quote the current price and get explicit approval before **every** paid generation — image or video, not just video. A quote is not approval, and each rerun needs its own.
    - ✅ Use real reference images for faces, logos, and brand marks instead of describing them in text.
    - ✅ Read the model's reference file in `references/` before calling it — model IDs and request shapes have already changed once in this skill's lifetime (Interactions API migration, `gemini-3-pro-image-preview` shutdown).
    - ❌ Don't generate "on brand" from an empty or nonexistent reference set — bootstrap the folder and stop until it has at least one real image.
    - ❌ Don't claim a generation is exactly reproducible — no model here documents a seed parameter. Reuse the exact prompt and references instead of promising identical output.
    - ❌ Don't run generations in parallel or reconstruct a prompt from memory when the original's sidecar still has the exact text.
    
    ## Limitations
    
    - Covers Google Gemini models only; there is no multi-provider routing to other image/video generators.
    - Requires a Google AI Studio API key (`GEMINI_API_KEY`) and, outside Antigravity's native tool fallback, the `google-genai` Python package.
    - No model documents a seed or reproducibility guarantee; reruns are best-effort via the saved prompt and references, not identical output.
    - Model IDs and pricing are Google's to change; the reference files carry the model IDs verified at the time this skill was last updated, and each links to the live docs to re-verify against.
    - This skill does not replace environment-specific validation, testing, or expert review of generated assets.
    - Stop and ask for clarification if a required reference image, permission, or the API key is missing.
    
    ## Security & Safety Notes
    
    - **Network** — Generation and file-transfer calls go to `generativelanguage.googleapis.com`; checking current docs or pricing contacts `ai.google.dev`, and an explicitly approved package install contacts the configured PyPI index. Never send prompts or reference media to any other endpoint.
    - **Secrets** — `GEMINI_API_KEY` is only ever read from the environment or a workspace `.env` the user already set up; it is never logged, printed, or written into a sidecar, prompt, or committed file. The skill never creates or edits `.env`, `.env.example`, or `.gitignore` itself.
    - **File writes** — skill-authored project outputs are confined to the workspace's `generations/` folder (including `generations/refs/`, REST request/response files, and `sets.json`); nothing is written outside the current project except an explicitly approved package installation in its selected environment.
    - **Package installs** — only the official `google-genai` PyPI package, and only when missing; never installed silently or alongside any other package.
    - **Cost** — every call spends real money against the user's Google AI Studio billing; that, plus filesystem writes, is why this skill is `risk: critical` rather than `safe`.
    - Treat any change that would add a new network endpoint, a new package install, or a write outside `generations/` as a design decision for the user to approve, not something to do quietly.
    
    ## Common Pitfalls
    
    - **Problem:** Requesting "on brand" generation before any reference images exist.
      **Solution:** Create `generations/refs/<name>/`, tell the user its path, and wait for at least one image before generating.
    - **Problem:** Varying an existing image by re-describing it from memory.
      **Solution:** Read the original's sidecar for its exact prompt and references, and change only the requested delta.
    - **Problem:** Running an image or video generation without a cost quote.
      **Solution:** Always quote the current per-unit price from the live pricing page and get explicit approval before submitting any paid call.
    - **Problem:** Calling a model ID from memory instead of the reference file.
      **Solution:** Model IDs shift (e.g. `gemini-3-pro-image-preview` was shut down and replaced by `gemini-3-pro-image`) — always read `references/<model>.md` first.
    
    ## Related Skills
    
    - `@image-generator` - Nano Banana Pro image generation and editing without the multi-model routing, reference-set library, or cost-gate workflow.
    - `@nanobanana-ppt-skills` - AI-powered PPT generation with document analysis and styled images.
    - `@2slides-ppt-generator` - Presentation generation via 2slides API.
    

Comments (0)

Sign in to join the conversation.

No comments yet.

Reviews (0)

No reviews yet.

Related