Claude Skill

invoking-gemini

Invokes Google Gemini models for structured outputs, image generation, text-to-speech narration, multi-modal tasks, and Google-specific features. Use when users request Gemini, image generation, Gemini TTS or a synthesized voice, structured JSON output, Google API integration, or

LLM Mart · 0 points · 0 views 0 listing impressions 0 install-command copies
Virus-scanned Reviewed automatically before listing.

Full trust report

Download oaustegard-claude-skills-plugins_ai-and-reasoning_skills_invoking-gemini-e39c726.zip · 43 KB
Part of oaustegard/claude-skills — 39 skills

Install

skills CLI npx skills add https://github.com/oaustegard/claude-skills/tree/main/plugins/ai-and-reasoning/skills/invoking-gemini
Claude Code claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install oaustegard-claude-skills@llmmart
Git git clone https://github.com/oaustegard/claude-skills.git

The skills CLI installs just this skill, for any of its supported agents. Claude Code installs the whole oaustegard/claude-skills collection as a plugin from our marketplace. Git is the plain clone.

README

invoking-gemini

Invokes Google Gemini models for structured outputs, multi-modal tasks, and Google-specific features. Use when users request Gemini, structured JSON output, Google API integration, or cost-effective parallel processing.

Skill manifest

Invoking Gemini

Delegate tasks to Google's Gemini models when they offer advantages over Claude.

When to Use Gemini

Image generation:

  • Blog header images, illustrations, diagrams
  • Style-guided image creation (risograph, editorial, etc.)
  • Text rendering in images

Speech (TTS):

  • Narration, voice-over, read-aloud with style direction per line
  • A custom voice designed from a written description
  • Two-speaker dialogue

Structured outputs:

  • JSON Schema validation with property ordering guarantees
  • Pydantic model compliance
  • Strict schema adherence (enum values, required fields)

Cost optimization:

  • Parallel batch processing (Gemini 3 Flash is lightweight)
  • High-volume simple tasks

Multi-modal tasks:

  • Image analysis with JSON output
  • Video processing
  • Audio transcription with structure

Setup

uv pip install requests pydantic

Credentials — Option A (recommended): Cloudflare AI Gateway

Source /mnt/project/proxy.env with CF_ACCOUNT_ID, CF_GATEWAY_ID, CF_API_TOKEN. Requests route through Cloudflare AI Gateway, bypassing IP blocks. Google API key stored in gateway via BYOK.

Credentials — Option B: Direct Google API

If no proxy.env, falls back to direct: GOOGLE_API_KEY.txt or API_CREDENTIALS.json.

Image Generation

Generate images using Gemini's native image models. This is the primary way to create illustrations, blog headers, diagrams, and visual content.

Quick Start

import sys
sys.path.append('/mnt/skills/user/invoking-gemini/scripts')
from gemini_client import generate_image

# One call — returns {"path": "...", "caption": "..."} or None
result = generate_image("A watercolor painting of a mountain lake at sunset")
print(result["path"])  # /mnt/user-data/outputs/gemini_image_1740000000.png

Function Signature

generate_image(
    prompt: str,                    # The image description
    output_path: str = None,        # Auto-generates if omitted
    model: str = "nano-banana-2",   # Default: fast. Use "image-pro" for quality
    temperature: float = 0.7,       # 0.5-0.7 for diagrams, 0.7-0.8 for illustrations
) -> dict | None
# Returns: {"path": "/mnt/user-data/outputs/gemini_image_*.png", "caption": str|None}
# Returns None on failure

Model Selection

Alias Model Best For Cost/image
"nano-banana-2" or "image" gemini-3.1-flash-image-preview Fast iteration, drafts $0.067
"image-pro" or "nano-banana-pro" gemini-3-pro-image-preview Published content, text rendering $0.134

Complete Blog Header Example

import sys
sys.path.append('/mnt/skills/user/invoking-gemini/scripts')
from gemini_client import generate_image

# 1. Compose prompt with style prefix + subject
style_prefix = (
    "Style: Risograph-inspired editorial illustration. "
    "Visible halftone dot texture and slight color misregistration between layers. "
    "Limited ink palette: deep indigo, warm coral, and sage green on off-white paper. "
    "Layered transparency where colors overlap creates rich secondary tones. "
    "Modern and professional — the aesthetic of an indie design studio, not a fantasy novel. "
    "Generous whitespace. No photorealism, no glow effects, no cyberpunk. No text or labels."
)
subject = "A raven perched on a stack of books, observing a network graph"
prompt = f"{style_prefix}\n\nSubject: {subject}. Wide landscape format, suitable as a blog header."

# 2. Generate (use image-pro for published content)
result = generate_image(prompt, model="image-pro", temperature=0.75)

if result:
    print(f"Saved: {result['path']}")
    # 3. Present to user
    # present_files([result["path"]])

Prompt Patterns

  • Style prefix + subject: Prepend a style description, then describe the subject
  • Be specific about style: "Risograph-inspired editorial illustration" not "a nice picture"
  • Include composition: "Wide landscape format" / "centered, high contrast"
  • Text rendering: "A poster with the text 'SALE' in bold red letters" (works well with image-pro)
  • Negative constraints: "No photorealism, no glow effects" to avoid defaults

Custom Output Path

result = generate_image(
    "A logo for a coffee shop called 'Bean There'",
    output_path="/mnt/user-data/outputs/coffee_logo.png"
)

Speech Generation (TTS)

Gemini 3.8 Flash TTS and Flash-Lite TTS went GA on 2026-09-23. Output is WAV, 24 kHz mono 16-bit, SynthID-watermarked.

from gemini_client import generate_speech, design_voice, list_voices

r = generate_speech("Odin kept two ravens. <short pause> Huginn was thought.",
                    output_path="/tmp/line.wav", voice="Algenib",
                    style="quiet and dry, unhurried")
# {'path': '/tmp/line.wav', 'seconds': 4.2, 'audio_tokens': 135} or None

v = design_voice("A low, dry, quietly amused male voice with a faint rasp. "
                 "Soft southern British accent.", "narrator", gender="male",
                 language_code="en-GB")      # {'id': 'voice_...', 'sample_path': ...}
generate_speech("...", voice=v["id"])

lows = list_voices(gender="male", pitch="low")   # library of 2,089 prebuilt voices
  • Voices: 30 studio voices (Charon, Kore, Algenib "gravelly", Enceladus "breathy", ...) plus 2,059 persona voices with ids like en-gb-storyteller-4. list_voices() returns accent, pitch, gender and a description for each; the accent filter needs the exact string ("Winchester English"), so filter accents on the returned field.
  • Style: pass style= (a speech_metadata annotation). Do not prefix the text with "Style: ..." — the 3.8 models read the prefix aloud, and systemInstruction is rejected. Inline events go in the text: <laugh>, <sigh>, <breath>, <short pause>; CAPITALS stress a word.
  • Designed voices are stored (1-year expiry, 200 per project). The description sets baseline delivery too: "thoughtful pauses" in it produced 1–1.8 s mid-line pauses that no per-line style removed.
  • The model can change words. Seen in a 29-line narration: "Hmm, I get things wrong", "tell them" for "tell him". Anything with subtitles or a fixed script needs an ASR check (faster-whisper medium.en) and a retake.
  • Cost: about 32 audio tokens per second of speech, $9/M through 2026-12-31 on 3.8 Flash TTS ($6/M Lite), so a minute is about $0.02.
  • Voice replication (cloning from a 30 s sample plus a recorded consent clip) is not wired in, and is unavailable in the EEA, UK, Switzerland, India, Illinois and Texas.

Basic Text Usage

import sys
sys.path.append('/mnt/skills/user/invoking-gemini/scripts')
from gemini_client import invoke_gemini

response = invoke_gemini(
    prompt="Explain quantum computing in 3 bullet points",
    model="flash",  # gemini-3.8-flash (default)
)
print(response)

Structured Output

Use Pydantic models for guaranteed JSON Schema compliance:

from gemini_client import invoke_with_structured_output
from pydantic import BaseModel, Field

class BookAnalysis(BaseModel):
    title: str
    genre: str = Field(description="Primary genre")
    key_themes: list[str] = Field(max_length=5)
    rating: int = Field(ge=1, le=5)

result = invoke_with_structured_output(
    prompt="Analyze the book '1984' by George Orwell",
    pydantic_model=BookAnalysis
)
print(result.title)  # "1984"

Nested models are supported. Gemini's responseSchema rejects $ref/$defs, which pydantic emits for every nested model, so the client inlines them before sending:

class Finding(BaseModel):
    claim: str
    confidence: Literal["high", "medium", "low"]
    note: str | None = None

class Analysis(BaseModel):
    findings: list[Finding]     # nested — inlined for you
    gaps: list[str]

Budget output generously. Thinking tokens count against max_output_tokens (default 32768). Too low and the JSON truncates mid-object, which surfaces as a pydantic parse error rather than a length error — the client now detects finishReason=MAX_TOKENS and says so explicitly.

Parallel Invocation

from gemini_client import invoke_parallel

results = invoke_parallel(
    prompts=["Summarize Hamlet", "Summarize Macbeth", "Summarize Othello"],
    model="lite",  # gemini-3.5-flash-lite — cheap/fast tier for batch
)

Available Models

The current frontier Flash is gemini-3.8-flash (GA 2026-09-02), the default and the flash alias. Google shipped three Flash generations in six weeks: 3.6 (2026-07-21), 3.7 (2026-08-13), 3.8 (2026-09-02). Each stays callable under a pinned alias (flash-3.7, flash-3.6, flash-3.5, flash-3), and none has a shutdown date. gemini-3.1-flash-lite-preview from earlier docs is gone (shut down 2026-05-25).

The Pro tier is off routing. gemini-3.1-pro-preview costs 2.7× the input and 3.2× the output of 3.8 Flash at today's rates and loses to the 3.5+ Flash line on the coding and agentic benchmarks that matter here. Do not target it; the pro alias now resolves to gemini-3.8-flash, and "maximum reasoning" means thinking_level='high' on Flash.

Text / Reasoning Models

Model Alias Input/1M Output/1M Context Notes
gemini-3.8-flash flash $0.75 → $1.50 $3.75 → $7.50 1M in / 64K out Default. GA 2026-09-02. Current frontier Flash. Vs 3.7: Terminal-Bench 2.1 90.8% vs 81.6%, SWE-Bench Pro 61.6% vs 60.4%, SWE-Atlas 51.9% vs 48.0%, HLE flat (45.4% vs 45.7%). Google says it "works harder" at higher effort, so expect more thinking tokens per task. thinking_level is low/medium/high only — minimal returns HTTP 400 and the client downgrades it to low. Default medium spent 79 thinking tokens on a one-word reply (measured 2026-09-03); pass low for non-reasoning tasks.
gemini-3.7-flash flash-3.7 $0.75 → $1.50 $3.75 → $7.50 1M / 64K GA 2026-08-13. DeepSWE v1.1 65.3% vs 49.0% on 3.6, Terminal-Bench 2.1 85.8%. Same minimal restriction as 3.8. Google keeps it "fully supported for efficiency-first workloads".
gemini-3.6-flash flash-3.6 $0.75 → $1.50 $3.75 → $7.50 1M / 64K GA 2026-07-21. ~17% fewer output tokens than 3.5 Flash. Last Flash that accepts thinking_level='minimal' (verified 2026-09-03).
gemini-3.5-flash flash-3.5 $1.50 $9.00 1M GA 2026-05-19. Google's model list now labels it "legacy". Accepts minimal. Costs more on output than 3.6–3.8.
gemini-3-flash-preview flash-3 $0.30 $2.50 1M Older preview Flash, kept for back compat. Google's listed migration target for it is gemini-3.6-flash; no shutdown date.
gemini-3.1-pro-preview — $2.00 (≤200K) / $4.00 $12.00 / $18.00 1M DEPRECATED from routing (2026-09-03). Price/quality dominated by 3.6+ Flash; 3.5 Flash already beat it on most coding/agentic benchmarks. ID stays callable for pinned code. pro now resolves to gemini-3.8-flash. 3.5 Pro was announced at I/O 2026-05-19 for June and is still absent from the API as of 2026-09-03; it gets the same price/quality test before any alias points at it.
gemini-3.5-flash-lite lite $0.30 $2.50 1M Cheap/bulk tier. GA 2026-07-21. Fastest 3.5-class (350 output tok/sec); beats gemini-3-flash on SWE-Bench Pro and OSWorld-Verified.
gemini-2.5-flash stable-flash $0.30 $2.50 1M DEPRECATED — 2025-era generation, do not route here.
gemini-2.5-flash-lite — $0.10 $0.40 1M DEPRECATED — cheaper, but a 2025-era generation. lite now resolves to gemini-3.5-flash-lite.
gemini-2.5-pro stable-pro $1.25 (≤200K) / $2.50 $10.00 / $20.00 1M DEPRECATED — 2025-era generation, do not route here.

$0.75 → $1.50 means introductory pricing: Google's pricing page (fetched 2026-09-03) lists 3.6, 3.7 and 3.8 Flash at $0.75 in / $3.75 out through 2026-12-31 and $1.50 / $7.50 from 2027-01-01. Context caching is $0.075 → $0.15; Batch is half of standard. Output prices include thinking tokens.

Image Models

Model Alias Input/1M Per Image
gemini-3.1-flash-image-preview image, nano-banana-2 $0.25 $0.067
gemini-3-pro-image-preview image-pro, nano-banana-pro $2.00 $0.134

Speech Models

Model Alias Input/1M Output/1M (audio) Notes
gemini-3.8-flash-tts tts $0.50 → $1.00 $9.00 → $18.00 GA 2026-09-23. Expressive, 130 languages, 2 speakers. Use generate_speech(), not invoke_gemini().
gemini-3.8-flash-lite-tts tts-lite $0.50 → $1.00 $6.00 → $12.00 GA 2026-09-23. Bulk / read-aloud, 101 languages. Replaces gemini-3.1-flash-tts-preview ($1 / $20).

Speech aliases live in SPEECH_ALIASES, not MODEL_ALIASES, so a text call can never resolve to an audio model.

See references/models.md for full details.

Thinking Budget (Gemini 3.x)

Gemini 3.x models reason before responding. The parameter changed in 2026: integer thinking_budget is gone; use string thinking_level ∈ {minimal, low, medium, high}. Default for 3.5–3.8 Flash is medium. For transcription / classification / extraction tasks, pass thinking_level='minimal' or the model will silently spend output tokens on reasoning (symptom: empty response with finishReason=MAX_TOKENS).

3.7 and 3.8 Flash reject minimal with HTTP 400 (Thinking level MINIMAL is not supported for this model); low is their floor. The client downgrades minimal to low on those two models and prints a note to stderr, so existing callers keep working. Measured on 3.8 (2026-09-03): low spent 0 thinking tokens on a one-word reply, the default medium spent 79. On 3.7, low still spent 45–88, and a max_output_tokens=50 call at low hit MAX_TOKENS and returned None, so budget output generously there. If a job needs a true no-thinking pass, pin flash-3.6 or lite, which still accept minimal.

response = invoke_gemini(
    prompt="Transcribe this image.",
    model="flash",
    image_path="/tmp/screenshot.png",
    max_output_tokens=4000,
    thinking_level="minimal",  # don't burn output budget on reasoning
)

Error Handling

response = invoke_gemini(prompt="...", model="flash")
if response is None:
    print("API call failed — check credentials")

result = generate_image("...")
if result is None:
    print("Image generation failed — check credentials or try again")

Common issues: Missing API key → see Setup. Rate limit → auto-retries with backoff. Network error → returns None.

Advanced Features

Custom Generation Config

response = invoke_gemini(
    prompt="Write a haiku",
    model="flash",                  # gemini-3.8-flash
    temperature=0.9,
    max_output_tokens=200,
    top_p=0.95,
    thinking_level="low",           # haiku is short; modest reasoning is fine
)

Multi-modal Input

from pydantic import BaseModel
from gemini_client import invoke_with_structured_output

class ImageDescription(BaseModel):
    objects: list[str]
    scene: str
    colors: list[str]

result = invoke_with_structured_output(
    prompt="Describe this image",
    pydantic_model=ImageDescription,
    image_path="/mnt/user-data/uploads/photo.jpg"
)

See references/advanced.md for more patterns.

Troubleshooting

"No credentials configured": Create /mnt/project/proxy.env with CF credentials, or add GOOGLE_API_KEY.txt.

CF Gateway 401/403: Verify CF_API_TOKEN has AI Gateway permissions. If not using BYOK, add GOOGLE_API_KEY to proxy.env.

Import errors: uv pip install requests pydantic

Image generation returns None: Check credentials. If persistent, try model="nano-banana-2" (more reliable than image-pro). Check for content policy blocks in error output.

Files (claude-skills)
  • references
    • advanced.md 10 KB
      # Advanced Patterns
      
      Advanced usage patterns for Gemini integration.
      
      ## Multi-Modal Processing
      
      ### Image Analysis with Structure
      
      ```python
      from pydantic import BaseModel, Field
      from gemini_client import invoke_with_structured_output
      
      class ProductAnalysis(BaseModel):
          product_name: str
          category: str
          colors: list[str]
          estimated_price_range: str
          condition: str = Field(description="new, used, or refurbished")
      
      result = invoke_with_structured_output(
          prompt="Analyze this product image. Identify the product, category, colors, and estimate price range.",
          pydantic_model=ProductAnalysis,
          image_path="/mnt/user-data/uploads/product.jpg"
      )
      
      print(f"Product: {result.product_name}")
      print(f"Category: {result.category}")
      print(f"Colors: {', '.join(result.colors)}")
      ```
      
      ### Batch Image Processing
      
      ```python
      from pathlib import Path
      from gemini_client import invoke_parallel
      
      image_dir = Path("/mnt/user-data/uploads/photos")
      image_files = list(image_dir.glob("*.jpg"))
      
      prompts = [
          f"Describe this image in one sentence"
          for _ in image_files
      ]
      
      # Note: parallel invocation doesn't support images directly
      # Process sequentially with progress
      for i, image_path in enumerate(image_files):
          result = invoke_gemini(
              prompt="Describe this image briefly",
              image_path=str(image_path)
          )
          print(f"[{i+1}/{len(image_files)}] {image_path.name}: {result}")
      ```
      
      ## Hybrid Workflows
      
      ### Claude Plans, Gemini Executes
      
      Use Claude's reasoning for planning, Gemini for structured execution:
      
      ```python
      # 1. Claude (you) analyzes requirements and creates extraction plan
      # 2. Gemini executes structured extractions
      
      from pydantic import BaseModel
      from gemini_client import invoke_with_structured_output
      
      class ContactInfo(BaseModel):
          name: str
          email: str
          phone: str
          company: str
      
      # Gemini extracts structured data
      documents = [...]  # List of document paths
      contacts = []
      
      for doc_path in documents:
          with open(doc_path) as f:
              doc_text = f.read()
      
          result = invoke_with_structured_output(
              prompt=f"Extract contact information from:\n\n{doc_text}",
              pydantic_model=ContactInfo
          )
          contacts.append(result)
      
      # 3. Claude synthesizes and analyzes results
      # ... your analysis code here ...
      ```
      
      ### Parallel Processing with Different Models
      
      ```python
      from concurrent.futures import ThreadPoolExecutor
      from gemini_client import invoke_gemini
      
      def process_with_model(text: str, model: str) -> str:
          return invoke_gemini(text, model=model)
      
      # Compare outputs from different models
      text = "Analyze the sentiment of this review: ..."
      
      with ThreadPoolExecutor(max_workers=3) as executor:
          futures = {
              executor.submit(process_with_model, text, "gemini-3-flash-preview"): "3-flash",
              executor.submit(process_with_model, text, "gemini-2.5-flash"): "2.5-flash",
              executor.submit(process_with_model, text, "gemini-2.5-pro"): "2.5-pro",
          }
      
          for future in as_completed(futures):
              model_name = futures[future]
              result = future.result()
              print(f"{model_name}: {result}")
      ```
      
      ## Complex Schema Patterns
      
      ### Nested Structures
      
      ```python
      from pydantic import BaseModel
      from typing import Optional
      
      class Address(BaseModel):
          street: str
          city: str
          state: str
          zip_code: str
          country: str = "USA"
      
      class Person(BaseModel):
          name: str
          age: Optional[int] = None
          email: str
          address: Address
      
      result = invoke_with_structured_output(
          prompt="Extract person info: John Doe, 30, john@example.com, 123 Main St, Springfield, IL 62701",
          pydantic_model=Person
      )
      ```
      
      ### Enums and Constraints
      
      ```python
      from pydantic import BaseModel, Field, validator
      from enum import Enum
      
      class Priority(str, Enum):
          LOW = "low"
          MEDIUM = "medium"
          HIGH = "high"
          URGENT = "urgent"
      
      class Task(BaseModel):
          title: str = Field(max_length=100)
          description: str
          priority: Priority
          estimated_hours: int = Field(ge=1, le=100)
          tags: list[str] = Field(max_length=10)
      
          @validator('tags')
          def validate_tags(cls, v):
              if len(v) > 10:
                  raise ValueError('Maximum 10 tags allowed')
              return v
      
      result = invoke_with_structured_output(
          prompt="Create a task: Fix login bug - Users can't login. This is urgent. Should take 4 hours. Tags: bug, auth, security",
          pydantic_model=Task
      )
      ```
      
      ### Lists and Arrays
      
      ```python
      from pydantic import BaseModel
      
      class Ingredient(BaseModel):
          name: str
          quantity: str
          unit: str
      
      class Recipe(BaseModel):
          title: str
          servings: int
          prep_time_minutes: int
          cook_time_minutes: int
          ingredients: list[Ingredient]
          instructions: list[str]
      
      recipe_text = """
      Make pasta carbonara.
      Serves 4. Prep: 10 min. Cook: 20 min.
      Ingredients: 400g spaghetti, 200g pancetta, 4 eggs, 100g parmesan, black pepper.
      Instructions:
      1. Boil pasta
      2. Fry pancetta
      3. Mix eggs and cheese
      4. Combine all
      """
      
      result = invoke_with_structured_output(
          prompt=f"Extract recipe from:\n{recipe_text}",
          pydantic_model=Recipe
      )
      ```
      
      ## Error Recovery
      
      ### Retry with Schema Relaxation
      
      ```python
      from pydantic import BaseModel
      
      class StrictData(BaseModel):
          field1: str
          field2: int
          field3: list[str]
      
      # Try strict schema first
      result = invoke_with_structured_output(prompt, StrictData)
      
      if not result:
          # Retry with relaxed schema
          class RelaxedData(BaseModel):
              field1: str
              field2: Optional[int] = None
              field3: Optional[list[str]] = []
      
          result = invoke_with_structured_output(prompt, RelaxedData)
      ```
      
      ### Validation and Correction
      
      ```python
      from pydantic import BaseModel, ValidationError
      
      result = invoke_with_structured_output(prompt, MyModel)
      
      if result:
          try:
              # Additional validation
              assert len(result.items) > 0, "Must have at least one item"
              assert result.total > 0, "Total must be positive"
          except AssertionError as e:
              print(f"Validation failed: {e}")
              # Retry with corrected prompt
              corrected_prompt = f"{prompt}\n\nIMPORTANT: {e}"
              result = invoke_with_structured_output(corrected_prompt, MyModel)
      ```
      
      ## Performance Optimization
      
      ### Prompt Caching
      
      For repeated prompts with same prefix:
      
      ```python
      # Share common context across requests
      base_context = """
      You are analyzing customer reviews for sentiment.
      Categories: positive, neutral, negative
      Format: JSON with 'sentiment' and 'confidence' fields
      """
      
      reviews = ["Review 1...", "Review 2...", "Review 3..."]
      
      for review in reviews:
          full_prompt = f"{base_context}\n\nReview: {review}"
          result = invoke_gemini(full_prompt)
      ```
      
      ### Batch Size Tuning
      
      ```python
      from gemini_client import invoke_parallel
      
      # Process in optimal batches
      all_prompts = [...]  # 1000 prompts
      batch_size = 50  # Tune based on rate limits
      
      results = []
      for i in range(0, len(all_prompts), batch_size):
          batch = all_prompts[i:i+batch_size]
          batch_results = invoke_parallel(batch, max_workers=10)
          results.extend(batch_results)
      
          # Rate limit breathing room
          if i + batch_size < len(all_prompts):
              time.sleep(2)
      ```
      
      ### Temperature Tuning
      
      ```python
      # Factual extraction: low temperature
      factual = invoke_gemini(
          "Extract the date from: Meeting scheduled for March 15th",
          temperature=0.1
      )
      
      # Creative generation: high temperature
      creative = invoke_gemini(
          "Write a creative tagline for a coffee shop",
          temperature=0.9
      )
      
      # Balanced: medium temperature
      balanced = invoke_gemini(
          "Summarize this article",
          temperature=0.7
      )
      ```
      
      ## Cost Optimization
      
      ### Token Counting
      
      ```python
      import google.generativeai as genai
      
      model = genai.GenerativeModel("gemini-3-flash-preview")
      
      # Count tokens before sending
      token_count = model.count_tokens("Your prompt here")
      print(f"Input tokens: {token_count.total_tokens}")
      
      # Estimate cost
      input_cost = (token_count.total_tokens / 1_000_000) * 0.50
      print(f"Estimated input cost: ${input_cost:.4f}")
      ```
      
      ### Prompt Compression
      
      ```python
      # Verbose (wasteful)
      verbose_prompt = """
      Please analyze the following text and extract the following information:
      - The name of the person
      - Their email address
      - Their phone number
      - Their company name
      
      Text:
      John Doe works at Acme Corp. Email: john@acme.com, Phone: 555-0100
      """
      
      # Concise (efficient)
      concise_prompt = """
      Extract: name, email, phone, company
      
      John Doe works at Acme Corp. Email: john@acme.com, Phone: 555-0100
      """
      
      # With structured output, schema provides the context
      result = invoke_with_structured_output(
          prompt="John Doe works at Acme Corp. Email: john@acme.com, Phone: 555-0100",
          pydantic_model=ContactInfo  # Schema explains structure
      )
      ```
      
      ## Integration with Other Tools
      
      ### Save Results to Database
      
      ```python
      import sqlite3
      from gemini_client import invoke_with_structured_output
      
      conn = sqlite3.connect('/home/claude/results.db')
      cursor = conn.cursor()
      
      cursor.execute('''
          CREATE TABLE IF NOT EXISTS contacts (
              name TEXT, email TEXT, phone TEXT, company TEXT
          )
      ''')
      
      documents = [...]
      for doc in documents:
          result = invoke_with_structured_output(doc, ContactInfo)
          if result:
              cursor.execute(
                  'INSERT INTO contacts VALUES (?, ?, ?, ?)',
                  (result.name, result.email, result.phone, result.company)
              )
      
      conn.commit()
      ```
      
      ### Export to CSV
      
      ```python
      import csv
      from gemini_client import invoke_with_structured_output
      
      results = []
      for item in items:
          result = invoke_with_structured_output(item, DataModel)
          if result:
              results.append(result.dict())
      
      # Write to CSV
      with open('/mnt/user-data/outputs/results.csv', 'w', newline='') as f:
          if results:
              writer = csv.DictWriter(f, fieldnames=results[0].keys())
              writer.writeheader()
              writer.writerows(results)
      ```
      
      ### Combine with Pandas
      
      ```python
      import pandas as pd
      from gemini_client import invoke_with_structured_output
      
      # Process data with Gemini, analyze with pandas
      data = []
      for text in texts:
          result = invoke_with_structured_output(text, StructuredData)
          if result:
              data.append(result.dict())
      
      df = pd.DataFrame(data)
      print(df.describe())
      print(df.groupby('category').size())
      ```
      
    • examples.md 21.1 KB
      # Examples
      
      Comprehensive examples of Gemini usage patterns.
      
      ## Example 1: Document Data Extraction
      
      Extract structured data from unstructured documents.
      
      ```python
      from pydantic import BaseModel, Field
      from gemini_client import invoke_with_structured_output
      from pathlib import Path
      
      class InvoiceData(BaseModel):
          invoice_number: str
          date: str
          vendor: str
          total_amount: float
          line_items: list[dict] = Field(description="List of items with description and price")
      
      invoice_dir = Path("/mnt/user-data/uploads/invoices")
      results = []
      
      for invoice_file in invoice_dir.glob("*.txt"):
          with open(invoice_file) as f:
              invoice_text = f.read()
      
          data = invoke_with_structured_output(
              prompt=f"Extract invoice data:\n\n{invoice_text}",
              pydantic_model=InvoiceData
          )
      
          if data:
              results.append({
                  'file': invoice_file.name,
                  'invoice_number': data.invoice_number,
                  'vendor': data.vendor,
                  'total': data.total_amount
              })
      
      # Save results
      import pandas as pd
      df = pd.DataFrame(results)
      df.to_csv('/mnt/user-data/outputs/invoice_summary.csv', index=False)
      ```
      
      ## Example 2: Batch Classification
      
      Classify large datasets efficiently.
      
      ```python
      from pydantic import BaseModel
      from enum import Enum
      from gemini_client import invoke_parallel, invoke_with_structured_output
      
      class Sentiment(str, Enum):
          POSITIVE = "positive"
          NEUTRAL = "neutral"
          NEGATIVE = "negative"
      
      class ReviewAnalysis(BaseModel):
          sentiment: Sentiment
          confidence: float = Field(ge=0.0, le=1.0)
          key_topics: list[str] = Field(max_length=5)
      
      # Load reviews
      import pandas as pd
      df = pd.read_csv('/mnt/user-data/uploads/reviews.csv')
      
      results = []
      for idx, row in df.iterrows():
          analysis = invoke_with_structured_output(
              prompt=f"Analyze this review: {row['review_text']}",
              pydantic_model=ReviewAnalysis,
              temperature=0.3  # Low temp for consistent classification
          )
      
          if analysis:
              results.append({
                  'review_id': row['id'],
                  'sentiment': analysis.sentiment.value,
                  'confidence': analysis.confidence,
                  'topics': ', '.join(analysis.key_topics)
              })
      
          # Progress
          if (idx + 1) % 10 == 0:
              print(f"Processed {idx + 1}/{len(df)} reviews")
      
      # Save results
      results_df = pd.DataFrame(results)
      results_df.to_csv('/mnt/user-data/outputs/sentiment_analysis.csv', index=False)
      
      # Summary statistics
      print("\nSentiment Distribution:")
      print(results_df['sentiment'].value_counts())
      print(f"\nAverage Confidence: {results_df['confidence'].mean():.2f}")
      ```
      
      ## Example 3: Multi-Modal Product Catalog
      
      Create structured product catalog from images.
      
      ```python
      from pydantic import BaseModel, Field
      from gemini_client import invoke_with_structured_output
      from pathlib import Path
      
      class Product(BaseModel):
          name: str
          category: str
          description: str = Field(max_length=200)
          primary_color: str
          additional_colors: list[str] = []
          estimated_price_tier: str = Field(description="budget, mid-range, or premium")
          key_features: list[str] = Field(max_length=5)
      
      product_images = Path("/mnt/user-data/uploads/products")
      catalog = []
      
      for img_path in product_images.glob("*.jpg"):
          product = invoke_with_structured_output(
              prompt="""
              Analyze this product image and provide:
              - Product name and category
              - Brief description
              - Colors visible
              - Estimated price tier (budget/mid-range/premium)
              - Key features
              """,
              pydantic_model=Product,
              image_path=str(img_path)
          )
      
          if product:
              catalog.append({
                  'image': img_path.name,
                  **product.dict()
              })
              print(f"✓ {img_path.name}: {product.name}")
      
      # Export catalog
      import json
      with open('/mnt/user-data/outputs/product_catalog.json', 'w') as f:
          json.dump(catalog, f, indent=2)
      
      print(f"\nProcessed {len(catalog)} products")
      ```
      
      ## Example 4: Resume Parser
      
      Extract structured data from resumes.
      
      ```python
      from pydantic import BaseModel, Field
      from typing import Optional
      
      class Education(BaseModel):
          degree: str
          institution: str
          year: Optional[str] = None
      
      class Experience(BaseModel):
          title: str
          company: str
          duration: str
          responsibilities: list[str]
      
      class Resume(BaseModel):
          name: str
          email: str
          phone: Optional[str] = None
          summary: str = Field(max_length=300)
          skills: list[str]
          education: list[Education]
          experience: list[Experience]
      
      resume_files = Path("/mnt/user-data/uploads/resumes")
      parsed_resumes = []
      
      for resume_file in resume_files.glob("*.txt"):
          with open(resume_file) as f:
              resume_text = f.read()
      
          parsed = invoke_with_structured_output(
              prompt=f"Parse this resume:\n\n{resume_text}",
              pydantic_model=Resume,
              temperature=0.2  # Low temp for accuracy
          )
      
          if parsed:
              parsed_resumes.append({
                  'file': resume_file.name,
                  'candidate': parsed.name,
                  'email': parsed.email,
                  'skills_count': len(parsed.skills),
                  'years_experience': len(parsed.experience),
                  'education_level': parsed.education[0].degree if parsed.education else 'None'
              })
      
      # Create summary
      import pandas as pd
      df = pd.DataFrame(parsed_resumes)
      df.to_csv('/mnt/user-data/outputs/resume_summary.csv', index=False)
      
      # Skill frequency analysis
      all_skills = []
      for resume in parsed_resumes:
          all_skills.extend(resume.get('skills', []))
      
      from collections import Counter
      skill_counts = Counter(all_skills)
      print("\nTop 10 Skills:")
      for skill, count in skill_counts.most_common(10):
          print(f"  {skill}: {count}")
      ```
      
      ## Example 5: Meeting Notes Summarization
      
      Batch process meeting notes into structured summaries.
      
      ```python
      from pydantic import BaseModel, Field
      from datetime import datetime
      
      class ActionItem(BaseModel):
          task: str
          assignee: str
          due_date: Optional[str] = None
          priority: str = Field(description="high, medium, or low")
      
      class MeetingSummary(BaseModel):
          meeting_date: str
          attendees: list[str]
          key_topics: list[str] = Field(max_length=5)
          decisions: list[str]
          action_items: list[ActionItem]
          next_meeting: Optional[str] = None
      
      notes_dir = Path("/mnt/user-data/uploads/meeting_notes")
      summaries = []
      
      for notes_file in sorted(notes_dir.glob("*.txt")):
          with open(notes_file) as f:
              notes = f.read()
      
          summary = invoke_with_structured_output(
              prompt=f"Summarize these meeting notes:\n\n{notes}",
              pydantic_model=MeetingSummary
          )
      
          if summary:
              summaries.append(summary)
              print(f"✓ {notes_file.name}: {len(summary.action_items)} action items")
      
      # Generate action items report
      all_action_items = []
      for summary in summaries:
          for item in summary.action_items:
              all_action_items.append({
                  'meeting_date': summary.meeting_date,
                  'task': item.task,
                  'assignee': item.assignee,
                  'due_date': item.due_date,
                  'priority': item.priority
              })
      
      import pandas as pd
      df = pd.DataFrame(all_action_items)
      df.to_csv('/mnt/user-data/outputs/action_items.csv', index=False)
      
      # Group by assignee
      print("\nAction Items by Assignee:")
      print(df.groupby('assignee').size().sort_values(ascending=False))
      ```
      
      ## Example 6: Parallel Translation
      
      Translate content to multiple languages in parallel.
      
      ```python
      from gemini_client import invoke_parallel
      
      content = """
      Welcome to our product! This innovative solution helps you
      manage your tasks efficiently and collaborate with your team.
      """
      
      languages = [
          "Spanish", "French", "German", "Italian", "Portuguese",
          "Japanese", "Korean", "Chinese", "Arabic", "Russian"
      ]
      
      prompts = [
          f"Translate to {lang} (output only the translation):\n\n{content}"
          for lang in languages
      ]
      
      translations = invoke_parallel(
          prompts=prompts,
          model="gemini-3-flash-preview",
          temperature=0.3,
          max_workers=10
      )
      
      # Create translation table
      results = []
      for lang, translation in zip(languages, translations):
          if translation:
              results.append({
                  'language': lang,
                  'translation': translation.strip()
              })
      
      # Export
      import pandas as pd
      df = pd.DataFrame(results)
      df.to_csv('/mnt/user-data/outputs/translations.csv', index=False)
      
      print(f"Translated to {len(results)} languages")
      ```
      
      ## Example 7: Code Documentation Generator
      
      Generate structured documentation from code.
      
      ```python
      from pydantic import BaseModel, Field
      
      class FunctionDoc(BaseModel):
          function_name: str
          description: str = Field(max_length=200)
          parameters: list[dict] = Field(description="List with name, type, description")
          return_type: str
          return_description: str
          example_usage: str
          complexity: str = Field(description="O(n), O(log n), etc.")
      
      # Read source files
      code_dir = Path("/mnt/user-data/uploads/source_code")
      documentation = []
      
      for code_file in code_dir.glob("*.py"):
          with open(code_file) as f:
              code = f.read()
      
          # Extract functions (simplified)
          import re
          functions = re.findall(r'def\s+(\w+)\s*\([^)]*\):[^}]*?(?=\ndef|\Z)', code, re.DOTALL)
      
          for func_code in functions[:5]:  # Limit to first 5 functions
              doc = invoke_with_structured_output(
                  prompt=f"Generate documentation for this Python function:\n\n{func_code}",
                  pydantic_model=FunctionDoc
              )
      
              if doc:
                  documentation.append(doc.dict())
      
      # Export as JSON
      import json
      with open('/mnt/user-data/outputs/api_documentation.json', 'w') as f:
          json.dump(documentation, f, indent=2)
      
      # Generate markdown
      with open('/mnt/user-data/outputs/API_DOCS.md', 'w') as f:
          f.write("# API Documentation\n\n")
          for doc in documentation:
              f.write(f"## {doc['function_name']}\n\n")
              f.write(f"{doc['description']}\n\n")
              f.write(f"**Returns:** `{doc['return_type']}` - {doc['return_description']}\n\n")
              f.write(f"**Complexity:** {doc['complexity']}\n\n")
              f.write(f"**Example:**\n```python\n{doc['example_usage']}\n```\n\n")
      ```
      
      ## Example 8: Financial Report Analysis
      
      Extract key metrics from financial reports.
      
      ```python
      from pydantic import BaseModel, Field
      
      class FinancialMetrics(BaseModel):
          company_name: str
          reporting_period: str
          revenue: float = Field(description="In millions")
          net_income: float = Field(description="In millions")
          profit_margin: float = Field(ge=0, le=100, description="Percentage")
          key_highlights: list[str] = Field(max_length=5)
          risks: list[str] = Field(max_length=3)
      
      reports_dir = Path("/mnt/user-data/uploads/financial_reports")
      metrics = []
      
      for report_file in reports_dir.glob("*.txt"):
          with open(report_file) as f:
              report_text = f.read()
      
          data = invoke_with_structured_output(
              prompt=f"""
              Extract financial metrics from this quarterly report.
              All monetary values should be in millions.
      
              {report_text}
              """,
              pydantic_model=FinancialMetrics,
              temperature=0.1  # Very low for numerical accuracy
          )
      
          if data:
              metrics.append(data.dict())
      
      # Analysis
      import pandas as pd
      df = pd.DataFrame(metrics)
      
      print("\nFinancial Summary:")
      print(f"Average Revenue: ${df['revenue'].mean():.2f}M")
      print(f"Average Net Income: ${df['net_income'].mean():.2f}M")
      print(f"Average Profit Margin: {df['profit_margin'].mean():.2f}%")
      
      df.to_csv('/mnt/user-data/outputs/financial_metrics.csv', index=False)
      ```
      
      ## Example 9: Survey Response Analysis
      
      Analyze open-ended survey responses.
      
      ```python
      from pydantic import BaseModel
      from enum import Enum
      
      class Satisfaction(str, Enum):
          VERY_SATISFIED = "very_satisfied"
          SATISFIED = "satisfied"
          NEUTRAL = "neutral"
          DISSATISFIED = "dissatisfied"
          VERY_DISSATISFIED = "very_dissatisfied"
      
      class SurveyAnalysis(BaseModel):
          satisfaction: Satisfaction
          main_sentiment: str = Field(max_length=100)
          mentioned_features: list[str] = Field(description="Features mentioned positively or negatively")
          pain_points: list[str] = Field(description="Problems or complaints")
          suggestions: list[str] = Field(description="Improvement suggestions")
      
      # Load survey data
      import pandas as pd
      df = pd.read_csv('/mnt/user-data/uploads/survey_responses.csv')
      
      analyses = []
      for idx, row in df.iterrows():
          analysis = invoke_with_structured_output(
              prompt=f"Analyze this survey response:\n\nQuestion: {row['question']}\nAnswer: {row['response']}",
              pydantic_model=SurveyAnalysis
          )
      
          if analysis:
              analyses.append({
                  'response_id': row['id'],
                  'satisfaction': analysis.satisfaction.value,
                  'sentiment': analysis.main_sentiment,
                  'features': ', '.join(analysis.mentioned_features),
                  'pain_points': ', '.join(analysis.pain_points),
                  'suggestions': ', '.join(analysis.suggestions)
              })
      
      results_df = pd.DataFrame(analyses)
      results_df.to_csv('/mnt/user-data/outputs/survey_analysis.csv', index=False)
      
      # Aggregate insights
      print("\nSatisfaction Distribution:")
      print(results_df['satisfaction'].value_counts())
      
      all_pain_points = [p for points in analyses for p in points.get('pain_points', [])]
      from collections import Counter
      print("\nTop Pain Points:")
      for pain, count in Counter(all_pain_points).most_common(5):
          print(f"  {pain}: {count}")
      ```
      
      ## Example 10: Hybrid Claude + Gemini Workflow
      
      Claude does complex reasoning, Gemini does structured extraction.
      
      ```python
      from pydantic import BaseModel
      from gemini_client import invoke_with_structured_output, invoke_parallel
      
      # Step 1: Claude (you) analyzes the dataset and determines categories
      # Assume you've identified key categories for classification
      
      class DataPoint(BaseModel):
          text: str
          category: str
          confidence: float
          key_terms: list[str]
      
      # Step 2: Gemini extracts structured data at scale
      raw_data = pd.read_csv('/mnt/user-data/uploads/raw_data.csv')
      
      structured_data = []
      batch_size = 50
      
      for i in range(0, len(raw_data), batch_size):
          batch = raw_data.iloc[i:i+batch_size]
      
          for idx, row in batch.iterrows():
              result = invoke_with_structured_output(
                  prompt=f"Classify and extract from: {row['text']}",
                  pydantic_model=DataPoint,
                  temperature=0.3
              )
      
              if result:
                  structured_data.append(result.dict())
      
          print(f"Processed {min(i+batch_size, len(raw_data))}/{len(raw_data)}")
      
      # Step 3: Claude analyzes the structured results
      df = pd.DataFrame(structured_data)
      
      # Your analysis here:
      # - Identify patterns
      # - Generate insights
      # - Create visualizations
      # - Produce final report
      
      print(f"\nProcessed {len(structured_data)} items")
      print(f"Average confidence: {df['confidence'].mean():.2f}")
      print("\nCategory distribution:")
      print(df['category'].value_counts())
      ```
      
      ## Example 11: Blog Header Image Generation
      
      Generate a styled blog header image with a single call.
      
      ```python
      import sys
      sys.path.append('/mnt/skills/user/invoking-gemini/scripts')
      from gemini_client import generate_image
      
      # Style prefix — prepend to any subject for consistent visual identity
      RISO_ILLUSTRATION = (
          "Style: Risograph-inspired editorial illustration. "
          "Visible halftone dot texture and slight color misregistration between layers. "
          "Limited ink palette: deep indigo, warm coral, and sage green on off-white paper. "
          "Layered transparency where colors overlap creates rich secondary tones. "
          "Modern and professional — the aesthetic of an indie design studio, not a fantasy novel. "
          "Generous whitespace. No photorealism, no glow effects, no cyberpunk. No text or labels."
      )
      
      RISO_DIAGRAM = (
          "Style: Risograph-inspired technical diagram. "
          "Visible halftone dot texture and slight color misregistration. "
          "Limited ink palette: deep indigo for primary shapes and text, "
          "warm coral for highlights and active elements, "
          "sage green for secondary elements and connections. "
          "Off-white paper background. Clean layout with generous spacing. "
          "Professional and readable."
      )
      
      # Generate illustration header
      subject = "A raven perched on a network graph, watching data flow between nodes"
      result = generate_image(
          f"{RISO_ILLUSTRATION}\n\nSubject: {subject}. Wide landscape format, suitable as a blog header.",
          model="image-pro",       # Use image-pro for published content
          temperature=0.75,
          output_path="/mnt/user-data/outputs/blog_header.png"
      )
      
      if result:
          print(f"Header saved: {result['path']}")
      else:
          print("Generation failed — retry or check credentials")
      ```
      
      Key patterns:
      - **Style prefix + subject composition**: The prefix sets visual rules, the subject describes content
      - **`image-pro` for published content**: Better quality and text rendering than default
      - **Temperature 0.7-0.8 for illustrations**: Allows creative variation while staying on-style
      - **Explicit output path**: Control where the file lands for downstream use
      
      ## Example 12: Technical Diagram Generation
      
      Generate a styled technical diagram with text labels.
      
      ```python
      from gemini_client import generate_image
      
      RISO_DIAGRAM = (
          "Style: Risograph-inspired technical diagram. "
          "Visible halftone dot texture and slight color misregistration. "
          "Limited ink palette: deep indigo for primary shapes and text, "
          "warm coral for highlights and active elements, "
          "sage green for secondary elements and connections. "
          "Off-white paper background. Clean layout with generous spacing. "
          "Professional and readable."
      )
      
      result = generate_image(
          f"{RISO_DIAGRAM}\n\n"
          "A flowchart showing: User Request → Claude Planning → Gemini Extraction → "
          "Claude Synthesis → Final Report. Highlight the Gemini step in coral. "
          "Wide landscape format.",
          model="image-pro",
          temperature=0.6,  # Lower temp for diagrams — more precise
      )
      ```
      
      ## Example 13: Batch Image Generation with Variants
      
      Generate multiple variants of the same concept for selection.
      
      ```python
      from gemini_client import generate_image
      
      subjects = [
          "A raven carrying a scroll through a library of glowing books",
          "A raven assembling puzzle pieces that form a constellation",
          "A raven observing its reflection in a pool of data streams",
      ]
      
      results = []
      for i, subject in enumerate(subjects):
          result = generate_image(
              f"Style: Risograph-inspired editorial illustration with deep indigo, "
              f"warm coral, sage green on off-white. No text.\n\n"
              f"Subject: {subject}. Wide landscape format.",
              model="nano-banana-2",   # Use fast model for drafts
              temperature=0.8,
              output_path=f"/mnt/user-data/outputs/variant_{i+1}.png"
          )
          if result:
              results.append(result["path"])
              print(f"Variant {i+1}: {result['path']}")
      
      # Present all variants for user to pick
      # present_files(results)
      ```
      
      ## Example 14: Image Generation with Structured Feedback Loop
      
      Generate an image, analyze it with Gemini vision, regenerate if needed.
      
      ```python
      from gemini_client import generate_image, invoke_with_structured_output
      from pydantic import BaseModel, Field
      
      class ImageQuality(BaseModel):
          has_text_artifacts: bool = Field(description="Unwanted text in the image")
          style_match: int = Field(ge=1, le=5, description="How well it matches risograph style")
          composition_score: int = Field(ge=1, le=5)
          issues: list[str] = Field(description="Problems to fix in re-generation")
      
      # Generate
      result = generate_image(
          "Style: Risograph editorial. Deep indigo, coral, sage green.\n\n"
          "Subject: A raven on a circuit board. Wide landscape.",
          model="image-pro",
          temperature=0.75,
      )
      
      if result:
          # Analyze with vision
          quality = invoke_with_structured_output(
              prompt="Evaluate this image for: unwanted text artifacts, "
                     "risograph style fidelity (halftone dots, misregistration, limited palette), "
                     "and composition quality.",
              pydantic_model=ImageQuality,
              image_path=result["path"]
          )
          print(f"Style match: {quality.style_match}/5")
          print(f"Issues: {quality.issues}")
      
          # Re-generate with fixes if needed
          if quality.style_match < 3 or quality.has_text_artifacts:
              fix_prompt = f"Fix: {', '.join(quality.issues)}. "
              # Regenerate with adjusted prompt...
      ```
      
      ## Best Practices from Examples
      
      **1. Temperature tuning:**
      - Factual extraction: 0.1-0.3
      - Classification: 0.3-0.5
      - Diagrams / technical images: 0.5-0.7
      - Creative tasks / illustrations: 0.7-0.9
      
      **2. Batch processing:**
      - Process in batches of 50-100
      - Add delays between batches for rate limits
      - Show progress to user
      
      **3. Error handling:**
      - Always check if result is None
      - Log failures for debugging
      - Consider retry logic for critical tasks
      
      **4. Schema design:**
      - Use Field descriptions for clarity
      - Add constraints (ge, le, max_length)
      - Use Enums for fixed categories
      
      **5. Output formats:**
      - CSV for tabular data
      - JSON for hierarchical data
      - Markdown for documentation
      - Database for large datasets
      
      **6. Image generation:**
      - Compose prompts as: style prefix + subject + format ("Wide landscape")
      - Use `image-pro` / `nano-banana-pro` for published content, `nano-banana-2` for drafts
      - Temperature 0.5-0.7 for diagrams, 0.7-0.8 for illustrations
      - Add negative constraints ("No photorealism, no glow effects") to avoid model defaults
      - Always check result is not None before using the path
      
    • models.md 20.7 KB
      # Gemini Models Reference
      
      Detailed information about available Gemini models (as of September 2026; speech models added 2026-09-24).
      
      ## Model Comparison
      
      ### Gemini 3.8 — Frontier Flash (GA, current default)
      
      #### gemini-3.8-flash
      
      **Status:** Generally available (released September 2, 2026)
      **Alias:** `flash` (the current default Flash)
      
      **Strengths:**
      - Google's "most intelligent Flash model", positioned for long-horizon
        software engineering, autonomous agents, and multi-step enterprise work
      - Vs 3.7 Flash (Google's own numbers): Terminal-Bench 2.1 90.8% vs 81.6%,
        SWE-Bench Pro 61.6% vs 60.4%, SWE-Atlas 51.9% vs 48.0%, τ³-bench Banking
        38.1% vs 30.9%, CharXiv 86.2% vs 84.5%; HLE-Verified 54.9%
      - Humanity's Last Exam is flat (45.4% vs 45.7%) — the gains are agentic and
        tool-use, not open-ended reasoning
      - Artificial Analysis Intelligence Index: 57 at `medium` (3.7 Flash 56,
        3.6 Flash 52); ~310 output tok/sec
      - Prompt-injection robustness improved (Gray Swan); CBRN and cyber-offense
        safeguards carried forward
      
      **Specifications:**
      - Context window: 1,048,576 tokens input / 65,536 tokens output
      - Multimodal input: text, image, video, audio, PDF; text output only
      - `thinking_level`: `low`, `medium` (default), `high`. **`minimal` is not
        supported and returns HTTP 400** ("Thinking level MINIMAL is not supported
        for this model", verified 2026-09-03). The client downgrades `minimal` to
        `low` on this model.
      - Google's release note: the model "works harder" on complex tasks — extra
        reasoning steps, iterative tool calls — so higher effort levels cost more
        tokens. Measured 2026-09-03: `low` spent 0 thinking tokens on a one-word
        reply; the default `medium` spent 79.
      - Supports caching, code execution, computer use (preview), file search,
        function calling, Maps and Search grounding, structured outputs, URL
        context, Batch / Flex / Priority inference. No audio generation, image
        generation, or Live API.
      - Google's migration notes for the 3.x line: `thinking_budget` is gone
        (use `thinking_level`); `temperature` / `top_p` / `top_k` /
        `candidate_count` are deprecated sampling params on this model.
      
      **Best for:**
      - Default Flash / sub-agent-delegation choice for most tasks
      - Agentic coding loops, terminal automation, multi-file projects
      - Finance/legal agent workflows (Vals Finance Agent V2, Harvey's Legal
        Agent Benchmark lead their Flash class)
      
      **Pricing:**
      - Input: $0.75 / 1M tokens through 2026-12-31; $1.50 from 2027-01-01
      - Output: $3.75 / 1M tokens through 2026-12-31; $7.50 from 2027-01-01
        (includes thinking tokens)
      - Context caching $0.075 → $0.15; Batch 50% off
      - 1M context window at base price (no surcharge tier)
      
      Shipped alongside **`gemini-3.8-flash-cyber`** — vulnerability discovery and
      patching (>70% on Google's internal 20-language vuln benchmark, 47.2% pass@1
      on CWE-Bench patching). Access is limited to Google's Fairwind Program
      (government authorities, critical-infrastructure operators, software
      maintainers), so this client cannot alias it.
      
      ---
      
      ### Gemini 3.7 — Prior Frontier Flash (GA)
      
      #### gemini-3.7-flash
      
      **Status:** Generally available (released August 13, 2026)
      **Alias:** `flash-3.7` (was `flash` for three weeks until 3.8 shipped)
      
      **Strengths:**
      - The coding jump in the 3.x Flash line: DeepSWE v1.1 65.3% vs 49.0% on
        3.6 Flash, FrontierCode 1.1 43.6% vs 34.4%, AutomationBench 30.4% vs
        17.0%, WebDev Arena Elo 1588 vs 1538
      - Terminal-Bench 2.1 85.8%, OSWorld-2.0 47.9%, GDM-MRCR v2 (128k) 97.0%
      - Google keeps it "fully supported for efficiency-first workloads". Measured
        2026-09-03 on a one-word prompt, though, `low` on 3.7 still spent 45–88
        thinking tokens where `low` on 3.8 spent 0, and a `max_output_tokens=50`
        call at `low` hit MAX_TOKENS. Budget output generously on this model.
      
      **Specifications:**
      - Context window: 1,048,576 tokens input / 65,536 tokens output
      - Multimodal input: text, image, video, audio, PDF
      - `thinking_level`: `low`, `medium` (default), `high`; **`minimal` returns
        HTTP 400** (verified 2026-09-03), same as 3.8
      
      **Pricing:**
      - Same schedule as 3.8: $0.75 / $3.75 through 2026-12-31, then $1.50 / $7.50
      
      ---
      
      ### Gemini 3.6 — Older Flash (GA)
      
      #### gemini-3.6-flash
      
      **Status:** Generally available (released July 21, 2026)
      **Alias:** `flash-3.6` (was `flash` until 3.7 shipped)
      
      **Strengths:**
      - Builds on 3.5 Flash for coding, knowledge work, and multimodal tasks
      - The last Flash that accepts `thinking_level='minimal'` (verified
        2026-09-03) — pin here for true no-thinking transcription/extraction
      - ~17% fewer output tokens than 3.5 Flash on the Artificial Analysis index
        (the headline efficiency win — addresses 3.5's verbosity)
      - Quality gains alongside efficiency: DeepSWE 49% vs 37%, MLE-Bench 63.9%
        vs 49.7%, OSWorld-Verified 83.0% vs 78.4%, GDPval-AA v2 1421 vs 1349
      - Built-in client-side Computer Use tool via the Gemini API (Preview)
      - Dynamic thinking on by default (configurable via `thinking_level`)
      
      **Specifications:**
      - Context window: ~1M tokens input / 65,536 tokens output
      - Multimodal: text, image, audio, video
      - Default `thinking_level`: `medium` — set explicitly to `minimal` for
        transcription/classification/extraction or the model will silently spend
        output tokens on reasoning
      - Enhanced Frontier Safety safeguards (CBRN, cyber-offense); model card
        notes a slight tone regression vs 3.5 Flash
      
      **Best for:**
      - Pinning prior-gen behavior, and `minimal`-thinking bulk work
      - Cost-sensitive high-volume agentic work (cheaper output than 3.5)
      
      **Pricing:**
      - Same schedule as 3.7 and 3.8 on Google's pricing page (fetched
        2026-09-03): $0.75 / $3.75 through 2026-12-31, then $1.50 / $7.50. The
        July table here said $1.50 / $7.50 flat; the intro rate now covers all three.
      - 1M context window at base price (no surcharge tier)
      
      Shipped alongside two sibling models, neither wired into this client's alias
      table:
      
      - `gemini-3.5-flash-lite` — GA. Fastest 3.5-class model (350 output tok/sec),
        $0.30 / $2.50. **This is now the `lite` alias target** (repointed 2026-07-21
        from gemini-2.5-flash-lite). It costs ~6x more on output than the 2.5 model it
        replaces; that was accepted deliberately — the 2.5 generation is retired
        regardless of price.
      - `gemini-3.5-flash-cyber` — vuln-finding, fine-tuned on 3.5 Flash; powers
        CodeMender. **NOT generally available**: access is limited to governments
        and trusted partners under a pilot program due to dual-use risk. It cannot
        simply be added as an alias.
      
      ---
      
      ### Gemini 3.5 — Legacy Flash (GA)
      
      #### gemini-3.5-flash
      
      **Status:** Generally available (released May 19, 2026 at Google I/O);
      Google's model list now labels it "legacy". No shutdown date.
      **Alias:** `flash-3.5` (was `flash` until 3.6 shipped)
      
      **Strengths:**
      - Frontier-class performance — beats Gemini 3.1 Pro on most coding and
        agentic benchmarks
      - Runs ~4× faster on output tokens than other frontier models
      - Frontier intelligence at sub-Pro pricing
      - Dynamic thinking on by default (configurable via `thinking_level`)
      
      **Specifications:**
      - Context window: ~1M tokens input
      - Multimodal: text, image, audio, video
      - Knowledge cutoff: January 2026
      - Default `thinking_level`: `medium` (was `high` on prior 3.x — set
        explicitly to `minimal` for transcription/classification/extraction or
        the model will silently spend output tokens on reasoning)
      
      **Best for:**
      - Pinning to prior-gen Flash behavior when 3.6–3.8 output differs
      - Agentic coding loops, terminal automation, multi-file projects
      - Multimodal document analysis where structure must be preserved
      
      **Pricing:**
      - Input: $1.50 / 1M tokens
      - Output: $9.00 / 1M tokens
      - 1M context window at base price (no surcharge tier)
      
      ---
      
      ### Gemini 3.x — Prior Preview Generation
      
      #### gemini-3-flash-preview
      
      **Status:** Preview (still callable — kept for back compat)
      **Alias:** `flash-3`
      
      The previous-generation Flash. Use when you need to pin behavior
      established before the 3.5 cutover. Google's deprecation page lists
      gemini-3.6-flash as its replacement, with no shutdown date. New code should
      target `flash` (gemini-3.8-flash) instead.
      
      **Pricing:**
      - Input: $0.30 / 1M tokens
      - Output: $2.50 / 1M tokens
      
      #### gemini-3.1-pro-preview
      
      **Status:** Preview. **DEPRECATED from routing 2026-09-03** (Oskar): "its
      Pareto efficiency is too poor compared to the later Flash models." At today's
      rates it costs 2.7× the input and 3.2× the output of 3.8 Flash (1.3× / 1.6× once
      the Flash intro pricing ends), and 3.5 Flash already beat it on most coding and
      agentic benchmarks. The ID stays callable for pinned code.
      **Alias:** none — `pro` now resolves to gemini-3.8-flash. For maximum
      reasoning use Flash with `thinking_level='high'`.
      
      **Strengths:**
      - Was the most capable Gemini Pro in the API
      - 1M context with tiered pricing above 200K
      
      **Specifications:**
      - Context window: ~1M tokens input
      - Long context surcharge: 2× above 200K input tokens
      - Multimodal: text, image, video, audio
      
      **Best for:**
      - Nothing in new code. Pinned callers only.
      
      **Pricing:**
      - Input: $2.00 / 1M tokens (≤200K), $4.00 (>200K)
      - Output: $12.00 / 1M tokens (≤200K), $18.00 (>200K) — the July table here
        said $24.00; Google's pricing page says $18.00 (fetched 2026-09-03)
      
      **Note:** Google announced `gemini-3.5-pro` at I/O on 2026-05-19 for June.
      As of 2026-09-03 it is not in the API model list or on the pricing page;
      DeepMind still lists it as "coming soon". When it ships it gets the same
      price/quality test against the current Flash before any alias points at it.
      
      ---
      
      ### Gemini 2.5 — DEPRECATED (retired 2026-07-21)
      
      ⚠️ The Gemini 2.5 text generation is **retired from routing**. A 2025-era
      generation; the cost saving does not justify the quality gap. Model IDs remain
      callable so pinned code does not hard-break, but do not target them in new work.
      The `lite` alias now resolves to `gemini-3.5-flash-lite`.
      
      #### gemini-2.5-flash
      
      **Status:** Stable, generally available
      **Alias:** `stable-flash`
      
      **Strengths:**
      - Production stability without preview-tier volatility
      - Solid price-performance for reasoning tasks
      - Empirically token-perfect on dense transcription benchmarks (May 2026)
      
      **Specifications:**
      - Context window: ~1M tokens input
      - Multimodal: text, image, video, audio
      
      **Best for:**
      - Production workloads where preview models are too volatile
      - High-volume tasks with a quality floor
      - Multimodal extraction when cost matters but accuracy can't slip
      
      **Pricing:**
      - Input: $0.30 / 1M tokens
      - Output: $2.50 / 1M tokens
      
      #### gemini-2.5-flash-lite
      
      **Status:** DEPRECATED (retired from routing 2026-07-21)
      **Alias:** none — `lite` now points at gemini-3.5-flash-lite
      
      **Strengths:**
      - **Cheapest major-provider production model** ($0.10 / $0.40)
      - Surprisingly capable on multimodal extraction — empirically transcribes
        dense tables on par with much pricier models
      - Fast: typically lowest latency in the lineup
      
      **Specifications:**
      - Context window: ~1M tokens input
      - Multimodal: text, image, video, audio
      
      **Best for:**
      - Ultra-budget batch processing
      - Routine triage tasks (zeitgeist runs, inbox review, bsky image
        transcription, classification, simple extraction)
      - Maximum throughput at minimum cost
      
      **Pricing:**
      - Input: $0.10 / 1M tokens
      - Output: $0.40 / 1M tokens
      
      #### gemini-2.5-pro
      
      **Status:** Stable, generally available
      **Alias:** `stable-pro`
      
      **Strengths:**
      - Pro-tier reasoning with production stability
      - Well-documented behavior across long-running deployments
      
      **Specifications:**
      - Context window: ~1M tokens input
      - Long context surcharge: 2× above 200K tokens
      - Multimodal: text, image, video, audio
      
      **Best for:**
      - Complex tasks requiring production stability
      - Long-document processing
      - Quality-critical workloads
      
      **Pricing:**
      - Input: $1.25 / 1M tokens (≤200K), $2.50 (>200K)
      - Output: $10.00 / 1M tokens (≤200K), $20.00 (>200K)
      
      ---
      
      ### Speech Generation Models (TTS)
      
      **Released 2026-09-23, GA** on the Gemini API and AI Studio (Gemini Enterprise
      in preview). Google's announcement claims #1 on Hume AI's Voice Design
      Benchmark (71.4) and the #1 and #2 spots on its Overall Quality Index.
      
      #### gemini-3.8-flash-tts
      
      - Flagship expressive TTS: acting, regional accents, long-form multi-turn
        stability. 130 languages, auto-detected.
      - Limits: 8,192 input tokens / 16,384 output tokens per request; up to two
        speakers.
      - Pricing: $0.50 in (text) / $9.00 out (audio) per 1M through 2026-12-31, then
        $1.00 / $18.00. Batch half. About 32 audio tokens per second of speech
        (measured 2026-09-24), so ~$0.02 per minute.
      
      #### gemini-3.8-flash-lite-tts
      
      - Cheaper workhorse for bulk narration, voice-agent cascades and read-aloud;
        101 languages. $0.50 / $6.00 per 1M through 2026-12-31, then $1.00 / $12.00.
      - Google's named replacement for `gemini-3.1-flash-tts-preview` ($1 / $20).
      
      #### API shape (verified through the CF gateway, 2026-09-24)
      
      - Use the **Interactions API**: `POST v1beta/interactions` with
        `{"model", "input": [{"type": "user_input", "content": [{"type": "text",
        "text", "annotations": [{"type": "speech_metadata", "style"}]}]}],
        "response_format": {"type": "audio", "mime_type": "audio/wav",
        "sample_rate": 24000}, "generation_config": {"speech_config": [{"voice"}]}}`.
        Audio is base64 at `steps[].content[]` where `type == "audio"`.
      - `generateContent` with `responseModalities: ["AUDIO"]` also returns audio,
        but a "Style: text" prefix is spoken aloud and `systemInstruction` returns
        HTTP 400 "Developer instruction is not enabled for this model".
      - Voices: 30 studio voices plus 2,059 persona voices (`GET v1beta/voices`,
        paged by `next_page_token`, max 1,000 per page). Designed voices come from
        `POST v1beta/voices` with `{"store": true, "voice": {"type": "prompted",
        "prompted": {"input": "<description>"}, ...}}` and return a `voice_...` id
        (1-year expiry, 200 per project) plus a `sample_audio` preview.
      - Output is watermarked with SynthID.
      - The model can paraphrase: it added "Hmm," and swapped pronouns in a scripted
        narration. Check scripted output with ASR.
      
      ### Image Generation Models
      
      **Updated 2026-05-28:** Nano Banana 2 and Nano Banana Pro reached general
      availability — announced GA on Vertex AI / Gemini Enterprise Agent Platform,
      where the GA model IDs drop the suffix (`gemini-3.1-flash-image`,
      `gemini-3-pro-image`).
      
      ⚠️ **Corrected 2026-07-21 (the previous note here was wrong).** The GA IDs
      `gemini-3.1-flash-image` and `gemini-3-pro-image` are **NOT** Vertex-only and do
      **NOT** 404 on the Developer API — they were released on this surface on
      2026-05-28 and were live-tested working through the CF gateway on 2026-07-21.
      The `-preview` IDs also still resolve (their announced 2026-06-25 shutdown
      appears to redirect rather than fail), so nothing is broken either way — but
      **new code should target the GA IDs**.
      
      Also available and not yet wired into this client: `gemini-3.1-flash-lite-image`
      (Nano Banana 2 Lite, GA) — the cheapest image tier, ~$0.034/image.
      
      #### nano-banana-2
      
      **Status:** GA on Vertex; Developer API still serves it as `-preview` (this client's surface)
      **API Model ID:** `gemini-3.1-flash-image-preview`
      **Alias:** `image`
      
      Fast generation/editing on the Gemini 3.1 Flash Image platform. Default image
      model in `generate_image()`. Capabilities on the Developer API:
      - Output resolutions: 512 (0.5K), 1K, 2K generally available; 4K in preview.
        512 is 3.1-Flash-only.
      - Up to 14 reference images (up to 10 high-fidelity objects + up to 4 characters).
      - Grounding with Google Search, plus Image Search grounding (3.1-Flash-only) —
        cannot search for images of people.
      - Thinking: `thinking_level` is {`minimal` (default), `high`}; thinking cannot be
        fully disabled and thinking tokens are billed.
      - Extra aspect ratios over 2.5 Flash Image: 1:4, 4:1, 1:8, 8:1.
      
      Note: the GA announcement's "video file as input prompt" capability is a Vertex
      preview feature. The Developer API does NOT accept video or audio input for
      image generation — don't route video here.
      
      #### nano-banana-pro
      
      **Status:** GA on Vertex; Developer API still serves it as `-preview` (this client's surface)
      **API Model ID:** `gemini-3-pro-image-preview`
      **Alias:** `image-pro`
      
      High-fidelity generation on the Gemini 3 Pro Image platform — legible stylized
      text rendering and professional asset production via advanced "thinking."
      Capabilities:
      - Output resolutions: 1K, 2K generally available; 4K in preview.
      - Up to 14 reference images (up to 6 high-fidelity objects + up to 5 characters).
      - Thinking always on (cannot be disabled).
      
      #### nano-banana
      
      **Status:** Stable, GA (unchanged)
      **API Model ID:** `gemini-2.5-flash-image`
      
      Production-grade stability on the Gemini 2.5 Flash Image platform. Works best
      with up to 3 input images.
      
      ---
      
      ## Model Selection Guide
      
      ```
      Default Flash (frontier)?              → gemini-3.8-flash (alias: flash)
      Maximum reasoning?                     → gemini-3.8-flash, thinking_level='high' (alias: pro)
      Pro tier?                              → off routing since 2026-09-03; see gemini-3.1-pro-preview
      Routine / bulk / cheap / fastest?      → gemini-3.5-flash-lite (alias: lite)
      No-thinking pass (minimal) on Flash?   → gemini-3.6-flash (alias: flash-3.6)
      Pin to prior frontier Flash (3.7)?     → gemini-3.7-flash (alias: flash-3.7)
      Pin to legacy Flash (3.5)?             → gemini-3.5-flash (alias: flash-3.5)
      Pin to older preview Flash?            → gemini-3-flash-preview (alias: flash-3)
      Image generation (fast)?               → nano-banana-2 (alias: image)
      Image generation (high-fidelity)?      → nano-banana-pro (alias: image-pro)
      ```
      
      ## Thinking Configuration (Gemini 3.x family)
      
      Gemini 3.x models reason before responding. By default, the model spends
      output tokens on reasoning, then on the visible answer. From Gemini 3.5
      Flash on, the default is `medium` — down from `high` on prior 3.x — and the
      parameter shape changed:
      
      - **Old:** integer `thinking_budget`
      - **New:** string enum `thinking_level` ∈ {`minimal`, `low`, `medium`, `high`}
      
      Which models take `minimal` (all verified live 2026-09-03):
      
      | Model | `minimal` |
      |---|---|
      | gemini-3.8-flash | HTTP 400 — client downgrades to `low` |
      | gemini-3.7-flash | HTTP 400 — client downgrades to `low` |
      | gemini-3.6-flash | accepted |
      | gemini-3.5-flash | accepted |
      | gemini-3.5-flash-lite | accepted |
      
      The Python client exposes this as `invoke_gemini(..., thinking_level="...")`.
      Pass `None` (default) to let the model use its built-in default.
      
      **When to set `thinking_level='minimal'`:**
      - Transcription, OCR, image-to-text
      - Classification, tagging, extraction with a fixed schema
      - Any task where the LLM doesn't need to reason — it just needs to emit
      
      **When to leave it as default or set higher:**
      - Code generation, debugging
      - Multi-step planning
      - Math, complex analysis
      
      **Why it matters:** A `max_output_tokens=50` request can return empty if
      thinking_level (default `medium` on 3.5–3.8) consumes all 50 tokens before
      emitting visible output. Symptom: response text is empty, `finishReason`
      is `MAX_TOKENS`. Fix: either raise `max_output_tokens` substantially or
      set `thinking_level='minimal'`.
      
      ## Multimodal Capabilities
      
      All text models support:
      - **Images:** JPEG, PNG, WebP, HEIC, HEIF
      - **Video:** MP4, MPEG, MOV, AVI, FLV, MPG, WebM, WMV, 3GPP
      - **Audio:** WAV, MP3, AIFF, AAC, OGG, FLAC
      
      **Audio input pricing:** Higher than text, typically ~$1.00 / 1M tokens
      on Flash-tier models.
      
      ## Deprecated / Retired Models
      
      | Model | Status | Migration Target |
      |---|---|---|
      | gemini-3-pro-preview | Retired (March 9, 2026) | gemini-3.1-pro-preview |
      | gemini-3-flash-preview | Callable, no shutdown date | gemini-3.6-flash (Google's listed target) |
      | gemini-3.1-flash-lite-preview | Retired (May 25, 2026) | gemini-3.1-flash-lite |
      | gemini-3.1-flash-lite | Shutdown May 7, 2027 | gemini-3.5-flash-lite |
      | gemini-3.1-flash-image-preview | Shutdown listed June 25, 2026 (still resolves) | gemini-3.1-flash-image |
      | gemini-3-pro-image-preview | Shutdown listed June 25, 2026 (still resolves) | gemini-3-pro-image |
      | gemini-2.5-flash-image (`nano-banana`) | Shutdown October 2, 2026 | gemini-3.1-flash-image |
      | gemini-2.0-flash-exp | Retired June 1, 2026 | gemini-3.6-flash |
      | gemini-2.0-flash | Retired June 1, 2026 | gemini-3.6-flash |
      | gemini-2.0-flash-lite | Retired June 1, 2026 | gemini-3.5-flash-lite |
      | gemini-1.5-pro | Retired (404) | gemini-2.5-pro |
      | gemini-1.5-flash | Retired (404) | gemini-3.6-flash |
      | gemini-1.0-* | Retired (404) | — |
      
      ## Cost Optimization Tips
      
      - **Batch API:** 50% discount on all paid models for async (≤24h) processing
      - **Context caching:** Up to 75–90% savings for repeated large prompts
      - **Long context:** Pro models charge 2× above 200K tokens — keep prompts concise
      - **Free tier:** Gemini app + AI Studio offer free access to Flash and Lite
        models with daily quotas; Pro is paid-only as of April 2026
      
      ## Rate Limits
      
      Vary by API tier (default free tier):
      - **Requests per minute:** 15
      - **Tokens per minute:** 1M
      - **Requests per day:** 1,500
      
      Client automatically handles rate limiting with exponential backoff.
      
  • scripts
    • gemini_client.py 46.5 KB
      #!/usr/bin/env python3
      """
      Gemini API Client
      
      Routes requests through Cloudflare AI Gateway when configured (preferred)
      or directly to Google's Generative Language API via the google-generativeai
      SDK (fallback).
      
      Credential priority:
      1. CF Gateway: proxy.env with CF_ACCOUNT_ID, CF_GATEWAY_ID, CF_API_TOKEN
      2. Direct API:  GOOGLE_API_KEY.txt or API_CREDENTIALS.json
      """
      
      import json
      import os
      import sys
      import time
      from pathlib import Path
      
      try:
          import requests
          HAS_REQUESTS = True
      except ImportError:
          HAS_REQUESTS = False
      
      try:
          import google.generativeai as genai
          HAS_GENAI = True
      except ImportError:
          HAS_GENAI = False
      
      try:
          from pydantic import BaseModel
          HAS_PYDANTIC = True
      except ImportError:
          HAS_PYDANTIC = False
          BaseModel = object  # type: ignore[assignment,misc]
      
      if not HAS_REQUESTS and not HAS_GENAI:
          print("Error: neither 'requests' nor 'google-generativeai' is installed.")
          print("Install with: uv pip install requests google-generativeai pydantic")
          import sys
          sys.exit(1)
      
      
      # ---------------------------------------------------------------------------
      # Model registry
      # ---------------------------------------------------------------------------
      
      # Text generation models
      MODELS = {
          # Gemini 3.8 — current frontier Flash (GA 2026-09-02)
          "gemini-3.8-flash": "gemini-3.8-flash",
          # Gemini 3.7 — prior frontier Flash (GA 2026-08-13)
          "gemini-3.7-flash": "gemini-3.7-flash",
          # Gemini 3.6 — older Flash (GA 2026-07-21)
          "gemini-3.6-flash": "gemini-3.6-flash",
          # Gemini 3.5 — prior frontier Flash (GA May 2026)
          "gemini-3.5-flash": "gemini-3.5-flash",
          # Gemini 3.x — preview (still callable, kept for back compat)
          "gemini-3-flash-preview": "gemini-3-flash-preview",
          # Gemini 3.1 Pro — DEPRECATED from routing 2026-09-03 (Oskar): its
          # price/quality is dominated by the 3.6+ Flash line. ID kept callable for
          # pinned code; `pro` no longer points here.
          "gemini-3.1-pro-preview": "gemini-3.1-pro-preview",
          # Gemini 3.5 Flash-Lite — cheap/bulk tier (GA 2026-07-21)
          "gemini-3.5-flash-lite": "gemini-3.5-flash-lite",
          # Gemini 2.5 — DEPRECATED 2026-07-21. A 2025-era generation; do NOT route
          # here. IDs kept callable so pinned code does not hard-break, but they are
          # no longer a recommended target and `lite` no longer points at 2.5.
          "gemini-2.5-flash": "gemini-2.5-flash",
          "gemini-2.5-flash-lite": "gemini-2.5-flash-lite",
          "gemini-2.5-pro": "gemini-2.5-pro",
      }
      
      # Image generation models — display name -> actual API model ID.
      # NOTE (2026-05-28): Nano Banana 2 / Pro went GA on Vertex (IDs there drop
      # the suffix: gemini-3.1-flash-image / gemini-3-pro-image). This client uses
      # the Gemini Developer API (google-ai-studio gateway), where both are STILL
      # served under the -preview IDs below (verified against the live docs).
      # Do NOT drop -preview here — the GA IDs are Vertex-only and 404 on this surface.
      IMAGE_MODELS = {
          "nano-banana-2": "gemini-3.1-flash-image-preview",
          "nano-banana-pro": "gemini-3-pro-image-preview",
          "nano-banana": "gemini-2.5-flash-image",
      }
      
      # Convenience aliases. `flash` points to the current frontier Flash (3.8, GA
      # 2026-09-02); `flash-3.7`, `flash-3.6`, `flash-3.5` and `flash-3` keep stable
      # handles on the prior Flash generations for code that pinned to them. `lite`
      # repointed 2026-07-21 from gemini-2.5-flash-lite to gemini-3.5-flash-lite
      # (BREAKING: ~6x output cost, $0.40 -> $2.50/M, in exchange for a
      # current-generation model).
      MODEL_ALIASES = {
          "flash": "gemini-3.8-flash",
          "flash-3.7": "gemini-3.7-flash",
          "flash-3.6": "gemini-3.6-flash",
          "flash-3.5": "gemini-3.5-flash",
          "flash-3": "gemini-3-flash-preview",
          # `pro` repointed 2026-09-03 from gemini-3.1-pro-preview to the frontier
          # Flash. The Pro tier is off routing (price/quality dominated by Flash);
          # "maximum reasoning" is Flash with thinking_level='high'.
          "pro": "gemini-3.8-flash",
          "lite": "gemini-3.5-flash-lite",
          # DEPRECATED aliases (Gemini 2.5). Retained for back compat only.
          "stable-flash": "gemini-2.5-flash",
          "stable-pro": "gemini-2.5-pro",
          "image": "nano-banana-2",
          "image-pro": "nano-banana-pro",
      }
      
      DEFAULT_MODEL = "gemini-3.8-flash"
      
      # Speech (TTS) models — GA 2026-09-23. Kept out of MODEL_ALIASES on purpose:
      # they return audio, not text, so invoke_gemini() must never resolve to them.
      SPEECH_MODELS = {
          "gemini-3.8-flash-tts": "gemini-3.8-flash-tts",            # expressive, 130 languages
          "gemini-3.8-flash-lite-tts": "gemini-3.8-flash-lite-tts",  # bulk / read-aloud, 101 languages
      }
      SPEECH_ALIASES = {"tts": "gemini-3.8-flash-tts", "tts-lite": "gemini-3.8-flash-lite-tts"}
      
      # Flash 3.7 and 3.8 reject thinking_level='minimal' with HTTP 400 ("Thinking
      # level MINIMAL is not supported for this model"); 3.6, 3.5 and 3.5-lite accept
      # it. All five verified live through the CF gateway on 2026-09-03.
      _MINIMAL_THINKING_UNSUPPORTED = frozenset({"gemini-3.7-flash", "gemini-3.8-flash"})
      
      # ---------------------------------------------------------------------------
      # Cloudflare AI Gateway constants
      # ---------------------------------------------------------------------------
      
      _CF_GATEWAY_BASE = "https://gateway.ai.cloudflare.com/v1"
      
      _PROXY_ENV_PATHS = [
          Path("/mnt/project/proxy.env"),
          Path("/mnt/user-data/proxy.env"),
          Path.home() / ".muninn" / "proxy.env",
      ]
      
      
      # ---------------------------------------------------------------------------
      # Credential loading
      # ---------------------------------------------------------------------------
      
      def _parse_env_file(path: Path) -> dict:
          """Parse a .env-format file into a dict, stripping quotes."""
          result = {}
          for line in path.read_text().splitlines():
              line = line.strip()
              if line and not line.startswith("#") and "=" in line:
                  key, _, value = line.partition("=")
                  result[key.strip()] = value.strip().strip('"').strip("'")
          return result
      
      
      def get_cf_credentials() -> dict | None:
          """
          Load Cloudflare AI Gateway credentials.
      
          Searches for proxy.env in well-known paths, then falls back to
          environment variables.
      
          Required keys: CF_ACCOUNT_ID, CF_GATEWAY_ID, CF_API_TOKEN
          Optional key:  GOOGLE_API_KEY (for non-BYOK setups)
      
          Returns:
              dict with credentials if fully configured, None otherwise
          """
          required = ("CF_ACCOUNT_ID", "CF_GATEWAY_ID", "CF_API_TOKEN")
      
          for env_path in _PROXY_ENV_PATHS:
              if env_path.exists():
                  try:
                      creds = _parse_env_file(env_path)
                      if all(creds.get(k) for k in required):
                          return creds
                  except OSError:
                      continue
      
          # Fall back to environment variables
          creds = {k: os.environ.get(k, "") for k in required}
          creds["GOOGLE_API_KEY"] = os.environ.get("GOOGLE_API_KEY", "")
          if all(creds.get(k) for k in required):
              return creds
      
          return None
      
      
      def get_google_api_key() -> str:
          """
          Get Google API key for direct (non-gateway) access.
      
          Priority order:
          1. Individual file: /mnt/project/GOOGLE_API_KEY.txt
          2. Combined file:   /mnt/project/API_CREDENTIALS.json
          3. Environment variable: GOOGLE_API_KEY
      
          Returns:
              str: Google API key
      
          Raises:
              ValueError: If no API key found in any source
          """
          # 1. Individual key file
          key_file = Path("/mnt/project/GOOGLE_API_KEY.txt")
          if key_file.exists():
              try:
                  key = key_file.read_text().strip()
                  if key:
                      return key
              except OSError as e:
                  raise ValueError(f"Found GOOGLE_API_KEY.txt but couldn't read it: {e}")
      
          # 2. Combined credentials file
          creds_file = Path("/mnt/project/API_CREDENTIALS.json")
          if creds_file.exists():
              try:
                  with open(creds_file) as f:
                      config = json.load(f)
                  key = config.get("google_api_key", "").strip()
                  if key:
                      return key
              except (json.JSONDecodeError, OSError) as e:
                  raise ValueError(f"Found API_CREDENTIALS.json but couldn't parse it: {e}")
      
          # 3. Environment variable
          key = os.environ.get("GOOGLE_API_KEY", "").strip()
          if key:
              return key
      
          raise ValueError(
              "No Google API key found!\n\n"
              "Option A (recommended): Configure Cloudflare AI Gateway\n"
              "  File: /mnt/project/proxy.env\n"
              "  Content:\n"
              "    CF_ACCOUNT_ID=<your-account-id>\n"
              "    CF_GATEWAY_ID=<your-gateway-id>\n"
              "    CF_API_TOKEN=<your-cf-api-token>\n\n"
              "Option B: Direct Google API\n"
              "  File: GOOGLE_API_KEY.txt  (content: AIzaSy...)\n"
              "  or\n"
              "  File: API_CREDENTIALS.json  (content: {\"google_api_key\": \"AIzaSy...\"})\n\n"
              "Get your Cloudflare token: https://dash.cloudflare.com/profile/api-tokens\n"
              "Get your Google key: https://console.cloud.google.com/apis/credentials"
          )
      
      
      # ---------------------------------------------------------------------------
      # Cloudflare AI Gateway — REST path
      # ---------------------------------------------------------------------------
      
      def _cf_request(
          model_id: str,
          contents: list,
          generation_config: dict,
          cf_creds: dict,
      ) -> dict:
          """
          POST a generateContent request via Cloudflare AI Gateway.
      
          Args:
              model_id: Gemini model ID (e.g., 'gemini-3-flash-preview')
              contents: Gemini REST API contents array
              generation_config: generationConfig dict (camelCase keys)
              cf_creds: dict with CF_ACCOUNT_ID, CF_GATEWAY_ID, CF_API_TOKEN
      
          Returns:
              Parsed JSON response dict
      
          Raises:
              requests.HTTPError: On non-2xx HTTP response
          """
          account_id = cf_creds["CF_ACCOUNT_ID"]
          gateway_id = cf_creds["CF_GATEWAY_ID"]
          api_token = cf_creds["CF_API_TOKEN"]
      
          url = (
              f"{_CF_GATEWAY_BASE}/{account_id}/{gateway_id}"
              f"/google-ai-studio/v1beta/models/{model_id}:generateContent"
          )
      
          # Include Google API key as query param for non-BYOK setups
          google_key = cf_creds.get("GOOGLE_API_KEY") or os.environ.get("GOOGLE_API_KEY", "")
          if google_key:
              url += f"?key={google_key}"
      
          payload: dict = {"contents": contents}
          if generation_config:
              payload["generationConfig"] = generation_config
      
          headers = {
              "Content-Type": "application/json",
              "cf-aig-authorization": f"Bearer {api_token}",
          }
      
          # Retry on 5xx / 429 / SSL / non-JSON proxy errors. The Claude.ai egress
          # proxy can return HTTP 503 with body 'DNS cache overflow' (text/plain),
          # most often on cold start; without this retry the caller sees an
          # opaque JSONDecodeError or HTTPError on what is actually a transient
          # proxy condition rather than a Gemini/CF-AI-Gateway failure.
          max_retries = 3
          base_delay = 0.5
          for attempt in range(max_retries):
              try:
                  response = requests.post(url, json=payload, headers=headers, timeout=120)
                  if response.status_code >= 500:
                      preview = (response.text or '')[:200]
                      raise RuntimeError(
                          f"HTTP {response.status_code} from CF AI Gateway "
                          f"(likely egress proxy, not Gemini): {preview!r}"
                      )
                  if 400 <= response.status_code < 500:
                      # raise_for_status() discards the body, and Gemini puts the ONLY
                      # useful diagnostic there — e.g. which schema keyword it rejected.
                      # A 4xx is deterministic, so mark it non-retriable too.
                      preview = (response.text or '')[:600]
                      raise _NonRetriableAPIError(
                          f"HTTP {response.status_code} from Gemini: {preview}"
                      )
                  response.raise_for_status()
                  return response.json()
              except (MediaInputError, _NonRetriableAPIError):
                  raise  # deterministic error — retrying cannot help
              except Exception as e:
                  if attempt == max_retries - 1:
                      raise
                  err = str(e)
                  retriable = (
                      '503' in err or '429' in err or 'Service Unavailable' in err
                      or 'DNS cache overflow' in err or 'Expecting value' in err
                      or 'JSONDecodeError' in err or 'SSL' in err or 'SSLError' in err
                      or 'HANDSHAKE_FAILURE' in err
                  )
                  if not retriable:
                      raise
                  delay = base_delay * (2 ** attempt)
                  print(
                      f"Warning: CF AI Gateway request failed "
                      f"(attempt {attempt + 1}/{max_retries}), retrying in {delay}s: {e}"
                  )
                  time.sleep(delay)
      
      
      def _extract_text(response: dict) -> str | None:
          """Extract generated text from a Gemini REST API response."""
          try:
              return response["candidates"][0]["content"]["parts"][0]["text"]
          except (KeyError, IndexError, TypeError):
              return None
      
      
      def _check_truncation(response: dict) -> None:
          """Raise if the model hit the output cap, instead of letting json.loads fail.
      
          A truncated structured response is invalid JSON, so the symptom is a
          JSONDecodeError / pydantic "EOF while parsing" that looks like a schema
          problem. finishReason says what actually happened.
          """
          try:
              reason = response["candidates"][0].get("finishReason")
          except (KeyError, IndexError, TypeError):
              return
          if reason in ("MAX_TOKENS", "MAX_OUTPUT_TOKENS"):
              raise _NonRetriableAPIError(
                  "Gemini hit maxOutputTokens before finishing the JSON object "
                  "(finishReason=MAX_TOKENS). Raise max_output_tokens — thinking "
                  "tokens count against the same budget."
              )
      
      
      class MediaInputError(ValueError):
          """Invalid media input. Deterministic — never worth retrying."""
      
      
      class _NonRetriableAPIError(RuntimeError):
          """A 4xx from Gemini. Deterministic (bad schema, bad request) — retrying
          only hides the response body, which is the one thing worth reading."""
      
      
      # Extensions mimetypes.guess_type() gets wrong or misses, per Gemini's accepted
      # media types. Audio/video matter here: a bad guess silently sends the wrong
      # mimeType and the model returns a confused answer instead of an error.
      _MEDIA_MIME_OVERRIDES = {
          ".m4a": "audio/mp4", ".aac": "audio/aac", ".flac": "audio/flac",
          ".ogg": "audio/ogg", ".opus": "audio/opus", ".mp3": "audio/mpeg",
          ".wav": "audio/wav", ".aiff": "audio/aiff",
          ".mp4": "video/mp4", ".mov": "video/quicktime", ".webm": "video/webm",
          ".heic": "image/heic", ".heif": "image/heif",
      }
      
      # generateContent inline payloads must fit the request; base64 inflates ~4/3.
      # Files above this need the Files API (not implemented in this client).
      _INLINE_MEDIA_MAX_BYTES = 15 * 1024 * 1024
      
      
      def _guess_media_mime(path: str) -> str:
          """Resolve a media mimeType, preferring explicit overrides over mimetypes."""
          import mimetypes
          ext = Path(path).suffix.lower()
          if ext in _MEDIA_MIME_OVERRIDES:
              return _MEDIA_MIME_OVERRIDES[ext]
          return mimetypes.guess_type(path)[0] or "application/octet-stream"
      
      
      def _build_contents(prompt: str, image_path: str | None) -> list:
          """Build the Gemini REST API 'contents' array.
      
          `image_path` accepts ANY supported media file — image, audio, or video —
          despite the legacy name. Audio input is verified working on this path
          (2026-07-21).
          """
          parts: list = [{"text": prompt}]
          if image_path:
              import base64
      
              size = Path(image_path).stat().st_size
              if size > _INLINE_MEDIA_MAX_BYTES:
                  raise MediaInputError(
                      f"{image_path} is {size/1e6:.1f}MB; inline media is capped at "
                      f"{_INLINE_MEDIA_MAX_BYTES/1e6:.0f}MB. Larger files need the "
                      f"Files API, which this client does not implement yet."
                  )
              image_data = Path(image_path).read_bytes()
              mime_type = _guess_media_mime(image_path)
              parts.append({
                  "inlineData": {
                      "mimeType": mime_type,
                      "data": base64.b64encode(image_data).decode(),
                  }
              })
          return [{"parts": parts}]
      
      
      #: Keywords Gemini's responseSchema rejects outright. Pydantic emits all of
      #: these; leaving any of them in produces a bare HTTP 400 with no usable body.
      _SCHEMA_STRIP_KEYS = ("$schema", "$defs", "title", "default", "additionalProperties",
                            "discriminator", "examples", "const")
      
      
      def _inline_refs(node, defs: dict):
          """Resolve $ref against $defs and strip Gemini-unsupported keywords, recursively.
      
          Gemini's responseSchema is NOT full JSON Schema: it rejects $ref/$defs
          entirely. Pydantic emits a $ref for every nested model, so any model with a
          nested model (the common case — List[SomeModel]) produced a bare HTTP 400
          that the retry loop then burned three attempts on before returning None.
          Flat single-level models worked, which is why this hid for so long.
          """
          if isinstance(node, dict):
              if "$ref" in node:
                  name = node["$ref"].split("/")[-1]
                  target = defs.get(name)
                  if target is None:
                      # Unresolvable ref — better a permissive object than a 400.
                      return {"type": "object"}
                  return _inline_refs(target, defs)
              out = {}
              for k, v in node.items():
                  if k in _SCHEMA_STRIP_KEYS:
                      continue
                  if k == "anyOf":
                      # Gemini has no anyOf. Optional[X] becomes [X, null]; take the
                      # first non-null branch and mark it nullable instead.
                      branches = [b for b in v if b.get("type") != "null"]
                      if branches:
                          resolved = _inline_refs(branches[0], defs)
                          if len(branches) < len(v):
                              resolved["nullable"] = True
                          out.update(resolved)
                      continue
                  out[k] = _inline_refs(v, defs)
              return out
          if isinstance(node, list):
              return [_inline_refs(x, defs) for x in node]
          return node
      
      
      def _pydantic_to_schema(model_class: type) -> dict:
          """Convert a Pydantic model class to a Gemini-compatible JSON schema dict.
      
          Handles nested models: pydantic's $defs/$ref indirection is inlined, because
          Gemini's responseSchema does not support it.
          """
          try:
              schema = model_class.model_json_schema()  # Pydantic v2
          except AttributeError:
              schema = model_class.schema()  # Pydantic v1
      
          defs = schema.get("$defs") or schema.get("definitions") or {}
          return _inline_refs(schema, defs)
      
      
      # ---------------------------------------------------------------------------
      # Direct SDK path helpers
      # ---------------------------------------------------------------------------
      
      def _initialize_direct_client() -> bool:
          """Configure google.generativeai SDK for direct API access."""
          if not HAS_GENAI:
              return False
          try:
              api_key = get_google_api_key()
              genai.configure(api_key=api_key)
              return True
          except ValueError as e:
              print(f"Error: {e}")
              return False
      
      
      def _build_genai_content(prompt: str, image_path: str | None):
          """Build content argument for google.generativeai SDK calls.
      
          NOTE: this direct-SDK fallback handles IMAGES only. Non-image media
          (audio/video) works on the CF Gateway path; PIL cannot open it, so fail
          with a clear message instead of an opaque UnidentifiedImageError.
          """
          if image_path and not _guess_media_mime(image_path).startswith("image/"):
              raise MediaInputError(
                  f"{image_path} is not an image ({_guess_media_mime(image_path)}). "
                  "Audio/video input requires the Cloudflare AI Gateway path — "
                  "configure proxy.env; the direct google-generativeai SDK fallback "
                  "supports images only."
              )
          if image_path:
              from PIL import Image  # type: ignore[import]
              return [prompt, Image.open(image_path)]
          return prompt
      
      
      # ---------------------------------------------------------------------------
      # Public API
      # ---------------------------------------------------------------------------
      
      def _resolve_model(model: str) -> str:
          """Resolve a model name or alias to its canonical API model ID.
      
          Handles chained resolution: alias → display name → API model ID.
      
          Args:
              model: Model name, alias, or direct ID
      
          Returns:
              Canonical API model ID string
      
          Raises:
              ValueError: If model is not recognized
          """
          # Resolve alias first (e.g., "image" → "nano-banana-2")
          resolved = MODEL_ALIASES.get(model, model)
      
          # Direct match in text models
          if resolved in MODELS:
              return MODELS[resolved]
          # Image models (display name → API model ID)
          if resolved in IMAGE_MODELS:
              return IMAGE_MODELS[resolved]
      
          all_names = list(MODELS) + list(MODEL_ALIASES) + list(IMAGE_MODELS)
          raise ValueError(f"Invalid model: {model}. Choose from {all_names}")
      
      
      # @lat: [[orchestration#Gemini Client]]
      def invoke_gemini(
          prompt: str,
          model: str = DEFAULT_MODEL,
          temperature: float = 0.7,
          max_output_tokens: int | None = None,
          top_p: float | None = None,
          top_k: int | None = None,
          image_path: str | None = None,
          thinking_level: str | None = None,
      ) -> str | None:
          """
          Invoke Gemini model with a text (or multi-modal) prompt.
      
          Routes through Cloudflare AI Gateway when proxy.env is configured;
          falls back to direct Google API via google-generativeai SDK.
      
          Args:
              prompt: The text prompt to send
              model: Model name or alias (default: gemini-3.6-flash).
                  Aliases: flash (3.6), flash-3.5 (prior Flash), flash-3 (prior
                  preview), pro, lite, stable-flash, stable-pro
              temperature: Sampling temperature (0.0–1.0)
              max_output_tokens: Maximum tokens in response. Note: with thinking
                  models (Gemini 3.x), thinking tokens consume part of this budget;
                  set generously or use thinking_level='minimal' for non-reasoning
                  tasks.
              top_p: Nucleus sampling parameter
              top_k: Top-k sampling parameter
              image_path: Optional path to a media file for multi-modal input.
                  Despite the name, accepts image, AUDIO, or video (audio verified
                  2026-07-21). Requires the CF Gateway path for non-image media.
                  Inline only — files >15MB need the Files API (not implemented).
              thinking_level: Reasoning budget for thinking models (Gemini 3.x).
                  One of 'minimal', 'low', 'medium', 'high'. Default (None) lets
                  the model use its built-in default — 'medium' for 3.x Flash
                  (incl. 3.8), which silently eats output budget. Set 'minimal' for tasks that
                  don't need reasoning (transcription, classification, extraction).
                  Flash 3.7 and 3.8 do not accept 'minimal' (HTTP 400); on those the
                  client downgrades it to 'low', the cheapest level they take, and
                  says so on stderr. Ignored by 2.5 models (which use a different
                  parameter).
      
          Returns:
              Response text if successful, None if error
          """
          model_id = _resolve_model(model)
          cf_creds = get_cf_credentials()
      
          # thinking_level is a Gemini 3.x feature. 2.5 (and earlier) models 400 on it.
          # Silently drop for non-3.x so callers can pass it uniformly without branching.
          if thinking_level is not None and not model_id.startswith("gemini-3"):
              thinking_level = None
          if thinking_level == "minimal" and model_id in _MINIMAL_THINKING_UNSUPPORTED:
              print(
                  f"Note: {model_id} does not support thinking_level='minimal'; using 'low'.",
                  file=sys.stderr,
              )
              thinking_level = "low"
      
          max_retries = 3
          for attempt in range(max_retries):
              try:
                  if cf_creds and HAS_REQUESTS:
                      # --- Cloudflare AI Gateway path ---
                      contents = _build_contents(prompt, image_path)
                      gen_cfg: dict = {"temperature": temperature}
                      if max_output_tokens:
                          gen_cfg["maxOutputTokens"] = max_output_tokens
                      if top_p is not None:
                          gen_cfg["topP"] = top_p
                      if top_k is not None:
                          gen_cfg["topK"] = top_k
                      if thinking_level is not None:
                          gen_cfg["thinkingConfig"] = {"thinkingLevel": thinking_level}
      
                      response = _cf_request(model_id, contents, gen_cfg, cf_creds)
                      return _extract_text(response)
      
                  else:
                      # --- Direct SDK path ---
                      if not _initialize_direct_client():
                          return None
      
                      gen_cfg_sdk = {"temperature": temperature}
                      if max_output_tokens:
                          gen_cfg_sdk["max_output_tokens"] = max_output_tokens
                      if top_p is not None:
                          gen_cfg_sdk["top_p"] = top_p
                      if top_k is not None:
                          gen_cfg_sdk["top_k"] = top_k
                      if thinking_level is not None:
                          # SDK uses snake_case; mapped to thinking_config later.
                          gen_cfg_sdk["thinking_config"] = {"thinking_level": thinking_level}
      
                      model_instance = genai.GenerativeModel(
                          model_name=model_id,
                          generation_config=gen_cfg_sdk,
                      )
                      content = _build_genai_content(prompt, image_path)
                      response = model_instance.generate_content(content)
                      return response.text
      
              except (MediaInputError, _NonRetriableAPIError):
                  raise  # deterministic error — retrying cannot help
              except Exception as e:
                  if attempt < max_retries - 1:
                      wait_time = 2 ** attempt
                      print(f"Retry {attempt + 1}/{max_retries} after {wait_time}s: {e}")
                      time.sleep(wait_time)
                  else:
                      print(f"Error invoking Gemini: {e}")
                      return None
      
          return None
      
      
      def generate_image(
          prompt: str,
          output_path: str | None = None,
          model: str = "nano-banana-2",
          temperature: float = 0.7,
      ) -> dict | None:
          """
          Generate an image using a Gemini image model.
      
          Sends a prompt with responseModalities ["IMAGE", "TEXT"] and saves the
          resulting image to disk.
      
          Args:
              prompt: Text prompt describing the desired image
              output_path: Where to save the PNG. If None, auto-generates under
                  /mnt/user-data/outputs/ (or /tmp/ if that doesn't exist)
              model: Image model name or alias (default: nano-banana-2).
                  Aliases: image, image-pro
              temperature: Sampling temperature (0.0–1.0)
      
          Returns:
              dict with keys 'path' (str) and 'caption' (str|None) on success,
              None on failure
          """
          import base64
      
          # Resolve model — must land in IMAGE_MODELS
          resolved = model
          if model in MODEL_ALIASES:
              resolved = MODEL_ALIASES[model]
          if resolved not in IMAGE_MODELS:
              raise ValueError(
                  f"Model '{model}' is not an image model. "
                  f"Use one of: {list(IMAGE_MODELS)} or aliases: image, image-pro"
              )
          model_id = IMAGE_MODELS[resolved]
      
          # Determine output path
          if output_path is None:
              ts = int(time.time())
              out_dir = Path("/mnt/user-data/outputs")
              if not out_dir.exists():
                  out_dir = Path("/tmp")
              output_path = str(out_dir / f"gemini_image_{ts}.png")
      
          cf_creds = get_cf_credentials()
      
          max_retries = 3
          for attempt in range(max_retries):
              try:
                  if cf_creds and HAS_REQUESTS:
                      # --- Cloudflare AI Gateway path ---
                      contents = [{"parts": [{"text": prompt}]}]
                      gen_cfg = {
                          "temperature": temperature,
                          "responseModalities": ["IMAGE", "TEXT"],
                      }
                      response = _cf_request(model_id, contents, gen_cfg, cf_creds)
      
                  elif HAS_GENAI:
                      # --- Direct SDK path ---
                      if not _initialize_direct_client():
                          return None
                      model_instance = genai.GenerativeModel(
                          model_name=model_id,
                          generation_config={
                              "temperature": temperature,
                              "response_modalities": ["IMAGE", "TEXT"],
                          },
                      )
                      response_obj = model_instance.generate_content(prompt)
                      # Convert SDK response to REST-like dict for unified extraction
                      response = _sdk_response_to_dict(response_obj)
      
                  else:
                      print("Error: no credentials configured and google-generativeai not installed")
                      return None
      
                  # Extract image and optional caption from response
                  image_data = None
                  caption = None
                  candidates = response.get("candidates", [])
                  if not candidates:
                      print("Error: no candidates in response")
                      return None
      
                  parts = candidates[0].get("content", {}).get("parts", [])
                  for part in parts:
                      if "inlineData" in part:
                          image_data = part["inlineData"].get("data")
                      elif "inline_data" in part:
                          image_data = part["inline_data"].get("data")
                      elif "text" in part:
                          caption = part["text"]
      
                  if not image_data:
                      print("Error: no image data in response")
                      return None
      
                  # Decode and save
                  Path(output_path).parent.mkdir(parents=True, exist_ok=True)
                  Path(output_path).write_bytes(base64.b64decode(image_data))
                  print(f"Image saved to {output_path}")
                  return {"path": output_path, "caption": caption}
      
              except (MediaInputError, _NonRetriableAPIError):
                  raise  # deterministic error — retrying cannot help
              except Exception as e:
                  if attempt < max_retries - 1:
                      wait_time = 2 ** attempt
                      print(f"Retry {attempt + 1}/{max_retries} after {wait_time}s: {e}")
                      time.sleep(wait_time)
                  else:
                      print(f"Error generating image: {e}")
                      return None
      
          return None
      
      
      def _rest_request(method: str, path: str, body: dict | None = None,
                        params: dict | None = None, timeout: int = 180) -> dict:
          """Call a v1beta REST path (e.g. 'interactions', 'voices') via the CF
          gateway, or directly with GOOGLE_API_KEY in a header (never in the URL).
          Retries 429/5xx; a 4xx raises with Google's error body."""
          cf = get_cf_credentials()
          if cf:
              url = (f"{_CF_GATEWAY_BASE}/{cf['CF_ACCOUNT_ID']}/{cf['CF_GATEWAY_ID']}"
                     f"/google-ai-studio/v1beta/{path}")
              headers = {"cf-aig-authorization": f"Bearer {cf['CF_API_TOKEN']}"}
              if cf.get("GOOGLE_API_KEY"):
                  headers["x-goog-api-key"] = cf["GOOGLE_API_KEY"]
          else:
              url = f"https://generativelanguage.googleapis.com/v1beta/{path}"
              headers = {"x-goog-api-key": get_google_api_key()}
          headers["Content-Type"] = "application/json"
          for attempt in range(4):
              r = requests.request(method, url, json=body, params=params, headers=headers, timeout=timeout)
              if r.status_code in (429, 500, 502, 503, 504) and attempt < 3:
                  time.sleep(1.0 * 2 ** attempt)
                  continue
              if r.status_code >= 400:
                  raise _NonRetriableAPIError(f"HTTP {r.status_code} from Gemini: {(r.text or '')[:600]}")
              return r.json()
          raise RuntimeError("unreachable")
      
      
      def generate_speech(
          text: str,
          output_path: str | None = None,
          voice: str = "Charon",
          style: str | None = None,
          model: str = "tts",
          sample_rate: int = 24000,
      ) -> dict | None:
          """
          Synthesize speech with Gemini 3.8 Flash TTS and save it as WAV.
      
          Uses the Interactions API. generateContent also returns audio for these
          models, but it has no style channel: a "Style: text" prefix is SPOKEN, and
          systemInstruction is refused ("Developer instruction is not enabled for
          this model"). Style goes in a speech_metadata annotation here.
      
          Args:
              text: What to say. Inline events work: <laugh>, <sigh>, <breath>,
                  <short pause>; CAPITALS stress a word.
              output_path: WAV path; defaults under /mnt/user-data/outputs/ or /tmp/.
              voice: A prebuilt name ("Charon", "Algenib", "en-gb-storyteller-4" —
                  see list_voices()) or a designed voice id ("voice_...").
              style: Turn-level delivery, e.g. "quiet and dry, unhurried".
              model: "tts" (gemini-3.8-flash-tts) or "tts-lite".
              sample_rate: Output rate in Hz (24000 default; mono 16-bit PCM).
      
          Returns:
              {'path', 'seconds', 'audio_tokens'} on success, None on failure.
          """
          import base64
          resolved = SPEECH_ALIASES.get(model, model)
          if resolved not in SPEECH_MODELS:
              print(f"Error: '{model}' is not a speech model. Known: {list(SPEECH_MODELS) + list(SPEECH_ALIASES)}",
                    file=sys.stderr)
              return None
          part = {"type": "text", "text": text}
          if style:
              part["annotations"] = [{"type": "speech_metadata", "style": style}]
          body = {"model": resolved,
                  "input": [{"type": "user_input", "content": [part]}],
                  "response_format": {"type": "audio", "mime_type": "audio/wav", "sample_rate": sample_rate},
                  "generation_config": {"speech_config": [{"voice": voice}]}}
          try:
              j = _rest_request("POST", "interactions", body)
              audio = [c for st in j.get("steps", []) for c in st.get("content", []) if c.get("type") == "audio"]
              if not audio:
                  print(f"Error: no audio in response: {json.dumps(j)[:300]}", file=sys.stderr)
                  return None
              data = base64.b64decode(audio[-1]["data"])
          except Exception as e:
              print(f"Error: speech generation failed: {e}", file=sys.stderr)
              return None
          if output_path is None:
              out_dir = Path("/mnt/user-data/outputs") if Path("/mnt/user-data/outputs").exists() else Path("/tmp")
              output_path = str(out_dir / f"gemini_speech_{int(time.time())}.wav")
          Path(output_path).write_bytes(data)
          # WAV body is 16-bit mono PCM after a 44-byte RIFF header
          return {"path": output_path, "seconds": max(0, len(data) - 44) / (2 * sample_rate),
                  "audio_tokens": j.get("usage", {}).get("total_output_tokens")}
      
      
      def design_voice(description: str, display_name: str, gender: str | None = None,
                       language_code: str = "en-US", model: str = "tts") -> dict | None:
          """
          Create a stored voice from a natural-language description (voice design).
      
          POST v1beta/voices with type "prompted". Prompted voices must be stored
          (store=true): the id ('voice_...') lasts one year, max 200 per project.
          The description shapes delivery too — "thoughtful pauses" in it produced
          1-1.8 s mid-line pauses that no per-line style removed (2026-09-24).
      
          Returns:
              {'id', 'expire_time', 'sample_path'} on success, None on failure.
          """
          import base64
          voice = {"model": SPEECH_ALIASES.get(model, model), "type": "prompted",
                   "display_name": display_name, "language_code": language_code,
                   "prompted": {"input": description}}
          if gender:
              voice["gender"] = gender
          try:
              j = _rest_request("POST", "voices", {"store": True, "voice": voice})
          except Exception as e:
              print(f"Error: voice design failed: {e}", file=sys.stderr)
              return None
          sample = (j.get("sample_audio") or {}).get("data")
          sample_path = None
          if sample:
              sample_path = f"/tmp/{j['id']}_sample.wav"
              Path(sample_path).write_bytes(base64.b64decode(sample))
          return {"id": j.get("id"), "expire_time": j.get("expire_time"), "sample_path": sample_path}
      
      
      def list_voices(**filters) -> list | None:
          """
          List the voice library (2,089 prebuilt voices as of 2026-09-24, plus this
          project's designed voices), paging through next_page_token.
      
          Filters are server-side query params: gender ('male'|'female'|'neutral'),
          pitch ('low'|'medium'|'high'), search (substring of the description).
          `accent` needs the exact string ('Winchester English', not 'British'), so
          filter accents client-side on the returned 'accent' field.
      
          Returns:
              list of voice dicts (id, display_name, language_code, accent, gender,
              pitch, persona, description), or None on failure.
          """
          out, token = [], None
          try:
              while True:
                  params = {"page_size": 1000, **filters}
                  if token:
                      params["page_token"] = token
                  j = _rest_request("GET", "voices", params=params, timeout=60)
                  out += j.get("voices", [])
                  token = j.get("next_page_token")
                  if not token:
                      return out
          except Exception as e:
              print(f"Error: listing voices failed: {e}", file=sys.stderr)
              return None
      
      
      def _sdk_response_to_dict(response_obj) -> dict:
          """Convert a google.generativeai SDK response to a REST-like dict.
      
          This allows generate_image() to use the same extraction logic for both
          the CF Gateway (REST) and direct SDK paths.
      
          Args:
              response_obj: GenerateContentResponse from the SDK
      
          Returns:
              dict matching the Gemini REST API response shape
          """
          import base64
      
          candidates = []
          for candidate in response_obj.candidates:
              parts = []
              for part in candidate.content.parts:
                  if hasattr(part, "text") and part.text:
                      parts.append({"text": part.text})
                  elif hasattr(part, "inline_data") and part.inline_data:
                      data = part.inline_data
                      parts.append({
                          "inlineData": {
                              "mimeType": data.mime_type,
                              "data": base64.b64encode(data.data).decode()
                              if isinstance(data.data, bytes)
                              else data.data,
                          }
                      })
              candidates.append({"content": {"parts": parts}})
          return {"candidates": candidates}
      
      
      def invoke_with_structured_output(
          prompt: str,
          pydantic_model: type,
          model: str = DEFAULT_MODEL,
          temperature: float = 0.7,
          image_path: str | None = None,
          max_output_tokens: int | None = 32768,
      ) -> object | None:
          """
          Invoke Gemini with structured (JSON schema) output using a Pydantic model.
      
          Args:
              prompt: The text prompt to send
              pydantic_model: Pydantic model class for response schema
              model: Model name or alias (default: DEFAULT_MODEL, gemini-3.8-flash).
                  Aliases: flash, pro, lite, stable-flash, stable-pro
              temperature: Sampling temperature (0.0–1.0)
              image_path: Optional path to a media file for multi-modal input.
                  Despite the name, accepts image, AUDIO, or video (audio verified
                  2026-07-21). Requires the CF Gateway path for non-image media.
                  Inline only — files >15MB need the Files API (not implemented).
      
          Returns:
              Instance of pydantic_model if successful, None if error
          """
          if not HAS_PYDANTIC:
              print("Error: pydantic not installed. Run: uv pip install pydantic")
              return None
      
          model_id = _resolve_model(model)
          cf_creds = get_cf_credentials()
      
          max_retries = 3
          for attempt in range(max_retries):
              try:
                  if cf_creds and HAS_REQUESTS:
                      # --- Cloudflare AI Gateway path ---
                      contents = _build_contents(prompt, image_path)
                      schema = _pydantic_to_schema(pydantic_model)
                      gen_cfg = {
                          "temperature": temperature,
                          "responseMimeType": "application/json",
                          "responseSchema": schema,
                      }
                      # Thinking tokens count against this budget, so a default sized
                      # for the visible answer truncates the JSON mid-object and
                      # surfaces as a confusing JSONDecodeError rather than a length
                      # error. Default generously.
                      if max_output_tokens:
                          gen_cfg["maxOutputTokens"] = max_output_tokens
                      response = _cf_request(model_id, contents, gen_cfg, cf_creds)
                      _check_truncation(response)
                      text = _extract_text(response)
                      if text:
                          json_data = json.loads(text)
                          return pydantic_model(**json_data)
      
                  else:
                      # --- Direct SDK path ---
                      if not _initialize_direct_client():
                          return None
      
                      model_instance = genai.GenerativeModel(
                          model_name=model_id,
                          generation_config={
                              "temperature": temperature,
                              "response_mime_type": "application/json",
                              "response_schema": pydantic_model,
                          },
                      )
                      content = _build_genai_content(prompt, image_path)
                      response = model_instance.generate_content(content)
                      json_data = json.loads(response.text)
                      return pydantic_model(**json_data)
      
              except (MediaInputError, _NonRetriableAPIError):
                  raise  # deterministic error — retrying cannot help
              except Exception as e:
                  if attempt < max_retries - 1:
                      wait_time = 2 ** attempt
                      print(f"Retry {attempt + 1}/{max_retries} after {wait_time}s: {e}")
                      time.sleep(wait_time)
                  else:
                      print(f"Error invoking Gemini with structured output: {e}")
                      return None
      
          return None
      
      
      def invoke_parallel(
          prompts: list,
          model: str = DEFAULT_MODEL,
          temperature: float = 0.7,
          max_workers: int = 5,
      ) -> list:
          """
          Invoke Gemini with multiple prompts in parallel.
      
          Args:
              prompts: List of text prompts to process
              model: Model name or alias (default: DEFAULT_MODEL, gemini-3.8-flash).
                  Aliases: flash, pro, lite, stable-flash, stable-pro
              temperature: Sampling temperature (0.0–1.0)
              max_workers: Maximum concurrent requests
      
          Returns:
              List of response strings (None for failed requests) in prompt order
          """
          from concurrent.futures import ThreadPoolExecutor, as_completed
      
          results: list = [None] * len(prompts)
      
          def _process(idx: int, prompt: str):
              return idx, invoke_gemini(prompt, model=model, temperature=temperature)
      
          with ThreadPoolExecutor(max_workers=max_workers) as executor:
              futures = {
                  executor.submit(_process, idx, prompt): idx
                  for idx, prompt in enumerate(prompts)
              }
              for future in as_completed(futures):
                  try:
                      idx, response = future.result()
                      results[idx] = response
                  except Exception as e:
                      idx = futures[future]
                      print(f"Error processing prompt {idx}: {e}")
                      results[idx] = None
      
          return results
      
      
      def get_available_models() -> dict:
          """Return dict of registered Gemini models grouped by category.
      
          Returns:
              dict with keys 'text', 'image', 'speech', 'aliases', 'speech_aliases'
          """
          return {
              "text": list(MODELS.keys()),
              "image": list(IMAGE_MODELS.keys()),
              "speech": list(SPEECH_MODELS.keys()),
              "aliases": dict(MODEL_ALIASES),
              "speech_aliases": dict(SPEECH_ALIASES),
          }
      
      
      def verify_setup() -> bool:
          """
          Verify that Gemini client is properly configured.
      
          Returns:
              True if at least one credential source is valid and a test call succeeds
          """
          cf_creds = get_cf_credentials()
          if cf_creds:
              print(f"Using Cloudflare AI Gateway (account: {cf_creds['CF_ACCOUNT_ID'][:8]}...)")
          elif HAS_GENAI:
              if not _initialize_direct_client():
                  return False
              print("Using direct Google API (google-generativeai SDK)")
          else:
              print("Error: no credentials configured and google-generativeai not installed")
              return False
      
          try:
              test_response = invoke_gemini("Say 'OK'", model=DEFAULT_MODEL)
              return test_response is not None
          except Exception as e:
              print(f"Setup verification failed: {e}")
              return False
      
      
      # ---------------------------------------------------------------------------
      # Self-test
      # ---------------------------------------------------------------------------
      
      if __name__ == "__main__":
          import sys
      
          print("Gemini Client Self-Test")
          print("=" * 50)
      
          cf = get_cf_credentials()
          if cf:
              print(f"Backend: Cloudflare AI Gateway ({cf['CF_ACCOUNT_ID'][:8]}.../{cf['CF_GATEWAY_ID']})")
          elif HAS_GENAI:
              print("Backend: Direct Google API (google-generativeai SDK)")
          else:
              print("ERROR: no credentials and google-generativeai not installed")
              sys.exit(1)
      
          print("\n1. Verifying setup...")
          if verify_setup():
              print("   ✓ Setup verified")
          else:
              print("   ✗ Setup failed")
              sys.exit(1)
      
          print("\n2. Available models:")
          available = get_available_models()
          for category, items in available.items():
              if isinstance(items, dict):
                  print(f"   {category}:")
                  for alias, target in items.items():
                      print(f"     {alias} → {target}")
              else:
                  print(f"   {category}: {', '.join(items)}")
      
          print("\n3. Testing basic invocation...")
          resp = invoke_gemini("What is 2+2? Answer in one word.", model=DEFAULT_MODEL)
          if resp:
              print(f"   Response: {resp.strip()}")
          else:
              print("   ✗ Invocation failed")
              sys.exit(1)
      
          if HAS_PYDANTIC:
              print("\n4. Testing structured output...")
              from pydantic import BaseModel as PM
              from pydantic import Field
      
              class MathAnswer(PM):
                  result: int = Field(description="The numerical result")
                  explanation: str = Field(description="Brief explanation")
      
              structured = invoke_with_structured_output(
                  prompt="What is 5+7? Provide result and explanation.",
                  pydantic_model=MathAnswer,
                  model=DEFAULT_MODEL,
              )
              if structured:
                  print(f"   Result: {structured.result}")
                  print(f"   Explanation: {structured.explanation}")
              else:
                  print("   ✗ Structured output failed")
      
          print("\n5. Testing parallel invocation...")
          test_prompts = [
              "Capital of France? One word.",
              "Capital of Japan? One word.",
              "Capital of Brazil? One word.",
          ]
          parallel_results = invoke_parallel(test_prompts, model=DEFAULT_MODEL)
          for prompt, result in zip(test_prompts, parallel_results):
              status = result.strip() if result else "Failed"
              print(f"   {prompt[:35]}... → {status}")
      
          print("\n" + "=" * 50)
          print("Self-test complete!")
      
  • CHANGELOG.md 10.9 KB
    # invoking-gemini - Changelog
    
    ## 2026-09-24
    
    ### Added — speech generation (Gemini 3.8 Flash TTS, GA 2026-09-23)
    - `generate_speech()` writes a WAV from text with a prebuilt or designed voice
      and an optional style; `design_voice()` creates a stored voice from a
      description; `list_voices()` pages the 2,089-voice library. Run through the
      CF gateway on 2026-09-24, they returned a 7.8 s WAV, a stored voice id with
      its sample, and 2,092 voices (the library plus this project's designs).
    - `SPEECH_MODELS` / `SPEECH_ALIASES` (`tts`, `tts-lite`) are kept apart from
      `MODEL_ALIASES`, so `invoke_gemini()` cannot resolve to an audio model.
    - The calls use the Interactions API. On `generateContent` a "Style:" prefix is
      spoken and `systemInstruction` is refused.
    - `_rest_request()` sends a direct-mode API key in the `x-goog-api-key` header
      rather than the URL.
    
    ## 2026-09-03
    
    ### ⚠️ BREAKING — `flash` alias and `DEFAULT_MODEL` repointed to gemini-3.8-flash
    - Gemini 3.8 Flash reached GA on 2026-09-02. It is now the registry default and
      the `flash` alias. Google also shipped 3.7 Flash on 2026-08-13, which this
      skill missed; both are added to `MODELS`.
    - New pinned aliases `flash-3.7` and `flash-3.6`. `flash-3.5` and `flash-3`
      are unchanged. Nothing is removed; no 3.x Flash has a shutdown date.
    - Price is the same as 3.6 on Google's current page ($0.75 in / $3.75 out
      through 2026-12-31, then $1.50 / $7.50), so the swap costs nothing per token.
      Google says 3.8 "works harder" at higher effort levels, so per-task thinking
      tokens may go up.
    
    ### ⚠️ BREAKING — `pro` alias repointed to gemini-3.8-flash; Pro tier off routing
    - Oskar, 2026-09-03: "We should never use 3.1 Pro, its Pareto efficiency is
      too poor compared to the later Flash models." `gemini-3.1-pro-preview` costs
      2.7× / 3.2× (in / out) what 3.8 Flash does at today's rates and the 3.5+
      Flash line already beat it on coding and agentic benchmarks. The ID stays in
      `MODELS` for pinned callers; the table marks it DEPRECATED. "Maximum
      reasoning" is now Flash with `thinking_level='high'`. A future 3.5 Pro gets
      the same test before any alias points at it.
    
    ### Changed — `thinking_level='minimal'` on 3.7 / 3.8 Flash
    - Both models return HTTP 400 for `minimal` (verified live 2026-09-03; 3.6,
      3.5 and 3.5-lite accept it). `invoke_gemini()` now downgrades `minimal` to
      `low` on those two model IDs and prints a one-line note to stderr, so callers
      that pass `minimal` uniformly (transcription, classification) keep working
      on the new default instead of burning a non-retriable 400 and returning
      `None`. Measured on 3.8: `low` spent 0 thinking tokens on a one-word reply,
      the default `medium` spent 79. On 3.7, `low` still spent 45–88 on the same
      prompt, so the downgrade is not free there — budget output generously.
    - For a true no-thinking pass, pin `flash-3.6` or `lite`.
    
    ### Fixed — stale figures in the model tables
    - gemini-3.1-pro-preview output above 200K is $18.00 / 1M, not $24.00.
    - gemini-3.6-flash pricing carried the flat $1.50 / $7.50 from its launch
      table; Google's pricing page now lists it on the same introductory schedule
      as 3.7 and 3.8.
    - 3.5 Pro was still described as "slated for June 2026"; it has not shipped
      as of 2026-09-03.
    - Deprecation table gains the 3.x preview rows and the 2026-10-02 shutdown of
      `gemini-2.5-flash-image` (the `nano-banana` alias target).
    - Two docstrings still named `gemini-3-flash-preview` as the default.
    
    ## 2026-07-29
    
    ### Fixed — nested Pydantic models raised a bare HTTP 400
    - `_pydantic_to_schema()` returned pydantic's `model_json_schema()` nearly
      verbatim, which emits `$defs` + `$ref` for every nested model. Gemini's
      `responseSchema` does not support `$ref`/`$defs`, so **any** model containing
      another model (`list[Finding]`, the common case) produced a 400 with no usable
      body, burned three retries, and returned `None`. Flat single-level models
      worked, which is why this went unnoticed. Refs are now inlined and the
      keywords Gemini rejects are stripped recursively rather than only at the top
      level (`title`, `default`, `additionalProperties`, `discriminator`, `examples`,
      `const`).
    - `Optional[X]` / `anyOf` is now translated to the non-null branch plus
      `nullable: true` instead of being passed through as `anyOf`, which Gemini also
      rejects.
    - 4xx responses now raise `_NonRetriableAPIError` carrying the **response body**.
      `raise_for_status()` discarded it, and Gemini puts the only useful diagnostic
      there — which schema keyword it refused. 4xx is deterministic, so it no longer
      burns the retry budget either.
    - `invoke_with_structured_output()` gained `max_output_tokens` (default 32768).
      Thinking tokens count against the output budget, so a small cap truncated the
      JSON mid-object and surfaced as a pydantic `EOF while parsing` that reads like
      a schema error. `finishReason=MAX_TOKENS` is now detected and reported as
      truncation.
    
    ## 2026-07-21
    
    ### Added — audio (and video) input
    - `image_path` now accepts **any** supported media file, not just images. Audio
      input verified working 2026-07-21 (3-beep WAV; `gemini-3.6-flash` returned the
      correct count and pitch direction). The param keeps its legacy name for
      back compat; docstrings now state what it really accepts.
    - Explicit mimeType overrides for audio/video extensions `mimetypes` guesses
      wrongly or misses (.m4a/.aac/.flac/.ogg/.opus/.mp3/.wav/.aiff/.mp4/.mov/.webm/
      .heic/.heif). A wrong guess previously sent a bad mimeType and produced a
      confused answer rather than an error.
    - New `MediaInputError` (subclass of `ValueError`) for deterministic input
      problems; retry loops re-raise it immediately instead of burning 3 attempts.
    - 15MB inline cap enforced with an actionable message. Larger files need the
      Files API, which this client still does not implement.
    - The direct google-generativeai SDK fallback now rejects non-image media with a
      clear message instead of an opaque PIL `UnidentifiedImageError`. Audio/video
      require the CF Gateway path.
    
    **Routing note:** send audio to `gemini-3.6-flash`. `gemini-3.5-flash-lite` was
    unreliable on the same clip — it reported three beeps then two on identical
    input, and got the pitch direction wrong both times.
    
    ### ⚠️ BREAKING — `lite` alias repointed
    - `MODEL_ALIASES['lite']`: `gemini-2.5-flash-lite` → `gemini-3.5-flash-lite`.
      Output cost goes from $0.40 to $2.50 per 1M (~6x) for any caller using
      `model="lite"`. Pin `gemini-2.5-flash-lite` by full ID if you need the old
      rate, though it is deprecated (below).
    - The Gemini 2.5 **text** generation is retired from routing: `gemini-2.5-flash`,
      `gemini-2.5-flash-lite`, `gemini-2.5-pro`. Model IDs stay callable and the
      `stable-flash` / `stable-pro` aliases still resolve, so nothing hard-breaks,
      but they are no longer recommended targets. Image model `nano-banana`
      (gemini-2.5-flash-image) is NOT affected.
    - Added `gemini-3.5-flash-lite` (GA 2026-07-21) as the cheap/bulk tier.
    
    - Gemini 3.6 Flash (`gemini-3.6-flash`) reached GA (2026-07-21). Added it to
      the model registry and made it the new `DEFAULT_MODEL`.
    - Repointed the `flash` alias from `gemini-3.5-flash` to `gemini-3.6-flash`.
      Added `flash-3.5` as a stable handle for the prior frontier Flash; `flash-3`
      still pins the older `gemini-3-flash-preview`.
    - Rationale: 3.6 Flash is ~half Sonnet's cost (in $1.50 / out $7.50 vs ~$3 /
      ~$15) and improves coding/agentic quality with ~17% fewer output tokens than
      3.5 Flash — making it the default for sub-agent delegation.
    - Updated SKILL.md and references/models.md tables (pricing, 64K output cap,
      benchmark deltas, tone-regression caveat).
    - Not touched: `gemini-3.5-flash-lite` / `gemini-3.5-flash-cyber` (shipped same
      day) are not yet aliased; two helper fns still hardcode a
      `gemini-3-flash-preview` default in their signatures (pre-existing).
    
    ## 2026-05-28
    - Nano Banana 2 (`gemini-3.1-flash-image-preview`) and Nano Banana Pro
      (`gemini-3-pro-image-preview`) reached GA on Vertex / Gemini Enterprise.
    - Kept the `-preview` model IDs: the Gemini Developer API surface this client
      uses still serves both under `-preview` (GA IDs without the suffix are
      Vertex-only and 404 here). Verified against the live image-generation docs.
    - Documented capabilities: 512/1K/2K GA + 4K preview, up to 14 reference
      images, Search + Image-Search grounding (3.1 Flash), thinking_level control.
    - Noted video-as-input is a Vertex preview only; not available on the Developer API.
    
    
    All notable changes to the `invoking-gemini` skill are documented in this file. The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.0.0/).
    
    ## [0.9.0] - 2026-09-24
    
    ### Other
    
    - invoking-gemini 0.9.0: add Gemini 3.8 Flash TTS (speech generation)
    
    ## [0.8.0] - 2026-09-03
    
    ### Other
    
    - invoking-gemini: Gemini 3.8 Flash is the default, add 3.7, retire Pro from routing, handle minimal-thinking 400 (#783)
    
    ## [0.7.0] - 2026-07-23
    
    ### Other
    
    - invoking-gemini: default to gemini-3.6-flash, retire Gemini 2.5 text models (#741)
    - invoking-gemini: Nano Banana 2/Pro GA — keep -preview IDs on Developer API surface (#676)
    
    ## [0.6.0] - 2026-05-23
    
    ### Added
    
    - add Gemini 3.5 Flash + thinking_level, fix stale model docs (#669)
    
    ### Fixed
    
    - retry on egress-proxy 503 ('DNS cache overflow') in remembering + invoking-gemini (#580)
    
    ### Other
    
    - Remove _MAP.md files, direct agents to tree-sitting for code navigation (#545)
    
    ## [0.5.0] - 2026-03-31
    
    ### Added
    
    - surface Image Generation, add examples (#520)
    - add mapping-features skill for behavioral web app documentation (#432)
    
    ### Other
    
    - Regenerate _MAP.md files after @lat: backlink insertion (#504)
    - Lattice v2: bidirectional source-anchored knowledge graph (#503)
    
    ## [0.3.1] - 2026-03-01
    
    ### Fixed
    
    - use camelCase keys for Gemini REST API inline data
    
    ## [0.3.0] - 2026-03-01
    
    ### Added
    
    - add image generation support + fix IMAGE_MODELS registry
    
    ## [0.3.0] - 2026-03-01
    
    ### Added
    
    - `generate_image()` function for native image generation via Gemini image models
    - `image` and `image-pro` model aliases for image generation
    - `nano-banana` (gemini-2.5-flash-image) to IMAGE_MODELS registry
    - Image generation documentation in SKILL.md with prompt patterns and examples
    
    ### Fixed
    
    - IMAGE_MODELS registry now maps display names to actual API model IDs
      (was mapping names to themselves, causing 404 errors)
    - `nano-banana-2` → `gemini-3.1-flash-image-preview`
    - `nano-banana-pro` → `gemini-3-pro-image-preview`
    - `nano-banana` → `gemini-2.5-flash-image`
    
    ## [0.2.0] - 2026-03-01
    
    ### Added
    
    - update invoking-gemini model registry to current Gemini lineup
    
    ## [0.1.0] - 2026-03-01
    
    ### Added
    
    - route invoking-gemini through Cloudflare AI Gateway
    - add line numbers, markdown ToC, and other files listing
    - add code maps and CLAUDE.md integration guidance
    - Delete VERSION files, complete migration to frontmatter
    - Migrate all 27 skills from VERSION files to frontmatter
    
    ### Changed
    
    - migrate API credential management to project knowledge files
    
    ### Fixed
    
    - limit markdown ToC to h1/h2 headings only
  • README.md 239 B
    # invoking-gemini
    
    Invokes Google Gemini models for structured outputs, multi-modal tasks, and Google-specific features. Use when users request Gemini, structured JSON output, Google API integration, or cost-effective parallel processing.
    
  • SKILL.md 16.1 KB
    ---
    name: invoking-gemini
    description: Invokes Google Gemini models for structured outputs, image generation, text-to-speech narration, multi-modal tasks, and Google-specific features. Use when users request Gemini, image generation, Gemini TTS or a synthesized voice, structured JSON output, Google API integration, or cost-effective parallel processing.
    metadata:
      version: 0.9.0
    ---
    
    # Invoking Gemini
    
    Delegate tasks to Google's Gemini models when they offer advantages over Claude.
    
    ## When to Use Gemini
    
    **Image generation:**
    - Blog header images, illustrations, diagrams
    - Style-guided image creation (risograph, editorial, etc.)
    - Text rendering in images
    
    **Speech (TTS):**
    - Narration, voice-over, read-aloud with style direction per line
    - A custom voice designed from a written description
    - Two-speaker dialogue
    
    **Structured outputs:**
    - JSON Schema validation with property ordering guarantees
    - Pydantic model compliance
    - Strict schema adherence (enum values, required fields)
    
    **Cost optimization:**
    - Parallel batch processing (Gemini 3 Flash is lightweight)
    - High-volume simple tasks
    
    **Multi-modal tasks:**
    - Image analysis with JSON output
    - Video processing
    - Audio transcription with structure
    
    ## Setup
    
    ```bash
    uv pip install requests pydantic
    ```
    
    **Credentials — Option A (recommended): Cloudflare AI Gateway**
    
    Source `/mnt/project/proxy.env` with `CF_ACCOUNT_ID`, `CF_GATEWAY_ID`, `CF_API_TOKEN`.
    Requests route through Cloudflare AI Gateway, bypassing IP blocks. Google API key stored in gateway via BYOK.
    
    **Credentials — Option B: Direct Google API**
    
    If no `proxy.env`, falls back to direct: `GOOGLE_API_KEY.txt` or `API_CREDENTIALS.json`.
    
    ## Image Generation
    
    Generate images using Gemini's native image models. This is the primary way to create illustrations, blog headers, diagrams, and visual content.
    
    ### Quick Start
    
    ```python
    import sys
    sys.path.append('/mnt/skills/user/invoking-gemini/scripts')
    from gemini_client import generate_image
    
    # One call — returns {"path": "...", "caption": "..."} or None
    result = generate_image("A watercolor painting of a mountain lake at sunset")
    print(result["path"])  # /mnt/user-data/outputs/gemini_image_1740000000.png
    ```
    
    ### Function Signature
    
    ```python
    generate_image(
        prompt: str,                    # The image description
        output_path: str = None,        # Auto-generates if omitted
        model: str = "nano-banana-2",   # Default: fast. Use "image-pro" for quality
        temperature: float = 0.7,       # 0.5-0.7 for diagrams, 0.7-0.8 for illustrations
    ) -> dict | None
    # Returns: {"path": "/mnt/user-data/outputs/gemini_image_*.png", "caption": str|None}
    # Returns None on failure
    ```
    
    ### Model Selection
    
    | Alias | Model | Best For | Cost/image |
    |-------|-------|----------|------------|
    | `"nano-banana-2"` or `"image"` | gemini-3.1-flash-image-preview | Fast iteration, drafts | $0.067 |
    | `"image-pro"` or `"nano-banana-pro"` | gemini-3-pro-image-preview | Published content, text rendering | $0.134 |
    
    ### Complete Blog Header Example
    
    ```python
    import sys
    sys.path.append('/mnt/skills/user/invoking-gemini/scripts')
    from gemini_client import generate_image
    
    # 1. Compose prompt with style prefix + subject
    style_prefix = (
        "Style: Risograph-inspired editorial illustration. "
        "Visible halftone dot texture and slight color misregistration between layers. "
        "Limited ink palette: deep indigo, warm coral, and sage green on off-white paper. "
        "Layered transparency where colors overlap creates rich secondary tones. "
        "Modern and professional — the aesthetic of an indie design studio, not a fantasy novel. "
        "Generous whitespace. No photorealism, no glow effects, no cyberpunk. No text or labels."
    )
    subject = "A raven perched on a stack of books, observing a network graph"
    prompt = f"{style_prefix}\n\nSubject: {subject}. Wide landscape format, suitable as a blog header."
    
    # 2. Generate (use image-pro for published content)
    result = generate_image(prompt, model="image-pro", temperature=0.75)
    
    if result:
        print(f"Saved: {result['path']}")
        # 3. Present to user
        # present_files([result["path"]])
    ```
    
    ### Prompt Patterns
    
    - **Style prefix + subject**: Prepend a style description, then describe the subject
    - **Be specific about style**: "Risograph-inspired editorial illustration" not "a nice picture"
    - **Include composition**: "Wide landscape format" / "centered, high contrast"
    - **Text rendering**: "A poster with the text 'SALE' in bold red letters" (works well with image-pro)
    - **Negative constraints**: "No photorealism, no glow effects" to avoid defaults
    
    ### Custom Output Path
    
    ```python
    result = generate_image(
        "A logo for a coffee shop called 'Bean There'",
        output_path="/mnt/user-data/outputs/coffee_logo.png"
    )
    ```
    
    ## Speech Generation (TTS)
    
    Gemini 3.8 Flash TTS and Flash-Lite TTS went GA on 2026-09-23. Output is WAV,
    24 kHz mono 16-bit, SynthID-watermarked.
    
    ```python
    from gemini_client import generate_speech, design_voice, list_voices
    
    r = generate_speech("Odin kept two ravens. <short pause> Huginn was thought.",
                        output_path="/tmp/line.wav", voice="Algenib",
                        style="quiet and dry, unhurried")
    # {'path': '/tmp/line.wav', 'seconds': 4.2, 'audio_tokens': 135} or None
    
    v = design_voice("A low, dry, quietly amused male voice with a faint rasp. "
                     "Soft southern British accent.", "narrator", gender="male",
                     language_code="en-GB")      # {'id': 'voice_...', 'sample_path': ...}
    generate_speech("...", voice=v["id"])
    
    lows = list_voices(gender="male", pitch="low")   # library of 2,089 prebuilt voices
    ```
    
    - **Voices:** 30 studio voices (`Charon`, `Kore`, `Algenib` "gravelly",
      `Enceladus` "breathy", ...) plus 2,059 persona voices with ids like
      `en-gb-storyteller-4`. `list_voices()` returns accent, pitch, gender and a
      description for each; the `accent` filter needs the exact string ("Winchester
      English"), so filter accents on the returned field.
    - **Style:** pass `style=` (a `speech_metadata` annotation). Do not prefix the
      text with "Style: ..." — the 3.8 models read the prefix aloud, and
      `systemInstruction` is rejected. Inline events go in the text: `<laugh>`,
      `<sigh>`, `<breath>`, `<short pause>`; CAPITALS stress a word.
    - **Designed voices** are stored (1-year expiry, 200 per project). The
      description sets baseline delivery too: "thoughtful pauses" in it produced
      1–1.8 s mid-line pauses that no per-line style removed.
    - **The model can change words.** Seen in a 29-line narration: "Hmm, I get
      things wrong", "tell them" for "tell him". Anything with subtitles or a
      fixed script needs an ASR check (faster-whisper `medium.en`) and a retake.
    - **Cost:** about 32 audio tokens per second of speech, $9/M through
      2026-12-31 on 3.8 Flash TTS ($6/M Lite), so a minute is about $0.02.
    - Voice replication (cloning from a 30 s sample plus a recorded consent clip) is
      not wired in, and is unavailable in the EEA, UK, Switzerland, India, Illinois
      and Texas.
    
    ## Basic Text Usage
    
    ```python
    import sys
    sys.path.append('/mnt/skills/user/invoking-gemini/scripts')
    from gemini_client import invoke_gemini
    
    response = invoke_gemini(
        prompt="Explain quantum computing in 3 bullet points",
        model="flash",  # gemini-3.8-flash (default)
    )
    print(response)
    ```
    
    ## Structured Output
    
    Use Pydantic models for guaranteed JSON Schema compliance:
    
    ```python
    from gemini_client import invoke_with_structured_output
    from pydantic import BaseModel, Field
    
    class BookAnalysis(BaseModel):
        title: str
        genre: str = Field(description="Primary genre")
        key_themes: list[str] = Field(max_length=5)
        rating: int = Field(ge=1, le=5)
    
    result = invoke_with_structured_output(
        prompt="Analyze the book '1984' by George Orwell",
        pydantic_model=BookAnalysis
    )
    print(result.title)  # "1984"
    ```
    
    **Nested models are supported.** Gemini's `responseSchema` rejects `$ref`/`$defs`,
    which pydantic emits for every nested model, so the client inlines them before
    sending:
    
    ```python
    class Finding(BaseModel):
        claim: str
        confidence: Literal["high", "medium", "low"]
        note: str | None = None
    
    class Analysis(BaseModel):
        findings: list[Finding]     # nested — inlined for you
        gaps: list[str]
    ```
    
    **Budget output generously.** Thinking tokens count against `max_output_tokens`
    (default 32768). Too low and the JSON truncates mid-object, which surfaces as a
    pydantic parse error rather than a length error — the client now detects
    `finishReason=MAX_TOKENS` and says so explicitly.
    
    ## Parallel Invocation
    
    ```python
    from gemini_client import invoke_parallel
    
    results = invoke_parallel(
        prompts=["Summarize Hamlet", "Summarize Macbeth", "Summarize Othello"],
        model="lite",  # gemini-3.5-flash-lite — cheap/fast tier for batch
    )
    ```
    
    ## Available Models
    
    The current frontier Flash is **gemini-3.8-flash** (GA 2026-09-02), the
    default and the `flash` alias. Google shipped three Flash generations in six
    weeks: 3.6 (2026-07-21), 3.7 (2026-08-13), 3.8 (2026-09-02). Each stays
    callable under a pinned alias (`flash-3.7`, `flash-3.6`, `flash-3.5`,
    `flash-3`), and none has a shutdown date. `gemini-3.1-flash-lite-preview` from
    earlier docs is gone (shut down 2026-05-25).
    
    The Pro tier is off routing. `gemini-3.1-pro-preview` costs 2.7× the input and
    3.2× the output of 3.8 Flash at today's rates and loses to the 3.5+ Flash line
    on the coding and agentic benchmarks that matter here. Do not target it; the
    `pro` alias now resolves to gemini-3.8-flash, and "maximum reasoning" means
    `thinking_level='high'` on Flash.
    
    ### Text / Reasoning Models
    
    | Model | Alias | Input/1M | Output/1M | Context | Notes |
    |-------|-------|----------|-----------|---------|-------|
    | gemini-3.8-flash | `flash` | $0.75 → $1.50 | $3.75 → $7.50 | 1M in / 64K out | **Default.** GA 2026-09-02. Current frontier Flash. Vs 3.7: Terminal-Bench 2.1 90.8% vs 81.6%, SWE-Bench Pro 61.6% vs 60.4%, SWE-Atlas 51.9% vs 48.0%, HLE flat (45.4% vs 45.7%). Google says it "works harder" at higher effort, so expect more thinking tokens per task. `thinking_level` is low/medium/high only — `minimal` returns HTTP 400 and the client downgrades it to `low`. Default `medium` spent 79 thinking tokens on a one-word reply (measured 2026-09-03); pass `low` for non-reasoning tasks. |
    | gemini-3.7-flash | `flash-3.7` | $0.75 → $1.50 | $3.75 → $7.50 | 1M / 64K | GA 2026-08-13. DeepSWE v1.1 65.3% vs 49.0% on 3.6, Terminal-Bench 2.1 85.8%. Same `minimal` restriction as 3.8. Google keeps it "fully supported for efficiency-first workloads". |
    | gemini-3.6-flash | `flash-3.6` | $0.75 → $1.50 | $3.75 → $7.50 | 1M / 64K | GA 2026-07-21. ~17% fewer output tokens than 3.5 Flash. Last Flash that accepts `thinking_level='minimal'` (verified 2026-09-03). |
    | gemini-3.5-flash | `flash-3.5` | $1.50 | $9.00 | 1M | GA 2026-05-19. Google's model list now labels it "legacy". Accepts `minimal`. Costs more on output than 3.6–3.8. |
    | gemini-3-flash-preview | `flash-3` | $0.30 | $2.50 | 1M | Older preview Flash, kept for back compat. Google's listed migration target for it is gemini-3.6-flash; no shutdown date. |
    | ~~gemini-3.1-pro-preview~~ | — | $2.00 (≤200K) / $4.00 | $12.00 / $18.00 | 1M | **DEPRECATED from routing (2026-09-03).** Price/quality dominated by 3.6+ Flash; 3.5 Flash already beat it on most coding/agentic benchmarks. ID stays callable for pinned code. `pro` now resolves to gemini-3.8-flash. 3.5 Pro was announced at I/O 2026-05-19 for June and is still absent from the API as of 2026-09-03; it gets the same price/quality test before any alias points at it. |
    | gemini-3.5-flash-lite | `lite` | $0.30 | $2.50 | 1M | **Cheap/bulk tier.** GA 2026-07-21. Fastest 3.5-class (350 output tok/sec); beats gemini-3-flash on SWE-Bench Pro and OSWorld-Verified. |
    | ~~gemini-2.5-flash~~ | `stable-flash` | $0.30 | $2.50 | 1M | **DEPRECATED** — 2025-era generation, do not route here. |
    | ~~gemini-2.5-flash-lite~~ | — | $0.10 | $0.40 | 1M | **DEPRECATED** — cheaper, but a 2025-era generation. `lite` now resolves to gemini-3.5-flash-lite. |
    | ~~gemini-2.5-pro~~ | `stable-pro` | $1.25 (≤200K) / $2.50 | $10.00 / $20.00 | 1M | **DEPRECATED** — 2025-era generation, do not route here. |
    
    `$0.75 → $1.50` means introductory pricing: Google's pricing page (fetched
    2026-09-03) lists 3.6, 3.7 and 3.8 Flash at $0.75 in / $3.75 out through
    2026-12-31 and $1.50 / $7.50 from 2027-01-01. Context caching is $0.075 → $0.15;
    Batch is half of standard. Output prices include thinking tokens.
    
    ### Image Models
    
    | Model | Alias | Input/1M | Per Image |
    |-------|-------|----------|-----------|
    | gemini-3.1-flash-image-preview | `image`, `nano-banana-2` | $0.25 | $0.067 |
    | gemini-3-pro-image-preview | `image-pro`, `nano-banana-pro` | $2.00 | $0.134 |
    
    ### Speech Models
    
    | Model | Alias | Input/1M | Output/1M (audio) | Notes |
    |-------|-------|----------|-------------------|-------|
    | gemini-3.8-flash-tts | `tts` | $0.50 → $1.00 | $9.00 → $18.00 | GA 2026-09-23. Expressive, 130 languages, 2 speakers. Use `generate_speech()`, not `invoke_gemini()`. |
    | gemini-3.8-flash-lite-tts | `tts-lite` | $0.50 → $1.00 | $6.00 → $12.00 | GA 2026-09-23. Bulk / read-aloud, 101 languages. Replaces gemini-3.1-flash-tts-preview ($1 / $20). |
    
    Speech aliases live in `SPEECH_ALIASES`, not `MODEL_ALIASES`, so a text call can
    never resolve to an audio model.
    
    See [references/models.md](references/models.md) for full details.
    
    ### Thinking Budget (Gemini 3.x)
    
    Gemini 3.x models reason before responding. The parameter changed in
    2026: integer `thinking_budget` is gone; use string `thinking_level`
    ∈ {`minimal`, `low`, `medium`, `high`}. Default for 3.5–3.8 Flash is
    `medium`. For transcription / classification / extraction tasks, pass
    `thinking_level='minimal'` or the model will silently spend output
    tokens on reasoning (symptom: empty response with
    `finishReason=MAX_TOKENS`).
    
    **3.7 and 3.8 Flash reject `minimal`** with HTTP 400 (`Thinking level MINIMAL
    is not supported for this model`); `low` is their floor. The client downgrades
    `minimal` to `low` on those two models and prints a note to stderr, so existing
    callers keep working. Measured on 3.8 (2026-09-03): `low` spent 0 thinking
    tokens on a one-word reply, the default `medium` spent 79. On 3.7, `low` still
    spent 45–88, and a `max_output_tokens=50` call at `low` hit MAX_TOKENS and
    returned None, so budget output generously there. If a job needs a true
    no-thinking pass, pin `flash-3.6` or `lite`, which still accept `minimal`.
    
    ```python
    response = invoke_gemini(
        prompt="Transcribe this image.",
        model="flash",
        image_path="/tmp/screenshot.png",
        max_output_tokens=4000,
        thinking_level="minimal",  # don't burn output budget on reasoning
    )
    ```
    
    ## Error Handling
    
    ```python
    response = invoke_gemini(prompt="...", model="flash")
    if response is None:
        print("API call failed — check credentials")
    
    result = generate_image("...")
    if result is None:
        print("Image generation failed — check credentials or try again")
    ```
    
    Common issues: Missing API key → see Setup. Rate limit → auto-retries with backoff. Network error → returns None.
    
    ## Advanced Features
    
    ### Custom Generation Config
    
    ```python
    response = invoke_gemini(
        prompt="Write a haiku",
        model="flash",                  # gemini-3.8-flash
        temperature=0.9,
        max_output_tokens=200,
        top_p=0.95,
        thinking_level="low",           # haiku is short; modest reasoning is fine
    )
    ```
    
    ### Multi-modal Input
    
    ```python
    from pydantic import BaseModel
    from gemini_client import invoke_with_structured_output
    
    class ImageDescription(BaseModel):
        objects: list[str]
        scene: str
        colors: list[str]
    
    result = invoke_with_structured_output(
        prompt="Describe this image",
        pydantic_model=ImageDescription,
        image_path="/mnt/user-data/uploads/photo.jpg"
    )
    ```
    
    See [references/advanced.md](references/advanced.md) for more patterns.
    
    ## Troubleshooting
    
    **"No credentials configured":** Create `/mnt/project/proxy.env` with CF credentials, or add `GOOGLE_API_KEY.txt`.
    
    **CF Gateway 401/403:** Verify `CF_API_TOKEN` has AI Gateway permissions. If not using BYOK, add `GOOGLE_API_KEY` to `proxy.env`.
    
    **Import errors:** `uv pip install requests pydantic`
    
    **Image generation returns None:** Check credentials. If persistent, try `model="nano-banana-2"` (more reliable than image-pro). Check for content policy blocks in error output.
    

Comments (0)

Sign in to join the conversation.

No comments yet.

Reviews (0)

No reviews yet.

Related