invoking-gemini
Invokes Google Gemini models for structured outputs, image generation, text-to-speech narration, multi-modal tasks, and Google-specific features. Use when users request Gemini, image generation, Gemini TTS or a synthesized voice, structured JSON output, Google API integration, or
Install
npx skills add https://github.com/oaustegard/claude-skills/tree/main/plugins/ai-and-reasoning/skills/invoking-gemini
claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install oaustegard-claude-skills@llmmart
git clone https://github.com/oaustegard/claude-skills.git
The skills CLI installs just this skill, for any of its supported agents. Claude Code installs the whole oaustegard/claude-skills collection as a plugin from our marketplace. Git is the plain clone.
README
invoking-gemini
Invokes Google Gemini models for structured outputs, multi-modal tasks, and Google-specific features. Use when users request Gemini, structured JSON output, Google API integration, or cost-effective parallel processing.
Skill manifest
Invoking Gemini
Delegate tasks to Google's Gemini models when they offer advantages over Claude.
When to Use Gemini
Image generation:
- Blog header images, illustrations, diagrams
- Style-guided image creation (risograph, editorial, etc.)
- Text rendering in images
Speech (TTS):
- Narration, voice-over, read-aloud with style direction per line
- A custom voice designed from a written description
- Two-speaker dialogue
Structured outputs:
- JSON Schema validation with property ordering guarantees
- Pydantic model compliance
- Strict schema adherence (enum values, required fields)
Cost optimization:
- Parallel batch processing (Gemini 3 Flash is lightweight)
- High-volume simple tasks
Multi-modal tasks:
- Image analysis with JSON output
- Video processing
- Audio transcription with structure
Setup
uv pip install requests pydantic
Credentials — Option A (recommended): Cloudflare AI Gateway
Source /mnt/project/proxy.env with CF_ACCOUNT_ID, CF_GATEWAY_ID, CF_API_TOKEN.
Requests route through Cloudflare AI Gateway, bypassing IP blocks. Google API key stored in gateway via BYOK.
Credentials — Option B: Direct Google API
If no proxy.env, falls back to direct: GOOGLE_API_KEY.txt or API_CREDENTIALS.json.
Image Generation
Generate images using Gemini's native image models. This is the primary way to create illustrations, blog headers, diagrams, and visual content.
Quick Start
import sys
sys.path.append('/mnt/skills/user/invoking-gemini/scripts')
from gemini_client import generate_image
# One call — returns {"path": "...", "caption": "..."} or None
result = generate_image("A watercolor painting of a mountain lake at sunset")
print(result["path"]) # /mnt/user-data/outputs/gemini_image_1740000000.png
Function Signature
generate_image(
prompt: str, # The image description
output_path: str = None, # Auto-generates if omitted
model: str = "nano-banana-2", # Default: fast. Use "image-pro" for quality
temperature: float = 0.7, # 0.5-0.7 for diagrams, 0.7-0.8 for illustrations
) -> dict | None
# Returns: {"path": "/mnt/user-data/outputs/gemini_image_*.png", "caption": str|None}
# Returns None on failure
Model Selection
| Alias | Model | Best For | Cost/image |
|---|---|---|---|
"nano-banana-2" or "image" |
gemini-3.1-flash-image-preview | Fast iteration, drafts | $0.067 |
"image-pro" or "nano-banana-pro" |
gemini-3-pro-image-preview | Published content, text rendering | $0.134 |
Complete Blog Header Example
import sys
sys.path.append('/mnt/skills/user/invoking-gemini/scripts')
from gemini_client import generate_image
# 1. Compose prompt with style prefix + subject
style_prefix = (
"Style: Risograph-inspired editorial illustration. "
"Visible halftone dot texture and slight color misregistration between layers. "
"Limited ink palette: deep indigo, warm coral, and sage green on off-white paper. "
"Layered transparency where colors overlap creates rich secondary tones. "
"Modern and professional — the aesthetic of an indie design studio, not a fantasy novel. "
"Generous whitespace. No photorealism, no glow effects, no cyberpunk. No text or labels."
)
subject = "A raven perched on a stack of books, observing a network graph"
prompt = f"{style_prefix}\n\nSubject: {subject}. Wide landscape format, suitable as a blog header."
# 2. Generate (use image-pro for published content)
result = generate_image(prompt, model="image-pro", temperature=0.75)
if result:
print(f"Saved: {result['path']}")
# 3. Present to user
# present_files([result["path"]])
Prompt Patterns
- Style prefix + subject: Prepend a style description, then describe the subject
- Be specific about style: "Risograph-inspired editorial illustration" not "a nice picture"
- Include composition: "Wide landscape format" / "centered, high contrast"
- Text rendering: "A poster with the text 'SALE' in bold red letters" (works well with image-pro)
- Negative constraints: "No photorealism, no glow effects" to avoid defaults
Custom Output Path
result = generate_image(
"A logo for a coffee shop called 'Bean There'",
output_path="/mnt/user-data/outputs/coffee_logo.png"
)
Speech Generation (TTS)
Gemini 3.8 Flash TTS and Flash-Lite TTS went GA on 2026-09-23. Output is WAV, 24 kHz mono 16-bit, SynthID-watermarked.
from gemini_client import generate_speech, design_voice, list_voices
r = generate_speech("Odin kept two ravens. <short pause> Huginn was thought.",
output_path="/tmp/line.wav", voice="Algenib",
style="quiet and dry, unhurried")
# {'path': '/tmp/line.wav', 'seconds': 4.2, 'audio_tokens': 135} or None
v = design_voice("A low, dry, quietly amused male voice with a faint rasp. "
"Soft southern British accent.", "narrator", gender="male",
language_code="en-GB") # {'id': 'voice_...', 'sample_path': ...}
generate_speech("...", voice=v["id"])
lows = list_voices(gender="male", pitch="low") # library of 2,089 prebuilt voices
- Voices: 30 studio voices (
Charon,Kore,Algenib"gravelly",Enceladus"breathy", ...) plus 2,059 persona voices with ids likeen-gb-storyteller-4.list_voices()returns accent, pitch, gender and a description for each; theaccentfilter needs the exact string ("Winchester English"), so filter accents on the returned field. - Style: pass
style=(aspeech_metadataannotation). Do not prefix the text with "Style: ..." — the 3.8 models read the prefix aloud, andsystemInstructionis rejected. Inline events go in the text:<laugh>,<sigh>,<breath>,<short pause>; CAPITALS stress a word. - Designed voices are stored (1-year expiry, 200 per project). The description sets baseline delivery too: "thoughtful pauses" in it produced 1–1.8 s mid-line pauses that no per-line style removed.
- The model can change words. Seen in a 29-line narration: "Hmm, I get
things wrong", "tell them" for "tell him". Anything with subtitles or a
fixed script needs an ASR check (faster-whisper
medium.en) and a retake. - Cost: about 32 audio tokens per second of speech, $9/M through 2026-12-31 on 3.8 Flash TTS ($6/M Lite), so a minute is about $0.02.
- Voice replication (cloning from a 30 s sample plus a recorded consent clip) is not wired in, and is unavailable in the EEA, UK, Switzerland, India, Illinois and Texas.
Basic Text Usage
import sys
sys.path.append('/mnt/skills/user/invoking-gemini/scripts')
from gemini_client import invoke_gemini
response = invoke_gemini(
prompt="Explain quantum computing in 3 bullet points",
model="flash", # gemini-3.8-flash (default)
)
print(response)
Structured Output
Use Pydantic models for guaranteed JSON Schema compliance:
from gemini_client import invoke_with_structured_output
from pydantic import BaseModel, Field
class BookAnalysis(BaseModel):
title: str
genre: str = Field(description="Primary genre")
key_themes: list[str] = Field(max_length=5)
rating: int = Field(ge=1, le=5)
result = invoke_with_structured_output(
prompt="Analyze the book '1984' by George Orwell",
pydantic_model=BookAnalysis
)
print(result.title) # "1984"
Nested models are supported. Gemini's responseSchema rejects $ref/$defs,
which pydantic emits for every nested model, so the client inlines them before
sending:
class Finding(BaseModel):
claim: str
confidence: Literal["high", "medium", "low"]
note: str | None = None
class Analysis(BaseModel):
findings: list[Finding] # nested — inlined for you
gaps: list[str]
Budget output generously. Thinking tokens count against max_output_tokens
(default 32768). Too low and the JSON truncates mid-object, which surfaces as a
pydantic parse error rather than a length error — the client now detects
finishReason=MAX_TOKENS and says so explicitly.
Parallel Invocation
from gemini_client import invoke_parallel
results = invoke_parallel(
prompts=["Summarize Hamlet", "Summarize Macbeth", "Summarize Othello"],
model="lite", # gemini-3.5-flash-lite — cheap/fast tier for batch
)
Available Models
The current frontier Flash is gemini-3.8-flash (GA 2026-09-02), the
default and the flash alias. Google shipped three Flash generations in six
weeks: 3.6 (2026-07-21), 3.7 (2026-08-13), 3.8 (2026-09-02). Each stays
callable under a pinned alias (flash-3.7, flash-3.6, flash-3.5,
flash-3), and none has a shutdown date. gemini-3.1-flash-lite-preview from
earlier docs is gone (shut down 2026-05-25).
The Pro tier is off routing. gemini-3.1-pro-preview costs 2.7× the input and
3.2× the output of 3.8 Flash at today's rates and loses to the 3.5+ Flash line
on the coding and agentic benchmarks that matter here. Do not target it; the
pro alias now resolves to gemini-3.8-flash, and "maximum reasoning" means
thinking_level='high' on Flash.
Text / Reasoning Models
| Model | Alias | Input/1M | Output/1M | Context | Notes |
|---|---|---|---|---|---|
| gemini-3.8-flash | flash |
$0.75 → $1.50 | $3.75 → $7.50 | 1M in / 64K out | Default. GA 2026-09-02. Current frontier Flash. Vs 3.7: Terminal-Bench 2.1 90.8% vs 81.6%, SWE-Bench Pro 61.6% vs 60.4%, SWE-Atlas 51.9% vs 48.0%, HLE flat (45.4% vs 45.7%). Google says it "works harder" at higher effort, so expect more thinking tokens per task. thinking_level is low/medium/high only — minimal returns HTTP 400 and the client downgrades it to low. Default medium spent 79 thinking tokens on a one-word reply (measured 2026-09-03); pass low for non-reasoning tasks. |
| gemini-3.7-flash | flash-3.7 |
$0.75 → $1.50 | $3.75 → $7.50 | 1M / 64K | GA 2026-08-13. DeepSWE v1.1 65.3% vs 49.0% on 3.6, Terminal-Bench 2.1 85.8%. Same minimal restriction as 3.8. Google keeps it "fully supported for efficiency-first workloads". |
| gemini-3.6-flash | flash-3.6 |
$0.75 → $1.50 | $3.75 → $7.50 | 1M / 64K | GA 2026-07-21. ~17% fewer output tokens than 3.5 Flash. Last Flash that accepts thinking_level='minimal' (verified 2026-09-03). |
| gemini-3.5-flash | flash-3.5 |
$1.50 | $9.00 | 1M | GA 2026-05-19. Google's model list now labels it "legacy". Accepts minimal. Costs more on output than 3.6–3.8. |
| gemini-3-flash-preview | flash-3 |
$0.30 | $2.50 | 1M | Older preview Flash, kept for back compat. Google's listed migration target for it is gemini-3.6-flash; no shutdown date. |
| — | $2.00 (≤200K) / $4.00 | $12.00 / $18.00 | 1M | DEPRECATED from routing (2026-09-03). Price/quality dominated by 3.6+ Flash; 3.5 Flash already beat it on most coding/agentic benchmarks. ID stays callable for pinned code. pro now resolves to gemini-3.8-flash. 3.5 Pro was announced at I/O 2026-05-19 for June and is still absent from the API as of 2026-09-03; it gets the same price/quality test before any alias points at it. |
|
| gemini-3.5-flash-lite | lite |
$0.30 | $2.50 | 1M | Cheap/bulk tier. GA 2026-07-21. Fastest 3.5-class (350 output tok/sec); beats gemini-3-flash on SWE-Bench Pro and OSWorld-Verified. |
stable-flash |
$0.30 | $2.50 | 1M | DEPRECATED — 2025-era generation, do not route here. | |
| — | $0.10 | $0.40 | 1M | DEPRECATED — cheaper, but a 2025-era generation. lite now resolves to gemini-3.5-flash-lite. |
|
stable-pro |
$1.25 (≤200K) / $2.50 | $10.00 / $20.00 | 1M | DEPRECATED — 2025-era generation, do not route here. |
$0.75 → $1.50 means introductory pricing: Google's pricing page (fetched
2026-09-03) lists 3.6, 3.7 and 3.8 Flash at $0.75 in / $3.75 out through
2026-12-31 and $1.50 / $7.50 from 2027-01-01. Context caching is $0.075 → $0.15;
Batch is half of standard. Output prices include thinking tokens.
Image Models
| Model | Alias | Input/1M | Per Image |
|---|---|---|---|
| gemini-3.1-flash-image-preview | image, nano-banana-2 |
$0.25 | $0.067 |
| gemini-3-pro-image-preview | image-pro, nano-banana-pro |
$2.00 | $0.134 |
Speech Models
| Model | Alias | Input/1M | Output/1M (audio) | Notes |
|---|---|---|---|---|
| gemini-3.8-flash-tts | tts |
$0.50 → $1.00 | $9.00 → $18.00 | GA 2026-09-23. Expressive, 130 languages, 2 speakers. Use generate_speech(), not invoke_gemini(). |
| gemini-3.8-flash-lite-tts | tts-lite |
$0.50 → $1.00 | $6.00 → $12.00 | GA 2026-09-23. Bulk / read-aloud, 101 languages. Replaces gemini-3.1-flash-tts-preview ($1 / $20). |
Speech aliases live in SPEECH_ALIASES, not MODEL_ALIASES, so a text call can
never resolve to an audio model.
See references/models.md for full details.
Thinking Budget (Gemini 3.x)
Gemini 3.x models reason before responding. The parameter changed in
2026: integer thinking_budget is gone; use string thinking_level
∈ {minimal, low, medium, high}. Default for 3.5–3.8 Flash is
medium. For transcription / classification / extraction tasks, pass
thinking_level='minimal' or the model will silently spend output
tokens on reasoning (symptom: empty response with
finishReason=MAX_TOKENS).
3.7 and 3.8 Flash reject minimal with HTTP 400 (Thinking level MINIMAL is not supported for this model); low is their floor. The client downgrades
minimal to low on those two models and prints a note to stderr, so existing
callers keep working. Measured on 3.8 (2026-09-03): low spent 0 thinking
tokens on a one-word reply, the default medium spent 79. On 3.7, low still
spent 45–88, and a max_output_tokens=50 call at low hit MAX_TOKENS and
returned None, so budget output generously there. If a job needs a true
no-thinking pass, pin flash-3.6 or lite, which still accept minimal.
response = invoke_gemini(
prompt="Transcribe this image.",
model="flash",
image_path="/tmp/screenshot.png",
max_output_tokens=4000,
thinking_level="minimal", # don't burn output budget on reasoning
)
Error Handling
response = invoke_gemini(prompt="...", model="flash")
if response is None:
print("API call failed — check credentials")
result = generate_image("...")
if result is None:
print("Image generation failed — check credentials or try again")
Common issues: Missing API key → see Setup. Rate limit → auto-retries with backoff. Network error → returns None.
Advanced Features
Custom Generation Config
response = invoke_gemini(
prompt="Write a haiku",
model="flash", # gemini-3.8-flash
temperature=0.9,
max_output_tokens=200,
top_p=0.95,
thinking_level="low", # haiku is short; modest reasoning is fine
)
Multi-modal Input
from pydantic import BaseModel
from gemini_client import invoke_with_structured_output
class ImageDescription(BaseModel):
objects: list[str]
scene: str
colors: list[str]
result = invoke_with_structured_output(
prompt="Describe this image",
pydantic_model=ImageDescription,
image_path="/mnt/user-data/uploads/photo.jpg"
)
See references/advanced.md for more patterns.
Troubleshooting
"No credentials configured": Create /mnt/project/proxy.env with CF credentials, or add GOOGLE_API_KEY.txt.
CF Gateway 401/403: Verify CF_API_TOKEN has AI Gateway permissions. If not using BYOK, add GOOGLE_API_KEY to proxy.env.
Import errors: uv pip install requests pydantic
Image generation returns None: Check credentials. If persistent, try model="nano-banana-2" (more reliable than image-pro). Check for content policy blocks in error output.
Files (claude-skills)
-
references
-
advanced.md 10 KB
# Advanced Patterns Advanced usage patterns for Gemini integration. ## Multi-Modal Processing ### Image Analysis with Structure ```python from pydantic import BaseModel, Field from gemini_client import invoke_with_structured_output class ProductAnalysis(BaseModel): product_name: str category: str colors: list[str] estimated_price_range: str condition: str = Field(description="new, used, or refurbished") result = invoke_with_structured_output( prompt="Analyze this product image. Identify the product, category, colors, and estimate price range.", pydantic_model=ProductAnalysis, image_path="/mnt/user-data/uploads/product.jpg" ) print(f"Product: {result.product_name}") print(f"Category: {result.category}") print(f"Colors: {', '.join(result.colors)}") ``` ### Batch Image Processing ```python from pathlib import Path from gemini_client import invoke_parallel image_dir = Path("/mnt/user-data/uploads/photos") image_files = list(image_dir.glob("*.jpg")) prompts = [ f"Describe this image in one sentence" for _ in image_files ] # Note: parallel invocation doesn't support images directly # Process sequentially with progress for i, image_path in enumerate(image_files): result = invoke_gemini( prompt="Describe this image briefly", image_path=str(image_path) ) print(f"[{i+1}/{len(image_files)}] {image_path.name}: {result}") ``` ## Hybrid Workflows ### Claude Plans, Gemini Executes Use Claude's reasoning for planning, Gemini for structured execution: ```python # 1. Claude (you) analyzes requirements and creates extraction plan # 2. Gemini executes structured extractions from pydantic import BaseModel from gemini_client import invoke_with_structured_output class ContactInfo(BaseModel): name: str email: str phone: str company: str # Gemini extracts structured data documents = [...] # List of document paths contacts = [] for doc_path in documents: with open(doc_path) as f: doc_text = f.read() result = invoke_with_structured_output( prompt=f"Extract contact information from:\n\n{doc_text}", pydantic_model=ContactInfo ) contacts.append(result) # 3. Claude synthesizes and analyzes results # ... your analysis code here ... ``` ### Parallel Processing with Different Models ```python from concurrent.futures import ThreadPoolExecutor from gemini_client import invoke_gemini def process_with_model(text: str, model: str) -> str: return invoke_gemini(text, model=model) # Compare outputs from different models text = "Analyze the sentiment of this review: ..." with ThreadPoolExecutor(max_workers=3) as executor: futures = { executor.submit(process_with_model, text, "gemini-3-flash-preview"): "3-flash", executor.submit(process_with_model, text, "gemini-2.5-flash"): "2.5-flash", executor.submit(process_with_model, text, "gemini-2.5-pro"): "2.5-pro", } for future in as_completed(futures): model_name = futures[future] result = future.result() print(f"{model_name}: {result}") ``` ## Complex Schema Patterns ### Nested Structures ```python from pydantic import BaseModel from typing import Optional class Address(BaseModel): street: str city: str state: str zip_code: str country: str = "USA" class Person(BaseModel): name: str age: Optional[int] = None email: str address: Address result = invoke_with_structured_output( prompt="Extract person info: John Doe, 30, john@example.com, 123 Main St, Springfield, IL 62701", pydantic_model=Person ) ``` ### Enums and Constraints ```python from pydantic import BaseModel, Field, validator from enum import Enum class Priority(str, Enum): LOW = "low" MEDIUM = "medium" HIGH = "high" URGENT = "urgent" class Task(BaseModel): title: str = Field(max_length=100) description: str priority: Priority estimated_hours: int = Field(ge=1, le=100) tags: list[str] = Field(max_length=10) @validator('tags') def validate_tags(cls, v): if len(v) > 10: raise ValueError('Maximum 10 tags allowed') return v result = invoke_with_structured_output( prompt="Create a task: Fix login bug - Users can't login. This is urgent. Should take 4 hours. Tags: bug, auth, security", pydantic_model=Task ) ``` ### Lists and Arrays ```python from pydantic import BaseModel class Ingredient(BaseModel): name: str quantity: str unit: str class Recipe(BaseModel): title: str servings: int prep_time_minutes: int cook_time_minutes: int ingredients: list[Ingredient] instructions: list[str] recipe_text = """ Make pasta carbonara. Serves 4. Prep: 10 min. Cook: 20 min. Ingredients: 400g spaghetti, 200g pancetta, 4 eggs, 100g parmesan, black pepper. Instructions: 1. Boil pasta 2. Fry pancetta 3. Mix eggs and cheese 4. Combine all """ result = invoke_with_structured_output( prompt=f"Extract recipe from:\n{recipe_text}", pydantic_model=Recipe ) ``` ## Error Recovery ### Retry with Schema Relaxation ```python from pydantic import BaseModel class StrictData(BaseModel): field1: str field2: int field3: list[str] # Try strict schema first result = invoke_with_structured_output(prompt, StrictData) if not result: # Retry with relaxed schema class RelaxedData(BaseModel): field1: str field2: Optional[int] = None field3: Optional[list[str]] = [] result = invoke_with_structured_output(prompt, RelaxedData) ``` ### Validation and Correction ```python from pydantic import BaseModel, ValidationError result = invoke_with_structured_output(prompt, MyModel) if result: try: # Additional validation assert len(result.items) > 0, "Must have at least one item" assert result.total > 0, "Total must be positive" except AssertionError as e: print(f"Validation failed: {e}") # Retry with corrected prompt corrected_prompt = f"{prompt}\n\nIMPORTANT: {e}" result = invoke_with_structured_output(corrected_prompt, MyModel) ``` ## Performance Optimization ### Prompt Caching For repeated prompts with same prefix: ```python # Share common context across requests base_context = """ You are analyzing customer reviews for sentiment. Categories: positive, neutral, negative Format: JSON with 'sentiment' and 'confidence' fields """ reviews = ["Review 1...", "Review 2...", "Review 3..."] for review in reviews: full_prompt = f"{base_context}\n\nReview: {review}" result = invoke_gemini(full_prompt) ``` ### Batch Size Tuning ```python from gemini_client import invoke_parallel # Process in optimal batches all_prompts = [...] # 1000 prompts batch_size = 50 # Tune based on rate limits results = [] for i in range(0, len(all_prompts), batch_size): batch = all_prompts[i:i+batch_size] batch_results = invoke_parallel(batch, max_workers=10) results.extend(batch_results) # Rate limit breathing room if i + batch_size < len(all_prompts): time.sleep(2) ``` ### Temperature Tuning ```python # Factual extraction: low temperature factual = invoke_gemini( "Extract the date from: Meeting scheduled for March 15th", temperature=0.1 ) # Creative generation: high temperature creative = invoke_gemini( "Write a creative tagline for a coffee shop", temperature=0.9 ) # Balanced: medium temperature balanced = invoke_gemini( "Summarize this article", temperature=0.7 ) ``` ## Cost Optimization ### Token Counting ```python import google.generativeai as genai model = genai.GenerativeModel("gemini-3-flash-preview") # Count tokens before sending token_count = model.count_tokens("Your prompt here") print(f"Input tokens: {token_count.total_tokens}") # Estimate cost input_cost = (token_count.total_tokens / 1_000_000) * 0.50 print(f"Estimated input cost: ${input_cost:.4f}") ``` ### Prompt Compression ```python # Verbose (wasteful) verbose_prompt = """ Please analyze the following text and extract the following information: - The name of the person - Their email address - Their phone number - Their company name Text: John Doe works at Acme Corp. Email: john@acme.com, Phone: 555-0100 """ # Concise (efficient) concise_prompt = """ Extract: name, email, phone, company John Doe works at Acme Corp. Email: john@acme.com, Phone: 555-0100 """ # With structured output, schema provides the context result = invoke_with_structured_output( prompt="John Doe works at Acme Corp. Email: john@acme.com, Phone: 555-0100", pydantic_model=ContactInfo # Schema explains structure ) ``` ## Integration with Other Tools ### Save Results to Database ```python import sqlite3 from gemini_client import invoke_with_structured_output conn = sqlite3.connect('/home/claude/results.db') cursor = conn.cursor() cursor.execute(''' CREATE TABLE IF NOT EXISTS contacts ( name TEXT, email TEXT, phone TEXT, company TEXT ) ''') documents = [...] for doc in documents: result = invoke_with_structured_output(doc, ContactInfo) if result: cursor.execute( 'INSERT INTO contacts VALUES (?, ?, ?, ?)', (result.name, result.email, result.phone, result.company) ) conn.commit() ``` ### Export to CSV ```python import csv from gemini_client import invoke_with_structured_output results = [] for item in items: result = invoke_with_structured_output(item, DataModel) if result: results.append(result.dict()) # Write to CSV with open('/mnt/user-data/outputs/results.csv', 'w', newline='') as f: if results: writer = csv.DictWriter(f, fieldnames=results[0].keys()) writer.writeheader() writer.writerows(results) ``` ### Combine with Pandas ```python import pandas as pd from gemini_client import invoke_with_structured_output # Process data with Gemini, analyze with pandas data = [] for text in texts: result = invoke_with_structured_output(text, StructuredData) if result: data.append(result.dict()) df = pd.DataFrame(data) print(df.describe()) print(df.groupby('category').size()) ``` -
examples.md 21.1 KB
# Examples Comprehensive examples of Gemini usage patterns. ## Example 1: Document Data Extraction Extract structured data from unstructured documents. ```python from pydantic import BaseModel, Field from gemini_client import invoke_with_structured_output from pathlib import Path class InvoiceData(BaseModel): invoice_number: str date: str vendor: str total_amount: float line_items: list[dict] = Field(description="List of items with description and price") invoice_dir = Path("/mnt/user-data/uploads/invoices") results = [] for invoice_file in invoice_dir.glob("*.txt"): with open(invoice_file) as f: invoice_text = f.read() data = invoke_with_structured_output( prompt=f"Extract invoice data:\n\n{invoice_text}", pydantic_model=InvoiceData ) if data: results.append({ 'file': invoice_file.name, 'invoice_number': data.invoice_number, 'vendor': data.vendor, 'total': data.total_amount }) # Save results import pandas as pd df = pd.DataFrame(results) df.to_csv('/mnt/user-data/outputs/invoice_summary.csv', index=False) ``` ## Example 2: Batch Classification Classify large datasets efficiently. ```python from pydantic import BaseModel from enum import Enum from gemini_client import invoke_parallel, invoke_with_structured_output class Sentiment(str, Enum): POSITIVE = "positive" NEUTRAL = "neutral" NEGATIVE = "negative" class ReviewAnalysis(BaseModel): sentiment: Sentiment confidence: float = Field(ge=0.0, le=1.0) key_topics: list[str] = Field(max_length=5) # Load reviews import pandas as pd df = pd.read_csv('/mnt/user-data/uploads/reviews.csv') results = [] for idx, row in df.iterrows(): analysis = invoke_with_structured_output( prompt=f"Analyze this review: {row['review_text']}", pydantic_model=ReviewAnalysis, temperature=0.3 # Low temp for consistent classification ) if analysis: results.append({ 'review_id': row['id'], 'sentiment': analysis.sentiment.value, 'confidence': analysis.confidence, 'topics': ', '.join(analysis.key_topics) }) # Progress if (idx + 1) % 10 == 0: print(f"Processed {idx + 1}/{len(df)} reviews") # Save results results_df = pd.DataFrame(results) results_df.to_csv('/mnt/user-data/outputs/sentiment_analysis.csv', index=False) # Summary statistics print("\nSentiment Distribution:") print(results_df['sentiment'].value_counts()) print(f"\nAverage Confidence: {results_df['confidence'].mean():.2f}") ``` ## Example 3: Multi-Modal Product Catalog Create structured product catalog from images. ```python from pydantic import BaseModel, Field from gemini_client import invoke_with_structured_output from pathlib import Path class Product(BaseModel): name: str category: str description: str = Field(max_length=200) primary_color: str additional_colors: list[str] = [] estimated_price_tier: str = Field(description="budget, mid-range, or premium") key_features: list[str] = Field(max_length=5) product_images = Path("/mnt/user-data/uploads/products") catalog = [] for img_path in product_images.glob("*.jpg"): product = invoke_with_structured_output( prompt=""" Analyze this product image and provide: - Product name and category - Brief description - Colors visible - Estimated price tier (budget/mid-range/premium) - Key features """, pydantic_model=Product, image_path=str(img_path) ) if product: catalog.append({ 'image': img_path.name, **product.dict() }) print(f"✓ {img_path.name}: {product.name}") # Export catalog import json with open('/mnt/user-data/outputs/product_catalog.json', 'w') as f: json.dump(catalog, f, indent=2) print(f"\nProcessed {len(catalog)} products") ``` ## Example 4: Resume Parser Extract structured data from resumes. ```python from pydantic import BaseModel, Field from typing import Optional class Education(BaseModel): degree: str institution: str year: Optional[str] = None class Experience(BaseModel): title: str company: str duration: str responsibilities: list[str] class Resume(BaseModel): name: str email: str phone: Optional[str] = None summary: str = Field(max_length=300) skills: list[str] education: list[Education] experience: list[Experience] resume_files = Path("/mnt/user-data/uploads/resumes") parsed_resumes = [] for resume_file in resume_files.glob("*.txt"): with open(resume_file) as f: resume_text = f.read() parsed = invoke_with_structured_output( prompt=f"Parse this resume:\n\n{resume_text}", pydantic_model=Resume, temperature=0.2 # Low temp for accuracy ) if parsed: parsed_resumes.append({ 'file': resume_file.name, 'candidate': parsed.name, 'email': parsed.email, 'skills_count': len(parsed.skills), 'years_experience': len(parsed.experience), 'education_level': parsed.education[0].degree if parsed.education else 'None' }) # Create summary import pandas as pd df = pd.DataFrame(parsed_resumes) df.to_csv('/mnt/user-data/outputs/resume_summary.csv', index=False) # Skill frequency analysis all_skills = [] for resume in parsed_resumes: all_skills.extend(resume.get('skills', [])) from collections import Counter skill_counts = Counter(all_skills) print("\nTop 10 Skills:") for skill, count in skill_counts.most_common(10): print(f" {skill}: {count}") ``` ## Example 5: Meeting Notes Summarization Batch process meeting notes into structured summaries. ```python from pydantic import BaseModel, Field from datetime import datetime class ActionItem(BaseModel): task: str assignee: str due_date: Optional[str] = None priority: str = Field(description="high, medium, or low") class MeetingSummary(BaseModel): meeting_date: str attendees: list[str] key_topics: list[str] = Field(max_length=5) decisions: list[str] action_items: list[ActionItem] next_meeting: Optional[str] = None notes_dir = Path("/mnt/user-data/uploads/meeting_notes") summaries = [] for notes_file in sorted(notes_dir.glob("*.txt")): with open(notes_file) as f: notes = f.read() summary = invoke_with_structured_output( prompt=f"Summarize these meeting notes:\n\n{notes}", pydantic_model=MeetingSummary ) if summary: summaries.append(summary) print(f"✓ {notes_file.name}: {len(summary.action_items)} action items") # Generate action items report all_action_items = [] for summary in summaries: for item in summary.action_items: all_action_items.append({ 'meeting_date': summary.meeting_date, 'task': item.task, 'assignee': item.assignee, 'due_date': item.due_date, 'priority': item.priority }) import pandas as pd df = pd.DataFrame(all_action_items) df.to_csv('/mnt/user-data/outputs/action_items.csv', index=False) # Group by assignee print("\nAction Items by Assignee:") print(df.groupby('assignee').size().sort_values(ascending=False)) ``` ## Example 6: Parallel Translation Translate content to multiple languages in parallel. ```python from gemini_client import invoke_parallel content = """ Welcome to our product! This innovative solution helps you manage your tasks efficiently and collaborate with your team. """ languages = [ "Spanish", "French", "German", "Italian", "Portuguese", "Japanese", "Korean", "Chinese", "Arabic", "Russian" ] prompts = [ f"Translate to {lang} (output only the translation):\n\n{content}" for lang in languages ] translations = invoke_parallel( prompts=prompts, model="gemini-3-flash-preview", temperature=0.3, max_workers=10 ) # Create translation table results = [] for lang, translation in zip(languages, translations): if translation: results.append({ 'language': lang, 'translation': translation.strip() }) # Export import pandas as pd df = pd.DataFrame(results) df.to_csv('/mnt/user-data/outputs/translations.csv', index=False) print(f"Translated to {len(results)} languages") ``` ## Example 7: Code Documentation Generator Generate structured documentation from code. ```python from pydantic import BaseModel, Field class FunctionDoc(BaseModel): function_name: str description: str = Field(max_length=200) parameters: list[dict] = Field(description="List with name, type, description") return_type: str return_description: str example_usage: str complexity: str = Field(description="O(n), O(log n), etc.") # Read source files code_dir = Path("/mnt/user-data/uploads/source_code") documentation = [] for code_file in code_dir.glob("*.py"): with open(code_file) as f: code = f.read() # Extract functions (simplified) import re functions = re.findall(r'def\s+(\w+)\s*\([^)]*\):[^}]*?(?=\ndef|\Z)', code, re.DOTALL) for func_code in functions[:5]: # Limit to first 5 functions doc = invoke_with_structured_output( prompt=f"Generate documentation for this Python function:\n\n{func_code}", pydantic_model=FunctionDoc ) if doc: documentation.append(doc.dict()) # Export as JSON import json with open('/mnt/user-data/outputs/api_documentation.json', 'w') as f: json.dump(documentation, f, indent=2) # Generate markdown with open('/mnt/user-data/outputs/API_DOCS.md', 'w') as f: f.write("# API Documentation\n\n") for doc in documentation: f.write(f"## {doc['function_name']}\n\n") f.write(f"{doc['description']}\n\n") f.write(f"**Returns:** `{doc['return_type']}` - {doc['return_description']}\n\n") f.write(f"**Complexity:** {doc['complexity']}\n\n") f.write(f"**Example:**\n```python\n{doc['example_usage']}\n```\n\n") ``` ## Example 8: Financial Report Analysis Extract key metrics from financial reports. ```python from pydantic import BaseModel, Field class FinancialMetrics(BaseModel): company_name: str reporting_period: str revenue: float = Field(description="In millions") net_income: float = Field(description="In millions") profit_margin: float = Field(ge=0, le=100, description="Percentage") key_highlights: list[str] = Field(max_length=5) risks: list[str] = Field(max_length=3) reports_dir = Path("/mnt/user-data/uploads/financial_reports") metrics = [] for report_file in reports_dir.glob("*.txt"): with open(report_file) as f: report_text = f.read() data = invoke_with_structured_output( prompt=f""" Extract financial metrics from this quarterly report. All monetary values should be in millions. {report_text} """, pydantic_model=FinancialMetrics, temperature=0.1 # Very low for numerical accuracy ) if data: metrics.append(data.dict()) # Analysis import pandas as pd df = pd.DataFrame(metrics) print("\nFinancial Summary:") print(f"Average Revenue: ${df['revenue'].mean():.2f}M") print(f"Average Net Income: ${df['net_income'].mean():.2f}M") print(f"Average Profit Margin: {df['profit_margin'].mean():.2f}%") df.to_csv('/mnt/user-data/outputs/financial_metrics.csv', index=False) ``` ## Example 9: Survey Response Analysis Analyze open-ended survey responses. ```python from pydantic import BaseModel from enum import Enum class Satisfaction(str, Enum): VERY_SATISFIED = "very_satisfied" SATISFIED = "satisfied" NEUTRAL = "neutral" DISSATISFIED = "dissatisfied" VERY_DISSATISFIED = "very_dissatisfied" class SurveyAnalysis(BaseModel): satisfaction: Satisfaction main_sentiment: str = Field(max_length=100) mentioned_features: list[str] = Field(description="Features mentioned positively or negatively") pain_points: list[str] = Field(description="Problems or complaints") suggestions: list[str] = Field(description="Improvement suggestions") # Load survey data import pandas as pd df = pd.read_csv('/mnt/user-data/uploads/survey_responses.csv') analyses = [] for idx, row in df.iterrows(): analysis = invoke_with_structured_output( prompt=f"Analyze this survey response:\n\nQuestion: {row['question']}\nAnswer: {row['response']}", pydantic_model=SurveyAnalysis ) if analysis: analyses.append({ 'response_id': row['id'], 'satisfaction': analysis.satisfaction.value, 'sentiment': analysis.main_sentiment, 'features': ', '.join(analysis.mentioned_features), 'pain_points': ', '.join(analysis.pain_points), 'suggestions': ', '.join(analysis.suggestions) }) results_df = pd.DataFrame(analyses) results_df.to_csv('/mnt/user-data/outputs/survey_analysis.csv', index=False) # Aggregate insights print("\nSatisfaction Distribution:") print(results_df['satisfaction'].value_counts()) all_pain_points = [p for points in analyses for p in points.get('pain_points', [])] from collections import Counter print("\nTop Pain Points:") for pain, count in Counter(all_pain_points).most_common(5): print(f" {pain}: {count}") ``` ## Example 10: Hybrid Claude + Gemini Workflow Claude does complex reasoning, Gemini does structured extraction. ```python from pydantic import BaseModel from gemini_client import invoke_with_structured_output, invoke_parallel # Step 1: Claude (you) analyzes the dataset and determines categories # Assume you've identified key categories for classification class DataPoint(BaseModel): text: str category: str confidence: float key_terms: list[str] # Step 2: Gemini extracts structured data at scale raw_data = pd.read_csv('/mnt/user-data/uploads/raw_data.csv') structured_data = [] batch_size = 50 for i in range(0, len(raw_data), batch_size): batch = raw_data.iloc[i:i+batch_size] for idx, row in batch.iterrows(): result = invoke_with_structured_output( prompt=f"Classify and extract from: {row['text']}", pydantic_model=DataPoint, temperature=0.3 ) if result: structured_data.append(result.dict()) print(f"Processed {min(i+batch_size, len(raw_data))}/{len(raw_data)}") # Step 3: Claude analyzes the structured results df = pd.DataFrame(structured_data) # Your analysis here: # - Identify patterns # - Generate insights # - Create visualizations # - Produce final report print(f"\nProcessed {len(structured_data)} items") print(f"Average confidence: {df['confidence'].mean():.2f}") print("\nCategory distribution:") print(df['category'].value_counts()) ``` ## Example 11: Blog Header Image Generation Generate a styled blog header image with a single call. ```python import sys sys.path.append('/mnt/skills/user/invoking-gemini/scripts') from gemini_client import generate_image # Style prefix — prepend to any subject for consistent visual identity RISO_ILLUSTRATION = ( "Style: Risograph-inspired editorial illustration. " "Visible halftone dot texture and slight color misregistration between layers. " "Limited ink palette: deep indigo, warm coral, and sage green on off-white paper. " "Layered transparency where colors overlap creates rich secondary tones. " "Modern and professional — the aesthetic of an indie design studio, not a fantasy novel. " "Generous whitespace. No photorealism, no glow effects, no cyberpunk. No text or labels." ) RISO_DIAGRAM = ( "Style: Risograph-inspired technical diagram. " "Visible halftone dot texture and slight color misregistration. " "Limited ink palette: deep indigo for primary shapes and text, " "warm coral for highlights and active elements, " "sage green for secondary elements and connections. " "Off-white paper background. Clean layout with generous spacing. " "Professional and readable." ) # Generate illustration header subject = "A raven perched on a network graph, watching data flow between nodes" result = generate_image( f"{RISO_ILLUSTRATION}\n\nSubject: {subject}. Wide landscape format, suitable as a blog header.", model="image-pro", # Use image-pro for published content temperature=0.75, output_path="/mnt/user-data/outputs/blog_header.png" ) if result: print(f"Header saved: {result['path']}") else: print("Generation failed — retry or check credentials") ``` Key patterns: - **Style prefix + subject composition**: The prefix sets visual rules, the subject describes content - **`image-pro` for published content**: Better quality and text rendering than default - **Temperature 0.7-0.8 for illustrations**: Allows creative variation while staying on-style - **Explicit output path**: Control where the file lands for downstream use ## Example 12: Technical Diagram Generation Generate a styled technical diagram with text labels. ```python from gemini_client import generate_image RISO_DIAGRAM = ( "Style: Risograph-inspired technical diagram. " "Visible halftone dot texture and slight color misregistration. " "Limited ink palette: deep indigo for primary shapes and text, " "warm coral for highlights and active elements, " "sage green for secondary elements and connections. " "Off-white paper background. Clean layout with generous spacing. " "Professional and readable." ) result = generate_image( f"{RISO_DIAGRAM}\n\n" "A flowchart showing: User Request → Claude Planning → Gemini Extraction → " "Claude Synthesis → Final Report. Highlight the Gemini step in coral. " "Wide landscape format.", model="image-pro", temperature=0.6, # Lower temp for diagrams — more precise ) ``` ## Example 13: Batch Image Generation with Variants Generate multiple variants of the same concept for selection. ```python from gemini_client import generate_image subjects = [ "A raven carrying a scroll through a library of glowing books", "A raven assembling puzzle pieces that form a constellation", "A raven observing its reflection in a pool of data streams", ] results = [] for i, subject in enumerate(subjects): result = generate_image( f"Style: Risograph-inspired editorial illustration with deep indigo, " f"warm coral, sage green on off-white. No text.\n\n" f"Subject: {subject}. Wide landscape format.", model="nano-banana-2", # Use fast model for drafts temperature=0.8, output_path=f"/mnt/user-data/outputs/variant_{i+1}.png" ) if result: results.append(result["path"]) print(f"Variant {i+1}: {result['path']}") # Present all variants for user to pick # present_files(results) ``` ## Example 14: Image Generation with Structured Feedback Loop Generate an image, analyze it with Gemini vision, regenerate if needed. ```python from gemini_client import generate_image, invoke_with_structured_output from pydantic import BaseModel, Field class ImageQuality(BaseModel): has_text_artifacts: bool = Field(description="Unwanted text in the image") style_match: int = Field(ge=1, le=5, description="How well it matches risograph style") composition_score: int = Field(ge=1, le=5) issues: list[str] = Field(description="Problems to fix in re-generation") # Generate result = generate_image( "Style: Risograph editorial. Deep indigo, coral, sage green.\n\n" "Subject: A raven on a circuit board. Wide landscape.", model="image-pro", temperature=0.75, ) if result: # Analyze with vision quality = invoke_with_structured_output( prompt="Evaluate this image for: unwanted text artifacts, " "risograph style fidelity (halftone dots, misregistration, limited palette), " "and composition quality.", pydantic_model=ImageQuality, image_path=result["path"] ) print(f"Style match: {quality.style_match}/5") print(f"Issues: {quality.issues}") # Re-generate with fixes if needed if quality.style_match < 3 or quality.has_text_artifacts: fix_prompt = f"Fix: {', '.join(quality.issues)}. " # Regenerate with adjusted prompt... ``` ## Best Practices from Examples **1. Temperature tuning:** - Factual extraction: 0.1-0.3 - Classification: 0.3-0.5 - Diagrams / technical images: 0.5-0.7 - Creative tasks / illustrations: 0.7-0.9 **2. Batch processing:** - Process in batches of 50-100 - Add delays between batches for rate limits - Show progress to user **3. Error handling:** - Always check if result is None - Log failures for debugging - Consider retry logic for critical tasks **4. Schema design:** - Use Field descriptions for clarity - Add constraints (ge, le, max_length) - Use Enums for fixed categories **5. Output formats:** - CSV for tabular data - JSON for hierarchical data - Markdown for documentation - Database for large datasets **6. Image generation:** - Compose prompts as: style prefix + subject + format ("Wide landscape") - Use `image-pro` / `nano-banana-pro` for published content, `nano-banana-2` for drafts - Temperature 0.5-0.7 for diagrams, 0.7-0.8 for illustrations - Add negative constraints ("No photorealism, no glow effects") to avoid model defaults - Always check result is not None before using the path -
models.md 20.7 KB
# Gemini Models Reference Detailed information about available Gemini models (as of September 2026; speech models added 2026-09-24). ## Model Comparison ### Gemini 3.8 — Frontier Flash (GA, current default) #### gemini-3.8-flash **Status:** Generally available (released September 2, 2026) **Alias:** `flash` (the current default Flash) **Strengths:** - Google's "most intelligent Flash model", positioned for long-horizon software engineering, autonomous agents, and multi-step enterprise work - Vs 3.7 Flash (Google's own numbers): Terminal-Bench 2.1 90.8% vs 81.6%, SWE-Bench Pro 61.6% vs 60.4%, SWE-Atlas 51.9% vs 48.0%, τ³-bench Banking 38.1% vs 30.9%, CharXiv 86.2% vs 84.5%; HLE-Verified 54.9% - Humanity's Last Exam is flat (45.4% vs 45.7%) — the gains are agentic and tool-use, not open-ended reasoning - Artificial Analysis Intelligence Index: 57 at `medium` (3.7 Flash 56, 3.6 Flash 52); ~310 output tok/sec - Prompt-injection robustness improved (Gray Swan); CBRN and cyber-offense safeguards carried forward **Specifications:** - Context window: 1,048,576 tokens input / 65,536 tokens output - Multimodal input: text, image, video, audio, PDF; text output only - `thinking_level`: `low`, `medium` (default), `high`. **`minimal` is not supported and returns HTTP 400** ("Thinking level MINIMAL is not supported for this model", verified 2026-09-03). The client downgrades `minimal` to `low` on this model. - Google's release note: the model "works harder" on complex tasks — extra reasoning steps, iterative tool calls — so higher effort levels cost more tokens. Measured 2026-09-03: `low` spent 0 thinking tokens on a one-word reply; the default `medium` spent 79. - Supports caching, code execution, computer use (preview), file search, function calling, Maps and Search grounding, structured outputs, URL context, Batch / Flex / Priority inference. No audio generation, image generation, or Live API. - Google's migration notes for the 3.x line: `thinking_budget` is gone (use `thinking_level`); `temperature` / `top_p` / `top_k` / `candidate_count` are deprecated sampling params on this model. **Best for:** - Default Flash / sub-agent-delegation choice for most tasks - Agentic coding loops, terminal automation, multi-file projects - Finance/legal agent workflows (Vals Finance Agent V2, Harvey's Legal Agent Benchmark lead their Flash class) **Pricing:** - Input: $0.75 / 1M tokens through 2026-12-31; $1.50 from 2027-01-01 - Output: $3.75 / 1M tokens through 2026-12-31; $7.50 from 2027-01-01 (includes thinking tokens) - Context caching $0.075 → $0.15; Batch 50% off - 1M context window at base price (no surcharge tier) Shipped alongside **`gemini-3.8-flash-cyber`** — vulnerability discovery and patching (>70% on Google's internal 20-language vuln benchmark, 47.2% pass@1 on CWE-Bench patching). Access is limited to Google's Fairwind Program (government authorities, critical-infrastructure operators, software maintainers), so this client cannot alias it. --- ### Gemini 3.7 — Prior Frontier Flash (GA) #### gemini-3.7-flash **Status:** Generally available (released August 13, 2026) **Alias:** `flash-3.7` (was `flash` for three weeks until 3.8 shipped) **Strengths:** - The coding jump in the 3.x Flash line: DeepSWE v1.1 65.3% vs 49.0% on 3.6 Flash, FrontierCode 1.1 43.6% vs 34.4%, AutomationBench 30.4% vs 17.0%, WebDev Arena Elo 1588 vs 1538 - Terminal-Bench 2.1 85.8%, OSWorld-2.0 47.9%, GDM-MRCR v2 (128k) 97.0% - Google keeps it "fully supported for efficiency-first workloads". Measured 2026-09-03 on a one-word prompt, though, `low` on 3.7 still spent 45–88 thinking tokens where `low` on 3.8 spent 0, and a `max_output_tokens=50` call at `low` hit MAX_TOKENS. Budget output generously on this model. **Specifications:** - Context window: 1,048,576 tokens input / 65,536 tokens output - Multimodal input: text, image, video, audio, PDF - `thinking_level`: `low`, `medium` (default), `high`; **`minimal` returns HTTP 400** (verified 2026-09-03), same as 3.8 **Pricing:** - Same schedule as 3.8: $0.75 / $3.75 through 2026-12-31, then $1.50 / $7.50 --- ### Gemini 3.6 — Older Flash (GA) #### gemini-3.6-flash **Status:** Generally available (released July 21, 2026) **Alias:** `flash-3.6` (was `flash` until 3.7 shipped) **Strengths:** - Builds on 3.5 Flash for coding, knowledge work, and multimodal tasks - The last Flash that accepts `thinking_level='minimal'` (verified 2026-09-03) — pin here for true no-thinking transcription/extraction - ~17% fewer output tokens than 3.5 Flash on the Artificial Analysis index (the headline efficiency win — addresses 3.5's verbosity) - Quality gains alongside efficiency: DeepSWE 49% vs 37%, MLE-Bench 63.9% vs 49.7%, OSWorld-Verified 83.0% vs 78.4%, GDPval-AA v2 1421 vs 1349 - Built-in client-side Computer Use tool via the Gemini API (Preview) - Dynamic thinking on by default (configurable via `thinking_level`) **Specifications:** - Context window: ~1M tokens input / 65,536 tokens output - Multimodal: text, image, audio, video - Default `thinking_level`: `medium` — set explicitly to `minimal` for transcription/classification/extraction or the model will silently spend output tokens on reasoning - Enhanced Frontier Safety safeguards (CBRN, cyber-offense); model card notes a slight tone regression vs 3.5 Flash **Best for:** - Pinning prior-gen behavior, and `minimal`-thinking bulk work - Cost-sensitive high-volume agentic work (cheaper output than 3.5) **Pricing:** - Same schedule as 3.7 and 3.8 on Google's pricing page (fetched 2026-09-03): $0.75 / $3.75 through 2026-12-31, then $1.50 / $7.50. The July table here said $1.50 / $7.50 flat; the intro rate now covers all three. - 1M context window at base price (no surcharge tier) Shipped alongside two sibling models, neither wired into this client's alias table: - `gemini-3.5-flash-lite` — GA. Fastest 3.5-class model (350 output tok/sec), $0.30 / $2.50. **This is now the `lite` alias target** (repointed 2026-07-21 from gemini-2.5-flash-lite). It costs ~6x more on output than the 2.5 model it replaces; that was accepted deliberately — the 2.5 generation is retired regardless of price. - `gemini-3.5-flash-cyber` — vuln-finding, fine-tuned on 3.5 Flash; powers CodeMender. **NOT generally available**: access is limited to governments and trusted partners under a pilot program due to dual-use risk. It cannot simply be added as an alias. --- ### Gemini 3.5 — Legacy Flash (GA) #### gemini-3.5-flash **Status:** Generally available (released May 19, 2026 at Google I/O); Google's model list now labels it "legacy". No shutdown date. **Alias:** `flash-3.5` (was `flash` until 3.6 shipped) **Strengths:** - Frontier-class performance — beats Gemini 3.1 Pro on most coding and agentic benchmarks - Runs ~4× faster on output tokens than other frontier models - Frontier intelligence at sub-Pro pricing - Dynamic thinking on by default (configurable via `thinking_level`) **Specifications:** - Context window: ~1M tokens input - Multimodal: text, image, audio, video - Knowledge cutoff: January 2026 - Default `thinking_level`: `medium` (was `high` on prior 3.x — set explicitly to `minimal` for transcription/classification/extraction or the model will silently spend output tokens on reasoning) **Best for:** - Pinning to prior-gen Flash behavior when 3.6–3.8 output differs - Agentic coding loops, terminal automation, multi-file projects - Multimodal document analysis where structure must be preserved **Pricing:** - Input: $1.50 / 1M tokens - Output: $9.00 / 1M tokens - 1M context window at base price (no surcharge tier) --- ### Gemini 3.x — Prior Preview Generation #### gemini-3-flash-preview **Status:** Preview (still callable — kept for back compat) **Alias:** `flash-3` The previous-generation Flash. Use when you need to pin behavior established before the 3.5 cutover. Google's deprecation page lists gemini-3.6-flash as its replacement, with no shutdown date. New code should target `flash` (gemini-3.8-flash) instead. **Pricing:** - Input: $0.30 / 1M tokens - Output: $2.50 / 1M tokens #### gemini-3.1-pro-preview **Status:** Preview. **DEPRECATED from routing 2026-09-03** (Oskar): "its Pareto efficiency is too poor compared to the later Flash models." At today's rates it costs 2.7× the input and 3.2× the output of 3.8 Flash (1.3× / 1.6× once the Flash intro pricing ends), and 3.5 Flash already beat it on most coding and agentic benchmarks. The ID stays callable for pinned code. **Alias:** none — `pro` now resolves to gemini-3.8-flash. For maximum reasoning use Flash with `thinking_level='high'`. **Strengths:** - Was the most capable Gemini Pro in the API - 1M context with tiered pricing above 200K **Specifications:** - Context window: ~1M tokens input - Long context surcharge: 2× above 200K input tokens - Multimodal: text, image, video, audio **Best for:** - Nothing in new code. Pinned callers only. **Pricing:** - Input: $2.00 / 1M tokens (≤200K), $4.00 (>200K) - Output: $12.00 / 1M tokens (≤200K), $18.00 (>200K) — the July table here said $24.00; Google's pricing page says $18.00 (fetched 2026-09-03) **Note:** Google announced `gemini-3.5-pro` at I/O on 2026-05-19 for June. As of 2026-09-03 it is not in the API model list or on the pricing page; DeepMind still lists it as "coming soon". When it ships it gets the same price/quality test against the current Flash before any alias points at it. --- ### Gemini 2.5 — DEPRECATED (retired 2026-07-21) ⚠️ The Gemini 2.5 text generation is **retired from routing**. A 2025-era generation; the cost saving does not justify the quality gap. Model IDs remain callable so pinned code does not hard-break, but do not target them in new work. The `lite` alias now resolves to `gemini-3.5-flash-lite`. #### gemini-2.5-flash **Status:** Stable, generally available **Alias:** `stable-flash` **Strengths:** - Production stability without preview-tier volatility - Solid price-performance for reasoning tasks - Empirically token-perfect on dense transcription benchmarks (May 2026) **Specifications:** - Context window: ~1M tokens input - Multimodal: text, image, video, audio **Best for:** - Production workloads where preview models are too volatile - High-volume tasks with a quality floor - Multimodal extraction when cost matters but accuracy can't slip **Pricing:** - Input: $0.30 / 1M tokens - Output: $2.50 / 1M tokens #### gemini-2.5-flash-lite **Status:** DEPRECATED (retired from routing 2026-07-21) **Alias:** none — `lite` now points at gemini-3.5-flash-lite **Strengths:** - **Cheapest major-provider production model** ($0.10 / $0.40) - Surprisingly capable on multimodal extraction — empirically transcribes dense tables on par with much pricier models - Fast: typically lowest latency in the lineup **Specifications:** - Context window: ~1M tokens input - Multimodal: text, image, video, audio **Best for:** - Ultra-budget batch processing - Routine triage tasks (zeitgeist runs, inbox review, bsky image transcription, classification, simple extraction) - Maximum throughput at minimum cost **Pricing:** - Input: $0.10 / 1M tokens - Output: $0.40 / 1M tokens #### gemini-2.5-pro **Status:** Stable, generally available **Alias:** `stable-pro` **Strengths:** - Pro-tier reasoning with production stability - Well-documented behavior across long-running deployments **Specifications:** - Context window: ~1M tokens input - Long context surcharge: 2× above 200K tokens - Multimodal: text, image, video, audio **Best for:** - Complex tasks requiring production stability - Long-document processing - Quality-critical workloads **Pricing:** - Input: $1.25 / 1M tokens (≤200K), $2.50 (>200K) - Output: $10.00 / 1M tokens (≤200K), $20.00 (>200K) --- ### Speech Generation Models (TTS) **Released 2026-09-23, GA** on the Gemini API and AI Studio (Gemini Enterprise in preview). Google's announcement claims #1 on Hume AI's Voice Design Benchmark (71.4) and the #1 and #2 spots on its Overall Quality Index. #### gemini-3.8-flash-tts - Flagship expressive TTS: acting, regional accents, long-form multi-turn stability. 130 languages, auto-detected. - Limits: 8,192 input tokens / 16,384 output tokens per request; up to two speakers. - Pricing: $0.50 in (text) / $9.00 out (audio) per 1M through 2026-12-31, then $1.00 / $18.00. Batch half. About 32 audio tokens per second of speech (measured 2026-09-24), so ~$0.02 per minute. #### gemini-3.8-flash-lite-tts - Cheaper workhorse for bulk narration, voice-agent cascades and read-aloud; 101 languages. $0.50 / $6.00 per 1M through 2026-12-31, then $1.00 / $12.00. - Google's named replacement for `gemini-3.1-flash-tts-preview` ($1 / $20). #### API shape (verified through the CF gateway, 2026-09-24) - Use the **Interactions API**: `POST v1beta/interactions` with `{"model", "input": [{"type": "user_input", "content": [{"type": "text", "text", "annotations": [{"type": "speech_metadata", "style"}]}]}], "response_format": {"type": "audio", "mime_type": "audio/wav", "sample_rate": 24000}, "generation_config": {"speech_config": [{"voice"}]}}`. Audio is base64 at `steps[].content[]` where `type == "audio"`. - `generateContent` with `responseModalities: ["AUDIO"]` also returns audio, but a "Style: text" prefix is spoken aloud and `systemInstruction` returns HTTP 400 "Developer instruction is not enabled for this model". - Voices: 30 studio voices plus 2,059 persona voices (`GET v1beta/voices`, paged by `next_page_token`, max 1,000 per page). Designed voices come from `POST v1beta/voices` with `{"store": true, "voice": {"type": "prompted", "prompted": {"input": "<description>"}, ...}}` and return a `voice_...` id (1-year expiry, 200 per project) plus a `sample_audio` preview. - Output is watermarked with SynthID. - The model can paraphrase: it added "Hmm," and swapped pronouns in a scripted narration. Check scripted output with ASR. ### Image Generation Models **Updated 2026-05-28:** Nano Banana 2 and Nano Banana Pro reached general availability — announced GA on Vertex AI / Gemini Enterprise Agent Platform, where the GA model IDs drop the suffix (`gemini-3.1-flash-image`, `gemini-3-pro-image`). ⚠️ **Corrected 2026-07-21 (the previous note here was wrong).** The GA IDs `gemini-3.1-flash-image` and `gemini-3-pro-image` are **NOT** Vertex-only and do **NOT** 404 on the Developer API — they were released on this surface on 2026-05-28 and were live-tested working through the CF gateway on 2026-07-21. The `-preview` IDs also still resolve (their announced 2026-06-25 shutdown appears to redirect rather than fail), so nothing is broken either way — but **new code should target the GA IDs**. Also available and not yet wired into this client: `gemini-3.1-flash-lite-image` (Nano Banana 2 Lite, GA) — the cheapest image tier, ~$0.034/image. #### nano-banana-2 **Status:** GA on Vertex; Developer API still serves it as `-preview` (this client's surface) **API Model ID:** `gemini-3.1-flash-image-preview` **Alias:** `image` Fast generation/editing on the Gemini 3.1 Flash Image platform. Default image model in `generate_image()`. Capabilities on the Developer API: - Output resolutions: 512 (0.5K), 1K, 2K generally available; 4K in preview. 512 is 3.1-Flash-only. - Up to 14 reference images (up to 10 high-fidelity objects + up to 4 characters). - Grounding with Google Search, plus Image Search grounding (3.1-Flash-only) — cannot search for images of people. - Thinking: `thinking_level` is {`minimal` (default), `high`}; thinking cannot be fully disabled and thinking tokens are billed. - Extra aspect ratios over 2.5 Flash Image: 1:4, 4:1, 1:8, 8:1. Note: the GA announcement's "video file as input prompt" capability is a Vertex preview feature. The Developer API does NOT accept video or audio input for image generation — don't route video here. #### nano-banana-pro **Status:** GA on Vertex; Developer API still serves it as `-preview` (this client's surface) **API Model ID:** `gemini-3-pro-image-preview` **Alias:** `image-pro` High-fidelity generation on the Gemini 3 Pro Image platform — legible stylized text rendering and professional asset production via advanced "thinking." Capabilities: - Output resolutions: 1K, 2K generally available; 4K in preview. - Up to 14 reference images (up to 6 high-fidelity objects + up to 5 characters). - Thinking always on (cannot be disabled). #### nano-banana **Status:** Stable, GA (unchanged) **API Model ID:** `gemini-2.5-flash-image` Production-grade stability on the Gemini 2.5 Flash Image platform. Works best with up to 3 input images. --- ## Model Selection Guide ``` Default Flash (frontier)? → gemini-3.8-flash (alias: flash) Maximum reasoning? → gemini-3.8-flash, thinking_level='high' (alias: pro) Pro tier? → off routing since 2026-09-03; see gemini-3.1-pro-preview Routine / bulk / cheap / fastest? → gemini-3.5-flash-lite (alias: lite) No-thinking pass (minimal) on Flash? → gemini-3.6-flash (alias: flash-3.6) Pin to prior frontier Flash (3.7)? → gemini-3.7-flash (alias: flash-3.7) Pin to legacy Flash (3.5)? → gemini-3.5-flash (alias: flash-3.5) Pin to older preview Flash? → gemini-3-flash-preview (alias: flash-3) Image generation (fast)? → nano-banana-2 (alias: image) Image generation (high-fidelity)? → nano-banana-pro (alias: image-pro) ``` ## Thinking Configuration (Gemini 3.x family) Gemini 3.x models reason before responding. By default, the model spends output tokens on reasoning, then on the visible answer. From Gemini 3.5 Flash on, the default is `medium` — down from `high` on prior 3.x — and the parameter shape changed: - **Old:** integer `thinking_budget` - **New:** string enum `thinking_level` ∈ {`minimal`, `low`, `medium`, `high`} Which models take `minimal` (all verified live 2026-09-03): | Model | `minimal` | |---|---| | gemini-3.8-flash | HTTP 400 — client downgrades to `low` | | gemini-3.7-flash | HTTP 400 — client downgrades to `low` | | gemini-3.6-flash | accepted | | gemini-3.5-flash | accepted | | gemini-3.5-flash-lite | accepted | The Python client exposes this as `invoke_gemini(..., thinking_level="...")`. Pass `None` (default) to let the model use its built-in default. **When to set `thinking_level='minimal'`:** - Transcription, OCR, image-to-text - Classification, tagging, extraction with a fixed schema - Any task where the LLM doesn't need to reason — it just needs to emit **When to leave it as default or set higher:** - Code generation, debugging - Multi-step planning - Math, complex analysis **Why it matters:** A `max_output_tokens=50` request can return empty if thinking_level (default `medium` on 3.5–3.8) consumes all 50 tokens before emitting visible output. Symptom: response text is empty, `finishReason` is `MAX_TOKENS`. Fix: either raise `max_output_tokens` substantially or set `thinking_level='minimal'`. ## Multimodal Capabilities All text models support: - **Images:** JPEG, PNG, WebP, HEIC, HEIF - **Video:** MP4, MPEG, MOV, AVI, FLV, MPG, WebM, WMV, 3GPP - **Audio:** WAV, MP3, AIFF, AAC, OGG, FLAC **Audio input pricing:** Higher than text, typically ~$1.00 / 1M tokens on Flash-tier models. ## Deprecated / Retired Models | Model | Status | Migration Target | |---|---|---| | gemini-3-pro-preview | Retired (March 9, 2026) | gemini-3.1-pro-preview | | gemini-3-flash-preview | Callable, no shutdown date | gemini-3.6-flash (Google's listed target) | | gemini-3.1-flash-lite-preview | Retired (May 25, 2026) | gemini-3.1-flash-lite | | gemini-3.1-flash-lite | Shutdown May 7, 2027 | gemini-3.5-flash-lite | | gemini-3.1-flash-image-preview | Shutdown listed June 25, 2026 (still resolves) | gemini-3.1-flash-image | | gemini-3-pro-image-preview | Shutdown listed June 25, 2026 (still resolves) | gemini-3-pro-image | | gemini-2.5-flash-image (`nano-banana`) | Shutdown October 2, 2026 | gemini-3.1-flash-image | | gemini-2.0-flash-exp | Retired June 1, 2026 | gemini-3.6-flash | | gemini-2.0-flash | Retired June 1, 2026 | gemini-3.6-flash | | gemini-2.0-flash-lite | Retired June 1, 2026 | gemini-3.5-flash-lite | | gemini-1.5-pro | Retired (404) | gemini-2.5-pro | | gemini-1.5-flash | Retired (404) | gemini-3.6-flash | | gemini-1.0-* | Retired (404) | — | ## Cost Optimization Tips - **Batch API:** 50% discount on all paid models for async (≤24h) processing - **Context caching:** Up to 75–90% savings for repeated large prompts - **Long context:** Pro models charge 2× above 200K tokens — keep prompts concise - **Free tier:** Gemini app + AI Studio offer free access to Flash and Lite models with daily quotas; Pro is paid-only as of April 2026 ## Rate Limits Vary by API tier (default free tier): - **Requests per minute:** 15 - **Tokens per minute:** 1M - **Requests per day:** 1,500 Client automatically handles rate limiting with exponential backoff.
-
-
scripts
-
gemini_client.py 46.5 KB
#!/usr/bin/env python3 """ Gemini API Client Routes requests through Cloudflare AI Gateway when configured (preferred) or directly to Google's Generative Language API via the google-generativeai SDK (fallback). Credential priority: 1. CF Gateway: proxy.env with CF_ACCOUNT_ID, CF_GATEWAY_ID, CF_API_TOKEN 2. Direct API: GOOGLE_API_KEY.txt or API_CREDENTIALS.json """ import json import os import sys import time from pathlib import Path try: import requests HAS_REQUESTS = True except ImportError: HAS_REQUESTS = False try: import google.generativeai as genai HAS_GENAI = True except ImportError: HAS_GENAI = False try: from pydantic import BaseModel HAS_PYDANTIC = True except ImportError: HAS_PYDANTIC = False BaseModel = object # type: ignore[assignment,misc] if not HAS_REQUESTS and not HAS_GENAI: print("Error: neither 'requests' nor 'google-generativeai' is installed.") print("Install with: uv pip install requests google-generativeai pydantic") import sys sys.exit(1) # --------------------------------------------------------------------------- # Model registry # --------------------------------------------------------------------------- # Text generation models MODELS = { # Gemini 3.8 — current frontier Flash (GA 2026-09-02) "gemini-3.8-flash": "gemini-3.8-flash", # Gemini 3.7 — prior frontier Flash (GA 2026-08-13) "gemini-3.7-flash": "gemini-3.7-flash", # Gemini 3.6 — older Flash (GA 2026-07-21) "gemini-3.6-flash": "gemini-3.6-flash", # Gemini 3.5 — prior frontier Flash (GA May 2026) "gemini-3.5-flash": "gemini-3.5-flash", # Gemini 3.x — preview (still callable, kept for back compat) "gemini-3-flash-preview": "gemini-3-flash-preview", # Gemini 3.1 Pro — DEPRECATED from routing 2026-09-03 (Oskar): its # price/quality is dominated by the 3.6+ Flash line. ID kept callable for # pinned code; `pro` no longer points here. "gemini-3.1-pro-preview": "gemini-3.1-pro-preview", # Gemini 3.5 Flash-Lite — cheap/bulk tier (GA 2026-07-21) "gemini-3.5-flash-lite": "gemini-3.5-flash-lite", # Gemini 2.5 — DEPRECATED 2026-07-21. A 2025-era generation; do NOT route # here. IDs kept callable so pinned code does not hard-break, but they are # no longer a recommended target and `lite` no longer points at 2.5. "gemini-2.5-flash": "gemini-2.5-flash", "gemini-2.5-flash-lite": "gemini-2.5-flash-lite", "gemini-2.5-pro": "gemini-2.5-pro", } # Image generation models — display name -> actual API model ID. # NOTE (2026-05-28): Nano Banana 2 / Pro went GA on Vertex (IDs there drop # the suffix: gemini-3.1-flash-image / gemini-3-pro-image). This client uses # the Gemini Developer API (google-ai-studio gateway), where both are STILL # served under the -preview IDs below (verified against the live docs). # Do NOT drop -preview here — the GA IDs are Vertex-only and 404 on this surface. IMAGE_MODELS = { "nano-banana-2": "gemini-3.1-flash-image-preview", "nano-banana-pro": "gemini-3-pro-image-preview", "nano-banana": "gemini-2.5-flash-image", } # Convenience aliases. `flash` points to the current frontier Flash (3.8, GA # 2026-09-02); `flash-3.7`, `flash-3.6`, `flash-3.5` and `flash-3` keep stable # handles on the prior Flash generations for code that pinned to them. `lite` # repointed 2026-07-21 from gemini-2.5-flash-lite to gemini-3.5-flash-lite # (BREAKING: ~6x output cost, $0.40 -> $2.50/M, in exchange for a # current-generation model). MODEL_ALIASES = { "flash": "gemini-3.8-flash", "flash-3.7": "gemini-3.7-flash", "flash-3.6": "gemini-3.6-flash", "flash-3.5": "gemini-3.5-flash", "flash-3": "gemini-3-flash-preview", # `pro` repointed 2026-09-03 from gemini-3.1-pro-preview to the frontier # Flash. The Pro tier is off routing (price/quality dominated by Flash); # "maximum reasoning" is Flash with thinking_level='high'. "pro": "gemini-3.8-flash", "lite": "gemini-3.5-flash-lite", # DEPRECATED aliases (Gemini 2.5). Retained for back compat only. "stable-flash": "gemini-2.5-flash", "stable-pro": "gemini-2.5-pro", "image": "nano-banana-2", "image-pro": "nano-banana-pro", } DEFAULT_MODEL = "gemini-3.8-flash" # Speech (TTS) models — GA 2026-09-23. Kept out of MODEL_ALIASES on purpose: # they return audio, not text, so invoke_gemini() must never resolve to them. SPEECH_MODELS = { "gemini-3.8-flash-tts": "gemini-3.8-flash-tts", # expressive, 130 languages "gemini-3.8-flash-lite-tts": "gemini-3.8-flash-lite-tts", # bulk / read-aloud, 101 languages } SPEECH_ALIASES = {"tts": "gemini-3.8-flash-tts", "tts-lite": "gemini-3.8-flash-lite-tts"} # Flash 3.7 and 3.8 reject thinking_level='minimal' with HTTP 400 ("Thinking # level MINIMAL is not supported for this model"); 3.6, 3.5 and 3.5-lite accept # it. All five verified live through the CF gateway on 2026-09-03. _MINIMAL_THINKING_UNSUPPORTED = frozenset({"gemini-3.7-flash", "gemini-3.8-flash"}) # --------------------------------------------------------------------------- # Cloudflare AI Gateway constants # --------------------------------------------------------------------------- _CF_GATEWAY_BASE = "https://gateway.ai.cloudflare.com/v1" _PROXY_ENV_PATHS = [ Path("/mnt/project/proxy.env"), Path("/mnt/user-data/proxy.env"), Path.home() / ".muninn" / "proxy.env", ] # --------------------------------------------------------------------------- # Credential loading # --------------------------------------------------------------------------- def _parse_env_file(path: Path) -> dict: """Parse a .env-format file into a dict, stripping quotes.""" result = {} for line in path.read_text().splitlines(): line = line.strip() if line and not line.startswith("#") and "=" in line: key, _, value = line.partition("=") result[key.strip()] = value.strip().strip('"').strip("'") return result def get_cf_credentials() -> dict | None: """ Load Cloudflare AI Gateway credentials. Searches for proxy.env in well-known paths, then falls back to environment variables. Required keys: CF_ACCOUNT_ID, CF_GATEWAY_ID, CF_API_TOKEN Optional key: GOOGLE_API_KEY (for non-BYOK setups) Returns: dict with credentials if fully configured, None otherwise """ required = ("CF_ACCOUNT_ID", "CF_GATEWAY_ID", "CF_API_TOKEN") for env_path in _PROXY_ENV_PATHS: if env_path.exists(): try: creds = _parse_env_file(env_path) if all(creds.get(k) for k in required): return creds except OSError: continue # Fall back to environment variables creds = {k: os.environ.get(k, "") for k in required} creds["GOOGLE_API_KEY"] = os.environ.get("GOOGLE_API_KEY", "") if all(creds.get(k) for k in required): return creds return None def get_google_api_key() -> str: """ Get Google API key for direct (non-gateway) access. Priority order: 1. Individual file: /mnt/project/GOOGLE_API_KEY.txt 2. Combined file: /mnt/project/API_CREDENTIALS.json 3. Environment variable: GOOGLE_API_KEY Returns: str: Google API key Raises: ValueError: If no API key found in any source """ # 1. Individual key file key_file = Path("/mnt/project/GOOGLE_API_KEY.txt") if key_file.exists(): try: key = key_file.read_text().strip() if key: return key except OSError as e: raise ValueError(f"Found GOOGLE_API_KEY.txt but couldn't read it: {e}") # 2. Combined credentials file creds_file = Path("/mnt/project/API_CREDENTIALS.json") if creds_file.exists(): try: with open(creds_file) as f: config = json.load(f) key = config.get("google_api_key", "").strip() if key: return key except (json.JSONDecodeError, OSError) as e: raise ValueError(f"Found API_CREDENTIALS.json but couldn't parse it: {e}") # 3. Environment variable key = os.environ.get("GOOGLE_API_KEY", "").strip() if key: return key raise ValueError( "No Google API key found!\n\n" "Option A (recommended): Configure Cloudflare AI Gateway\n" " File: /mnt/project/proxy.env\n" " Content:\n" " CF_ACCOUNT_ID=<your-account-id>\n" " CF_GATEWAY_ID=<your-gateway-id>\n" " CF_API_TOKEN=<your-cf-api-token>\n\n" "Option B: Direct Google API\n" " File: GOOGLE_API_KEY.txt (content: AIzaSy...)\n" " or\n" " File: API_CREDENTIALS.json (content: {\"google_api_key\": \"AIzaSy...\"})\n\n" "Get your Cloudflare token: https://dash.cloudflare.com/profile/api-tokens\n" "Get your Google key: https://console.cloud.google.com/apis/credentials" ) # --------------------------------------------------------------------------- # Cloudflare AI Gateway — REST path # --------------------------------------------------------------------------- def _cf_request( model_id: str, contents: list, generation_config: dict, cf_creds: dict, ) -> dict: """ POST a generateContent request via Cloudflare AI Gateway. Args: model_id: Gemini model ID (e.g., 'gemini-3-flash-preview') contents: Gemini REST API contents array generation_config: generationConfig dict (camelCase keys) cf_creds: dict with CF_ACCOUNT_ID, CF_GATEWAY_ID, CF_API_TOKEN Returns: Parsed JSON response dict Raises: requests.HTTPError: On non-2xx HTTP response """ account_id = cf_creds["CF_ACCOUNT_ID"] gateway_id = cf_creds["CF_GATEWAY_ID"] api_token = cf_creds["CF_API_TOKEN"] url = ( f"{_CF_GATEWAY_BASE}/{account_id}/{gateway_id}" f"/google-ai-studio/v1beta/models/{model_id}:generateContent" ) # Include Google API key as query param for non-BYOK setups google_key = cf_creds.get("GOOGLE_API_KEY") or os.environ.get("GOOGLE_API_KEY", "") if google_key: url += f"?key={google_key}" payload: dict = {"contents": contents} if generation_config: payload["generationConfig"] = generation_config headers = { "Content-Type": "application/json", "cf-aig-authorization": f"Bearer {api_token}", } # Retry on 5xx / 429 / SSL / non-JSON proxy errors. The Claude.ai egress # proxy can return HTTP 503 with body 'DNS cache overflow' (text/plain), # most often on cold start; without this retry the caller sees an # opaque JSONDecodeError or HTTPError on what is actually a transient # proxy condition rather than a Gemini/CF-AI-Gateway failure. max_retries = 3 base_delay = 0.5 for attempt in range(max_retries): try: response = requests.post(url, json=payload, headers=headers, timeout=120) if response.status_code >= 500: preview = (response.text or '')[:200] raise RuntimeError( f"HTTP {response.status_code} from CF AI Gateway " f"(likely egress proxy, not Gemini): {preview!r}" ) if 400 <= response.status_code < 500: # raise_for_status() discards the body, and Gemini puts the ONLY # useful diagnostic there — e.g. which schema keyword it rejected. # A 4xx is deterministic, so mark it non-retriable too. preview = (response.text or '')[:600] raise _NonRetriableAPIError( f"HTTP {response.status_code} from Gemini: {preview}" ) response.raise_for_status() return response.json() except (MediaInputError, _NonRetriableAPIError): raise # deterministic error — retrying cannot help except Exception as e: if attempt == max_retries - 1: raise err = str(e) retriable = ( '503' in err or '429' in err or 'Service Unavailable' in err or 'DNS cache overflow' in err or 'Expecting value' in err or 'JSONDecodeError' in err or 'SSL' in err or 'SSLError' in err or 'HANDSHAKE_FAILURE' in err ) if not retriable: raise delay = base_delay * (2 ** attempt) print( f"Warning: CF AI Gateway request failed " f"(attempt {attempt + 1}/{max_retries}), retrying in {delay}s: {e}" ) time.sleep(delay) def _extract_text(response: dict) -> str | None: """Extract generated text from a Gemini REST API response.""" try: return response["candidates"][0]["content"]["parts"][0]["text"] except (KeyError, IndexError, TypeError): return None def _check_truncation(response: dict) -> None: """Raise if the model hit the output cap, instead of letting json.loads fail. A truncated structured response is invalid JSON, so the symptom is a JSONDecodeError / pydantic "EOF while parsing" that looks like a schema problem. finishReason says what actually happened. """ try: reason = response["candidates"][0].get("finishReason") except (KeyError, IndexError, TypeError): return if reason in ("MAX_TOKENS", "MAX_OUTPUT_TOKENS"): raise _NonRetriableAPIError( "Gemini hit maxOutputTokens before finishing the JSON object " "(finishReason=MAX_TOKENS). Raise max_output_tokens — thinking " "tokens count against the same budget." ) class MediaInputError(ValueError): """Invalid media input. Deterministic — never worth retrying.""" class _NonRetriableAPIError(RuntimeError): """A 4xx from Gemini. Deterministic (bad schema, bad request) — retrying only hides the response body, which is the one thing worth reading.""" # Extensions mimetypes.guess_type() gets wrong or misses, per Gemini's accepted # media types. Audio/video matter here: a bad guess silently sends the wrong # mimeType and the model returns a confused answer instead of an error. _MEDIA_MIME_OVERRIDES = { ".m4a": "audio/mp4", ".aac": "audio/aac", ".flac": "audio/flac", ".ogg": "audio/ogg", ".opus": "audio/opus", ".mp3": "audio/mpeg", ".wav": "audio/wav", ".aiff": "audio/aiff", ".mp4": "video/mp4", ".mov": "video/quicktime", ".webm": "video/webm", ".heic": "image/heic", ".heif": "image/heif", } # generateContent inline payloads must fit the request; base64 inflates ~4/3. # Files above this need the Files API (not implemented in this client). _INLINE_MEDIA_MAX_BYTES = 15 * 1024 * 1024 def _guess_media_mime(path: str) -> str: """Resolve a media mimeType, preferring explicit overrides over mimetypes.""" import mimetypes ext = Path(path).suffix.lower() if ext in _MEDIA_MIME_OVERRIDES: return _MEDIA_MIME_OVERRIDES[ext] return mimetypes.guess_type(path)[0] or "application/octet-stream" def _build_contents(prompt: str, image_path: str | None) -> list: """Build the Gemini REST API 'contents' array. `image_path` accepts ANY supported media file — image, audio, or video — despite the legacy name. Audio input is verified working on this path (2026-07-21). """ parts: list = [{"text": prompt}] if image_path: import base64 size = Path(image_path).stat().st_size if size > _INLINE_MEDIA_MAX_BYTES: raise MediaInputError( f"{image_path} is {size/1e6:.1f}MB; inline media is capped at " f"{_INLINE_MEDIA_MAX_BYTES/1e6:.0f}MB. Larger files need the " f"Files API, which this client does not implement yet." ) image_data = Path(image_path).read_bytes() mime_type = _guess_media_mime(image_path) parts.append({ "inlineData": { "mimeType": mime_type, "data": base64.b64encode(image_data).decode(), } }) return [{"parts": parts}] #: Keywords Gemini's responseSchema rejects outright. Pydantic emits all of #: these; leaving any of them in produces a bare HTTP 400 with no usable body. _SCHEMA_STRIP_KEYS = ("$schema", "$defs", "title", "default", "additionalProperties", "discriminator", "examples", "const") def _inline_refs(node, defs: dict): """Resolve $ref against $defs and strip Gemini-unsupported keywords, recursively. Gemini's responseSchema is NOT full JSON Schema: it rejects $ref/$defs entirely. Pydantic emits a $ref for every nested model, so any model with a nested model (the common case — List[SomeModel]) produced a bare HTTP 400 that the retry loop then burned three attempts on before returning None. Flat single-level models worked, which is why this hid for so long. """ if isinstance(node, dict): if "$ref" in node: name = node["$ref"].split("/")[-1] target = defs.get(name) if target is None: # Unresolvable ref — better a permissive object than a 400. return {"type": "object"} return _inline_refs(target, defs) out = {} for k, v in node.items(): if k in _SCHEMA_STRIP_KEYS: continue if k == "anyOf": # Gemini has no anyOf. Optional[X] becomes [X, null]; take the # first non-null branch and mark it nullable instead. branches = [b for b in v if b.get("type") != "null"] if branches: resolved = _inline_refs(branches[0], defs) if len(branches) < len(v): resolved["nullable"] = True out.update(resolved) continue out[k] = _inline_refs(v, defs) return out if isinstance(node, list): return [_inline_refs(x, defs) for x in node] return node def _pydantic_to_schema(model_class: type) -> dict: """Convert a Pydantic model class to a Gemini-compatible JSON schema dict. Handles nested models: pydantic's $defs/$ref indirection is inlined, because Gemini's responseSchema does not support it. """ try: schema = model_class.model_json_schema() # Pydantic v2 except AttributeError: schema = model_class.schema() # Pydantic v1 defs = schema.get("$defs") or schema.get("definitions") or {} return _inline_refs(schema, defs) # --------------------------------------------------------------------------- # Direct SDK path helpers # --------------------------------------------------------------------------- def _initialize_direct_client() -> bool: """Configure google.generativeai SDK for direct API access.""" if not HAS_GENAI: return False try: api_key = get_google_api_key() genai.configure(api_key=api_key) return True except ValueError as e: print(f"Error: {e}") return False def _build_genai_content(prompt: str, image_path: str | None): """Build content argument for google.generativeai SDK calls. NOTE: this direct-SDK fallback handles IMAGES only. Non-image media (audio/video) works on the CF Gateway path; PIL cannot open it, so fail with a clear message instead of an opaque UnidentifiedImageError. """ if image_path and not _guess_media_mime(image_path).startswith("image/"): raise MediaInputError( f"{image_path} is not an image ({_guess_media_mime(image_path)}). " "Audio/video input requires the Cloudflare AI Gateway path — " "configure proxy.env; the direct google-generativeai SDK fallback " "supports images only." ) if image_path: from PIL import Image # type: ignore[import] return [prompt, Image.open(image_path)] return prompt # --------------------------------------------------------------------------- # Public API # --------------------------------------------------------------------------- def _resolve_model(model: str) -> str: """Resolve a model name or alias to its canonical API model ID. Handles chained resolution: alias → display name → API model ID. Args: model: Model name, alias, or direct ID Returns: Canonical API model ID string Raises: ValueError: If model is not recognized """ # Resolve alias first (e.g., "image" → "nano-banana-2") resolved = MODEL_ALIASES.get(model, model) # Direct match in text models if resolved in MODELS: return MODELS[resolved] # Image models (display name → API model ID) if resolved in IMAGE_MODELS: return IMAGE_MODELS[resolved] all_names = list(MODELS) + list(MODEL_ALIASES) + list(IMAGE_MODELS) raise ValueError(f"Invalid model: {model}. Choose from {all_names}") # @lat: [[orchestration#Gemini Client]] def invoke_gemini( prompt: str, model: str = DEFAULT_MODEL, temperature: float = 0.7, max_output_tokens: int | None = None, top_p: float | None = None, top_k: int | None = None, image_path: str | None = None, thinking_level: str | None = None, ) -> str | None: """ Invoke Gemini model with a text (or multi-modal) prompt. Routes through Cloudflare AI Gateway when proxy.env is configured; falls back to direct Google API via google-generativeai SDK. Args: prompt: The text prompt to send model: Model name or alias (default: gemini-3.6-flash). Aliases: flash (3.6), flash-3.5 (prior Flash), flash-3 (prior preview), pro, lite, stable-flash, stable-pro temperature: Sampling temperature (0.0–1.0) max_output_tokens: Maximum tokens in response. Note: with thinking models (Gemini 3.x), thinking tokens consume part of this budget; set generously or use thinking_level='minimal' for non-reasoning tasks. top_p: Nucleus sampling parameter top_k: Top-k sampling parameter image_path: Optional path to a media file for multi-modal input. Despite the name, accepts image, AUDIO, or video (audio verified 2026-07-21). Requires the CF Gateway path for non-image media. Inline only — files >15MB need the Files API (not implemented). thinking_level: Reasoning budget for thinking models (Gemini 3.x). One of 'minimal', 'low', 'medium', 'high'. Default (None) lets the model use its built-in default — 'medium' for 3.x Flash (incl. 3.8), which silently eats output budget. Set 'minimal' for tasks that don't need reasoning (transcription, classification, extraction). Flash 3.7 and 3.8 do not accept 'minimal' (HTTP 400); on those the client downgrades it to 'low', the cheapest level they take, and says so on stderr. Ignored by 2.5 models (which use a different parameter). Returns: Response text if successful, None if error """ model_id = _resolve_model(model) cf_creds = get_cf_credentials() # thinking_level is a Gemini 3.x feature. 2.5 (and earlier) models 400 on it. # Silently drop for non-3.x so callers can pass it uniformly without branching. if thinking_level is not None and not model_id.startswith("gemini-3"): thinking_level = None if thinking_level == "minimal" and model_id in _MINIMAL_THINKING_UNSUPPORTED: print( f"Note: {model_id} does not support thinking_level='minimal'; using 'low'.", file=sys.stderr, ) thinking_level = "low" max_retries = 3 for attempt in range(max_retries): try: if cf_creds and HAS_REQUESTS: # --- Cloudflare AI Gateway path --- contents = _build_contents(prompt, image_path) gen_cfg: dict = {"temperature": temperature} if max_output_tokens: gen_cfg["maxOutputTokens"] = max_output_tokens if top_p is not None: gen_cfg["topP"] = top_p if top_k is not None: gen_cfg["topK"] = top_k if thinking_level is not None: gen_cfg["thinkingConfig"] = {"thinkingLevel": thinking_level} response = _cf_request(model_id, contents, gen_cfg, cf_creds) return _extract_text(response) else: # --- Direct SDK path --- if not _initialize_direct_client(): return None gen_cfg_sdk = {"temperature": temperature} if max_output_tokens: gen_cfg_sdk["max_output_tokens"] = max_output_tokens if top_p is not None: gen_cfg_sdk["top_p"] = top_p if top_k is not None: gen_cfg_sdk["top_k"] = top_k if thinking_level is not None: # SDK uses snake_case; mapped to thinking_config later. gen_cfg_sdk["thinking_config"] = {"thinking_level": thinking_level} model_instance = genai.GenerativeModel( model_name=model_id, generation_config=gen_cfg_sdk, ) content = _build_genai_content(prompt, image_path) response = model_instance.generate_content(content) return response.text except (MediaInputError, _NonRetriableAPIError): raise # deterministic error — retrying cannot help except Exception as e: if attempt < max_retries - 1: wait_time = 2 ** attempt print(f"Retry {attempt + 1}/{max_retries} after {wait_time}s: {e}") time.sleep(wait_time) else: print(f"Error invoking Gemini: {e}") return None return None def generate_image( prompt: str, output_path: str | None = None, model: str = "nano-banana-2", temperature: float = 0.7, ) -> dict | None: """ Generate an image using a Gemini image model. Sends a prompt with responseModalities ["IMAGE", "TEXT"] and saves the resulting image to disk. Args: prompt: Text prompt describing the desired image output_path: Where to save the PNG. If None, auto-generates under /mnt/user-data/outputs/ (or /tmp/ if that doesn't exist) model: Image model name or alias (default: nano-banana-2). Aliases: image, image-pro temperature: Sampling temperature (0.0–1.0) Returns: dict with keys 'path' (str) and 'caption' (str|None) on success, None on failure """ import base64 # Resolve model — must land in IMAGE_MODELS resolved = model if model in MODEL_ALIASES: resolved = MODEL_ALIASES[model] if resolved not in IMAGE_MODELS: raise ValueError( f"Model '{model}' is not an image model. " f"Use one of: {list(IMAGE_MODELS)} or aliases: image, image-pro" ) model_id = IMAGE_MODELS[resolved] # Determine output path if output_path is None: ts = int(time.time()) out_dir = Path("/mnt/user-data/outputs") if not out_dir.exists(): out_dir = Path("/tmp") output_path = str(out_dir / f"gemini_image_{ts}.png") cf_creds = get_cf_credentials() max_retries = 3 for attempt in range(max_retries): try: if cf_creds and HAS_REQUESTS: # --- Cloudflare AI Gateway path --- contents = [{"parts": [{"text": prompt}]}] gen_cfg = { "temperature": temperature, "responseModalities": ["IMAGE", "TEXT"], } response = _cf_request(model_id, contents, gen_cfg, cf_creds) elif HAS_GENAI: # --- Direct SDK path --- if not _initialize_direct_client(): return None model_instance = genai.GenerativeModel( model_name=model_id, generation_config={ "temperature": temperature, "response_modalities": ["IMAGE", "TEXT"], }, ) response_obj = model_instance.generate_content(prompt) # Convert SDK response to REST-like dict for unified extraction response = _sdk_response_to_dict(response_obj) else: print("Error: no credentials configured and google-generativeai not installed") return None # Extract image and optional caption from response image_data = None caption = None candidates = response.get("candidates", []) if not candidates: print("Error: no candidates in response") return None parts = candidates[0].get("content", {}).get("parts", []) for part in parts: if "inlineData" in part: image_data = part["inlineData"].get("data") elif "inline_data" in part: image_data = part["inline_data"].get("data") elif "text" in part: caption = part["text"] if not image_data: print("Error: no image data in response") return None # Decode and save Path(output_path).parent.mkdir(parents=True, exist_ok=True) Path(output_path).write_bytes(base64.b64decode(image_data)) print(f"Image saved to {output_path}") return {"path": output_path, "caption": caption} except (MediaInputError, _NonRetriableAPIError): raise # deterministic error — retrying cannot help except Exception as e: if attempt < max_retries - 1: wait_time = 2 ** attempt print(f"Retry {attempt + 1}/{max_retries} after {wait_time}s: {e}") time.sleep(wait_time) else: print(f"Error generating image: {e}") return None return None def _rest_request(method: str, path: str, body: dict | None = None, params: dict | None = None, timeout: int = 180) -> dict: """Call a v1beta REST path (e.g. 'interactions', 'voices') via the CF gateway, or directly with GOOGLE_API_KEY in a header (never in the URL). Retries 429/5xx; a 4xx raises with Google's error body.""" cf = get_cf_credentials() if cf: url = (f"{_CF_GATEWAY_BASE}/{cf['CF_ACCOUNT_ID']}/{cf['CF_GATEWAY_ID']}" f"/google-ai-studio/v1beta/{path}") headers = {"cf-aig-authorization": f"Bearer {cf['CF_API_TOKEN']}"} if cf.get("GOOGLE_API_KEY"): headers["x-goog-api-key"] = cf["GOOGLE_API_KEY"] else: url = f"https://generativelanguage.googleapis.com/v1beta/{path}" headers = {"x-goog-api-key": get_google_api_key()} headers["Content-Type"] = "application/json" for attempt in range(4): r = requests.request(method, url, json=body, params=params, headers=headers, timeout=timeout) if r.status_code in (429, 500, 502, 503, 504) and attempt < 3: time.sleep(1.0 * 2 ** attempt) continue if r.status_code >= 400: raise _NonRetriableAPIError(f"HTTP {r.status_code} from Gemini: {(r.text or '')[:600]}") return r.json() raise RuntimeError("unreachable") def generate_speech( text: str, output_path: str | None = None, voice: str = "Charon", style: str | None = None, model: str = "tts", sample_rate: int = 24000, ) -> dict | None: """ Synthesize speech with Gemini 3.8 Flash TTS and save it as WAV. Uses the Interactions API. generateContent also returns audio for these models, but it has no style channel: a "Style: text" prefix is SPOKEN, and systemInstruction is refused ("Developer instruction is not enabled for this model"). Style goes in a speech_metadata annotation here. Args: text: What to say. Inline events work: <laugh>, <sigh>, <breath>, <short pause>; CAPITALS stress a word. output_path: WAV path; defaults under /mnt/user-data/outputs/ or /tmp/. voice: A prebuilt name ("Charon", "Algenib", "en-gb-storyteller-4" — see list_voices()) or a designed voice id ("voice_..."). style: Turn-level delivery, e.g. "quiet and dry, unhurried". model: "tts" (gemini-3.8-flash-tts) or "tts-lite". sample_rate: Output rate in Hz (24000 default; mono 16-bit PCM). Returns: {'path', 'seconds', 'audio_tokens'} on success, None on failure. """ import base64 resolved = SPEECH_ALIASES.get(model, model) if resolved not in SPEECH_MODELS: print(f"Error: '{model}' is not a speech model. Known: {list(SPEECH_MODELS) + list(SPEECH_ALIASES)}", file=sys.stderr) return None part = {"type": "text", "text": text} if style: part["annotations"] = [{"type": "speech_metadata", "style": style}] body = {"model": resolved, "input": [{"type": "user_input", "content": [part]}], "response_format": {"type": "audio", "mime_type": "audio/wav", "sample_rate": sample_rate}, "generation_config": {"speech_config": [{"voice": voice}]}} try: j = _rest_request("POST", "interactions", body) audio = [c for st in j.get("steps", []) for c in st.get("content", []) if c.get("type") == "audio"] if not audio: print(f"Error: no audio in response: {json.dumps(j)[:300]}", file=sys.stderr) return None data = base64.b64decode(audio[-1]["data"]) except Exception as e: print(f"Error: speech generation failed: {e}", file=sys.stderr) return None if output_path is None: out_dir = Path("/mnt/user-data/outputs") if Path("/mnt/user-data/outputs").exists() else Path("/tmp") output_path = str(out_dir / f"gemini_speech_{int(time.time())}.wav") Path(output_path).write_bytes(data) # WAV body is 16-bit mono PCM after a 44-byte RIFF header return {"path": output_path, "seconds": max(0, len(data) - 44) / (2 * sample_rate), "audio_tokens": j.get("usage", {}).get("total_output_tokens")} def design_voice(description: str, display_name: str, gender: str | None = None, language_code: str = "en-US", model: str = "tts") -> dict | None: """ Create a stored voice from a natural-language description (voice design). POST v1beta/voices with type "prompted". Prompted voices must be stored (store=true): the id ('voice_...') lasts one year, max 200 per project. The description shapes delivery too — "thoughtful pauses" in it produced 1-1.8 s mid-line pauses that no per-line style removed (2026-09-24). Returns: {'id', 'expire_time', 'sample_path'} on success, None on failure. """ import base64 voice = {"model": SPEECH_ALIASES.get(model, model), "type": "prompted", "display_name": display_name, "language_code": language_code, "prompted": {"input": description}} if gender: voice["gender"] = gender try: j = _rest_request("POST", "voices", {"store": True, "voice": voice}) except Exception as e: print(f"Error: voice design failed: {e}", file=sys.stderr) return None sample = (j.get("sample_audio") or {}).get("data") sample_path = None if sample: sample_path = f"/tmp/{j['id']}_sample.wav" Path(sample_path).write_bytes(base64.b64decode(sample)) return {"id": j.get("id"), "expire_time": j.get("expire_time"), "sample_path": sample_path} def list_voices(**filters) -> list | None: """ List the voice library (2,089 prebuilt voices as of 2026-09-24, plus this project's designed voices), paging through next_page_token. Filters are server-side query params: gender ('male'|'female'|'neutral'), pitch ('low'|'medium'|'high'), search (substring of the description). `accent` needs the exact string ('Winchester English', not 'British'), so filter accents client-side on the returned 'accent' field. Returns: list of voice dicts (id, display_name, language_code, accent, gender, pitch, persona, description), or None on failure. """ out, token = [], None try: while True: params = {"page_size": 1000, **filters} if token: params["page_token"] = token j = _rest_request("GET", "voices", params=params, timeout=60) out += j.get("voices", []) token = j.get("next_page_token") if not token: return out except Exception as e: print(f"Error: listing voices failed: {e}", file=sys.stderr) return None def _sdk_response_to_dict(response_obj) -> dict: """Convert a google.generativeai SDK response to a REST-like dict. This allows generate_image() to use the same extraction logic for both the CF Gateway (REST) and direct SDK paths. Args: response_obj: GenerateContentResponse from the SDK Returns: dict matching the Gemini REST API response shape """ import base64 candidates = [] for candidate in response_obj.candidates: parts = [] for part in candidate.content.parts: if hasattr(part, "text") and part.text: parts.append({"text": part.text}) elif hasattr(part, "inline_data") and part.inline_data: data = part.inline_data parts.append({ "inlineData": { "mimeType": data.mime_type, "data": base64.b64encode(data.data).decode() if isinstance(data.data, bytes) else data.data, } }) candidates.append({"content": {"parts": parts}}) return {"candidates": candidates} def invoke_with_structured_output( prompt: str, pydantic_model: type, model: str = DEFAULT_MODEL, temperature: float = 0.7, image_path: str | None = None, max_output_tokens: int | None = 32768, ) -> object | None: """ Invoke Gemini with structured (JSON schema) output using a Pydantic model. Args: prompt: The text prompt to send pydantic_model: Pydantic model class for response schema model: Model name or alias (default: DEFAULT_MODEL, gemini-3.8-flash). Aliases: flash, pro, lite, stable-flash, stable-pro temperature: Sampling temperature (0.0–1.0) image_path: Optional path to a media file for multi-modal input. Despite the name, accepts image, AUDIO, or video (audio verified 2026-07-21). Requires the CF Gateway path for non-image media. Inline only — files >15MB need the Files API (not implemented). Returns: Instance of pydantic_model if successful, None if error """ if not HAS_PYDANTIC: print("Error: pydantic not installed. Run: uv pip install pydantic") return None model_id = _resolve_model(model) cf_creds = get_cf_credentials() max_retries = 3 for attempt in range(max_retries): try: if cf_creds and HAS_REQUESTS: # --- Cloudflare AI Gateway path --- contents = _build_contents(prompt, image_path) schema = _pydantic_to_schema(pydantic_model) gen_cfg = { "temperature": temperature, "responseMimeType": "application/json", "responseSchema": schema, } # Thinking tokens count against this budget, so a default sized # for the visible answer truncates the JSON mid-object and # surfaces as a confusing JSONDecodeError rather than a length # error. Default generously. if max_output_tokens: gen_cfg["maxOutputTokens"] = max_output_tokens response = _cf_request(model_id, contents, gen_cfg, cf_creds) _check_truncation(response) text = _extract_text(response) if text: json_data = json.loads(text) return pydantic_model(**json_data) else: # --- Direct SDK path --- if not _initialize_direct_client(): return None model_instance = genai.GenerativeModel( model_name=model_id, generation_config={ "temperature": temperature, "response_mime_type": "application/json", "response_schema": pydantic_model, }, ) content = _build_genai_content(prompt, image_path) response = model_instance.generate_content(content) json_data = json.loads(response.text) return pydantic_model(**json_data) except (MediaInputError, _NonRetriableAPIError): raise # deterministic error — retrying cannot help except Exception as e: if attempt < max_retries - 1: wait_time = 2 ** attempt print(f"Retry {attempt + 1}/{max_retries} after {wait_time}s: {e}") time.sleep(wait_time) else: print(f"Error invoking Gemini with structured output: {e}") return None return None def invoke_parallel( prompts: list, model: str = DEFAULT_MODEL, temperature: float = 0.7, max_workers: int = 5, ) -> list: """ Invoke Gemini with multiple prompts in parallel. Args: prompts: List of text prompts to process model: Model name or alias (default: DEFAULT_MODEL, gemini-3.8-flash). Aliases: flash, pro, lite, stable-flash, stable-pro temperature: Sampling temperature (0.0–1.0) max_workers: Maximum concurrent requests Returns: List of response strings (None for failed requests) in prompt order """ from concurrent.futures import ThreadPoolExecutor, as_completed results: list = [None] * len(prompts) def _process(idx: int, prompt: str): return idx, invoke_gemini(prompt, model=model, temperature=temperature) with ThreadPoolExecutor(max_workers=max_workers) as executor: futures = { executor.submit(_process, idx, prompt): idx for idx, prompt in enumerate(prompts) } for future in as_completed(futures): try: idx, response = future.result() results[idx] = response except Exception as e: idx = futures[future] print(f"Error processing prompt {idx}: {e}") results[idx] = None return results def get_available_models() -> dict: """Return dict of registered Gemini models grouped by category. Returns: dict with keys 'text', 'image', 'speech', 'aliases', 'speech_aliases' """ return { "text": list(MODELS.keys()), "image": list(IMAGE_MODELS.keys()), "speech": list(SPEECH_MODELS.keys()), "aliases": dict(MODEL_ALIASES), "speech_aliases": dict(SPEECH_ALIASES), } def verify_setup() -> bool: """ Verify that Gemini client is properly configured. Returns: True if at least one credential source is valid and a test call succeeds """ cf_creds = get_cf_credentials() if cf_creds: print(f"Using Cloudflare AI Gateway (account: {cf_creds['CF_ACCOUNT_ID'][:8]}...)") elif HAS_GENAI: if not _initialize_direct_client(): return False print("Using direct Google API (google-generativeai SDK)") else: print("Error: no credentials configured and google-generativeai not installed") return False try: test_response = invoke_gemini("Say 'OK'", model=DEFAULT_MODEL) return test_response is not None except Exception as e: print(f"Setup verification failed: {e}") return False # --------------------------------------------------------------------------- # Self-test # --------------------------------------------------------------------------- if __name__ == "__main__": import sys print("Gemini Client Self-Test") print("=" * 50) cf = get_cf_credentials() if cf: print(f"Backend: Cloudflare AI Gateway ({cf['CF_ACCOUNT_ID'][:8]}.../{cf['CF_GATEWAY_ID']})") elif HAS_GENAI: print("Backend: Direct Google API (google-generativeai SDK)") else: print("ERROR: no credentials and google-generativeai not installed") sys.exit(1) print("\n1. Verifying setup...") if verify_setup(): print(" ✓ Setup verified") else: print(" ✗ Setup failed") sys.exit(1) print("\n2. Available models:") available = get_available_models() for category, items in available.items(): if isinstance(items, dict): print(f" {category}:") for alias, target in items.items(): print(f" {alias} → {target}") else: print(f" {category}: {', '.join(items)}") print("\n3. Testing basic invocation...") resp = invoke_gemini("What is 2+2? Answer in one word.", model=DEFAULT_MODEL) if resp: print(f" Response: {resp.strip()}") else: print(" ✗ Invocation failed") sys.exit(1) if HAS_PYDANTIC: print("\n4. Testing structured output...") from pydantic import BaseModel as PM from pydantic import Field class MathAnswer(PM): result: int = Field(description="The numerical result") explanation: str = Field(description="Brief explanation") structured = invoke_with_structured_output( prompt="What is 5+7? Provide result and explanation.", pydantic_model=MathAnswer, model=DEFAULT_MODEL, ) if structured: print(f" Result: {structured.result}") print(f" Explanation: {structured.explanation}") else: print(" ✗ Structured output failed") print("\n5. Testing parallel invocation...") test_prompts = [ "Capital of France? One word.", "Capital of Japan? One word.", "Capital of Brazil? One word.", ] parallel_results = invoke_parallel(test_prompts, model=DEFAULT_MODEL) for prompt, result in zip(test_prompts, parallel_results): status = result.strip() if result else "Failed" print(f" {prompt[:35]}... → {status}") print("\n" + "=" * 50) print("Self-test complete!")
-
-
CHANGELOG.md 10.9 KB
# invoking-gemini - Changelog ## 2026-09-24 ### Added — speech generation (Gemini 3.8 Flash TTS, GA 2026-09-23) - `generate_speech()` writes a WAV from text with a prebuilt or designed voice and an optional style; `design_voice()` creates a stored voice from a description; `list_voices()` pages the 2,089-voice library. Run through the CF gateway on 2026-09-24, they returned a 7.8 s WAV, a stored voice id with its sample, and 2,092 voices (the library plus this project's designs). - `SPEECH_MODELS` / `SPEECH_ALIASES` (`tts`, `tts-lite`) are kept apart from `MODEL_ALIASES`, so `invoke_gemini()` cannot resolve to an audio model. - The calls use the Interactions API. On `generateContent` a "Style:" prefix is spoken and `systemInstruction` is refused. - `_rest_request()` sends a direct-mode API key in the `x-goog-api-key` header rather than the URL. ## 2026-09-03 ### ⚠️ BREAKING — `flash` alias and `DEFAULT_MODEL` repointed to gemini-3.8-flash - Gemini 3.8 Flash reached GA on 2026-09-02. It is now the registry default and the `flash` alias. Google also shipped 3.7 Flash on 2026-08-13, which this skill missed; both are added to `MODELS`. - New pinned aliases `flash-3.7` and `flash-3.6`. `flash-3.5` and `flash-3` are unchanged. Nothing is removed; no 3.x Flash has a shutdown date. - Price is the same as 3.6 on Google's current page ($0.75 in / $3.75 out through 2026-12-31, then $1.50 / $7.50), so the swap costs nothing per token. Google says 3.8 "works harder" at higher effort levels, so per-task thinking tokens may go up. ### ⚠️ BREAKING — `pro` alias repointed to gemini-3.8-flash; Pro tier off routing - Oskar, 2026-09-03: "We should never use 3.1 Pro, its Pareto efficiency is too poor compared to the later Flash models." `gemini-3.1-pro-preview` costs 2.7× / 3.2× (in / out) what 3.8 Flash does at today's rates and the 3.5+ Flash line already beat it on coding and agentic benchmarks. The ID stays in `MODELS` for pinned callers; the table marks it DEPRECATED. "Maximum reasoning" is now Flash with `thinking_level='high'`. A future 3.5 Pro gets the same test before any alias points at it. ### Changed — `thinking_level='minimal'` on 3.7 / 3.8 Flash - Both models return HTTP 400 for `minimal` (verified live 2026-09-03; 3.6, 3.5 and 3.5-lite accept it). `invoke_gemini()` now downgrades `minimal` to `low` on those two model IDs and prints a one-line note to stderr, so callers that pass `minimal` uniformly (transcription, classification) keep working on the new default instead of burning a non-retriable 400 and returning `None`. Measured on 3.8: `low` spent 0 thinking tokens on a one-word reply, the default `medium` spent 79. On 3.7, `low` still spent 45–88 on the same prompt, so the downgrade is not free there — budget output generously. - For a true no-thinking pass, pin `flash-3.6` or `lite`. ### Fixed — stale figures in the model tables - gemini-3.1-pro-preview output above 200K is $18.00 / 1M, not $24.00. - gemini-3.6-flash pricing carried the flat $1.50 / $7.50 from its launch table; Google's pricing page now lists it on the same introductory schedule as 3.7 and 3.8. - 3.5 Pro was still described as "slated for June 2026"; it has not shipped as of 2026-09-03. - Deprecation table gains the 3.x preview rows and the 2026-10-02 shutdown of `gemini-2.5-flash-image` (the `nano-banana` alias target). - Two docstrings still named `gemini-3-flash-preview` as the default. ## 2026-07-29 ### Fixed — nested Pydantic models raised a bare HTTP 400 - `_pydantic_to_schema()` returned pydantic's `model_json_schema()` nearly verbatim, which emits `$defs` + `$ref` for every nested model. Gemini's `responseSchema` does not support `$ref`/`$defs`, so **any** model containing another model (`list[Finding]`, the common case) produced a 400 with no usable body, burned three retries, and returned `None`. Flat single-level models worked, which is why this went unnoticed. Refs are now inlined and the keywords Gemini rejects are stripped recursively rather than only at the top level (`title`, `default`, `additionalProperties`, `discriminator`, `examples`, `const`). - `Optional[X]` / `anyOf` is now translated to the non-null branch plus `nullable: true` instead of being passed through as `anyOf`, which Gemini also rejects. - 4xx responses now raise `_NonRetriableAPIError` carrying the **response body**. `raise_for_status()` discarded it, and Gemini puts the only useful diagnostic there — which schema keyword it refused. 4xx is deterministic, so it no longer burns the retry budget either. - `invoke_with_structured_output()` gained `max_output_tokens` (default 32768). Thinking tokens count against the output budget, so a small cap truncated the JSON mid-object and surfaced as a pydantic `EOF while parsing` that reads like a schema error. `finishReason=MAX_TOKENS` is now detected and reported as truncation. ## 2026-07-21 ### Added — audio (and video) input - `image_path` now accepts **any** supported media file, not just images. Audio input verified working 2026-07-21 (3-beep WAV; `gemini-3.6-flash` returned the correct count and pitch direction). The param keeps its legacy name for back compat; docstrings now state what it really accepts. - Explicit mimeType overrides for audio/video extensions `mimetypes` guesses wrongly or misses (.m4a/.aac/.flac/.ogg/.opus/.mp3/.wav/.aiff/.mp4/.mov/.webm/ .heic/.heif). A wrong guess previously sent a bad mimeType and produced a confused answer rather than an error. - New `MediaInputError` (subclass of `ValueError`) for deterministic input problems; retry loops re-raise it immediately instead of burning 3 attempts. - 15MB inline cap enforced with an actionable message. Larger files need the Files API, which this client still does not implement. - The direct google-generativeai SDK fallback now rejects non-image media with a clear message instead of an opaque PIL `UnidentifiedImageError`. Audio/video require the CF Gateway path. **Routing note:** send audio to `gemini-3.6-flash`. `gemini-3.5-flash-lite` was unreliable on the same clip — it reported three beeps then two on identical input, and got the pitch direction wrong both times. ### ⚠️ BREAKING — `lite` alias repointed - `MODEL_ALIASES['lite']`: `gemini-2.5-flash-lite` → `gemini-3.5-flash-lite`. Output cost goes from $0.40 to $2.50 per 1M (~6x) for any caller using `model="lite"`. Pin `gemini-2.5-flash-lite` by full ID if you need the old rate, though it is deprecated (below). - The Gemini 2.5 **text** generation is retired from routing: `gemini-2.5-flash`, `gemini-2.5-flash-lite`, `gemini-2.5-pro`. Model IDs stay callable and the `stable-flash` / `stable-pro` aliases still resolve, so nothing hard-breaks, but they are no longer recommended targets. Image model `nano-banana` (gemini-2.5-flash-image) is NOT affected. - Added `gemini-3.5-flash-lite` (GA 2026-07-21) as the cheap/bulk tier. - Gemini 3.6 Flash (`gemini-3.6-flash`) reached GA (2026-07-21). Added it to the model registry and made it the new `DEFAULT_MODEL`. - Repointed the `flash` alias from `gemini-3.5-flash` to `gemini-3.6-flash`. Added `flash-3.5` as a stable handle for the prior frontier Flash; `flash-3` still pins the older `gemini-3-flash-preview`. - Rationale: 3.6 Flash is ~half Sonnet's cost (in $1.50 / out $7.50 vs ~$3 / ~$15) and improves coding/agentic quality with ~17% fewer output tokens than 3.5 Flash — making it the default for sub-agent delegation. - Updated SKILL.md and references/models.md tables (pricing, 64K output cap, benchmark deltas, tone-regression caveat). - Not touched: `gemini-3.5-flash-lite` / `gemini-3.5-flash-cyber` (shipped same day) are not yet aliased; two helper fns still hardcode a `gemini-3-flash-preview` default in their signatures (pre-existing). ## 2026-05-28 - Nano Banana 2 (`gemini-3.1-flash-image-preview`) and Nano Banana Pro (`gemini-3-pro-image-preview`) reached GA on Vertex / Gemini Enterprise. - Kept the `-preview` model IDs: the Gemini Developer API surface this client uses still serves both under `-preview` (GA IDs without the suffix are Vertex-only and 404 here). Verified against the live image-generation docs. - Documented capabilities: 512/1K/2K GA + 4K preview, up to 14 reference images, Search + Image-Search grounding (3.1 Flash), thinking_level control. - Noted video-as-input is a Vertex preview only; not available on the Developer API. All notable changes to the `invoking-gemini` skill are documented in this file. The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.0.0/). ## [0.9.0] - 2026-09-24 ### Other - invoking-gemini 0.9.0: add Gemini 3.8 Flash TTS (speech generation) ## [0.8.0] - 2026-09-03 ### Other - invoking-gemini: Gemini 3.8 Flash is the default, add 3.7, retire Pro from routing, handle minimal-thinking 400 (#783) ## [0.7.0] - 2026-07-23 ### Other - invoking-gemini: default to gemini-3.6-flash, retire Gemini 2.5 text models (#741) - invoking-gemini: Nano Banana 2/Pro GA — keep -preview IDs on Developer API surface (#676) ## [0.6.0] - 2026-05-23 ### Added - add Gemini 3.5 Flash + thinking_level, fix stale model docs (#669) ### Fixed - retry on egress-proxy 503 ('DNS cache overflow') in remembering + invoking-gemini (#580) ### Other - Remove _MAP.md files, direct agents to tree-sitting for code navigation (#545) ## [0.5.0] - 2026-03-31 ### Added - surface Image Generation, add examples (#520) - add mapping-features skill for behavioral web app documentation (#432) ### Other - Regenerate _MAP.md files after @lat: backlink insertion (#504) - Lattice v2: bidirectional source-anchored knowledge graph (#503) ## [0.3.1] - 2026-03-01 ### Fixed - use camelCase keys for Gemini REST API inline data ## [0.3.0] - 2026-03-01 ### Added - add image generation support + fix IMAGE_MODELS registry ## [0.3.0] - 2026-03-01 ### Added - `generate_image()` function for native image generation via Gemini image models - `image` and `image-pro` model aliases for image generation - `nano-banana` (gemini-2.5-flash-image) to IMAGE_MODELS registry - Image generation documentation in SKILL.md with prompt patterns and examples ### Fixed - IMAGE_MODELS registry now maps display names to actual API model IDs (was mapping names to themselves, causing 404 errors) - `nano-banana-2` → `gemini-3.1-flash-image-preview` - `nano-banana-pro` → `gemini-3-pro-image-preview` - `nano-banana` → `gemini-2.5-flash-image` ## [0.2.0] - 2026-03-01 ### Added - update invoking-gemini model registry to current Gemini lineup ## [0.1.0] - 2026-03-01 ### Added - route invoking-gemini through Cloudflare AI Gateway - add line numbers, markdown ToC, and other files listing - add code maps and CLAUDE.md integration guidance - Delete VERSION files, complete migration to frontmatter - Migrate all 27 skills from VERSION files to frontmatter ### Changed - migrate API credential management to project knowledge files ### Fixed - limit markdown ToC to h1/h2 headings only -
README.md 239 B
# invoking-gemini Invokes Google Gemini models for structured outputs, multi-modal tasks, and Google-specific features. Use when users request Gemini, structured JSON output, Google API integration, or cost-effective parallel processing. -
SKILL.md 16.1 KB
--- name: invoking-gemini description: Invokes Google Gemini models for structured outputs, image generation, text-to-speech narration, multi-modal tasks, and Google-specific features. Use when users request Gemini, image generation, Gemini TTS or a synthesized voice, structured JSON output, Google API integration, or cost-effective parallel processing. metadata: version: 0.9.0 --- # Invoking Gemini Delegate tasks to Google's Gemini models when they offer advantages over Claude. ## When to Use Gemini **Image generation:** - Blog header images, illustrations, diagrams - Style-guided image creation (risograph, editorial, etc.) - Text rendering in images **Speech (TTS):** - Narration, voice-over, read-aloud with style direction per line - A custom voice designed from a written description - Two-speaker dialogue **Structured outputs:** - JSON Schema validation with property ordering guarantees - Pydantic model compliance - Strict schema adherence (enum values, required fields) **Cost optimization:** - Parallel batch processing (Gemini 3 Flash is lightweight) - High-volume simple tasks **Multi-modal tasks:** - Image analysis with JSON output - Video processing - Audio transcription with structure ## Setup ```bash uv pip install requests pydantic ``` **Credentials — Option A (recommended): Cloudflare AI Gateway** Source `/mnt/project/proxy.env` with `CF_ACCOUNT_ID`, `CF_GATEWAY_ID`, `CF_API_TOKEN`. Requests route through Cloudflare AI Gateway, bypassing IP blocks. Google API key stored in gateway via BYOK. **Credentials — Option B: Direct Google API** If no `proxy.env`, falls back to direct: `GOOGLE_API_KEY.txt` or `API_CREDENTIALS.json`. ## Image Generation Generate images using Gemini's native image models. This is the primary way to create illustrations, blog headers, diagrams, and visual content. ### Quick Start ```python import sys sys.path.append('/mnt/skills/user/invoking-gemini/scripts') from gemini_client import generate_image # One call — returns {"path": "...", "caption": "..."} or None result = generate_image("A watercolor painting of a mountain lake at sunset") print(result["path"]) # /mnt/user-data/outputs/gemini_image_1740000000.png ``` ### Function Signature ```python generate_image( prompt: str, # The image description output_path: str = None, # Auto-generates if omitted model: str = "nano-banana-2", # Default: fast. Use "image-pro" for quality temperature: float = 0.7, # 0.5-0.7 for diagrams, 0.7-0.8 for illustrations ) -> dict | None # Returns: {"path": "/mnt/user-data/outputs/gemini_image_*.png", "caption": str|None} # Returns None on failure ``` ### Model Selection | Alias | Model | Best For | Cost/image | |-------|-------|----------|------------| | `"nano-banana-2"` or `"image"` | gemini-3.1-flash-image-preview | Fast iteration, drafts | $0.067 | | `"image-pro"` or `"nano-banana-pro"` | gemini-3-pro-image-preview | Published content, text rendering | $0.134 | ### Complete Blog Header Example ```python import sys sys.path.append('/mnt/skills/user/invoking-gemini/scripts') from gemini_client import generate_image # 1. Compose prompt with style prefix + subject style_prefix = ( "Style: Risograph-inspired editorial illustration. " "Visible halftone dot texture and slight color misregistration between layers. " "Limited ink palette: deep indigo, warm coral, and sage green on off-white paper. " "Layered transparency where colors overlap creates rich secondary tones. " "Modern and professional — the aesthetic of an indie design studio, not a fantasy novel. " "Generous whitespace. No photorealism, no glow effects, no cyberpunk. No text or labels." ) subject = "A raven perched on a stack of books, observing a network graph" prompt = f"{style_prefix}\n\nSubject: {subject}. Wide landscape format, suitable as a blog header." # 2. Generate (use image-pro for published content) result = generate_image(prompt, model="image-pro", temperature=0.75) if result: print(f"Saved: {result['path']}") # 3. Present to user # present_files([result["path"]]) ``` ### Prompt Patterns - **Style prefix + subject**: Prepend a style description, then describe the subject - **Be specific about style**: "Risograph-inspired editorial illustration" not "a nice picture" - **Include composition**: "Wide landscape format" / "centered, high contrast" - **Text rendering**: "A poster with the text 'SALE' in bold red letters" (works well with image-pro) - **Negative constraints**: "No photorealism, no glow effects" to avoid defaults ### Custom Output Path ```python result = generate_image( "A logo for a coffee shop called 'Bean There'", output_path="/mnt/user-data/outputs/coffee_logo.png" ) ``` ## Speech Generation (TTS) Gemini 3.8 Flash TTS and Flash-Lite TTS went GA on 2026-09-23. Output is WAV, 24 kHz mono 16-bit, SynthID-watermarked. ```python from gemini_client import generate_speech, design_voice, list_voices r = generate_speech("Odin kept two ravens. <short pause> Huginn was thought.", output_path="/tmp/line.wav", voice="Algenib", style="quiet and dry, unhurried") # {'path': '/tmp/line.wav', 'seconds': 4.2, 'audio_tokens': 135} or None v = design_voice("A low, dry, quietly amused male voice with a faint rasp. " "Soft southern British accent.", "narrator", gender="male", language_code="en-GB") # {'id': 'voice_...', 'sample_path': ...} generate_speech("...", voice=v["id"]) lows = list_voices(gender="male", pitch="low") # library of 2,089 prebuilt voices ``` - **Voices:** 30 studio voices (`Charon`, `Kore`, `Algenib` "gravelly", `Enceladus` "breathy", ...) plus 2,059 persona voices with ids like `en-gb-storyteller-4`. `list_voices()` returns accent, pitch, gender and a description for each; the `accent` filter needs the exact string ("Winchester English"), so filter accents on the returned field. - **Style:** pass `style=` (a `speech_metadata` annotation). Do not prefix the text with "Style: ..." — the 3.8 models read the prefix aloud, and `systemInstruction` is rejected. Inline events go in the text: `<laugh>`, `<sigh>`, `<breath>`, `<short pause>`; CAPITALS stress a word. - **Designed voices** are stored (1-year expiry, 200 per project). The description sets baseline delivery too: "thoughtful pauses" in it produced 1–1.8 s mid-line pauses that no per-line style removed. - **The model can change words.** Seen in a 29-line narration: "Hmm, I get things wrong", "tell them" for "tell him". Anything with subtitles or a fixed script needs an ASR check (faster-whisper `medium.en`) and a retake. - **Cost:** about 32 audio tokens per second of speech, $9/M through 2026-12-31 on 3.8 Flash TTS ($6/M Lite), so a minute is about $0.02. - Voice replication (cloning from a 30 s sample plus a recorded consent clip) is not wired in, and is unavailable in the EEA, UK, Switzerland, India, Illinois and Texas. ## Basic Text Usage ```python import sys sys.path.append('/mnt/skills/user/invoking-gemini/scripts') from gemini_client import invoke_gemini response = invoke_gemini( prompt="Explain quantum computing in 3 bullet points", model="flash", # gemini-3.8-flash (default) ) print(response) ``` ## Structured Output Use Pydantic models for guaranteed JSON Schema compliance: ```python from gemini_client import invoke_with_structured_output from pydantic import BaseModel, Field class BookAnalysis(BaseModel): title: str genre: str = Field(description="Primary genre") key_themes: list[str] = Field(max_length=5) rating: int = Field(ge=1, le=5) result = invoke_with_structured_output( prompt="Analyze the book '1984' by George Orwell", pydantic_model=BookAnalysis ) print(result.title) # "1984" ``` **Nested models are supported.** Gemini's `responseSchema` rejects `$ref`/`$defs`, which pydantic emits for every nested model, so the client inlines them before sending: ```python class Finding(BaseModel): claim: str confidence: Literal["high", "medium", "low"] note: str | None = None class Analysis(BaseModel): findings: list[Finding] # nested — inlined for you gaps: list[str] ``` **Budget output generously.** Thinking tokens count against `max_output_tokens` (default 32768). Too low and the JSON truncates mid-object, which surfaces as a pydantic parse error rather than a length error — the client now detects `finishReason=MAX_TOKENS` and says so explicitly. ## Parallel Invocation ```python from gemini_client import invoke_parallel results = invoke_parallel( prompts=["Summarize Hamlet", "Summarize Macbeth", "Summarize Othello"], model="lite", # gemini-3.5-flash-lite — cheap/fast tier for batch ) ``` ## Available Models The current frontier Flash is **gemini-3.8-flash** (GA 2026-09-02), the default and the `flash` alias. Google shipped three Flash generations in six weeks: 3.6 (2026-07-21), 3.7 (2026-08-13), 3.8 (2026-09-02). Each stays callable under a pinned alias (`flash-3.7`, `flash-3.6`, `flash-3.5`, `flash-3`), and none has a shutdown date. `gemini-3.1-flash-lite-preview` from earlier docs is gone (shut down 2026-05-25). The Pro tier is off routing. `gemini-3.1-pro-preview` costs 2.7× the input and 3.2× the output of 3.8 Flash at today's rates and loses to the 3.5+ Flash line on the coding and agentic benchmarks that matter here. Do not target it; the `pro` alias now resolves to gemini-3.8-flash, and "maximum reasoning" means `thinking_level='high'` on Flash. ### Text / Reasoning Models | Model | Alias | Input/1M | Output/1M | Context | Notes | |-------|-------|----------|-----------|---------|-------| | gemini-3.8-flash | `flash` | $0.75 → $1.50 | $3.75 → $7.50 | 1M in / 64K out | **Default.** GA 2026-09-02. Current frontier Flash. Vs 3.7: Terminal-Bench 2.1 90.8% vs 81.6%, SWE-Bench Pro 61.6% vs 60.4%, SWE-Atlas 51.9% vs 48.0%, HLE flat (45.4% vs 45.7%). Google says it "works harder" at higher effort, so expect more thinking tokens per task. `thinking_level` is low/medium/high only — `minimal` returns HTTP 400 and the client downgrades it to `low`. Default `medium` spent 79 thinking tokens on a one-word reply (measured 2026-09-03); pass `low` for non-reasoning tasks. | | gemini-3.7-flash | `flash-3.7` | $0.75 → $1.50 | $3.75 → $7.50 | 1M / 64K | GA 2026-08-13. DeepSWE v1.1 65.3% vs 49.0% on 3.6, Terminal-Bench 2.1 85.8%. Same `minimal` restriction as 3.8. Google keeps it "fully supported for efficiency-first workloads". | | gemini-3.6-flash | `flash-3.6` | $0.75 → $1.50 | $3.75 → $7.50 | 1M / 64K | GA 2026-07-21. ~17% fewer output tokens than 3.5 Flash. Last Flash that accepts `thinking_level='minimal'` (verified 2026-09-03). | | gemini-3.5-flash | `flash-3.5` | $1.50 | $9.00 | 1M | GA 2026-05-19. Google's model list now labels it "legacy". Accepts `minimal`. Costs more on output than 3.6–3.8. | | gemini-3-flash-preview | `flash-3` | $0.30 | $2.50 | 1M | Older preview Flash, kept for back compat. Google's listed migration target for it is gemini-3.6-flash; no shutdown date. | | ~~gemini-3.1-pro-preview~~ | — | $2.00 (≤200K) / $4.00 | $12.00 / $18.00 | 1M | **DEPRECATED from routing (2026-09-03).** Price/quality dominated by 3.6+ Flash; 3.5 Flash already beat it on most coding/agentic benchmarks. ID stays callable for pinned code. `pro` now resolves to gemini-3.8-flash. 3.5 Pro was announced at I/O 2026-05-19 for June and is still absent from the API as of 2026-09-03; it gets the same price/quality test before any alias points at it. | | gemini-3.5-flash-lite | `lite` | $0.30 | $2.50 | 1M | **Cheap/bulk tier.** GA 2026-07-21. Fastest 3.5-class (350 output tok/sec); beats gemini-3-flash on SWE-Bench Pro and OSWorld-Verified. | | ~~gemini-2.5-flash~~ | `stable-flash` | $0.30 | $2.50 | 1M | **DEPRECATED** — 2025-era generation, do not route here. | | ~~gemini-2.5-flash-lite~~ | — | $0.10 | $0.40 | 1M | **DEPRECATED** — cheaper, but a 2025-era generation. `lite` now resolves to gemini-3.5-flash-lite. | | ~~gemini-2.5-pro~~ | `stable-pro` | $1.25 (≤200K) / $2.50 | $10.00 / $20.00 | 1M | **DEPRECATED** — 2025-era generation, do not route here. | `$0.75 → $1.50` means introductory pricing: Google's pricing page (fetched 2026-09-03) lists 3.6, 3.7 and 3.8 Flash at $0.75 in / $3.75 out through 2026-12-31 and $1.50 / $7.50 from 2027-01-01. Context caching is $0.075 → $0.15; Batch is half of standard. Output prices include thinking tokens. ### Image Models | Model | Alias | Input/1M | Per Image | |-------|-------|----------|-----------| | gemini-3.1-flash-image-preview | `image`, `nano-banana-2` | $0.25 | $0.067 | | gemini-3-pro-image-preview | `image-pro`, `nano-banana-pro` | $2.00 | $0.134 | ### Speech Models | Model | Alias | Input/1M | Output/1M (audio) | Notes | |-------|-------|----------|-------------------|-------| | gemini-3.8-flash-tts | `tts` | $0.50 → $1.00 | $9.00 → $18.00 | GA 2026-09-23. Expressive, 130 languages, 2 speakers. Use `generate_speech()`, not `invoke_gemini()`. | | gemini-3.8-flash-lite-tts | `tts-lite` | $0.50 → $1.00 | $6.00 → $12.00 | GA 2026-09-23. Bulk / read-aloud, 101 languages. Replaces gemini-3.1-flash-tts-preview ($1 / $20). | Speech aliases live in `SPEECH_ALIASES`, not `MODEL_ALIASES`, so a text call can never resolve to an audio model. See [references/models.md](references/models.md) for full details. ### Thinking Budget (Gemini 3.x) Gemini 3.x models reason before responding. The parameter changed in 2026: integer `thinking_budget` is gone; use string `thinking_level` ∈ {`minimal`, `low`, `medium`, `high`}. Default for 3.5–3.8 Flash is `medium`. For transcription / classification / extraction tasks, pass `thinking_level='minimal'` or the model will silently spend output tokens on reasoning (symptom: empty response with `finishReason=MAX_TOKENS`). **3.7 and 3.8 Flash reject `minimal`** with HTTP 400 (`Thinking level MINIMAL is not supported for this model`); `low` is their floor. The client downgrades `minimal` to `low` on those two models and prints a note to stderr, so existing callers keep working. Measured on 3.8 (2026-09-03): `low` spent 0 thinking tokens on a one-word reply, the default `medium` spent 79. On 3.7, `low` still spent 45–88, and a `max_output_tokens=50` call at `low` hit MAX_TOKENS and returned None, so budget output generously there. If a job needs a true no-thinking pass, pin `flash-3.6` or `lite`, which still accept `minimal`. ```python response = invoke_gemini( prompt="Transcribe this image.", model="flash", image_path="/tmp/screenshot.png", max_output_tokens=4000, thinking_level="minimal", # don't burn output budget on reasoning ) ``` ## Error Handling ```python response = invoke_gemini(prompt="...", model="flash") if response is None: print("API call failed — check credentials") result = generate_image("...") if result is None: print("Image generation failed — check credentials or try again") ``` Common issues: Missing API key → see Setup. Rate limit → auto-retries with backoff. Network error → returns None. ## Advanced Features ### Custom Generation Config ```python response = invoke_gemini( prompt="Write a haiku", model="flash", # gemini-3.8-flash temperature=0.9, max_output_tokens=200, top_p=0.95, thinking_level="low", # haiku is short; modest reasoning is fine ) ``` ### Multi-modal Input ```python from pydantic import BaseModel from gemini_client import invoke_with_structured_output class ImageDescription(BaseModel): objects: list[str] scene: str colors: list[str] result = invoke_with_structured_output( prompt="Describe this image", pydantic_model=ImageDescription, image_path="/mnt/user-data/uploads/photo.jpg" ) ``` See [references/advanced.md](references/advanced.md) for more patterns. ## Troubleshooting **"No credentials configured":** Create `/mnt/project/proxy.env` with CF credentials, or add `GOOGLE_API_KEY.txt`. **CF Gateway 401/403:** Verify `CF_API_TOKEN` has AI Gateway permissions. If not using BYOK, add `GOOGLE_API_KEY` to `proxy.env`. **Import errors:** `uv pip install requests pydantic` **Image generation returns None:** Check credentials. If persistent, try `model="nano-banana-2"` (more reliable than image-pro). Check for content policy blocks in error output.
Comments (0)
Sign in to join the conversation.
Reviews (0)
No reviews yet.
No comments yet.