# Reality Fabricator — Complete Technical Documentation > AI Agent Platform & Creative Developer API — 30+ endpoints for images, video, VTuber avatars, segmentation, lip-sync Last updated: 2026-07-27 ## Overview Reality Fabricator is an AI agent platform for building autonomous AI agents that think independently and take real-world actions. It also provides a **Creative Developer API** with 30+ endpoints covering image generation (100+ models, 6 providers), video generation, VTuber avatar creation with expressions, lip-sync talking-head animation, and advanced computer vision (SAM2, GroundedSAM, BiSeNet face parsing, background removal). All accessible via a single API key. Website: https://rfab.ai API Base URL: https://api.rfab.ai Contact: support@rfab.ai --- ## Autonomous AI Agents ### What They Are Autonomous AI agents are self-directed AI entities that run continuously, thinking independently and taking actions without user intervention. Unlike traditional chatbots that only respond when prompted, these agents have their own goals, persistent memories, and can initiate actions on their own schedule. ### Thinking Cycle - **Timed Mode**: Agent thinks every X seconds (configurable: 30s, 1min, 5min, etc.) - **Continuous Mode**: Agent thinks as fast as possible (minimum 3.5s between thoughts) ### 7-Step Process 1. Load agent personality, goals, and memories from database 2. Build prompt with personality, goals, recent memories (last 50 events), available actions 3. Call selected AI model via streaming 4. Stream response to frontend in real-time 5. Parse thought for action commands (CALL_PHONE, SEND_EMAIL, SEARCH_WEB, etc.) 6. Execute detected actions; each action adds results to agent memory 7. Schedule next thinking cycle ### 70+ Real-World Actions **Communication (8)**: - CALL_PHONE — Make real phone calls via Vapi + OpenAI Realtime API - SEND_EMAIL — Send emails via AWS SES - SEND_SMS — Text messages via Twilio - SEND_WHATSAPP — WhatsApp messages via Twilio - SEARCH_EMAIL, SEARCH_SMS, SEARCH_WHATSAPP — Search message history - ASK_USER, SHOW_USER — Request input or display information **Coding & Development (40+)**: - Full IDE capabilities: LSP-powered code understanding, semantic search - Multi-file editing, Git integration, testing, deployment - EDIT_FUNCTION, RENAME_SYMBOL, RUN_TESTS, DEPLOY_WEB_APP, GIT_COMMIT, MULTI_FILE_EDIT - Natural language browser automation via Playwright **Research (3)**: - SEARCH_WEB — Internet search via Tavily API - READ_URL — Read and extract content from web pages - SEARCH_MEMORY — Search agent's own memory store **Memory & Learning (5)**: - ADD_MEMORY, SEARCH_MEMORY, CONSOLIDATE_MEMORY - SET_GOAL, COMPLETE_GOAL **Visual (2)**: - GENERATE_IMAGE — Create images with 8+ models - TAKE_SCREENSHOT — Capture browser screenshots **Integration (10+)**: - ADD_TO_CONTACT_BOOK, SEARCH_CONTACT_BOOK - ADD_TO_CALENDAR, SEARCH_CALENDAR - Discord integration - RF Bridge local desktop integration ### Phone Call Integration Agent outputs CALL_PHONE → Backend calls Vapi API with phone number, voice selection (8 voices: Alloy, Echo, Shimmer, Coral, Sage, Ash, Ballad, Verse), agent personality, and context → Vapi initiates call using OpenAI Realtime API → Voice conversation occurs → Call ends → Transcript added to agent memory → Agent continues thinking. Voice quality: Sub-200ms latency, voice interruption support, emotional nuance. ### Email Integration **Sending**: Agent outputs SEND_EMAIL → AI generates email (recipient, subject, body) → Sent via AWS SES → Stored in agent memory. **Receiving**: Email to agent@rfab.ai → AWS SES → S3 bucket → Lambda function → Backend webhook → Added to agent memory → Agent processes on next cycle. ### Memory System Types: thought, action_result, user_message, call_initiated, call_completed, email_sent, email_received, sms_sent, sms_received, web_search_result, goal_set, goal_completed. Storage: PostgreSQL, indexed by agent ID. Last 50 memories loaded per thinking cycle. Searchable via SEARCH_MEMORY action. --- ## AI Models (15+) Last updated: 2026-07-19. Rates are token multipliers vs. the platform baseline. ### Anthropic Claude (3+) - **Claude Fable 5** — Mythos-class flagship, 1M context, context caching, 4.0x tokens - **Claude Opus 4.7** — 200K context, context caching, 2.04x tokens - **Claude Sonnet 4.6** — 200K context, context caching, 1.26x tokens ### OpenAI (6) - **GPT-5.6 Sol / Terra / Luna** — 1M context, 2.3x / 1.15x / 0.46x tokens - **GPT-5.4** — 1M context, context caching, 1.19x tokens - **GPT-5.4 Mini** — 1M context, context caching, 0.345x tokens - **GPT-4.1** — 1M context, context caching, 0.68x tokens ### Google Gemini (4+) - **Gemini 3.1 Pro Preview** — Multimodal, 1M context, 0.95x tokens - **Gemini 3 Flash Preview** — Multimodal, 1M context, 0.23x tokens (+ web-search variant) - **Gemini 2.5 Pro** — Multimodal, 1M context, 0.725x tokens (+ web-search variant) - **Gemini 2.5 Flash** — Multimodal, 1M context, 0.18x tokens (+ web-search variant) ### xAI Grok (3+) - **Grok 4.5** — 500K context, built-in web search, best for stories, 0.56x tokens - **Grok 4.20** — 2M context, built-in web search, 0.3x tokens - **Grok Build 0.1** — 256K context, fast coding model, 0.24x tokens ### Mistral (3) - **Mistral Medium 3.5** — 128K context, 0.63x tokens - **Mistral Large 2512** — 128K context, 0.16x tokens - **Mistral Small 4** — 128K context, 0.051x tokens ### DeepSeek (3) - **DeepSeek V4 Pro** — 1M context / 384K output, 0.096x tokens - **DeepSeek V4 Flash** — 1M context / 384K output, 0.017x tokens - **DeepSeek V3.2** — 131K context, 0.025x tokens ### Free Models (0x tokens) - **Llama 3.3 70B**, **Nemotron 3 Super 120B**, **Gemma 4 31B**, **MythoMax 13B** ### Alloy Mode Multi-model intelligence pools. Agents cycle through multiple models (e.g., Claude Sonnet 4.6 → Grok 4.5 → Gemini 3 Flash) with automatic exclusion of failing models and round-robin rotation. Only Reality Fabricator offers this. ### Image Generation (8+) - Flux Schnell, Flux Dev, Flux Pro - SDXL (Stable Diffusion XL) - Ideogram V3 (text-in-images) - Anime Art Diffusion - Realistic Vision V5 - Dreamshaper --- ## Pricing - **Signup**: 50,000 free tokens (no credit card required) - **Daily**: 25,000 free tokens for all users, every day - **Free models**: Llama 3.3 70B, Nemotron 3 Super, Gemma 4 31B, MythoMax 13B — 0x tokens - **Paid tokens**: Transparent cost-based pricing (API cost + minimal markup) - **Crypto**: BTC, ETH, USDC accepted via Coinbase Commerce ### Token Rates - Budget (0.02x–0.25x): DeepSeek V4, Gemini 3 Flash, Mistral Large 2512, Grok Build - Standard (0.3x–1.3x): Grok 4.5, GPT-5.4, Gemini 3.1 Pro, Claude Sonnet 4.6 - Premium (2x–2.5x): Claude Opus 4.7, GPT-5.6 Sol - Elite (4.0x): Claude Fable 5 (Mythos-class) - Free (0x): Llama 3.3 70B, Nemotron 3 Super, Gemma 4 31B, MythoMax 13B --- ## Technical Architecture ### Backend - Node.js v22.x, Express 5 - PostgreSQL with Prisma ORM - Redis (caching, sessions, event streams) - WebSocket streaming (express-ws) - JWT + Google OAuth + Passport.js authentication - Event sourcing architecture ### Frontend - Angular 17, TypeScript - NgRx state management (Redux pattern) - IndexedDB with AES-256-GCM client-side encryption - PWA with Service Workers and offline support - PrimeNG UI components ### External Integrations - **Phone**: Vapi + OpenAI Realtime API - **Email**: AWS SES + Lambda + S3 - **SMS/WhatsApp**: Twilio - **Web Search**: Tavily API - **Image Generation**: Replicate, ModelsLab, AWS Bedrock - **Payments**: Stripe + Coinbase Commerce - **Voice**: Deepgram STT/TTS ### Security - AES-256-GCM encrypted local storage - JWT tokens with secure httpOnly cookies - TLS 1.3 for all traffic - Helmet.js security headers - Per-user and per-IP rate limiting - Comprehensive input validation and sanitization - GDPR compliant --- ## Language Support 6 fully supported languages with complete UI translation and native AI responses: - English (en) - Spanish (es) - Portuguese (pt) - French (fr) - German (de) - Polish (pl) --- ## Platform Features ### Plain Language Action Chaining Agents chain actions through natural reasoning without coding: SEARCH_WEB → CONSOLIDATE_MEMORY → SEND_EMAIL No scripts, JSON, or configuration required. ### ReZero Context Compression AI-powered context compression that summarizes conversation history, enabling infinite-length conversations without memory loss or context window limits. Works together with Story Memory (below): the adventure's memory book carries over to the continued adventure, so granular facts survive compression verbatim. ### Story Memory & Lorebooks (World Guide + Adventure Journal) Persistent, player-owned memory for AI adventures — the equivalent of SillyTavern/NovelAI/AI Dungeon lorebooks and World Info, plus an automatic write path: - **World Guide**: Adventure authors ship a lorebook with a published scenario; players adopt it automatically. Large guides are chunked and embedded, and each turn the engine retrieves only the most relevant entries within a token budget (semantic RAG, not keyword triggers). Guides can be built from uploaded documents. - **Story Memory**: During play, the narrative engine records short durable facts — names, debts, promises, relationship changes, open plot threads — into a per-adventure memory book. A fact written on one turn is retrievable the next. Engines can only write to memory slots they declare, and only with explicit player consent. - **Supersede, not accumulate**: A note that updates an old fact retires the stale entry from retrieval instead of stacking contradictory duplicates; newer facts outrank older ones when they conflict. - **Timeline-consistent**: Rerolling a response, deleting a message, or editing an earlier message (restarting the story from that point) automatically retires notes written by the abandoned timeline and restores any older fact a retired note had displaced. The AI never remembers events that no longer happened. - **Adventure Journal**: Every remembered fact is visible in a chronological, provenance-stamped journal — review, delete, undo, or add manual notes. Returning players get a "Previously on…" recap assembled from recent entries. - **Pinned memories**: Any journal entry (or in-flow recall chip) can be pinned "never forget" — pinned entries bypass ranked retrieval and are always injected first within the slot budget. - **Author's Note**: A per-story steering note (Story Setup, game sidebar) — the player's own instruction on tone, pacing, or style, injected verbatim at the end of the prompt every turn. Blank = nothing injected. The equivalent of AI Dungeon's Author's Note; complements the account-wide Player Persona (who you play as). - **Survives ReZero**: Memory lives in a book, not the context window — facts survive arbitrarily many context compressions. - **Shareable memory schemas**: Engine authors can ship a memory scaffold (e.g. "## Characters / ## Debts & Promises / ## Open Threads") that pre-structures the memory book; memory-capable engines carry a 📝 Memory badge. ### RF Bridge Desktop companion app (Windows/macOS) that connects local AI models (via Ollama) and local file system access to the platform. ### Self-Hosting Full Docker support for self-hosting. Open architecture with JSON-configurable narrative engine. --- ## Developer API — Complete Referencecontinue Reality Fabricator provides a comprehensive REST API with 30+ endpoints for AI-powered creative media. One API key gives access to everything: image generation, video, VTuber avatars, computer vision, lip-sync, and more. All through a simple REST interface designed for AI agents, scripts, and integrations. ### Quick Start (3 minutes) 1. **Sign up free** at https://rfab.ai (Google login, no credit card required) 2. You receive **50,000 free tokens** instantly + **25,000 tokens daily** (no expiry) 3. Go to **Settings → Developer API** section 4. Click **Create Key**, give it a name (e.g. "Cascade", "My Game Pipeline") 5. Click the key card to view your full key anytime — copy it 6. Use the key in the `X-API-Key` HTTP header on any API endpoint ### Authentication All API requests require the `X-API-Key` header. The key format is `rfab_...`: ```bash curl -X POST https://api.rfab.ai/api/image-generation/generate \ -H "X-API-Key: rfab_your_key_here" \ -H "Content-Type: application/json" \ -d '{"prompt": "a cyberpunk cityscape at sunset", "model_id": "flux-schnell"}' ``` ```python import requests API_KEY = "rfab_your_key_here" BASE = "https://api.rfab.ai" HEADERS = {"X-API-Key": API_KEY, "Content-Type": "application/json"} # Generate an image r = requests.post(f"{BASE}/api/image-generation/generate", headers=HEADERS, json={ "prompt": "a cyberpunk cityscape at sunset", "model_id": "flux-schnell" }) print(r.json()["imageUrl"]) # Permanent S3 URL ``` ```javascript // Node.js const API_KEY = "rfab_your_key_here"; const BASE = "https://api.rfab.ai"; const res = await fetch(`${BASE}/api/image-generation/generate`, { method: "POST", headers: { "X-API-Key": API_KEY, "Content-Type": "application/json" }, body: JSON.stringify({ prompt: "a cyberpunk cityscape at sunset", model_id: "flux-schnell" }) }); const { imageUrl } = await res.json(); console.log(imageUrl); // Permanent S3 URL ``` ### Error Handling All errors return JSON with `success: false` and an `error` field: ```json { "success": false, "error": "Invalid or revoked API key", "code": "INVALID_API_KEY" } ``` | HTTP Status | Code | Meaning | |---|---|---| | 400 | `BAD_REQUEST` | Missing or invalid parameters | | 401 | `INVALID_API_KEY` | Key is invalid, revoked, or expired | | 401 | `MISSING_AUTH_TOKEN` | No X-API-Key header provided | | 403 | `ACCOUNT_DISABLED` | Account is banned or inactive | | 403 | `INSUFFICIENT_TOKENS` | Not enough tokens for this operation | | 429 | `RATE_LIMITED` | Too many requests — wait and retry | | 500 | `INTERNAL_ERROR` | Server error — retry with exponential backoff | **Retry strategy for AI agents:** On 429 or 500, wait 2s then retry (max 3 attempts). On 403 insufficient tokens, stop and notify user. ### API Key Management (Programmatic) Manage API keys via the REST API itself (accepts JWT auth or an existing API key): - `POST /api/user/api-keys` — Create key: `{ "name": "My Key", "scopes": [...], "expiresInDays": 90 }` - `GET /api/user/api-keys` — List all keys with full decrypted values - `DELETE /api/user/api-keys/:id` — Revoke a key Default scopes: `image:generate`, `image:list-models`, `image:img2img`, `image:inpaint`, `video:generate`, `phone:call` --- ### 1. Image Generation (100+ models, 6 providers) #### Text-to-Image `POST /api/image-generation/generate` Request: ```json { "prompt": "pixel art spaceship, transparent background, 64x64", "negativePrompt": "blurry, low quality", "modelId": "flux-schnell", "width": 512, "height": 512, "imageCount": 1, "steps": 20, "guidanceScale": 7.5, "seed": null } ``` **Style parameter (recommended for AI agents):** Instead of choosing a `modelId`, pass a `style` string and the API auto-selects the best model: ```json { "prompt": "retro spaceship cockpit, CRT monitors, 80s sci-fi", "style": "pixel-art" } ``` Available styles: `pixel-art`, `retro`, `game-assets`, `8bit`, `16bit`, `anime`, `anime-nsfw`, `photorealistic`, `realistic`, `artistic`, `fantasy`, `cartoon`, `instruction-following`, `precise`, `text-rendering`, `premium`, `fast`, `budget`, `furry`, `character-consistency` The `style` parameter is ignored if `modelId` is also provided. **alphaRegion parameter (transparent cutouts):** Generate images with a transparent rectangular region cut out. Essential for game UI frames, HUD overlays, cockpit viewports, etc.: ```json { "prompt": "sci-fi cockpit frame, dark metal, CRT monitors, retro 80s", "style": "pixel-art", "width": 1280, "height": 800, "alphaRegion": { "x": 0.05, "y": 0.05, "width": 0.7, "height": 0.85, "cornerRadius": 0.02, "feather": 0.01 } } ``` - Coordinates can be **normalized (0-1)** or **pixel values**. If all four values are ≤1, they're treated as normalized. - `cornerRadius`: Optional rounded corners for the cutout (normalized or pixels). Default: 0 - `feather`: Optional edge softening/blur (normalized or pixels). Default: 0 - Response includes `alphaRegionApplied: true` and the `imageUrl` is a PNG data URI with transparency - **Use case:** Generate a cockpit frame with a transparent viewport, then layer it over game content Response: ```json { "success": true, "images": ["https://rfab-media.s3.amazonaws.com/.../image.png"], "imageUrl": "https://rfab-media.s3.amazonaws.com/.../image.png", "modelUsed": "replicate:black-forest-labs/flux-schnell", "prompt": "pixel art spaceship, transparent background, 64x64", "seed": 12345, "generationTime": 3.2, "provider": "replicate", "savedImageId": "img_..." } ``` **All returned S3 URLs are permanent and publicly accessible.** You can download, display, or pass them to any other endpoint. #### Image-to-Image `POST /api/image-generation/img2img` Transform an existing image with a prompt. Use for style transfer, refinement, or variation. Request: ```json { "imageUrl": "https://... (source image URL or data URI)", "prompt": "convert to watercolor painting style", "modelId": "flux-schnell", "strength": 0.7, "width": 512, "height": 512, "steps": 20, "guidanceScale": 7.5, "negativePrompt": "blurry", "characterReferenceUrl": "https://... (optional: reference image for character consistency)" } ``` **Base64 alternative:** Instead of `imageUrl`, pass `imageBase64` with a raw base64 string or data URI. Useful when the source image exists only in memory (e.g. programmatically generated masks, screenshots): ```json { "imageBase64": "data:image/png;base64,iVBORw0KGgo...", "prompt": "convert to pixel art style" } ``` - `strength` (0.0–1.0): How much to transform. 0.3 = subtle changes, 0.9 = near complete regeneration. Default: 0.7 - `characterReferenceUrl`: Optional — provide a character reference image for style/identity consistency Response: Same format as text-to-image. #### Inpainting `POST /api/image-generation/inpaint` Edit specific regions of an image using a mask. The masked area is regenerated according to the prompt. Request: ```json { "imageUrl": "https://... (source image)", "maskUrl": "https://... (black & white mask — white = area to inpaint)", "prompt": "a glowing magical sword", "modelId": "replicate:black-forest-labs/flux-fill-dev", "strength": 0.85, "steps": 25, "guidanceScale": 7.5, "seed": null } ``` **Base64 alternative:** Pass `imageBase64` and/or `maskBase64` instead of URLs. This is essential for programmatic workflows where you generate masks in code (e.g. with PIL/Pillow or Canvas) without uploading them first: ```json { "imageBase64": "data:image/png;base64,iVBORw0KGgo...", "maskBase64": "data:image/png;base64,iVBORw0KGgo...", "prompt": "transparent viewport hole, empty space" } ``` - `maskUrl` / `maskBase64`: White pixels = area to regenerate, black pixels = area to keep - Generate masks programmatically using the segmentation endpoints below, or create them in code - Accepts raw base64 strings or full `data:image/png;base64,...` data URIs Response: Same format as text-to-image. #### AI Prompt Enhancement `POST /api/image-generation/enhance-prompt` Use AI to improve a rough prompt into a detailed, optimized prompt for better generation results. Request: ```json { "prompt": "a knight", "style": "fantasy art" } ``` Response: ```json { "success": true, "enhancedPrompt": "a noble knight in gleaming silver plate armor, standing heroically atop a windswept cliff, dramatic sunset lighting, detailed fantasy art style, volumetric fog, epic composition" } ``` #### List Models - `GET /api/image-generation/models` — Returns all available image models with capabilities, tags, and style mappings - `GET /api/image-generation/img2img-models` — Models that support image-to-image Response (models): ```json { "success": true, "models": [ { "model_id": "pixel-art-diffusion-xl", "name": "Pixel Art XL", "description": "Dedicated SDXL checkpoint for pixel art style images and sprites", "category": "Image Generation", "nsfw_capable": true, "recommended": false, "tags": ["pixel-art", "retro", "game-assets", "sprites", "8bit", "16bit", "indie-game"], "best_for": "Retro pixel art, game sprites, 8/16-bit style assets. Best model for indie game pixel art.", "supportsImg2Img": true, "maxWidth": 768, "maxHeight": 768, "capabilities": { "negativePrompt": true, "size": true, "steps": true, "guidanceScale": true, "scheduler": true } } ], "availableStyles": [ { "style": "pixel-art", "model_id": "pixel-art-diffusion-xl" }, { "style": "photorealistic", "model_id": "runware:civitai:133005@782002" }, { "style": "anime", "model_id": "nuke-colormax-anime" }, { "style": "instruction-following", "model_id": "replicate:google/nano-banana-pro" } ] } ``` **AI Agent model selection guide:** - **Don't know which model to use?** Pass the `style` parameter on `/generate` instead of `modelId` - **Need pixel art / game sprites?** → `style: "pixel-art"` or `modelId: "pixel-art-diffusion-xl"` - **Need precise instruction following?** → `style: "instruction-following"` or `modelId: "replicate:google/nano-banana-pro"` - **Need character consistency across images?** → `style: "character-consistency"` or `modelId: "replicate:black-forest-labs/flux-kontext-pro"` - **Need fast/cheap iteration?** → `style: "fast"` or `modelId: "atlascloud:flux-schnell"` - **Need photorealism?** → `style: "photorealistic"` or `modelId: "runware:civitai:133005@782002"` - **Filter models by tag:** Use the `tags` array in the response to programmatically find models matching your needs #### Providers & Models | Provider | Key Models | Best For | |----------|-----------|----------| | **Replicate** | Nano Banana Pro (best instruction following), Flux Kontext Pro (character consistency), Flux Dev (quality), GPT Image 2 (text rendering), SDXL | General purpose, precise compositions | | **ModelsLab** | 100+ community models: anime, realistic, pixel art, niche | Specialized styles, pixel-art-diffusion-xl | | **Runware** | FLUX.2 Pro (#1 ranked), Pony Diffusion XL, Juggernaut XL, DreamShaper XL, CivitAI models | Premium quality, anime, photorealism | | **Atlas Cloud** | Flux Schnell (cheapest), Flux Dev (uncensored), Wan 2.7 | NSFW-capable, budget batch generation | | **Grok (xAI)** | Grok Imagine, Grok Imagine Quality 2K | Fast, good text rendering, budget | | **AWS Bedrock** | Amazon Nova Canvas | Enterprise reliability | #### Model ID Format - `pixel-art-diffusion-xl` — ModelsLab pixel art model - `replicate:owner/model-name` — Full Replicate model path - `runware:civitai:ID@VERSION` — CivitAI models via Runware - `atlascloud:model-name` — Atlas Cloud uncensored models - `grok:model-name` — xAI Grok models - `amazon.nova-canvas-v1:0` — Bedrock - Call `GET /models` to discover all available model IDs with tags and best_for descriptions #### Automatic Failover If a provider is unavailable, the API automatically falls back: Replicate → ModelsLab → AWS Bedrock (Nova Canvas). No client-side retry logic needed. --- ### 2. Video Generation #### Image-to-Video `POST /api/image-generation/generate-video` Animate a still image into a short video clip. Supports various motion styles. Request: ```json { "imageUrl": "https://... (source image URL)", "prompt": "camera slowly zooms in, cinematic motion, subtle animation", "videoModelId": "kling-2.5", "duration": 5, "resolution": "720p", "aspect_ratio": "16:9" } ``` - `duration`: 5 or 10 seconds - `resolution`: "720p" or "1080p" - `aspect_ratio`: "16:9", "9:16", "1:1" - `videoModelId`: "kling-2.5", "seedance-2.0", or Atlas Cloud models Response (video is always asynchronous — you get a job, not a URL): ```json { "success": true, "async": true, "jobId": "a1b2c3d4-...", "status": "processing", "message": "Video generation started. Poll /api/image-generation/job/ for status." } ``` **Polling:** `GET /api/image-generation/job/:jobId` every 3–5 seconds until `status` is `"completed"` (the `result` object then carries the same payload a synchronous response would have, e.g. `result.videoUrl`) or `"failed"` (`error` has the reason). Jobs expire from the poll endpoint 2 hours after creation. Typical video generation takes 30–120 seconds. Image generation supports the same flow when you pass `"async": true` — recommended for slow models like gpt-image-2, which can outlive proxy timeouts on the synchronous path. #### List Video Models `GET /api/image-generation/i2v-models` Returns available video models with resolution and duration capabilities. #### Recipe: Transparent Looping Animation (loading mascots, game sprites, stream overlays) Chain two endpoints + one local ffmpeg command to get a seamless, transparent, looping character animation (alpha-channel `.webm`). This is the exact pipeline behind rfab.ai's own catgirl loading mascots (40+ clips in production). **Step 1 — Generate the character on a solid chroma key** (`POST /api/image-generation/generate`): ```json { "prompt": "Full-body anime character, centered, facing the viewer. Absolutely NOTHING green anywhere on the character. The entire background is one solid flat pure green #00FF00, no gradient, no shadow on the background.", "modelId": "openai:gpt-image-2", "width": 1024, "height": 1536 } ``` Pick a key color that never appears on your character (pure green `#00FF00` suits pink/cyan palettes; use magenta `#FF00FF` for green-heavy characters). **Step 2 — Animate it** (`POST /api/image-generation/generate-video`): ```json { "imageUrl": "(imageUrl from step 1)", "prompt": "She dances energetically in place, staying centered and fully in frame. Camera completely locked, no zoom, no cuts. The solid pure green #00FF00 background stays flat and empty.", "videoModelId": "gemini:gemini-omni-flash-preview", "duration": 4, "audio": false } ``` `gemini:gemini-omni-flash-preview` is the recommended model for loops — it holds a locked camera and a flat background reliably. Always restate the key color and "camera locked" in the motion prompt. **Step 3 — Key it out and make it loop** (local ffmpeg, any modern build): ```bash ffmpeg -i motion.mp4 -filter_complex \ "[0:v]fps=24,scale=-2:400,chromakey=0x00FF00:0.22:0.08,despill=type=green,format=yuva420p,split[a][b];[b]reverse[r];[a][r]concat=n=2:v=1[out]" \ -map "[out]" -c:v libvpx-vp9 -pix_fmt yuva420p -b:v 0 -crf 38 -auto-alt-ref 0 -an out.webm ``` - The forward+reverse ("ping-pong") concat makes any clip loop seamlessly. - `-auto-alt-ref 0` is **required** — without it libvpx silently drops the alpha plane. - Result plays with real transparency in `